← Finance · Black Swans

Finance · Black Swans

The (ab)Normal Normal, Part I: Where Two Giraffes Can Have a Horse

Part 1 of a series on the use, and misuse, of probability distributions in modelling financial data.

All models are wrong, but some are completely wrong!

Mathematics is the universal language of nature. Statistics is the most widely used discipline within mathematics. A statistical model is a description of a system using mathematical concepts. Raw data is boring and confusing. However, raw data presented statistically (and visually) can be interesting, educational, and on occasion, even inspiring. Within the field of statistics, there is no (probability) model more well-known than the Gaussian distribution. In fact, it is so well-known that it is often called the Normal distribution. It is also known as the Bell curve, since it looks like a bell; though other probability distributions, like the Student's-t distribution, also look like a bell.

The Normal distribution is the curve of nature. We see it all over nature (height, blood pressure, particle diffusion, to name a few). If there is only one distribution of nature that people should understand, it is the Normal distribution. It was developed over two centuries, across three European countries, in stages, starting from Jacob Bernoulli (1654 - 1705) to Abraham de Moivre (1667 - 1754) to Pierre Simon Laplace (1749 - 1827) to Adrien-Marie Legendre (1752 - 1833) and finally Johann Carl Friedrich Gauss (1777 - 1855; after whom it is named); eventually maturing into the following equation (known as its probability density function):

Let x be a continuous random variable. x follows a Gaussian distribution of arithmetic mean μ and standard deviation σ its probability density function is given by:

If the elegance of a mathematical equation lies in its simplicity, then it is difficult to find a more elegant equation than the one that describes the Normal distribution. In physics, Einstein presented his theory of Special relativity through the simplicity of two variables—e (energy) and m (mass), combined with one constant c (the speed of light). The Normal distribution provides us with a road map for predicting the future of a random variable x through two parameters, as well—µ (mean) and σ (standard deviation), combined with two constants e (Euler's number) and π (pi).

Graphically, a Normal distribution looks even more attractive than its parent mathematical equation:

−4σ−3σ−2σ−1σμ1σ2σ3σ4σ68.26%95.44% within ±2σ · 99.74% within ±3σ
The Normal distribution curve

A symmetric single-humped curve, centered on its mean/median/mode (all the same), compartmentalized into chunks via standard deviations (describing its width), with most of the observations clustered around the central peak, stretching from -∞ to +∞, which despite its efforts is never quite able to reach the x-axis in either direction (asymptotic to the x-axis). 68.26% of the values of a Gaussian random variable are between a positive standard deviation and a negative standard deviation (also known as sigma) of one (between -1 and 1 on the graph) in relation to its mean. 95.44% of the values of a Gaussian random variable are between two positive standard deviations and two negative standard deviations relative to their mean. 99.74% of the values of a Gaussian random variable fall between three positive standard deviations and three standard deviations. 99.9937% of the values fall within four standard deviations.

Beyond four sigma is where it starts getting interesting. This is the world of tails (of the probability distribution), where we start seeing outliers - extreme values that fall a long way outside of the other observations. Probability distributions are difficult to understand, to begin with. Understanding the tails of probability distributions is even more difficult. Correctly relating and (most importantly) quantifying the impact of the outliers in the tail to events in our daily lives has proven nearly impossible; unless we are dealing with a Normal distribution (as opposed to other distributions, e.g. a Cauchy distribution). This is because, with the Normal distribution, we can (more or less) ignore the outliers, and simply concentrate on the events within the confidence level (defined as the portion of the curve between +/- 3 sigma or +/- 4 sigma). Which is also why enormous intellectual bandwidth has been spent, over the past century, to map not only physical phenomenon but also social phenomenon onto the Normal distribution; even if they do not quite map neatly onto a Normal distribution. If it becomes too difficult to understand (and model) what is really going on, just chop and change it, throw out a couple of irritating outliers (or better yet, assume they will never occur or, even better, if they do occur, assume their impact will be limited) till it all matches a model you already understand (the Normal distribution), and, happily (and rather conveniently) move along from there.

All of this works fine if the outliers are events that actually may never occur and / or even if our model miscalculates their probability of occurrence, their impact is harmless. But what if the outlier is the asteroid that destroyed the dinosaurs 70 million years ago? Or even worse, caused the S&P to crash on Black Monday (the odds of which, as per the Normal distribution financial models, were even more extreme than the dinosaur-eliminating asteroid)? And what if our Normal distribution model has incorrectly estimated its frequency of occurrence at once every 70 million years? What if we correctly calculated its occurrence, but completely miscalculated its impact, assuming it would kill only the dinosaurs on the coastal area of the Gulf of Mexico; not on the whole world? Worst of all, what if we should not have been using the Normal distribution to begin with to model the phenomenon of possible asteroid collisions with earth? Before delving into the details (and shortfalls) of the above approach, it is essential to discuss why we need probability distributions (Normal or otherwise) to begin with.

We need probability distributions because we want to predict the future; specifically in stochastic domains - domains where we cannot define a mathematical equation to calculate a precise answer. The alternative to a stochastic domain is the deterministic domain. In a deterministic domain, the output (answer) is fully determined by parameter values and initial conditions. What will be the exact location of the moon five hundred and two years, four hours, twenty minutes and two seconds? We can predict (calculate) the exact answer to this question via physics equation(s), hence, we do not require a probability distribution. However, certain phenomenon cannot be calculated deterministically. They incorporate inherent randomness. The same set of parameters and initial conditions can result in, not a single exact output, but a set of different possible outputs with certain probabilities (of occurrence). What will be the height of a fully grown male in the USA? This falls in the stochastic domain; modelled elegantly by the Normal distribution centered on a mean of 5'10" with a standard deviation of 4 inches. There is a 68.26% chance the height of the adult male will be between 5'6" and 6'2".

The Normal distribution makes the stochastic world exceedingly simple. We just need to come up with two values - µ and σ -, and throw them into the probability density function to get the Normal distribution graph. How do we get these two values? We get the values from the historical data. There are two issues regarding the historical data that should jump out for everyone in such an approach: the historical data must be accurate and it must be complete. How do we know if the historical data is accurate? We know it is accurate because it is recorded, empirically, that is, we know the height of adult males in the USA. Based on that information, we can model the past into a probability distribution (defining the past). How do we know it is complete, that is, how do we know the historical data available to us covers all possible occurrences? We kind of know, and we kind of don't, that is, for some cases we know and for some we don't; though we easily convince ourselves that we know for most, if not all, cases.

This error in judgement requires understanding the difference between bounded randomness and unbounded randomness. Bounded randomness places a boundary around the possible outcomes. As a simple example, there are only two possible outcomes for a coin toss and only six for a roll of dice. Nature evolves on the basis of bounded randomness with significantly higher outcomes than a coin toss; but bounded, still the same. The offspring of two giraffes will always be a giraffe with some level of randomness due to the combination of the parent giraffe genes. If this were a deterministic process, we would get a clone of the parent giraffes. The randomness allows some level of change (evolution) based on survival of the fittest. However, nature bounds the randomness, that is, the offspring of two giraffes will not, in fact cannot, be a horse. Such domains are stochastic in the physical domain. In this domain, an American adult male, let’s say Bill Gates cannot be 14 miles tall, since all other American adult males have a mean height of 5'10" tall.

On the other hand, there are certain domains that follow unbounded randomness. These are stochastic in the non-physical domain, also known as the informational domain. In this domain the (informational) offspring of two (informational) giraffes can be an (informational) horse. A better description would be that nearly all the (informational) offsprings will be giraffes, and then, completely unexpectedly, a (informational) horse will appear (and the horse could be so large in size that it may completely dominate all the giraffes). Such a horse would actually be a swan; a Black Swan. This is the domain of social sciences (including finance); where Bill Gates can have a net worth of 70 billion dollars, even if everyone else has a mean net worth 210 thousand dollars.

Why not simply analyse historical data, without constructing a probability distribution? This is because our motivation is not that of a historian. Our motivation is that of an oracle. We want to model the past to predict the future. A probability distribution is needed because the future can only be predicted if it is assumed the future operates under a known pattern, that is, a well-defined probability distribution. From where will we obtain such a probability distribution? We cannot get it from the future, since we do not know what will happen in the future (in fact, that is what we are trying to figure out). We (obviously) define the distribution from the past data, extrapolate the probability distribution into the future, and then calculate the future data from the distribution. This leads to a catch-22 situation—the distribution (of the future) is defined by the data (from the past) and the data (for the future) is defined by the distribution (of the past). If one is incorrect (the data or the distribution), the other will be incorrect, automatically. If both are incorrect, the results will be disastrous. To achieve this, we must ensure the probability distribution of the past will always be followed, in an identical manner, in the future. This involves four assumptions:

  1. The data from the past is accurate
  2. The data from the past is complete
  3. The data from the past actually follows the probability distribution that we think it follows
  4. The data from the past will continue to follow the same probability distribution in the future

Based on the above, can we simply assume the Normal distribution for various phenomenon? If not, then how do we decide when we should use the Normal distribution or use another probability distribution to describe a phenomenon? There is a step wise approach that should be followed to reach such a conclusion. As mentioned, the first step is to figure out the accuracy and completeness of the historical data. This is done via the first important theorem of probability: Law of Large Numbers (LLN), which states:

The sample mean (X̄) converges to the true population mean (µ) as the sample size (n) approaches infinity, with probability one.

It is through LLN that we reach our second important theorem of probability—the Central Limit Theorem (CLT), which states:

The distribution of sample means approximates a Normal distribution as the sample size gets larger (assuming that all samples are identical in size), regardless of population distribution shape.

In simple English, LLN implies that as the number of events / observations / items in the sample size increase, the mean of the sample gets closer and closer to the mean of the population. Toss a coin an infinite number of times, and the number of heads and tails will be equal (actually, they will move towards being equal as the coin tosses move towards infinite). CLT implies that if you take a large group of samples of equal size (e.g. an NBA basketball player taking 30 sets of 100 free throws), and then plot the mean values of each sample, the result will be a Normal distribution; higher the number of sets, the closer the approximation to an exact Normal distribution. LLN is easy to understand, while CLT is somewhat less intuitive.

It is through CLT that we reach the Gaussian distribution. This answers the question on how we decide whether we should use the Normal distribution. Firstly, we must have enough historical data to ensure LLN applies, i.e. the historical data must converge towards the population mean, that is, it must have a finite mean. Secondly, we must ensure the probability distribution generated from that data is conducive to CLT. To ensure the later, we need to add an additional caveat to CLT by replacing the term regardless in the above definition: the probability distribution of the samples must have a finite variance. This implies that CLT does not apply for all phenomena. We may end up modelling a two-humped camel as a one-humped camel, if we only have enough data to see the first hump, thereby, violating LLN. We also may end up modelling a horse as a giraffe, thereby violating CLT.

What do these giraffes, horses, camels and (most importantly) swans have to do with each other (not to mention, with kangaroos)? This is covered in Part II (to be continued).

More essays in the Black Swans series are on their way.