Thursday, December 5, 2019

The Importance of a Log Scale

This short post was inspired by a financial crises dataset posted on Kaggle. Kaggle is a great resource for the data sciences, hosting competitions and curated datasets. I browse Kaggle from time to time as a way to spark creativity and see how others are interpreting data. When I saw this Global Financial Stability dataset, I was curious to see how the computer science dominated user base of Kaggle would catch some of the nuances of this dataset. 

The first graph below is a plot of inflation over time for Zimbabwe, this popular post by Aditya Keshri creates an identical graph, although I use a subset of 2000-2008 while his uses Zimbabwe's entire history. Initial interpretation of the graph shows how absurd Zimbabwe's hyperinflation peaked at in 2008 with an inflation rate of 21,989,695%. 

This initial interpretation is the problem with the visualization, as the number is so large that it entirely blows out the linear scale. The growth of inflation between 2006 and 2007 is an imperceptible increase on the graph. Seriously, if I wasn't lampshading that slight increase in the slope would anyone even notice it? The reality of the situation is that this imperceptible increase was actually a growth of the inflation rate from 1,281% to 66,279%, a 52.7 times increase. This is the strength of a log scale, and the weakness of programmers haphazardly graphing data without understanding the data. A log scale excels in portraying large ranges of quantities due to each distance on the scale representing an order of magnitude, rather than a linear growth of the value.