I had an interesting experience recently. Someone made a map (and before anyone starts...this was not someone at the company I work for). It was going viral. I contacted said author privately and had a good conversation about the issue. I'd faced a similar issue about 6 months ago and was able to give him some advice on how to tackle it. He then published a blog that paraphrased the conversation and explained how important it was to blah blah blah yackity schmackity (you get the idea). His blog on the issue is currently going viral. His 'insight' is taking his worth to new levels. Good luck to him, but he didn't even mention any conversation we had or the fact that he was both wrong and lacked understanding prior to someone reaching out to help. It wasn't necessary to acknowledge my help but the fact that he didn't somehow feels bad.
Yesterday I was asked to help an academic friend of mine with another map. I experienced a few technical issues and couldn't quite get done what I had hoped but I offered a few pieces of technical advice which he then went off and implemented (expertly). When he published a blog post with the resulting work he was gracious to acknowledge my very small part in the process. It wasn't at all necessary but it felt good. It's nice to be appreciated.
In the academic world, doing research and publishing ideas is based on learning, doing and reporting. This is normally done objectively, with peer review and with proper attribution for work that went before including citations for previously published work and so on. It takes time. Often too much time... but you progress your research interests. You begin collaborating with like-minded people and the total often becomes more than the sum of the parts. You become an acknowledged expert in a niche area and you make a name for yourself based on your expertise. Your friends listen and it feeds future collaborations. You become interested in quality over quantity. The art of an academic is keeping your friends close and working with them.
In the corporate world, it's more important to come first, be seen to be first, even if the ideas are old and someone got there before you anyway. You learn where to steal ideas and how to then pass them off as your own. You act as if you're engaged and interested but really the only thing you're interested in is bleeding someone of ideas and then pretending you know it all to the next person you meet. This doesn't take much time and you can iterate through various different types of work with speed and efficiency. You become an expert at being a non-expert in the next best idea. You make a name for yourself regardless of any expertise and everyone listens. You become interested in quantity over quality. The art of being corporate is keeping your friends apart so they never find out they can bypass you entirely.
Having moved from academia to the corporate world I brought many good friends along with me. Finding out who the new ones might be still poses a challenge every now and then.
Thursday, 25 July 2013
Monday, 24 June 2013
Sharing or Showing Off...you can't have it both ways.
Interesting debate emerging today with the discovery that software games producers Naughty Dog have seemingly used a map from the interwebs without seeking permission or arranging for royalties to be paid to the owner of the work. Check out the Tumblr and comments by the map-maker Cameron Booth (@transitmap) whose map it is below (reproduced under fair use...I think!):
Cameron built a new map to his design of the MBTA rapid transit system. Looks remarkably like an official map, takes a hefty dose of design cues from the official one and many others derived, ultimately from Harry Beck's iconic 1933 London Underground map (which of course, itself has design cues you can trace back even further that questions his own innovation or lack thereof).
Mr Booth is unhappy his map has been re-used without his permission or the exchange of money. On one hand I can very well understand his upset. It's his work. It's been stolen and re-purposed without anyone even having the courtesy to give him a call. On the other hand, isn't this a case of pot calling kettle black? His map wouldn't exist without those that went before it. He made his map and shared it, hoping for it to be picked up on a wave of internet 'likes' to either become something MBTA might want, as a showcase for his talents or just to prompt discussion on the design of transit maps. Presumably he enjoyed all the positive comments his very well crafted version elicited but when it's used in a way he hadn't envisaged the reaction is the exact opposite.
The simple point here is no-one is unique and if you put your work out there it's bound to get bad press and be mis-used from time to time even if it doesn't deserve it. No-one's work is that innovative it transcends all others so we have to be a little less precious about what we have a right to claim ownership of. We can recognize design elements from many different works if we look close enough in most maps. A good cartographer will learn what works, what looks good, what communicates and fashion something that brings together those elements in a clean, clear and well structured product. Those not as capable tend to produce clumsy work, unable to bring together ideas from a range of sources. Then there are the blatant rip-offs. All work is a function of what has gone before and contributes to a continuum. We can normally see a very clear lineage in maps and only in very very rare cases can we genuinely look at something and say it's truly original. That being the case we need to apply the same caution when sharing our own work.
I also had a dose of this today. I've been preparing maps as part of my day job so that I can show those who use a certain software platform how to create high quality cartography, often going beyond the defaults and using custom tools and techniques (I'm not going to name them because this isn't a blog about my day job). I posted a screen shot of part of one such map and got what seemed a less than enthusiastic response from someone who appeared to be questioning my right to produce the map because it has a similar look and feel to one he had also created (again, no names because this is written without his input). I happen to be an admirer of his work and yes, I had seen the map before....but does this mean I have blatantly copied it? No. Different data, different results and a different purpose. Where they are similar is in their use of the same general technique (hex binning), the same colour background (white...like many maps) and a yellow-blue colour scheme for the symbols (again, like many maps). Their version is a single scale, mine a multiscale. They wrote a blog explaining how it was built using a certain set of tools. I'll be using my map to demo a different set of tools.
No-one holds the rights to cartography and by demonstrating how such maps can be created I take the view that I'm sharing 'best practice'. I give shout-outs to other work that I consider to demonstrate best practice in cartography (whether they use the same software stack or not) and encourage people to look at such examples to develop an appreciation of the current cream of the cartographic crop. I could have used a different colour scheme and maybe that would have been sufficiently different not to cause a reaction but where do we draw the line? Should I never again create a choropleth map using an orange sequential hue scheme again (actually...probably not but for other reasons!)?
It seems to me that the internet is breeding a mindset which on the face of it seems open and sharing but when you appear to get too close to what someone else has shared it can easily kick off. Cartography isn't copyrighted. Neither are white backgrounds or colours. Neither are hexagons...as I've had to write about before. Give credit where credit is due and cite your inspiration and ideas. By sharing best practice we're encouraging people to work towards making better maps which does a service to cartography as a whole. By having choice they can make use of a variety of software stacks and do their work in a way that suits them. But we cannot take the view that having publicised a map and a technique that we then get irked if someone else benefits from the work. I wonder sometimes if the Open Source movement are partly to blame for this cultural mindset. The idea of sharing (normally for free) is fantastically altruistic but ultimately we all survive off the quality of our own work and while striving to be seen above the crowd we can often get caught up in our own hype. Do good work. Share it. Be happy if people like it and make use of it for positive reasons. If you share your work and its good, don't be surprised when others find it useful and make use of it for the benefit of others. If you feel like your work has been ripped-off then step back and wonder if it's possible that the ideas can actually be traced a little further back. Chances are you weren't the first...and I won't be the last. We simply cannot have it both ways.
Thursday, 20 June 2013
3 billion tweets on a map
It's been hard this past few days to find anyone willing to say a bad word about the new Locals and Tourists map from that magic combination...MapBox, Eric Fischer and Gnip. All capable of producing cool stuff. This is broadly what you get from their collaboration...a multiscale cached map of the world showing 3 billion dots each one red or blue:
The MapBox blog explains it in more detail but essentially it's every geotagged tweet since September 2011, categorised as either red to indicate a tweet from a tourist and blue, a tweet from a local.
It's impressive in scale and scope. It looks beautiful. The crowd on the interwebs are going gaga for it. I like it as a piece of map porn along with the multitude of other such works I'd put in that category.
Where I perhaps differ from some, though, is in casting a critical eye on the actuality of the map and bypassing the hype and rhetoric so here's a few thoughts...
The map claims to show 'incredible new detail', 'demographic, cultural and social patterns down to city level'. It apparently allows you to 'explore stories of space, language and access to technology'. OK, if they'd left it at "here's a cool looking map, whaddya think" I'd have bought it but c'mon.
The points are just geotagged tweets. Best estimates suggest only 1% of tweets are geotagged. Some 13% of them don't have exact coordinates. It might be reasonable to suggest we're looking at a sample of all tweeters but is the sample the same or sufficiently similar across all places to make the assumption we can compare like for like? We don't know.
And what of the demographics of Twitter Use? 67% of all internet users use social media. People who live in cities spend more time on social media. Only 16% of those who use social media use Twitter and they are most likely to be adults aged between 18-29...and male.
In short, the data itself is full of error, bias and uncertainty. You can't make any sensible inferences from a map like this. You certainly cannot claim to be able to explore the map in the way suggested.
And then we get to the classification of the data. Tweets are red if the user has posted for one consecutive month in the same place and blue for users whose tweets are normally elsewhere. And from that, they've decided they are able to identify locals and tourists. Spurious to say the least. There are probably hundreds of ways that this rule can be broken and someone move from the red to the blue or vice versa. It may be a convenient hook for the map title but as a classification it is based in fantasy.
Finally we get to the map...take a look and check out places you are familiar with. Does it stack up? More likely than not the answer is no. There appears a much larger element of so-called 'tourism' in places you'd hardly classify as tourist spots and where there are tourist honeypots or towns...there's no discernable difference. Roads are predominantly comprised of red dots as are motorway service stations and anywhere else that people travel through.
The map is largely empty in areas you'd expect there to be lower levels of mobile phone usage (most of the developing world) but we know this already. The map doesn't tell us anything new. Does it give us an insight into large scale city level detail? Not really....it's patchy at best and why not just use a decent street map if that's what you need to find out.
Finally, I was curious about how the red and blue dots were displayed. Having just spent quite some time making a dasymetric dot density map of the 2012 Presidential election results I'm quite familiar with the problems of representing a lot of blue and red dots in close proximity. In order to give equal visual weighting to each colour you have to do a little bit of processing of the data (well, quite a bit actually). So what have they done with the twitter map? I re-engineered the map and played around with the tiles and it's clear that they've just overprinted the blue dots with the red dots. One layer on top of another. In visual terms, then, the map gives far more prominence to the red 'tourist' dots which, of course, makes for a more interesting map. Here's a quick comparison:
Two completely different maps. Two completely different impressions of the data...but users of the map don't get that choice; they get the map on the left and base their interpretations on that...incorrectly.
Let's get one thing clear...as a piece of map porn it's spectacular. It also demonstrates that technologically, the mapping world is really at the cutting edge of handling large data sets (not 'big data' per se...this is just a lot of data). The point of this blog, though, is to counter the common perception of this sort of work. There are big questions regarding bias, error and uncertainty and how we represent this sort of data for mass public consumption. Is it useful? Can we really find anything out from it? Or is it just 'art'? Maps have always lied it's true but far more eyes are cast over work these days simply because of its accessibility online and the way social media itself transmits the latest and greatest ideas...but the internet has no quality control, no quality assurance and as with everything out there, if you consume it without taking a moment to think about what it is you're looking at you just drink the Koolaid.
Enjoy it for what it is but don't be taken in by the rhetoric. It is what it is, nothing more...despite what the label says and what all those 'likes' and re-tweets infer.
The MapBox blog explains it in more detail but essentially it's every geotagged tweet since September 2011, categorised as either red to indicate a tweet from a tourist and blue, a tweet from a local.
It's impressive in scale and scope. It looks beautiful. The crowd on the interwebs are going gaga for it. I like it as a piece of map porn along with the multitude of other such works I'd put in that category.
Where I perhaps differ from some, though, is in casting a critical eye on the actuality of the map and bypassing the hype and rhetoric so here's a few thoughts...
The map claims to show 'incredible new detail', 'demographic, cultural and social patterns down to city level'. It apparently allows you to 'explore stories of space, language and access to technology'. OK, if they'd left it at "here's a cool looking map, whaddya think" I'd have bought it but c'mon.
The points are just geotagged tweets. Best estimates suggest only 1% of tweets are geotagged. Some 13% of them don't have exact coordinates. It might be reasonable to suggest we're looking at a sample of all tweeters but is the sample the same or sufficiently similar across all places to make the assumption we can compare like for like? We don't know.
And what of the demographics of Twitter Use? 67% of all internet users use social media. People who live in cities spend more time on social media. Only 16% of those who use social media use Twitter and they are most likely to be adults aged between 18-29...and male.
In short, the data itself is full of error, bias and uncertainty. You can't make any sensible inferences from a map like this. You certainly cannot claim to be able to explore the map in the way suggested.
And then we get to the classification of the data. Tweets are red if the user has posted for one consecutive month in the same place and blue for users whose tweets are normally elsewhere. And from that, they've decided they are able to identify locals and tourists. Spurious to say the least. There are probably hundreds of ways that this rule can be broken and someone move from the red to the blue or vice versa. It may be a convenient hook for the map title but as a classification it is based in fantasy.
Finally we get to the map...take a look and check out places you are familiar with. Does it stack up? More likely than not the answer is no. There appears a much larger element of so-called 'tourism' in places you'd hardly classify as tourist spots and where there are tourist honeypots or towns...there's no discernable difference. Roads are predominantly comprised of red dots as are motorway service stations and anywhere else that people travel through.
The map is largely empty in areas you'd expect there to be lower levels of mobile phone usage (most of the developing world) but we know this already. The map doesn't tell us anything new. Does it give us an insight into large scale city level detail? Not really....it's patchy at best and why not just use a decent street map if that's what you need to find out.
Finally, I was curious about how the red and blue dots were displayed. Having just spent quite some time making a dasymetric dot density map of the 2012 Presidential election results I'm quite familiar with the problems of representing a lot of blue and red dots in close proximity. In order to give equal visual weighting to each colour you have to do a little bit of processing of the data (well, quite a bit actually). So what have they done with the twitter map? I re-engineered the map and played around with the tiles and it's clear that they've just overprinted the blue dots with the red dots. One layer on top of another. In visual terms, then, the map gives far more prominence to the red 'tourist' dots which, of course, makes for a more interesting map. Here's a quick comparison:
Two completely different maps. Two completely different impressions of the data...but users of the map don't get that choice; they get the map on the left and base their interpretations on that...incorrectly.
Let's get one thing clear...as a piece of map porn it's spectacular. It also demonstrates that technologically, the mapping world is really at the cutting edge of handling large data sets (not 'big data' per se...this is just a lot of data). The point of this blog, though, is to counter the common perception of this sort of work. There are big questions regarding bias, error and uncertainty and how we represent this sort of data for mass public consumption. Is it useful? Can we really find anything out from it? Or is it just 'art'? Maps have always lied it's true but far more eyes are cast over work these days simply because of its accessibility online and the way social media itself transmits the latest and greatest ideas...but the internet has no quality control, no quality assurance and as with everything out there, if you consume it without taking a moment to think about what it is you're looking at you just drink the Koolaid.
Enjoy it for what it is but don't be taken in by the rhetoric. It is what it is, nothing more...despite what the label says and what all those 'likes' and re-tweets infer.
Monday, 10 June 2013
Choropleth totals: one more shot
Regular readers might have come across a few comments I've made about the problems of mapping totals using choropleth maps here and there. I bore myself silly sometimes but in a week that's seen the all-knowing National Security Agency PRISM tracking tool report surveillance in one of the worst culprits I have ever seen it's worth giving it one more shot to try and banish the erroneous practice.
The Guardian's reporting of the recording and analysing of voluminous data included a few screen shots of the maps used to show countries' exposure to surveillance. Here's one for the purposes of discussion:
The colours range from green, those with the fewest number of items of surveillance, to red, the most. Putting aside the fact the colour scheme is not easy to interpret (is light green least...or is it dark green?) and that the use of a Mercator projection distorts the size of areas giving false prominence to more northerly and southerly latitudes...let's focus on the data.
So there have been 2,892,343,446 pieces of intelligence data gathered from US computer networks over a 30-day period ending in March 2013. The US is shown the same as Germany, Saudi Arabia, Kenya and Iraq. The map quite clearly leaves readers with the impression that surveillance is pretty similar across these countries.
Let's assume that each of these countries has 3 billion pieces of information collected (a fair assumption of the mapping of totals above). If we calculate the number of pieces of surveillance data per capita we get an entirely different picture
USA, 313 million people = an average 9.5 pieces of surveillance data per person
Saudi Arabia, 28 million people = an average 106.8 pieces of surveillance data per person
Kenya, 42 million people = an average 72 pieces of surveillance data per person
Iraq, 33 million people = an average 91 pieces of surveillance data per person
Mapping these rates would give an accurate picture and one which, crucially, allows us to visually compare one country against another across the map. Without normalising our data to a consistent denominator the map is utterly useless. We can't make any sensible interpretations of the information. So, as it turns out the level of surveillance in the US is an order of 10 times less than Saudi Arabia. That is not a story the above map even vaguely illustrates. So it's worrying if these maps are genuinely being used to inform national security don't you think?
In trying to come up with a way for mere mortals to understand, my better half Linda Beale (@lindabeale) came up with a great analogy. We spent some time over the weekend making some slides to explain the issue...using alcohol, conveniently colour coded green to match the NSA map. It struck us that if we cannot use logic to explain it then let's dumb this down to something most people can identify with...size of a drink.
So here goes with one more shot...have a look at the following picture:
Two glasses of the same size. It should be fairly easy to see that the glass on the right is holding double that of the left. In fact, the right glass is filled with two shots of Creme de Menthe and the right one has only one (or...his and hers if you will!). Now let's look at adding a third glass:
The glass on the right is a different size and shape so how much Creme de Menthe does it hold? Trickier question eh? Is it one shot, two shots or somewhere in between? In fact it's one shot but it doesn't look the same as the glass on the left and it looks significantly less than half of the glass holding the double. Now let's take an aerial view:
Hmm...they all look quite similar but one's a hexagon. Here then is the problem...a choropleth map is a container for data. Looking at a choropleth map is similar to looking down on a whole array of glasses filled with liquids of different quantities and it's impossible to make any sense of the amount of liquid...the totals are a function of the size and shape of the vessel in which they are contained. When the sizes of areas vary substantially the problems of estimating quantities gets even harder unless you adjust for the differences caused by the containers so you can make a sensible assessment of the relative amount of liquid in each glass:
The Guardian's reporting of the recording and analysing of voluminous data included a few screen shots of the maps used to show countries' exposure to surveillance. Here's one for the purposes of discussion:
The colours range from green, those with the fewest number of items of surveillance, to red, the most. Putting aside the fact the colour scheme is not easy to interpret (is light green least...or is it dark green?) and that the use of a Mercator projection distorts the size of areas giving false prominence to more northerly and southerly latitudes...let's focus on the data.
So there have been 2,892,343,446 pieces of intelligence data gathered from US computer networks over a 30-day period ending in March 2013. The US is shown the same as Germany, Saudi Arabia, Kenya and Iraq. The map quite clearly leaves readers with the impression that surveillance is pretty similar across these countries.
Let's assume that each of these countries has 3 billion pieces of information collected (a fair assumption of the mapping of totals above). If we calculate the number of pieces of surveillance data per capita we get an entirely different picture
USA, 313 million people = an average 9.5 pieces of surveillance data per person
Saudi Arabia, 28 million people = an average 106.8 pieces of surveillance data per person
Kenya, 42 million people = an average 72 pieces of surveillance data per person
Iraq, 33 million people = an average 91 pieces of surveillance data per person
Mapping these rates would give an accurate picture and one which, crucially, allows us to visually compare one country against another across the map. Without normalising our data to a consistent denominator the map is utterly useless. We can't make any sensible interpretations of the information. So, as it turns out the level of surveillance in the US is an order of 10 times less than Saudi Arabia. That is not a story the above map even vaguely illustrates. So it's worrying if these maps are genuinely being used to inform national security don't you think?
In trying to come up with a way for mere mortals to understand, my better half Linda Beale (@lindabeale) came up with a great analogy. We spent some time over the weekend making some slides to explain the issue...using alcohol, conveniently colour coded green to match the NSA map. It struck us that if we cannot use logic to explain it then let's dumb this down to something most people can identify with...size of a drink.
So here goes with one more shot...have a look at the following picture:
Two glasses of the same size. It should be fairly easy to see that the glass on the right is holding double that of the left. In fact, the right glass is filled with two shots of Creme de Menthe and the right one has only one (or...his and hers if you will!). Now let's look at adding a third glass:
The glass on the right is a different size and shape so how much Creme de Menthe does it hold? Trickier question eh? Is it one shot, two shots or somewhere in between? In fact it's one shot but it doesn't look the same as the glass on the left and it looks significantly less than half of the glass holding the double. Now let's take an aerial view:
Hmm...they all look quite similar but one's a hexagon. Here then is the problem...a choropleth map is a container for data. Looking at a choropleth map is similar to looking down on a whole array of glasses filled with liquids of different quantities and it's impossible to make any sense of the amount of liquid...the totals are a function of the size and shape of the vessel in which they are contained. When the sizes of areas vary substantially the problems of estimating quantities gets even harder unless you adjust for the differences caused by the containers so you can make a sensible assessment of the relative amount of liquid in each glass:
In this example, the top left glass contains one shot, top right, two shots. Both these glasses are the same size so we have a consistent basis for comparison and the darker colour suggests more liquid...a correct visual interpretation. What about the big glass? Actually it contains one shot but because it's spread out across a much wider area the colour is diluted and it appears that the glass holds far less than either of the other two. This would be an incorrect assumption. It's the same as the top left glass and for us to interpret that correctly we need to see the colours the same. In map terms, we need to show similarity of the character of areas using similar symbols so our eyes and brains interpret things properly.
Let's be clear, this isn't some sort of quirky thing that cartographers do to finesse a map, it's not just 'best practice'. It's fundamental data analysis to support proper mapping of data. It's non-negotiable (and it's not even hard to do...). Your map will, quite simply, be utterly meaningless without your data being normalized.
Consider John Snow's famous map of the 1854 cholera outbreak in Soho, London. His map is widely regarded as a classic. He mapped deaths as dots and was able to make an inference as to the potential cause of the cholera outbreak. If he had mapped the same data as totals on a choropleth the map would have been one of the most useless maps ever made; it's unlikely he'd have made any reasonable assumption about the mode of transmission of cholera or had the foresight to trace it to the Broad Street pump. So why do so many people persist in making their maps useless? It strikes me as an odd thing to do but hopefully this little explanation will help those who can think about their mapping in a slightly different way.
Got it? Cheers, it was fun...
Subscribe to:
Posts (Atom)




