Showing posts with label google. Show all posts
Showing posts with label google. Show all posts

Monday, February 28, 2011

New open source server options for Ubisense myWorld

We have been busy working away on various aspects of Ubisense myWorld. One of the biggest enhancements is behind the scenes, with support for new server options, so that we can run in the cloud or in house.

Up to this point we’ve been working with Arc2Earth, which runs on top of Google App Engine, and both these platforms have worked very well for us, and were a great way of getting an initial system up and running quickly. We see a lot of benefits to running in the cloud, as I’ve talked about on several occasions.

However, a number of our customers, including large utilities and telecom companies, have said that they really like what we’re doing with myWorld, but they would be more comfortable with a solution where the server can run in house. So to support this we have added a new server architecture based on the open source products MapFish and PostGIS. As many of you will know, PostGIS is a very robust spatial database, built on top of PostgreSQL. I have used this on a few projects including whereyougonnabe and have always been very impressed with its functionality and performance. MapFish provides services using data from PostGIS (or from other spatial data sources, including Oracle Spatial, MySQL and Spatialite), using a very similar REST API to that used by Arc2Earth, so that made the migration straightforward and means we can support both server options with a largely common set of code. We’re just using server side components of MapFish, not its client side components (though we might consider using those in the future).

This new server code can run on various operating systems, including Linux and Windows (and Mac!). Customers can run the server in house, while we can now offer services using many different cloud infrastructure providers. We’re currently using Amazon, which has been working well, but it’s good to have alternatives available. We've continued to be pleased with PostGIS, and MapFish too based on our experience so far.

Stay tuned for more news on other cool new functionality on the front end of myWorld coming soon!

Thursday, January 27, 2011

Geospatial in the cloud

As mentioned previously, earlier in the week myself, Brian Timoney and Chris Helm did a set of presentations and demos on geospatial technology in the cloud, to the Boulder Denver Geospatial Technologists group. We were aiming to give a quick taste of a variety of interesting geo-things currently happening in the cloud, and we did it as six slots of about ten minutes each, and apart from my introductory opening slot these were all demos:
  • Peter: Why the cloud?
  • Brian: Google Fusion Tables
  • Chris: the OpenGeo stack on Amazon (PostGIS, GeoServer, OpenLayers, etc)
  • Peter: Ubisense myWorld and Arc2Earth
  • Chris: GeoCommons
  • Peter: OpenStreetMap
We got a lot of good feedback on the session. Here's the video (for best quality click through to vimeo and watch in HD):

Geo in the cloud from Peter Batty on Vimeo.

Here are links to the demos we used, or related sites:
And finally, here are my slides on slideshare:

Monday, January 24, 2011

Google Chrome notebook review

Today I was very happy to receive out of the blue a Google Chrome notebook. It's very black and stealth-bomber-like! I've written a short review of it over on my new posterous blog, where I'll be writing about (mainly) non-geo-things, including general tech, photography, videography, travel and anything else that comes to mind!

Wednesday, December 1, 2010

Ubisense myWorld featured on Google blogs

Today I had a guest post about Ubisense myWorld published on the Google Geo Developers and Google Enterprise blogs, check it out! Thanks to Mano Marks of Google for working with me on this.

Thursday, September 16, 2010

Using Google Maps to broaden the reach of GIS: Ubisense myWorld

Readers of my blog will know that for several years I have been advocating that Google Maps and other "neogeography" systems have a strong role to play in more traditional GIS application areas. In recent months I've been quite busy working on making this a reality, and last week at the Smallworld User conference in Baltimore we announced a new product called Ubisense myWorld. I'm really excited about what we've come up with. Check out the video demo, and read on below for more information (demo video best viewed in full screen, HD, no scaling).
Our initial focus has been on using myWorld with GE Smallworld data (for those not familiar with GE Smallworld, that's where I used to work a few jobs back, and they are still the market leader for GIS in utilities and telecom). However, we are using the excellent Arc2Earth Cloud as the repository for our geospatial data in the cloud (running on Google App Engine), which has comprehensive support for uploading ESRI data, so we can integrate with ESRI too. And much of our functionality is quite applicable to other application areas. So if you might be interested in using myWorld outside a Smallworld environment, let me know!

We've had a really strong emphasis on usability - our aim is that people should be able to use the system with no training. We display all the asset data as raster tiles, which is essential for scalability - Smallworld data models tend to be among the most dense of any GIS applications, with very detailed network data, and tables often running into millions of records. But you can select anything with a single click on the map, which sends a query to the spatial database hosted in Arc2Earth (using a REST API, for those interested in that sort of thing) to get the relevant attribute information. We've found performance to be excellent with this approach.

We've implemented a very fast Google style single box search across the whole geospatial database (well, whichever tables and columns the administrator designates), so users don't need to know any technical details about table names, field names or query structures to find what they are interested in - they can just start typing a pole number, asset id, customer name, etc and we have an autocomplete capability that will show them a list of options to choose from. This uses Google's App Engine datastore, which is based on BigTable, the same technology that powers Google Search, so is obviously very fast and scalable.

Another cool feature, which I haven't seen elsewhere yet done in the same way that we've done it, is tight integration with Google Street View. When you click on an item on the map, like a pole or a customer, we calculate the best street view to show you what that looks like (where one exists - it doesn't in all cases, as not everywhere is covered by street view). So far in our testing, the data has matched up better than I thought it would, with the automatic calculation working well in most cases. However, in some cases there will be mismatches between the data in the GIS and the data in street view, so we make it easy to adjust the view and save it, so that next time you click on the same item it will remember the adjusted view. This street view capability is one of the main reasons we decided to use Google Maps rather than the various other options out there.

We also see a lot of potential for the idea of "enterprise mashups" - being able to easily pull data from other systems within an enterprise that also contain spatial data, like outage management, vehicle tracking, work management, and customer information systems, as well as data from external sources like traffic, weather and more.

Since the application is just based on JavaScript with no plug-ins, it also runs really nicely on the iPad, and on smart phones like the iPhone and Android (with a modified user interface for the phones, to accommodate the smaller screen size). We think that all of these devices have great potential for use in the field.

We are just scratching the surface with what we're showing so far - we have a long list of ideas for more things we want to do, while at the same time maintaining a very strong focus on keeping the user interface super simple. I'll post more in due course about some more detailed aspects of what we've been doing.

You also can see some of my broader perspectives behind what we're doing in the presentation I did at the Smallworld user conference in Baltimore last week:

Smallworld and Google: the best of both worlds from Peter Batty on Vimeo.

Wednesday, May 19, 2010

Google's approach to user generated map updates not working?

My friend Greg Johnson found this interesting story saying that Google is hiring 300 people for a year to work "to improve the accuracy of Google Maps", though the commentary is rather uninformed (IMHO!). It doesn't discuss the fact that Google ditched Tele Atlas in the US 7 months ago to use their own data, and were widely perceived as having taken quite a step back in terms of data quality, as reported by various people including me, James, Matt, and Maitri. Google's vision seemed to be that they would improve the data quality over time by allowing users to report errors, but I had questioned whether typical users would be motivated to submit error reports, when it was easier to just switch to using Bing or MapQuest or whoever, who used more proven data from Tele Atlas or NAVTEQ. And most people interested in doing their own mapping are more likely to use OpenStreetMap, so they and others can use the raw data they have created (Google's equivalent, Mapmaker, is only available in some countries, and only lets you use raster map tiles derived from the data you have created rather than the raw data, and only under the terms of the Google Maps API which has various restrictions).

The article says that Google is paying $14.50 an hour, so a back of the envelope calculation for 300 people for a year says that they will be spending around $8.5m on labor alone (excluding overheads), which is not a huge deal for Google, but not insignificant either. Perhaps there is some other grand new plan behind this, but I have to think that this indicates that Google has realized they have a lot of work to do to improve their map data.

Monday, November 9, 2009

Transit routing on iPhone maps is cool!

I have been meaning to blog for some time about how useful I find the transit information in Google Maps on the iPhone. It's been around for a while, but I have been using it quite a lot recently and haven't blogged about it before. For those who haven't used it, when you calculate directions you can pick one of three options: driving, public transit, or walking (you also have the same options on Google Maps online). When you choose public transit, it shows you the next available trip, as follows:

iPhone Maps Transit information

You can click on the clock icon to see later trips and alternative options. This is one example where the application knowing your current location really adds to the convenience of getting the information - you just choose where you want to go, from a search or your contacts, and then the default is to show you how to get there from your current location. Even if you know the route, being able to easily find the time of the next bus is a great convenience. Another nice aspect is that the GPS tracks the bus location as you travel, and shows the scheduled arrival time, making it easy to figure out where to get off, which is another potential source of stress when you're traveling on a route you don't know.

iPhone Maps Transit

I think that one of the main inhibitors to people using public transit when they're not used to doing it is just the effort of figuring out their options. In a lot of US cities like Denver the culture really isn't to use transit - the default option is just to take the car. We've had some new light rail lines opening over the past few years, with more being built, and a few more people use those, but very few people I know really think about using the bus here. But we actually have a very good bus system, even though many people don't realize it. I am fortunate to live downtown so most of the time I just walk when we go out in the evening, but increasingly if we go further afield I try to take the bus or the light rail, and a strong contributing factor to that is the convenience of figuring out the transit routes (which are often ones I haven't taken before, or at least don't take regularly) on the iPhone.

This is also available in Google Maps Mobile on other platforms apart from the iPhone. But it's only available for some cities - it depends whether the transit agency has made their data available to Google. So it works for example in Denver, and even in Cropston, the small village in Leicestershire in the UK where I grew up, but not (at the moment) in London or Washington DC.

Anyway, I think that mobile multi-modal transit routing applications like iPhone maps and others have great potential to encourage people to use public transit more. I encourage you to try it, and leave the car at home! Other iPhone applications that can supplement this include things like Taxi Magic, which lets you call a taxi to your current location, and car sharing schemes like zipcar, which now has a cool iPhone app (of which one of the niftiest features is that you can use your iPhone to unlock the car, or even honk its horn to help you find it!).

By the way, I only found out fairly recently how to capture iPhone screen shots - in case you don't know, to do this you hold down the "Home" button and at the same time press and release the power button, and this will save an image of the current screen in your Camera Roll photos. I found this out from the 9 year old son of Dale Lutz!

Thursday, November 5, 2009

Was the Google Maps data change a big mistake?

So the discussions about the great Google map data change in the US rage on, and we are seeing more and more reports of significant data quality issues. I wrote about how Central City Parkway was completely missing, and I reported this to Google to see how the change process would work. I posted later about how it had been partially fixed, with a new geometry visible but not routable, and with the wrong road name and classification. The latest state (on November 5, after reporting the error on October 9), is that is now routable, but it still has the wrong road classification, being shown as a minor road rather than a major highway. This means that if you calculate the best route from Denver to Central City, Google gets it wrong and doesn't go on Central City Parkway, choosing some much smaller mountain roads instead, which take a lot longer. Microsoft Bing, Yahoo, MapQuest and CloudMade (using OpenStreetMap) all calculate the correct route using Central City Parkway. Another substantial error I found recently is that if you search for one of the companies I work for, Enspiria Solutions (by name or address), the location returned was about a mile off. This has now been partially but not entirely fixed (after I reported it).

Steve Citron-Pousty recently wrote about whether Google made the data change too soon. He talked about how his wife has always used Google Maps, but it has got her lost four times in the past week and she has now switched to using MapQuest or Bing. And Google got Steve lost in the Bay Area last week too. Maitri says she is "splitting up with Google Maps" over issues in Ohio, as "there is no excuse for such shoddy mapping when MapQuest and Yahoo do exceptional work in this area the first time around". She links to an article in a Canton, Ohio, newspaper about how the town was mis-named after the recent data change (we call it Canton, Google calls it Colesville). James Fee pointed out an error with Google showing a lake that hadn't been there for 25 years or so. Matt Ball did a round-up discussion on the importance of trusted data. The well known tech journalist Walt Mossberg reviews the new Motorola Droid phone (which uses the new Google data for navigation), and in passing says when reviewing the navigation application "but it also gave me a couple of bad directions, such as sending me the wrong way at a fork in the road". And then in news which is presumably unrelated technically (being the in the UK), there was a lot of coverage of a story about how Google Maps contained a completely fictitious town called Argleton - which even though a separate issue does not help the public perception of the reliability of Google Maps data.

Update: see quite a few more stories about data issues in the comments below.

So anyway, this is a long and maybe somewhat boring list, but I think that it is worth getting a feel for the number of stories that are appearing about map data errors. As anyone in the geo world knows, all maps have errors, and it's hard to do a really rigorous analysis on Google's current dataset versus others. But I think there is strong evidence that the new Google dataset in the US is a significant step down in quality from what they had before, and from what Microsoft, Yahoo and MapQuest have (via Tele Atlas or NAVTEQ).

Google clearly hopes to clean up the data fairly quickly by having users notify them of errors. But looking at the situation, I think that they may have a few challenges with this. One is just that the number of errors seems to be pretty large. But more importantly, I think the question for Google is whether consumers will be motivated to help them fix up the data, when there are plenty of good free alternatives available. If Google gives you the wrong answer once maybe you let it slide, and perhaps you notice the link to inform them of the problem and maybe fill it out. But if it happens a couple of times, is the average consumer likely to keep informing Google of errors, or just say "*&#% this, I'm switching to MapQuest/Bing/Yahoo"?

Google has made some reasonable progress with Google MapMaker (its crowdsourced system for letting people create their own map data) in certain parts of the world, but these are generally places where there are not good alternative maps already, or people may be unaware of alternatives like OpenStreetMap. So in those cases, people have a clearer motivation to contribute their time to making updates. People who contribute time to OpenStreetMap have a range of motivations, but in general for most of them it is important that the data is open and freely available, which is not the case with Google (at least not so much, I won't get into the details of that discussion here). Most if not all the people I know who contribute effort to OpenStreetMap (myself included) would not be inclined to contribute significant updates to Google (except for some experiments to see how good or bad the update process is).

Consumer confidence is a fickle thing, and you probably don't need too many stories in the newspapers of mishaps due to bad data, or more than a couple of direct experiences of getting lost yourself due to bad data, to switch to a different provider (especially when you are choosing between different free systems - you have a bit more incentive to stick with a navigation system and try to make it work if you've spent a few hundred dollars on it).

The risks are even higher with real time turn by turn directions - no matter how many caveats you put on it, you are likely to get some drivers who follow the directions from the system and don't notice relevant road signs. You only need a couple of accidents because people drove the wrong way up one way streets because of bad data to damage consumer confidence even further.

So I think it will be very interesting over the next few months to see whether the data quality issues are bad enough to result in significant numbers of users moving away from Google Maps in the US or not - and whether Google will get significant uptake in terms of number of people contributing error reports in the US (beyond the initial wave of curiosity-driven updates just to test if the process works). Obviously the answer to the second question is likely to have a big influence on the first. Stay tuned ...

Monday, October 26, 2009

Talk on "The Geospatial Revolution" in Minnesota

Here is a video of my recent keynote talk at the Minnesota GIS/LIS conference in Duluth, which was an excellent event. There were about 500 people there, which is great in the current economic climate. It was mainly a "traditional GIS" audience, and I got a lot of good feedback on the talk which was nice.

I talk about current trends in the industry in three main areas: moving to the mainstream (at last!); a real time, multimedia view of the world; and crowdsourcing. There's a lot of the same material that I presented in my talk with the same title at AGI GeoCommunity (which doesn't have an online video), but this one also has additional content (~50 minutes versus 30 minutes).

Click through to vimeo for a larger video, and if you click on "HD" you will get the full high definition version!! I used a different approach to produce this video compared to previous presentation videos, using a separate camera and a different layout for combining the slides and video. I like the way this came out - I'll do a separate blog post soon with some tips on how to video presentations, I think.

The Geospatial Revolution (Minnesota) from Peter Batty on Vimeo.

You can also view the slides here:

Sunday, October 25, 2009

Google Maps data correction - a strange semi-update

I reported previously that I found that Google Maps was missing Central City Parkway after their change in street data provider (they are now providing their own street data rather than using Tele Atlas). I reported the error to Google and said I would report back here when it was fixed - Google is aiming to fix errors within 30 days. This evening Tom Churchill commented on my previous post to say that Central City Parkway was now present on the map - I thought it was odd that it had been fixed but I hadn't received an email to let me know, as promised. I reported the problem 17 days ago and received confirmation that it was an error and they would work on it 13 days ago.

Anyway, when I went to check out the updated map, this is what I saw (Google is on the right, OpenStreetMap on the left):

Partial update to Central City Parkway in Google Maps

This compares to the previous comparison screenshot I did, which looks like this:

Central City Parkway missing from Google Maps

So there is now a road on the map that follows the path of Central City Parkway, when there wasn't before. But it's drawn as a minor road when it should be a major highway (a "divided highway" or "dual carriageway" depending on where you come from!), and it has no name shown on the map. And when I try to get Google Maps to route along it, it stubbornly refuses to do so, even if I try to drag the route to force it along there (when dragging, it doesn't allow me to drop on this street):

Central City Parkway not routing yet in Google Maps

So anyway, we seem to have a curious semi-update - there's a new geometry there that wasn't there previously, along the route of Central City Parkway, but with no name, the wrong road classification, and you can't route along it. Seems odd that a partial update like this should find its way into the live database ... I guess the process is still slightly "beta" :O !! Will keep an eye on it and report back on further progress!

Wednesday, October 14, 2009

A black hole isn't "evil", but ...

I loved this quote from Paul Ramsey, commenting on Paul Bisset's blog post about the "Google data earthquake":
Right, a black hole isn’t “evil”, but that doesn’t change the fact that it massively distorts the shape of space-time everywhere it goes, which can be a bummer for any object in its immediate neighbourhood.
That summarizes rather nicely concerns I've expressed in recent posts.

Tuesday, October 13, 2009

More on the "Google data earthquake"

Following on from my previous post about Google shaking up the geospatial data industry, Steve Coast invited me and James Fee to join him for a discussion on the topic. James' blog post on the topic has 138 comments at the time of writing, which is a good indication of the interest in this change! You can listen to the podcast on the "Google data earthquake" here.

One topic I talk about in the call which I didn't cover in my previous post is where the Google street data comes from (they haven't said anything about this). To me it looks like a mixture of data they have captured from their StreetView cars, which seems to be good quality, and then probably TIGER data, which is much lower quality, where their cars haven't driven. Quite a few people have reported finding errors in street data that weren't there previously since the change. I found that Central City Parkway was missing from the map entirely, which is a pretty major highway that was completed in 2004. You can see this below, with OpenStreetMap on the left, and Google Maps on the right (screen shot using GeoFabrik's nice map compare tool):

Central City Parkway missing from Google Maps

I've reported the error to Google, so it will be interesting to see how quickly it gets fixed - and in general, how quickly they are able to fix up the apparently lower quality data in areas they haven't driven yet (though this assessment is not based on anything scientific).

Wednesday, October 7, 2009

Google shakes up the geospatial data industry

Well, the big news of the day is that Google has dumped Tele Atlas as the main data provider for Google Maps in the US, and is providing its own map data from a variety of sources (presumably also including its own Streetview teams). They've also added the ability to point out errors in the map, another addition to the crowdsourcing techniques they've been using. The announcement has caused a flurry of discussion of course. James raises questions about various aspects of the data (especially parcels). Steve speculates that the same thing will happen in Europe and that the beneficiary there will probably be AND.

The new Google data certainly adds details in some places, from a quick random sampling - for example check out Commons Park in downtown Denver using the nice GeoFabrik Map Compare tool. None of those paths were there previously in Google. Still not quite as good as OpenStreetMap in this case though :).

This does dramatically reshape the geospatial data industry though. Previously there were two commercial providers with a detailed routable database of roads in the US, NAVTEQ (owned by Nokia) and Tele Atlas (owned by TomTom), now at a stroke there is a third in Google. OpenStreetMap is a fourth provider of course, not quite up with the other three in terms of coverage and routing quality in the US yet, but getting there very quickly.

This raises lots of interesting questions:
  • Will Google sell its data to providers of third party navigation systems and compete with Tele Atlas and NAVTEQ? Or indeed will they sell/license it to others who could use it (users of GIS, etc)?
  • Will Google Maps on the iPhone (and other mobile devices) get real time turn by turn directions? This was previously prohibited by licensing terms from Tele Atlas and NAVTEQ. Existing real time navigation systems using data from these two providers generally cost in the region of $100. Will Google add this to the free maps offering? Or sell a version that does real time turn by turn directions?
  • Will Google contribute any of this data to open data initiatives like OpenStreetMap? Or make it available to USGS for the US National Map? In the past they have cited licensing constraints from their data providers as a reason for not being more open with their geospatial data, that reason largely goes away now (though we don't know all the new data providers and their terms). I'm not holding my breath on this one, but we can hope!
  • Will this negatively or positively impact OpenStreetMap? Previously in areas with active communities, OpenStreetMap had significantly more detail, and more current data, than Google - this appears to move Google forward in that regard. But will Google taking another step towards total world domination encourage more people to want an open alternative?
So anyway, definitely a very interesting development for the geospatial data industry (albeit one that has been on the cards for a little while). It will take a little while to understand the full implications. I'm sure Tele Atlas is glad they are no longer an independent company, I wouldn't like to have seen how their stock price would have dropped today otherwise :O !!

Update: some more discussion and a link to a podcast on this topic featuring Steve Coast, James Fee and me in this follow up post.

Wednesday, September 16, 2009

Google PowerMeter accidentally wipes out small industry on the way to changing the world??

I spent the last few days at the Autovation conference in Denver, which is focused on Smart Metering and the Smart Grid, an area that I am becoming increasingly interested in and one where Enspiria is doing a lot of work (where I work part time as Chief Technology Advisor). It was a very interesting conference - there is certainly lots of activity and energy in the space, especially since the stimulus bill committed $4.3bn to Smart Grid projects - and utilities need to match this funding, so close to $9bn will be spent over the next couple of years. That's a large enough sum to even interest the likes of Google and Microsoft in electric utility applications, something they haven't been into previously.

So it was interesting that the closing speaker at the conference was Ed Hu from Google (who is incidentally a former astronaut, who has spent six months on the space station!). He is responsible for their PowerMeter initiative. This was announced earlier this year and I had previously skimmed articles on it, but have to admit that I hadn't grasped the full significance of it until Ed's talk yesterday. You can see a short description of what it's all about in this one minute video:


In summary, it lets you see detailed information about your home's power consumption, enabling you to change your behavior to reduce consumption. Ed draws an analogy with the fuel consumption readout in a Toyota Prius, which encourages you to modify your driving style to maximize your fuel consumption (I can vouch for this). He says that in their trials so far, people using PowerMeter have typically got anywhere from 5-15% savings on their electricity bill. As I talked about in my previous Smart Grid video, one reason for having smart meters is to enable customers to have access to this type of information, in order to encourage them to change their behavior and reduce electricity consumption. This has various benefits including reduction in carbon emissions. Usage data for Google PowerMeter can be obtained either via your local utility, if they have installed smart meters and choose to offer the Google service (which requires them to interface their Meter Data Management System, MDMS, to Google), or alternatively you will be able to buy devices to install in your home and read consumption directly. Ed said that if they could get the same level of usage reductions as they got in their pilot, from 6 million users, this would be equivalent to the reduction in carbon emissions due to all the hybrid cars currently on the road.

This initiative is being run by google.org, the philanthropic arm of Google, and the system is offered free to both consumers and utilities (and will continue to be free, Ed said). He also said that he'd been told personally by the CEO of Google, Eric Schmidt, that the aim of the project was "to change the world".

I'm excited in many ways to see Google getting involved - certainly they understand how to build applications to engage consumers, and this is not something that electric utilities generally have expertise in. And having Google involved certainly could significantly accelerate the development of this aspect of the Smart Grid. However, it's potentially very disruptive for a number of existing companies in this space, who are trying to do much the same thing. Companies like Greenbox and Tendril seem to have strong overlap with what Google is doing here. In this story at earth2tech shortly after the initial announcement of PowerMeter, both try to put a somewhat positive spin on Google getting involved in the space, but it will be hard for them or others to compete with the core Google offering, especially since it is going to be free. Perhaps they can find niches that are complementary to what Google is doing, but they and other companies in this space seem in a somewhat precarious position to me.

The one other player in this space that I haven't mentioned, who are probably in a stronger position to compete with Google, is Microsoft, who have a relatively similar offering called Microsoft Hohm. Earth2tech compares the two. They say that Microsoft intends to charge utilities for their offering eventually - and also says that Microsoft intends to move into the space of controlling devices too. Someone at Autovation asked Ed if Google was planning to move into control of devices too, in addition to the data display they are doing now, and he indicated that this was very likely - though he said they wanted to focus on getting the display part right first. While he wasn't specific, if they did provide the ability to control devices for the consumer it is a logical step to provide that to the utility too, which gets them into the whole market area called Demand Response, potentially overlapping even more with existing companies.

After seeing all the disruption that Google and Microsoft have brought to the geospatial industry, it is interesting to see them moving aggressively into the consumer-related aspects of the Smart Grid. Google providing completely free enterprise solutions through its philanthropic arm is in some ways commendable and in other ways concerning (in terms of the ability of others to provide competition, and the risk that they could just wipe out multiple companies). One of the most common questions web entrepreneurs get asked when presenting to investors is "what if Google decides to do what you're doing?", and it seems as though people will need to be asking that question of companies in ever more diverse fields! It will be very interesting to see how all this develops over the next year or two, and how the existing companies in the space respond.

Wednesday, April 8, 2009

Google App Engine and BigTable - VERY interesting!

Every so often you come across a radically different approach to a certain class of data processing problem which makes you completely rethink what you knew before about how best to develop applications in that space. Systems which I would put in this category over my career include:
  • Smallworld VMDS (early 90s), for its approach to handling long transactions and graphically intensive applications in a database
  • Apama (early 2000s), for its approach to streaming analytics on real time data streams, by indexing queries instead of indexing data
  • Netezza (last year - for me), for its approach to data warehousing using "SQL in hardware" in its smart disk readers, together with extreme parallel processing
I would now add to that list Google's BigTable for incredibly scalable request-oriented web applications (update: Barry Hunter pointed out in the comments that the App Engine datastore and BigTable are not the same thing - datastore is built on top of the lower level BigTable, and adds extra capabilities. I haven't updated the whole post but in most cases where I say BigTable, it should say datastore. Thanks Barry!). I know I'm a bit behind the times on this - Google has been using it internally for years, and it was first made available for external use with the release of Google's App Engine last year. I had read about App Engine but hadn't got around to looking at it in any detail until last weekend, when for some reason I read a few more detailed articles, downloaded the App Engine development environment and ran through the tutorial and played around with it a bit.

There are too many interesting things to talk about in this regard for one post, so I'll spread them over several. And I should add the caveat that everything I say here is based on a few hours of poking around App Engine and BigTable, so it is quite possible I have missed or misunderstood certain things - if anyone with more experience in the environment has thoughts I would be very interested to hear them.

In general I was very impressed with App Engine - in less than an hour I was able to run through the getting started tutorial, which included getting a local development environment set up, going through 5 or 6 iterations of a simple web application, including user authentication, database setup, etc, and deploying several iterations of the application online. It takes care of a huge amount for you, including the ability to automatically scale to zillions of users. We could throw away large portions of our code for whereyougonnabe if we moved to App Engine, something which I am now seriously considering.

But for the rest of this post I'd like to talk about BigTable, the "database" behind App Engine. Google stresses that it isn't a traditional database - this paper, from several years ago, describes it as a "distributed storage system". It can handle petabytes of data spread across thousands of servers and is used by many Google applications, including search and Google Earth. So clearly BigTable is enormously scalable.

However, it also has some limitations on the types of queries it allows, which at first glance for someone used to a traditional relational database management system seem incredibly restrictive. Some of these restrictions seem to have good technical reasons and some seem a bit arbitrary. For example:
  • A query cannot return more than 1000 rows
  • You cannot use an inequality operator (<, <=, >=, >, !=) on more than one property (aka "field") in a query - so you can do
    SELECT * FROM Person WHERE birth_year >= :min
    AND birth_year <= :max
    but not
    SELECT * FROM Person WHERE birth_year >= :min_year
    AND height >= :min_height
  • If a query has both a filter with an inequality comparison and one or more sort orders, the query must include a sort order for the property used in the inequality, and the sort order must appear before sort orders on other properties.
  • And more - see the full list.
While these sort of constraints impose some challenges, the positive side is that as far as I can see, you can't write an inefficient query using BigTable (if anyone has a counterexample to this statement - based as I said on a couple of hours exposure to BigTable - please let me know!). It changes your whole approach to the problem. A traditional relational DBMS makes it very easy to ask whatever question you want (generally speaking), but you may then need quite a lot of work in terms of indexing, tuning, even data model redesign, to make the answer to that question come back quickly. It's easy to do sloppy data model and query design. With BigTable you may need to think more up front about how to fit your problem into the constraints it imposes, but if you can then you are guaranteed (I think!) that it will run quickly and scale.

There's an interesting example of this type of redesign in this post, which shows how you can redesign a query on a date range, where the obvious approach is to have two fields storing start_date and end_date, and run a query which includes an inequality operator against both fields - something which BigTable does not allow. The interesting solution given here is to use one field containing a list of (two) dates, which BigTable does allow (and most traditional DBMSs don't). This is a real world example of a query which is pretty inefficient if you do it in the obvious way in a traditional database (I have seen performance issues for this type of query in the development of whereyougonnabe) - BigTable forces you to structure the data in a different way which ends up being far more efficient.

I am still in two minds about the restriction of not allowing inequality operators on more than one field. This clearly guarantees that the query can run quickly, but restricts you from answering certain questions. Most database management systems would use the approach of having a "primary filter" and a "secondary filter" for a compound query like this - the system uses the primary filter to efficiently retrieve candidate records from the database which satisfy the first condition, and then you test each of those against the second condition to decide whether to return them. This is very common in spatial queries, where you return candidate records based on a bounding box search which can be done very efficiently, and then you compare candidate records returned against a more precise polygon to decide if they should be included in the result set. But this also adds complexity - it is non-trivial to decide which one of multiple clauses to use as the primary filter (this is what a query optimizer does), and it is quite possible that you end up scanning large portions of a table, which seems to be one of the things that BigTable wants to avoid.

Nonetheless, technically it would be easy for Google to implement a secondary filter capability, so I can only assume it is a conscious decision to omit this, to force you to design your data structures and queries in a way which only scan a small portion of a table (as the restriction of 1000 records returned does) - so ensuring the scalability of the application. I would be curious as to whether some of these restrictions, like the 1000 record limit, apply to internal Google applications also, or just to the public site where App Engine runs (in order to stop applications consuming too many resources). When App Engine first came out it was free with quotas, but Google now has a system for charging based on system usage (CPU, bandwidth, etc) once you go above certain limits - so it will be interesting to see if they lift some of these restrictions at some point or not.

But in general it's an interesting philosophical approach to impose certain artificial restrictions to make you design things in a certain way (in this case, for efficiency and scalability). Twitter imposing a message length of 140 characters is limiting in certain ways, but imposes a certain communication style which is key to how it works. The Ignite and Pecha-Kucha presentation formats impose restrictive and artificial constraints on how you present a topic (20 slides which auto-advance after 15 or 20 seconds respectively), but they force you to really think about how to present your subject matter concisely.

With BigTable I think there is an interesting mix of constraints which have clear technical reasons (they can't be easily overcome) and those which don't (they could be easily overcome - like secondary inequality filters and the 1000 record limit). Whether there is really a conscious philosophy here or whether the approach is just to avoid overloading resources on the public site (or a mix of both), I am intrigued by this idea of having a system where seemingly any query you can write is "guaranteed" to run fast and be extremely scalable (not formally guaranteed by Google, but it seems to me that this should be the case).

Of course one key question for me and for readers of this blog is how well does BigTable handle geospatial data - especially since a standard bounding box query involves inequality operators on multiple fields, which is not allowed. BigTable does support a simple "GeoPt" data type, but doesn't support spatial queries out of the box. I have seen some examples using a geohash (which is claimed on Wikipedia to be a recent invention, which as a referencing scheme may be true, but as an indexing scheme it is just a simple form of the good old quadtree index which has been in use since the 1980s - see the end of this old paper of mine). The examples I have seen so far using a geohash are simple and just "approximate" - they will give good results in some cases but incorrect results in others. I have several ideas for using simple quadtree or grid indexes which I will experiment with, but I'll save the discussion on spatial data in BigTable for another post in the not too distant future.

Sunday, December 21, 2008

Google Earth used to "discover lost Eden" in Africa

In a welcome change to the recent rush of stories on the theme of "terrorists use Google Earth" (I'm still not sure of the ongoing fascination with that theme versus "terrorists use cell phones" or "terrorists use boats"), there's a story in today's Observer in the UK about how Google Earth was used to "discover a lost Eden" in Africa:
It was one of the few places on the planet that remained unmapped and unexplored, but now Mount Mabu has started to yield its secrets to the world.

Until a few years ago this giant forest in the mountainous north of Mozambique was known only to local villagers; it did not feature on maps nor, it is believed, in scientific collections or literature. But after "finding" the forest on a Google Earth internet map, a British-led team of scientists has returned from what is thought to be the first full-scale expedition into the canopy. Below the trees, which rise 45m above the ground, they discovered land filled with astonishingly rich biodiversity. It was one of the few places on the planet that remained unmapped and unexplored, but now Mount Mabu has started to yield its secrets to the world.
Check out the whole story.

Wednesday, October 8, 2008

Which city is this airport in? (An exercise in using geonames)

Today I spent some time working updating on updating the airport dataset we use with whereyougonnabe. I thought that it was an interesting process to iteratively come up with a reasonable algorithm for what I needed to do, so I thought I would explain it here. I'd be interested in any feedback both on improvements to the technical approach (or entirely different approaches), and also on what the "right answer" should be from a user perspective in some of the questionable cases I mention.

As I've mentioned previously, we use Google Local Search for geocoding in general, but in our experience this really doesn't work well for geocoding airports based on the three letter airport codes. When we import travel itineraries from Tripit, the three letter airport codes are the best way to unambiguously identify which airports you will be in, so we need a reliable way of geocoding those.

We have been using the geonames dataset to do this, a rich open source database, which includes cities, airports and other points of interest from around the world. There are about 6.6 million points in the whole dataset, of which about 23000 are airports, and of those about 3500 have three letter IATA codes. Coverage seems to be fairly complete, but we have found a few airports with IATA codes missing - Victoria, BC, is one (code YYJ), and today I found that Knoxville was missing its code (TYS). I updated Victoria in the source data online, and it shows up in the online map but for some reason still does not appear in the downloaded data, I have that on my list of things to look into.

In general, the geonames airport records usually have null data in the city field. This is something I really wanted for each airport, so I can show a short high level description like "departing from Denver" or "arriving in London". Often the city name is included in the airport name, but not always - for example, the main airport in Las Vegas is just named "McCarran International Airport", which would not be an immediately obvious location to many people. And even where the city is included, the name is often longer than I want for a short description, for example "Ronald Reagan Washington National Airport".

For our first quick pass at populating the city for each airport, we used the Yahoo geocoding service, as it has a reverse geocoding function (and we thought we would try it out). You just pass this a coordinate and it gives you a city name (and other data) back. This is the data we are using in the live version of whereyougonnabe today. However, I noticed that this doesn't always give me what I would like - for example, it tells me that London Heathrow airport is in Ashford, and Vancouver airport is in Richmond. This may be technically correct (or not, I don't know for sure), but I'd really find it more useful to see the main city nearby in my high level description - so it will say I am arriving in London or Vancouver, in these two examples.

One nice aspect of the geonames dataset is that it includes the population of (many) cities, so we can use this in determining which is the largest nearby city. So in my first pass at solving the problem, I ran a PostGIS query for each airport to look for "large" cities nearby. Somewhat arbitrarily I initially specified population thresholds of 250000, 100000 and 0, and distance thresholds of 5, 25 and 50 miles. I would first search for cities with population > 250000 within 5 miles, and if I didn't find any I would extend to 25 miles and then 50 miles. If I still hadn't found anything, I would decrease the population to the next threshold and try again. If I found more than one in a given category I chose the closest, in this initial approach (the largest is obviously another alternative, but neither will be the right answer in all cases).

This worked reasonably well, returning London and Vancouver for the examples already mentioned. But it gave some wrong answers too, including Denver Airport which was associated with Aurora (a large suburb of Denver), and for a lot of small town airports it returned a larger town nearby, which in some cases might have been reasonable but in many cases wasn't.

So I decided to extend my algorithm to include matching on names - if there was a city called 'Denver' near 'Denver International Airport' then that is the one I was after. This required a little bit of SQL trickery. If I knew the city name and wanted to search for airports this is easy, you just use a SQL clause like:

where name like '%Denver%'

But it was the other way around, the value I had was a superstring of the value in the field I was querying. I found that in PostgreSQL you can do a statement like the following:

where 'Denver International Airport' like '%' || name || '%'

(|| is the string concatenation in PostgreSQL, so if the value in the field "name" is "Denver", the expression evaluates to 'Denver International Airport' like '%Denver%', which is true). I wouldn't really want to be running that sort of clause alone against the very large geonames table, for performance reasons, but since I also had clauses based on location and population to use as primary filters, performance was fine. And in fact to eliminate issues of upper versus lower case, the clause actually looked like:

where 'denver international airport' like lower('%' || name || '%')

So this would match the city names I found to any substring of the airport name (whether the city name was at the beginning or not). If I found a matching city name within a specified distance (I am currently using 40 miles), I use that, if not then I fall back to using the first approach to look for a "large city" nearby. On my first pass through with this approach, I was able to match names for 2356 of the 3562 airports with an IATA code.

I noticed that some cases were not matching when they looked like they should, where the relevant city names had accented characters. I realized that in some cases, the airport name had been entered with non-accented characters, so it was not matching the version of the city name which did have accented characters. The geonames data includes both an accented and non-accented (ascii) version of each city name, so I extended my where clause to do the wild card matching above on either the (accented) name field or the (non-accented) asciiname field. This increased the number of matching names to 2674, which was a pretty good improvement.

I noticed that in a few places, there were two potential "city" names in the airport name - in particular London Gatwick was matching on a "city" called Gatwick, rather than London, and Amsterdam Schiphol was matching a "city" called Schiphol. Looking at maps of the local area, I'm actually not convinced there is a city (or even a village) called Gatwick, and on further examination the geoname record shows its population as zero, and the same turned out to be true of Schiphol. But regardless, I decided to handle this by checking whether there was more than one city name match with the airport name, within the search radius, and if so to choose the largest city found (rather than the closest). There are some airports which serve two real cities also, and have both in their title, so this is a reasonable approach in those cases (no offense intended to the smaller cities!). This change returned me London instead of Gatwick and Amsterdam instead of Schiphol, as I wanted.

I was now getting pretty close, but I stumbled on a couple of small local airports where I was choosing a really really small town which happened to be closest, rather than a not quite so small town nearby which was really the obvious choice. For example, for Garfield County Airport in the mountains of Colorado, I was picking the town of Antlers (population listed as 0, though I think it might be a few more than that) rather than Rifle, population 7897! This was fixed just by added one more value to my population thresholds (I used 5000).

Obviously tweaking the population and distance thresholds would change some results, but from a quick skim I think the current results are pretty reasonable. One example which I am not sure whether to regard as right is JeffCo airport, a small airport which is in Broomfield, just outside Denver. The current thresholds associate this with Denver, which is probably reasonable, but you could argue that Broomfield is large enough to be named (45000).

If you're interested in checking out the full list, to see if your favorite airport is assigned to the city you think it should be associated with, you can check out this text file (the encoding of accented characters is garbled here, though it is correct in our database - but I haven't taken the time to figure out how to handle this using basic line printing in Java). Let me know if you think anything is wrong! The format of a line in the report is as follows:

13: GRZ Graz-Thalerhof-Flughafen -- 8073 Feldkirchen to Graz (Graz), true

The IATA airport code is first (GRZ), then the airport name from geonames (Graz-Thalerof-Fluhafen), then the city name we got from Yahoo reverse geocoding (8073 Feldkirchen), then the associated city name generated by our algorithm (Graz), followed by the ascii version of the name in parentheses (the same in this case), and lastly the "true" indicates that a match was found on the city name, rather than just using proximity and population.

So anyway, I thought this was an interesting example of how you can use the geonames dataset. I would be happy to contribute these city names back to geonames if people think this would be valuable - it is on my list to look into how best to do this, but if anyone has particular suggestions on this front let me know.

Tuesday, September 2, 2008

First experience with Google Chrome

Whereyougonnabe makes heavy use of JavaScript in the browser, using the Ext JS library, which provides a very rich and dynamic user experience but also tests the limits of browser compatibility. This is especially true when running inside Facebook, which provides another level of complexity and has a knack of showing up obscure low level differences between browsers - we have hit several tricky issues in handling sessions and security recently. FireFox has generally given us the fewest issues. Early on we had some problems with Safari, and today we have had a day of major frustration as we were about to launch a new release, everything was running fine on FireFox and Safari, and on Internet Explorer if you ran our application outside a Facebook frame, but inside Facebook Internet Explorer has suddenly started abruptly refusing to load pages with no helpful error messages - we're still bashing our heads against the wall trying to track this one down.

So I have to say my initial reaction to seeing that Google has announced a new browser called Chrome was "oh no, not another browser to support".  It's currently Windows only which I also wasn't too keen on. But nevertheless I fired up my Windows laptop for the second time in a week to try it out. And I was pleasantly surprised and impressed that it worked first time with whereyougonnabe with no apparent problems :) !!  And subjectively it seemed fast too. We'll be testing it more thoroughly of course, but first impressions are encouraging in terms of compatibility with complex Javascript applications, which Google claim as a major design aim for Chrome.

Thursday, August 21, 2008

More on Google Local Search

As I recently discussed, we have decided to use Google Local Search rather than Google Geocoding (well, rather than a hybrid approach) for whereyougonnabe. That post got some interesting comments, including one from Pamela Fox of Google who said that Local Search will soon be using Tele Atlas data (rather than NAVTEQ according to one of the other posts, though she didn't say that explicitly), and the implication was that this was the same source data as the Google Geocoder. One would hope that this may mean that they are bringing the two services together, which would be a lot nicer for users. I also ran into Martin May from Brightkite yesterday at the Techstars investor day (which was excellent) and he mentioned that they are making the same move, which was interesting - it turned out we had been looking at a lot of common issues at the same time.

Anyway, I thought I would follow up with a little more discussion on some of the benefits we see of using Google Local Search, as well as some outstanding issues. One aspect that I really like is that you can give it a search point to provide context, and this is very useful for our application. We are especially leveraging this in our calendar synchronization. A typical (and recommended) way of using the system when you are traveling is to create a high level activity saying that, for example, you are in New York City for four days next week, and then subsequently enter more specific activities for those days. When you add an activity in your calendar, we look at the location string in the context of where we think you are on that date. So for example, if I just said I was having lunch at 1100 Broadway, it would locate me at 11 Broadway in New York during those four days, or at 1100 Broadway in Denver if I was at home. And as I discussed in the previous post, the feature we like most about Local Search is that you can search for business names, and again these take that location context, which makes it very easy to specify many locations. Some examples that I have used successfully at home in Denver include "Union Station", "Vesta Dipping Grill" (or just "Vesta"), "Enspiria" (a company where I have an advisory role"), and even "Dr Slota" (my dentist). It's very convenient to be able to use names like this rather than having to look up addresses.

So that's great, but there are still a few issues. One is that Local Search really doesn't work well for airports - if you enter a 3 letter airport code it is very hit or miss whether it finds it. I'm pretty sure this used to work a lot better, though I haven't tracked down specific tests to prove that. But we plan to do our own handling of this case (unless we see an improvement really soon), using geonames (which we already use for locating airports in our tripit interface, but we haven't hooked this into our general geocoding).

I reported over a year ago that I felt there were some issues with Google Local Search on the iPhone, so I thought it would be interesting to revisit some of the same problem areas I reported on then. These problems generally revolved around local search being too inclusive in the data it incorporates, which is good in terms of finding some sort of result, but the risk is that you get bad data back in some cases. Given the good results I had found in my initial testing this time round, I expected that I would probably see some improvement, so I was a little surprised to find that in general most of the same problems were still there. You can see these by trying the search terms in Google Maps (and in MapQuest for comparison).

My test cases were as follows, centered around my home at 1792 Wynkoop St, Denver:
  • King Soopers (a local supermarket chain): last time in the top 10 on Google, there were three entries with an incomplete address (just Denver, CO), all of which showed as being closer than the closest real King Soopers, and there was one (wrong) entry which said it was an "unverified listing". This time, there was just one incomplete address like this, so that was a definite improvement, but unfortunately this appeared top of the list so the closest King Soopers was not returned as the "best guess" from the local search API. There were two other bogus looking entries in the top 4, for a total of 3 apparently bad results in the top 10. On MapQuest, 10 legitimate addresses are returned, though in some cases they appear to have multiple addresses for the same store (e.g. entry points on two different streets) - but this is not a serious error as you would still find a King Soopers. This was the same as last time. I tried typing "King Soopers Speer" (the name of the street where the store is located) and that returned me the correct location as the top result.
  • Tattered Cover (a well known Denver book store with 3 locations). MapQuest returns just 3 results, all correct. Google Maps returns 9 entries in its initial list, of which three have incomplete addresses (duplicates of entries which do have complete addresses, but they show up at entirely different locations on the map), one is a store which closed two years ago, and one is a duplicate of a correct store. The old store does say removal requested, but it seems surprising that this closed two years ago and is still there.
  • Searches for Office Depot and Home Depot were more successful, with no obvious errors - hooray :) !! This was the same as last time.
  • Searching for grocery in Google Maps online last time returned Market 41, a nightclub which closed a year before, and four entries with incomplete addresses. This time there were only two which looked bad (including the King Soopers with an incomplete address we saw before). The same search on MapQuest had one incomplete address, the rest of the results looked reasonable. So a definite improvement in this case.
So in summary, Google showed a little improvement over last year but not a huge change - they still could use some more quality control to remove bad data entries. But as I've discussed, overall we've obtained very good results for our application needs, so we are looking forward to rolling out the new Local Search based capabilities shortly.

Wednesday, August 13, 2008

Google Local Search better than Google Geocoding?

Google offers two (apparently) unrelated solutions for doing geocoding (converting a text string into a latitude-longitude and a structured address): the Google Geocoding API (which is part of the Google Maps API) and the Google Local Search API (which has JavaScript and non-JavaScript variants). This post discusses our experiences with both and the results of recent testing we have been doing, which has led us to a decision to switch from our current hybrid approach to using only local search.

Initially we were attracted to using Google Local Search for whereyougonnabe, as it lets you search for places of interest like "Wynkoop Brewing Company" or "Intergraph Huntsville", rather than having to know an address in all cases - whereas the geocoding API only works with addresses.

However, in our initial testing with local search we found a number of cases where it returned a location and address string successfully, but did not properly parse the address into its constituents correctly (for example, returning a blank city). For our application it was important to be able to separate out the city from the rest of the address. In our initial testing, the geocoding API seemed to do a better job of correctly parsing out the address. In addition, it returned a number indicating how accurate the geocoding was (for example street level or city level). So we ended up implementing a rather ugly hybrid solution, where we first used local search to allow us to search for places of interest in addition to addresses, and then passed the address string which was returned to the geocoding API to try to structure it more consistently. In most cases this was transparent to the user, but in a number of cases we hit problems where the second call would not be able to find the address returned by the first call, or where the address would get mysteriously "moved". With so much to do and so little time :) we elected to go with this rather unsatisfactory solution to start with, but I have recently been revisiting this area, and doing some more detailed testing of the two options.

Before I get into more details of the testing, I should just comment on a couple of other things. One is that a key reason we are revisiting our geocoding in general is that we recently introduced the ability to import activities from external calendar systems, which means we need to call geocoding functionality from our server from a periodic background process, whereas before that we just called it from the browser on each user's client machine. This is sigifnicant because the Google geocoding API restricts you to 15,000 calls per IP address per day, which is not an issue at all if you do all your geocoding on the client but quickly becomes an issue if you need to do server based geocoding. Interestingly, Google Local Search has a different approach, with no specified transaction limits (which is not to say that they could not introduce some of course, but it's a much better starting point than the hard limit on the geocoding API).

Secondly, a natural question to ask is whether we have looked at other solutions beyond these two, and the answer is yes, though not in huge detail. Most of the other solutions out there do not have the ability to handle points of interest in addition to addresses, which is a big issue for us. Microsoft Virtual Earth looks like the strongest competitor in this regard, but it seems that we would need to pay in order to use that, and we need to talk to Microsoft to figure out how much, which we haven't done yet - and obviously if we can get a free solution which works well we would prefer that. Several solutions suffer from lack of global coverage and/or even lower transaction limits than Google. We are using the open source geonames database for an increasing number of things, which I'll talk about more in a future post, but that won't do address level geocoding and points of interest are more limited than in Google (currently at least).

Anyway, on to the main point of the post :) ! I tested quite a variety of addresses and points of interest on both services. Some of these were fairly random, a number were addresses which specific users had reported to us as causing problems in their use of whereyougonnabe. In almost all the specific problem cases, we found that the issue was with the geocoding API rather than local search. The following output shows a sample of our test cases (mainly those where one or other had some sort of problem):

I=Input, GC=geocoding API result, LS=local search API result, C=comments

I: 1792 Wynkoop St, Denver
GC: 1792 Wynkoop St, Denver, CO 80202, USA
LS: 1792 Wynkoop St, Denver, CO 80202
C: LS did not include country in the address string returned, but it was included separately in the country field. The postcode field was not set in LS, even though it appeared in the address string. Other fields including the city and state (called "region") were broken out correctly. This was typical with other US addresses - LS did not set the zip code / postal code, but otherwise generally broke out the address components correctly.

I: 42 Latimer Road, Cropston
GC: 42 Latimer Rd, Cropston, Leicestershire, LE7 7, UK
LS: null
C: The main case where GC fared better was with relatively incomplete addresses, like this one with the country and nearest large town omitted (Cropston is a very small village)

I: 72 Latimer Road, Cropston, England
GC: 72 Latimer Rd, Cropston, Leicestershire, LE7 7, UK
LS: 72 Latimer Rd, Cropston, Leicester, UK
C: Both worked in this case, with slight variations. GC included the postal code (a low accuracy version) and LS did not. LS included the local larger town Leicester which is technically part of the mailing address.

I: 8225 Sage Hill Rd, St Francisville, LA 70775
GC: null
LS: 8225 Sage Hill Rd, St Francisville, LA 70775
C: One of several examples from our users which didn't work in GC but did in LS.

I: Wynkoop Brewing Company Denver
GC: null
LS: 1634 18th St, Denver, CO
C: As expected, points of interest like this do not get found by GC

I: Kismet, NY
GC: Kismet Ct, Ridge, NY 11961, USA
LS: Kismet, Islip, NY
C: Another real example from a user, where GC returned an incorrect location about 45 miles away from the correct location, which was found by LS.

I: 1111 West Georgia Street, Vancouver, BC, Canada
GC: 1111 E Georgia St, Vancouver, BC, Canada
LS: 1111 W Georgia St, Vancouver, BC, Canada
C: Another real example from a user which GC gets wrong - strangely it switches the street from W Georgia St to E Georgia St, which moves the location about 2 miles from where it should be.

I: London E18
GC: London, Alser Strafle 63a, 1080 Josefstadt, Wien, Austria
LS: London E18, UK
C: Another real user example. London E18 is a common way of denoting an area of London (the E18 is the high level portion of the postcode). GC gets it completely wrong, relocating the user to Austria, but LS gets it right.


So in summary, in our test cases we found a lot more addresses that could not be found or were incorrectly located by the Google Geocoding API, but were correctly located by Google Local Search, than vice versa. It is hard to draw firm conclusions without doing much larger scale tests, and it is possible that there is a bias in our problem cases as our existing application may have tended to make more problems visible from the Geocoding API than from Local Search (it is hard to tell whether this is the case or not). But nevertheless, based on these tests we feel much more comfortable going with Local Search rather than the Geocoding API, especially given its compelling benefits in locating points of interest by name rather than address. These point of interest searches can also take a search point, which is also a very useful feature for us, which I will save discussion of for a future post. Advantages of the geocoding API are that it returns an accuracy indicator and generally returns the zip/postal code also, neither of which are true for Local Search. Neither of these were critical issues for us, though we would really like to have the accuracy indicator in local search also. For our application we are not too concerned about precisely how the addresses align with the base maps - one reason I have heard given in passing for Google having these two separate datasets is that they want a geocoding dataset which uses the same base data as Google Maps uses, to try to ensure that locations returned by geocoding align as well as possible with the maps. If this is a high priority for your application, you might want to test this aspect in more detail. It also appears that the transaction limits are more flexible for server based geocoding with Local Search.

So we'll be switching to an approach which uses only Google Local Search in an upcoming release of whereyougonnabe. I'll report back if anything happens to change our approach, and I'll also talk more about what we're doing with geonames relating to geocoding in a future post.