Selasa, 02 September 2008

Does anyone publish the Dataset of New Zealand Geographic Place Names already in XML form?

I've been playing with the Dataset of New Zealand Geographic Place Names which is a set of CSV files published by Toitū te whenua / Land Information New Zealand (LINZ). The data takes quite a bit of massaging, and I was wondering whether anyone else had already done the work of making acceptable XML out of the data rather than doing all the work myself.


I've attached the script I have so far, but it's not perfect. In particular:


  1. It doesn't include place names with Macrons
  2. It makes lots of ASCII-type assumptions
  3. Many of the element names are poorly named and map non-obviously to fields in the CSV files.
  4. The script isn't very generic and does little or no checking

Anyway, here's he script, hopefully it's successfully escaped. The basics are that it creates an sqlite database and streams the CSV files into it direct from the zip (which it expects to have been downloaded into the current directory). It then streams each point out using awk to transform it to XML.




#!/bin/bash
# script to import data from
# http://www.linz.govt.nz/placenames/search/place-names-dataset-download/index.aspx
# into an XML file.
# this script licensed under the GPL/BSD/Apache 2 licences

echo \(re\)creating the database, expect DROP errors the first time you run this
sqlite nzgeonames.db << EOF
DROP TABLE name;
CREATE TABLE name (id, name, east, north, pdescription, district, sheet, lat, long);

DROP TABLE district;
CREATE TABLE district (district, description);


DROP TABLE pdescription;
CREATE TABLE pdescription (pdescription, short, description);


DROP TABLE sheet;
CREATE TABLE sheet (edition, map, sheet);

VACUUM;
EOF

echo importing the names
unzip -p nznames_6Aug08.zip namedata.txt | sed 's/\r//' | sed 's/`/","/g' | awk -F^ '{print "INSERT INTO name VALUES (\"" $0 "\");"}' | sqlite nzgeonames.db

echo importing the districts
unzip -p nznames_6Aug08.zip landdist.txt | sed 's/\r//' | sed 's/`/","/g' | awk -F^ '{print "INSERT INTO district VALUES (\"" $0 "\");"}' | sqlite nzgeonames.db

echo importing the point descriptions \(expect two lines of errors\)
unzip -p nznames_6Aug08.zip pointdes.txt | sed 's/\r//' | sed 's/`/","/g' | awk -F^ '{print "INSERT INTO pdescription VALUES (\"" $0 "\");"}' | sed 's/:/","/' | sqlite nzgeonames.db

echo importing the sheet names
unzip -p nznames_6Aug08.zip sheetnam.txt | sed 's/\r//' | sed 's/`/","/g' | awk -F^ '{print "INSERT INTO sheet VALUES (\"" $0 "\");"}' | sqlite nzgeonames.db


# pick up the ugly duckling
sqlite nzgeonames << EOF
INSERT INTO pdescription VALUES ("MRFM","MARINE ROCK FORMATION","Marine Rock Formation");
EOF

echo exporting points as xml
echo "<document source=\"Sourced from Land Information New Zealand, [date]. Crown copyright reserved.\">" > nzgeonames.xml
sqlite nzgeonames.db "SELECT name.id, name.name, name.east, name.north, name.pdescription, name.district, name.sheet, name.lat, name.long, district.description, pdescription.short, pdescription.description AS descriptionA, sheet.edition, sheet.map FROM name, district, pdescription, sheet WHERE name.district = district.district AND name.pdescription = pdescription.pdescription AND name.sheet = sheet.sheet;" | awk -F\| '{print "<point><id>" $1 "</id><name>" $2 "</name><east>" $3 "</east><north>" $4 "</north><pdescription>" $5 "</pdescription><district>" $6 "</district><sheet>" $7 "</sheet><lat>" $8 "</lat><long>" $9 "</long><description>" $10 "</description><short>" $11 "</short><descriptionA>" $12 "</descriptionA> <edition>" $13 "</edition> <map>" $14 "</map> </point>"}' | sed 's/&/&amp;/' >> nzgeonames.xml
echo "</document>" >> nzgeonames.xml

echo formatting the points nicely
xmllint --format nzgeonames.xml > nzgeonames-formatted.xml


Senin, 01 September 2008

Library of Congress flickr experiment

While processing the photos from my parent's ruby wedding anniversary, I ran into the Library of Congress's flickr experiment.



I probably shouldn't have been, but I was astounded. It looks the the bastion of old-school cataloguing is coming to bathe in the fountain of social tagging.



This is part of a larger effort described at http://www.flickr.com/commons/


Selasa, 26 Agustus 2008

Saxon joy!

I've just moved to saxon from libxml for some XSLT stuff I'm doing, and I'm really loving it.

Not only does saxon run take much less memory, it also speaks XSLT 2.0.

Selasa, 05 Agustus 2008

moving back to google reader from bloglines

A couple of months ago I migrated to bloglines from google reader, not because I was necessarily unhappy with google reader, but because I was interested in seeing what else was available and how it might differ. I've just moved back to google reader.

OMPL just worked. I was able to move my RSS "reading list" from google reader to bloglines and back again with no fuss, no hassle and no duplication.

The advantages of google reader over bloglines are:
  1. AJAX - whereas bloglines marks all items on a page as read when you browse to it, google reader marks them as read when you scroll past them.
  2. Ordering - google entwines items from all feeds in time order, bloglines presents items feed by feed
  3. Better integration with other services
The advantages of bloglines over google reader are:
  1. Fast scanning of voluminous feeds
  2. Fast browsing (it seems _much_ faster when there are thousands of items)
  3. Less integration with other services
You'll notice that better integration is both a positive and a negative.

The fact that I have several google accounts and and only one of them is tied to my RSS reading means that there are tasks I can't multi-task between, even at the coarsest of levels and also means that contacts from the google account almost never get forwarded articles I discover via RSS.

The fact that my blogger.com account and my google reader accounts magically know about each other is great, as is being able to sign in once to a whole suite of tools.

In the end the reason for changing back was ordering. I read too many RSS feeds that cover the same topic for reading them out of order to make sense.

I've also just culled some of my RSS feeds, with the a prime criterion being the quality of their RSS. A number of web comics require one to click a link to read the strip and I no longer read them, but I still read Unshelved, which has the strip (and an ad) in the RSS.

Senin, 04 Agustus 2008

Decent editor for blogger.com?

Can someone recommend a decent replacement for the default editor for blogger.com?

Before it drives me insane...

KDE/Gnome Māori localisation on the rocks?

It looks like Maori localisation has been removed from the KDE 4.0 repository:

stuartyeates@stuartyeates:~/tmp/mi$ svn co svn://anonsvn.kde.org/home/kde/trunk/l10n-kde4/mi/messages
svn: URL 'svn://anonsvn.kde.org/home/kde/trunk/l10n-kde4/mi/messages' doesn't exist
stuartyeates@stuartyeates:~/tmp/mi$ svn co svn://anonsvn.kde.org/home/kde/trunk/l10n-kde4/mi/docmessages
svn: URL 'svn://anonsvn.kde.org/home/kde/trunk/l10n-kde4/mi/docmessages' doesn't exist
stuartyeates@stuartyeates:~/tmp/mi$ svn co svn://anonsvn.kde.org/home/kde/branches/stable/l10n-kde4/mi/messages
svn: URL 'svn://anonsvn.kde.org/home/kde/branches/stable/l10n-kde4/mi/messages' doesn't exist

Things don't look good for the upcoming 4.* releases, with the stats for translation at 0%: http://l10n.kde.org/stats/gui/trunk-kde4/team/

Gnome Māori localisation is not much better: stable at 1%: http://l10n.gnome.org/teams/mi

In the medium/long term there is hope that much of this localisation can be bootstrapped by application-centric localisation that appears to be thriving, particularly with respect to firefox, thunderbird and OOo.