Sunday, November 14, 2010

Wofo Si

Temple of the Reclinging Buddha (Wofo Si). Nice temple (filled with VERY FAT CATS), the Buddha was quite spectacular, 15 feet long, all bronze and the largest in China.

Hot Pot Spot

Stopped along a little alley way for lunch and had a hot pot for lunch. You picked out the pot you wanted (each and everyone slightly different) and then they cooked it on the propane powered burners.

You then went to go sit in this little room on the otherside of the alley:

2010-11-14-IMG_0211

Here we are sitting along the wall, it was packed when we got there. I promptly knocked over a big baking tray of rising dough, but picked it up and put it on the side (it was later used, so no harm done).

2010-11-14-IMG_0209

And here's the finished hot pots:

2010-11-14-IMG_0210

Atop Xianglu Peak

Chris and I, as the only two Western men caused a photo frenzy when we posed on top of Xianglu Peak. Keri took the photo because she didn't want to add to the frenzy.

a "Good Wall", not a "Great" Wall

Went to the top of Incense Burner Hill (Xianglu Peak, one of the highest points in Beijing), we took the cable car (ski lift) and didn't walk up. There was a nice wall though along the route, good, but not great.

Saturday, November 13, 2010

Along Ring Road 5, Beijing

Dusk settling in on Beijing as I road in a taxi from PEK to the Fragrant Hill Empark Hotel

Thursday, November 11, 2010

Wednesday, November 10, 2010

Sometimes less is more (when it comes to data)


2008-11-05-dscn6431
Originally uploaded by martin_kalfatovic
In a project I work on (the Biodiversity Heritage Library), we always say “BHL is important because it’s a complete (planned at least) of a type of data (biodiv lit)”. Generally speaking, we don’t data to support this assumption, but a firm called Infochimps is designing metrics and data analysis that can quantify this assumption; here’s an interesting post by Eric Hellman (on his "Go to Hellman" blog which you should all ready by the way) about scaling data value, Eric has a longer/fuller discussion, but the key point in relation to BHL  is the following:
Kromer has noticed that the price (or perhaps cost) of a partial data set follows a non-monotonic curve (see graphic). Small amounts of data are essentially free, but a peak value is reached when portions of the data set are extracted from the full data set.
Kromer has noticed that the price (or perhaps cost) of a partial data set follows a non-monotonic curve (see graphic). Small amounts of data are essentially free, but a peak value is reached when portions of the data set are extracted from the full data set. If we were discussing book metadata, for example, peak value might accrue for a set of the 100,000 top selling books.
There's much less value, according to Kromer, in having a large incomplete chunk of a data set. Data for 10,000,000 books, for example, would have less value than the 100,000 book data set, because it's not complete. Complete data sets become extremely expensive because of the logistics involved, and because of the value of having the complete set.
In the BHL context, substitute “Google” for 10m books and “BHL” for 100k books. The BHL data set, acquired at a higher unit cost than the Google data set, is of more “value” because of the coherency of the data (operations on a small, coherent set of data will return greater value than on large incoherent data sets). So, the current ~85,000 BHL volumes online could be of more value than the entire Google Books set.