At a data session during the Online News Association conference in October 2009, Aron Pilhofer of the New York Times demonstrated an early alpha of DocumentCloud, a project meant to help reporters present, share and search the volumes of documents an investigation produces.

The premise was that documents are data, but data without structure. DocumentCloud set out to add that structure and act as an analytical, publishing and search tool at once, with the twin goals of improving transparency and helping journalists find linked data buried inside document sets. PDFs, Pilhofer said, are a terrible way of putting documents online. The tool would be free so long as the uploader was willing to make the documents public.

The processing relied on OpenCalais, which reads the text and returns the companies, people and places named in it — entity extraction, followed by the step Pilhofer described as his new favourite word, disambiguation. Calais assigns an entity such as IBM a single identifier that then holds across every document, so a shared store contributed to by different news organisations would surface connections between documents that arrived from different sources.

In the demonstration, uploaded documents sat down the left of the screen; a search term narrowed the list, and a panel of topics narrowed it again until specific documents could be pulled out. At the time of the session DocumentCloud had 27 member organisations and was looking for more.