Showing posts with label Internet Archive. Show all posts
Showing posts with label Internet Archive. Show all posts

Saturday, December 19, 2009

When You Look Into a Digital Abyss, Does It Also Look Into You?: A Look at Historical Abundance on the Internet

It seems fitting that this blogging assignment on the abundance of materials available on the internet will be my last post of the semester. Assigned in September, it’s taken me nearly four months to get past more than a sketchy outline of what I wanted to discuss in this post. It’s not through lack of trying though – I’ve probably sat down at my computer to write it more than a half dozen times only to find myself quickly distracted by an errant thought or query which led me to an internet search for information which led me to a hyperlink for a related topic which led me to either another hyperlink or another question and so on and so on until the time that I’d set aside for writing that day had run out.

Over the course of the semester I have likened this process of traveling from hyperlink to hyperlink to that of Alice falling down the rabbit hole in Lewis Carroll’s Alice in Wonderland – you never know what weird and wonderful things you’ll see during your trip, nor do you end up quite where you expected to when all is said and done. I’m not complaining; I’ve come across hundreds of amazing websites, useful tools, interesting articles and countless pieces of fascinating information this way. In truth, I think that my education and knowledge base are probably significantly broader as a result.

But this “infinite archive,” [1] as the internet has been described in recent years, is something of a double-edged sword, especially for academic and public historians. In the words of Daniel J. Cohen and Roy Rosenzweig [2]:

The digital era seems likely to confront historians—who were more likely in the past to worry about the scarcity of surviving evidence from the past—with a new “problem” of abundance. A much deeper and denser historical record, especially one in digital form, seems like an incredible opportunity and gift. But its overwhelming size means that we will have to spend a lot of time looking at this particular gift horse in the mouth... (2005)

As Cohen and Rosenzweig point out, in the past historians had to build their craft on the sometimes tenuous threads of the surviving evidence, a harsh reality which certainly restricted historical research. I also think that the caveat “that they could get their hands on” needs to be added to the aforementioned quote. Historical records and artifacts did not rest in a single, central and easily accessible location. As a result, historical research was often regional, not necessarily in terms of topic but in terms of sources used to support the arguments of historians. It’s not unreasonable to think that historians in the past could become acquainted with all of the materials available to them on their subject in their area and still be missing that critical document or artifact that had the potential to either cement or refute their claims.

Thankfully, these constraints of time and distance have largely been removed by improvements in technology. With the growth of the internet and the sheer abundance of information that is now available online, academics in even the most remote regions of the world suddenly have access to primary and secondary sources they may have previously been unaware even existed. For instance, the internet and massive digitization efforts have made it possible for students of history, like our own Megan Arnott, to be able to research and write on the Medieval Era without ever having to leave her home in Canada. Further, through email, chat rooms, Twitter, Facebook, Google Wave and any other number of social networking tools, Megan (and historians, in general) are better equipped than ever to quickly and easily communicate and collaborate with colleagues around the world.

Unfortunately, this is where the “‘problem’ of abundance” that Cohen and Rosenzweig speak of comes in. Those same digitization efforts that seem amazing at the outset do have their drawbacks – even as early as 1996, the internet had become “a sprawling megalopolis that no one person could fully explore” [3]. With thousands, if not tens of thousands of new websites added to the internet each day, how does an historian keep up?

The short answer is, we can’t. Instead, we must do as Cohen and Rosenzweig suggest and work, as historians in concert with computer programmers, website designers, researchers, publishers, museum curators, librarians and archivists to create “sophisticated statistical and data mining tools to do some of the looking” for us [2]. And, while the internet currently acts as an amazing archive of human record, no singular database has yet been created in which all historical records within it are stored. Because of this, I believe that it is the job of public historians to support the creation and use of these data mining tools and smaller databases in order to help make as much of the digital record as possible both knowable and useful to academics and the public alike.

Another issue that must be considered in connection with the growing abundance of historical sources on the internet is that of authority within the study of humanities and social sciences. As we discussed in class earlier in the semester, while the constraints of time, distance and availability no doubt frustrated historians in the past, they also helped to make the professionalization of history what it is today. For instance, with the scarcity of information came a tight control on existing intellectual materials; museums, archives, and libraries, especially those housing extremely rare items, were often restricted to only the most respected historians in the field (and largely only those associated with a trusted institution). This resulted in the control of scholarship resting in the hands of the few. It likely also served to privilege certain interpretations of historical events.

But one of the more amazing “gifts” to come out of the abundance of materials on the internet is the growth of democratization of the historical process. As more items that were previously inaccessible to the general public are digitized and compiled in databases like the Internet Archive, more grassroots historians are popping up each day. They read, research and comment on what they see and learn, adding a unique perspective and voice to literature on a variety of historical topics. They now have not only the ability to share their interpretations of historical events, but now they also have the primary and secondary sources to support it.

Don’t get me wrong, even the internet hasn’t been able to overcome the entrenched intellectual hierarchy which privileges the contributions made to the field by professional scholars in universities and other elite institutions [3]. I’m not even arguing that it necessarily should. After all, academic historians have spent years training in proper research methodology and historiography. But neither should the amateur historian be dismissed out of hand. As someone who is untrained in this type of research, I believe that there is a lot that historians can learn from these amateurs about how those outside of the ivory tower interact with history, what they need and want from it, and perhaps where academic research needs to go in order to satisfy the interests of the wider public.

And that’s where I believe public historians come in. I think the task now falls to us to mediate between the two groups (the public and the academics) in order to ensure that the contributions of both sides are acknowledged, valued and made accessible to the other.

Sources:

[1] Turkel, William J. Digital History Hacks: Methodology for the Infinite Archive (2005-08). [Weblog.]

[2] Cohen, Daniel and Roy Rosenzweig, “Web of lies? Historical knowledge on the Internet,” First Monday, Volume 10, Number 12 - 5 December 2005, http://firstmonday.org/htbin/cgiwrap/bin/ojs/index.php/fm/article/view/1299/1219.

[3] Cohen, Daniel and Roy Rosenzweig, “Exploring the History Web,” Digital History: A Guide to Gathering, Preserving and Presenting to Past on the Web, http://chnm.gmu.edu/digitalhistory/exploring/1.php.

Saturday, November 21, 2009

Digitizing Books Faster Than the Speed of Copyright

The Internet is arguably the greatest thing to happen to reading since the invention of the printing press (I can’t help myself – here I have to give credit to the Koreans, not the Germans, for their invention of the first metal printing press in 1377). While some detractors might rail against the fact that the Internet is changing the way that we read, it’s also true that efforts in digitization have forever changed (in a positive way) how we access that content.

Unless you’ve been living under a rock for the last seven or eight years, you’re likely aware that websites such as Google Books, the Gutenberg Project and the Infinite Archive (to name but a prominent few) have been working to digitize millions of books that have entered the public domain since they were originally published. Our Digital History assignment for this week is to peruse the books section of the Eaton’s Fall and Winter Catalogue from 1913-1914, choose a half dozen books and attempt track down full online copies of them.

I decided to start my book search with Anna Sewell’s Black Beauty: the autobiography of a horse and Lewis Carroll’s Through the Looking Glass. My reasons for starting with these books were two-fold: first, I’ve always had a soft spot for classic children’s literature. Secondly, I have a sneaking suspicion that anything in the Western canon that has had its copyright protection expire is going to be relatively easy to find since, according to Choudhury et al., the more widely quoted the text, the more likely it is that a well-transcribed digital version exists.

Sure enough, Google Books returned a search for “Black Beauty” in 0.17 seconds, with the top result a full-view of its 1922 printing. A search for Carroll’s book was similarly simple. The novel was found easily enough in Google Books (as plain text only), but out of curiosity I also typed it into Google’s search engine just to see what I could come up with. Oddly, the second search result (even before that on Google Books) was for a full copy at Literature.org and the third was the complete text as found on Project Gutenberg’s website.

Still within the Western canon, my next book was Charles Dickens’ David Copperfield. Originally published in 1850, it was a snap to find a copy of an 1869 printing on Google Books. Presumably with a classic such as this, it would be just as easy to find it on a multitude of other websites as well.

After Dickens I decided to branch out and try to find some of the more obscure titles that were advertised in the Eaton’s Catalogue. I’d never heard of Rosa N. Carey’s Not Like Other Girls before, so I tried to look it up on Wikipedia to see if I could find a basic plot outline and original publish date. No such luck. As far as I can tell, Wikipedia doesn’t have this information available. But apparently this book isn’t nearly as obscure as I thought it was as I was able to find copies in Google Books, the Internet Archive and Open Library.

My fifth book was Robinson’s Book of Conundrums, a book that claims it’s a “veritable dispenser of cheerfulness and dispeller of the blues.” As much as I wanted to find this book, I couldn’t have done it if my life depended on it. A number of similarly titled books come up Google Books, but none that fit exactly. The search was made a bit more difficult by the fact that I was only able to make out the author’s last name (Allen) in the Eaton’s catalogue. Alas, maybe my blues just weren’t meant to be dispelled.

My sixth and final foray into the world of online books was to look for Fred T. Hodgson’s Modern Carpentry: A Practical Manual (Vol. 2) (1917). Despite (or maybe because of) being reprinted several times, only a snippet view is available of this book on Google Books. The Internet Archive, on the other hand, had links to two full-length digital copies of the book, one more effective than the other. The first version was plain text only, something that puzzled me a bit considering that a carpentry book without pictures and diagrams seems a little counter intuitive. But I.A. came through in the end and linked me to a page-by-page scanned version of the original text, diagrams and pictures included.

It’s no surprise to anyone who knows me that I love books. I like their look. I like their smell. Call me old-fashioned, but I like physically holding an object as I’m reading and being required to turn pages as I progress. In my day-to-day life I read a lot of content online, but I’m still somewhat of a sceptic when it comes to reading entire books using the Internet. It just doesn’t feel the same.

Having used the Google Books website before, I was relatively familiar with the site’s layout and the fact that their books are page-by-page scans where each page is presented individually to the reader who will then scroll through them “vertically” (for lack of a better word). With all of the money Google has to put into projects such as this, it isn’t surprising that their website has a slick, user-friendly design. That said, I still feel as if there is something missing from the reader’s experience when I view content in this format. Project Gutenberg’s “read online” option is positively stark in comparison to Google Books’ flashier site. It still allows you view pages individually, but its plain text approach is perhaps just a little too plain for my liking; it feels a little soulless.

Open Library, on the other hand, is positively “homey” in comparison to the other two sites. It doesn’t seem to have the same selection as some of the bigger websites in terms of number of titles available, but I enjoyed that they’ve designed their scans to look like real books. Readers get to see two pages presented side-by-side and animation gives the effect that you are flipping through the pages just as you would in a real book. On top of that you still get the visuals like wear and tear, stains on a page and writing in the margins that you would from any well-loved print book. This format was definitely friendlier to this user.

Still, it’s not the same.

And it’s because it’s not the same that I don’t think that ebooks and the digitization of existing books are going to make print books obsolete any time soon. They may eventually, but I doubt that it’s anything I need to be afraid of seeing in my lifetime. But who knows, as quickly as technology is moving, I may be eating my words by 2050. Of course, by then they’d be nothing more than words on a computer screen, so eating them might prove rather difficult.