Showing posts with label data mining. Show all posts
Showing posts with label data mining. Show all posts

Saturday, April 10, 2010

Visualization Tools Revisited


Six months ago I used Wordle to create a visual representation of my blog (see October's Wordle Wordle Wordle). By looking at the word cloud of my labels that was generated by Wordle, it was easy to see that I wrote quite frequently about public history, digital history, UWO, museums, human rights and the Holocaust. None of that is a big surprise when you consider that I'm a student at UWO in the Public History MA Program and that I am taking courses which deal with public history, digital history and museums and that I have a strong interest in human rights issues and the Holocaust.

In that entry, I wondered what a similar visualization would look like in six months or a year. Would it change drastically and reflect new interests and experiences? Or, would it remain the same, reflecting a constant interest in certain topics?

As you can see by looking at the word cloud at the top of this blog post, the visualization has gotten considerably more complex and diverse over the last six months (I have added dozens of new labels to the cloud). However, I would argue that it still largely reflects the initial trends of my blog. The single exception to this relative continuity is the inclusion of new labels which show an increased awareness of and interest in topics relating to local history and heritage (who knew Durham Region had such fascinating history?!).

That's a mildly interesting observation to make, but so what? What can a visual representation of my blog tell you that a quick skim of my labels list cannot? Sadly, I'm afraid the answer is probably nothing.

In October I had a cautious optimism for the use of digital visualization tools as techniques to help with the data mining process. Now, I am thoroughly unconvinced as to their usefulness. It might take me 2 or 3 seconds longer to scroll through the alphabetized list of labels that I have used in my blog since its inception than it does to glance at word cluster like the one above, but it tells me exactly the same thing. In fact, scrolling the list of labels actually saves me considerable time when you consider that I had to first go to the Wordle website and input my data before I could create the word cloud in question. Also, I can't help but notice that the picture that was produced by the generator is now so crowded with new, low-level labels that it has become an eye strain to try and extract any but the most obvious data from it.

In my humble opinion, Wordle and similar word/tag cloud generators may create pretty pictures, but they are a bust as far as their usefulness to the study of history goes.

But what about other visualization tools? Surely I can't paint all tools with the same brush, can I? You wouldn't think so. The esteemed Dan Cohen certainly doesn't. He used different visualization technologies in his research to map whether Americans prayed or watched CNN after hearing about the terrorist attacks on September 11th, 2001. Cohen stands by the fact that it was only through mapping the results of his survey onto a satellite image of the USA provided by Google Earth that he was able to observe that people in rural areas of the United States were far more likely to pray after hearing of the attacks than they were to watch CNN. Conversely, people in urban areas were far more likely to turn to CNN for information and solace than they were to pray.

So it was with this conflicting ringing endorsement by Dan Cohen (digital history guru) and growing personal skepticism that I attended the Historical Geographic Information Systems (Historical GIS) lecture given several weeks ago by Don LaFreniere, a colleague from the Geography Department. Don is a huge fan of GIS and the objective of his lecture was three-fold:

1) To teach us the basics of navigating GIS software
2) To introduce us, as historians, to the potential applications of GIS programs to our research questions and future careers in public history
3) To explain how he is currently using GIS visualization tools in his research process

Extremely knowledgeable and enthusiastic, Don did a fantastic job of patiently leading even the most technologically challenged of us through his tutorial. But even Don's enthusiasm couldn't convince me of the usefulness of visualization techniques. In order to get to the point of visualization Don and a team of undergraduates had to spend more than two years compiling a database of census materials from which they drew the information they mapped using the GIS software. For my money, it was the database that was truly the incredible feat, not the mapping we were able to do with the data points. Using the spreadsheets that he had created, we were able to search for specific criteria and filter the overall findings according to our desired specifications. At this point we were able to see all of the relevant details and make (in my opinion) the same observations that we were able to make after applying the visualization tools to the data set.

I don't know. Maybe I'm missing something. Maybe, as Sara Sirianni suggests in her blog post about the workshop, visualization techniques are the greatest thing since sliced bread. Maybe.

For now I remain skeptical.

Saturday, December 19, 2009

When You Look Into a Digital Abyss, Does It Also Look Into You?: A Look at Historical Abundance on the Internet

It seems fitting that this blogging assignment on the abundance of materials available on the internet will be my last post of the semester. Assigned in September, it’s taken me nearly four months to get past more than a sketchy outline of what I wanted to discuss in this post. It’s not through lack of trying though – I’ve probably sat down at my computer to write it more than a half dozen times only to find myself quickly distracted by an errant thought or query which led me to an internet search for information which led me to a hyperlink for a related topic which led me to either another hyperlink or another question and so on and so on until the time that I’d set aside for writing that day had run out.

Over the course of the semester I have likened this process of traveling from hyperlink to hyperlink to that of Alice falling down the rabbit hole in Lewis Carroll’s Alice in Wonderland – you never know what weird and wonderful things you’ll see during your trip, nor do you end up quite where you expected to when all is said and done. I’m not complaining; I’ve come across hundreds of amazing websites, useful tools, interesting articles and countless pieces of fascinating information this way. In truth, I think that my education and knowledge base are probably significantly broader as a result.

But this “infinite archive,” [1] as the internet has been described in recent years, is something of a double-edged sword, especially for academic and public historians. In the words of Daniel J. Cohen and Roy Rosenzweig [2]:

The digital era seems likely to confront historians—who were more likely in the past to worry about the scarcity of surviving evidence from the past—with a new “problem” of abundance. A much deeper and denser historical record, especially one in digital form, seems like an incredible opportunity and gift. But its overwhelming size means that we will have to spend a lot of time looking at this particular gift horse in the mouth... (2005)

As Cohen and Rosenzweig point out, in the past historians had to build their craft on the sometimes tenuous threads of the surviving evidence, a harsh reality which certainly restricted historical research. I also think that the caveat “that they could get their hands on” needs to be added to the aforementioned quote. Historical records and artifacts did not rest in a single, central and easily accessible location. As a result, historical research was often regional, not necessarily in terms of topic but in terms of sources used to support the arguments of historians. It’s not unreasonable to think that historians in the past could become acquainted with all of the materials available to them on their subject in their area and still be missing that critical document or artifact that had the potential to either cement or refute their claims.

Thankfully, these constraints of time and distance have largely been removed by improvements in technology. With the growth of the internet and the sheer abundance of information that is now available online, academics in even the most remote regions of the world suddenly have access to primary and secondary sources they may have previously been unaware even existed. For instance, the internet and massive digitization efforts have made it possible for students of history, like our own Megan Arnott, to be able to research and write on the Medieval Era without ever having to leave her home in Canada. Further, through email, chat rooms, Twitter, Facebook, Google Wave and any other number of social networking tools, Megan (and historians, in general) are better equipped than ever to quickly and easily communicate and collaborate with colleagues around the world.

Unfortunately, this is where the “‘problem’ of abundance” that Cohen and Rosenzweig speak of comes in. Those same digitization efforts that seem amazing at the outset do have their drawbacks – even as early as 1996, the internet had become “a sprawling megalopolis that no one person could fully explore” [3]. With thousands, if not tens of thousands of new websites added to the internet each day, how does an historian keep up?

The short answer is, we can’t. Instead, we must do as Cohen and Rosenzweig suggest and work, as historians in concert with computer programmers, website designers, researchers, publishers, museum curators, librarians and archivists to create “sophisticated statistical and data mining tools to do some of the looking” for us [2]. And, while the internet currently acts as an amazing archive of human record, no singular database has yet been created in which all historical records within it are stored. Because of this, I believe that it is the job of public historians to support the creation and use of these data mining tools and smaller databases in order to help make as much of the digital record as possible both knowable and useful to academics and the public alike.

Another issue that must be considered in connection with the growing abundance of historical sources on the internet is that of authority within the study of humanities and social sciences. As we discussed in class earlier in the semester, while the constraints of time, distance and availability no doubt frustrated historians in the past, they also helped to make the professionalization of history what it is today. For instance, with the scarcity of information came a tight control on existing intellectual materials; museums, archives, and libraries, especially those housing extremely rare items, were often restricted to only the most respected historians in the field (and largely only those associated with a trusted institution). This resulted in the control of scholarship resting in the hands of the few. It likely also served to privilege certain interpretations of historical events.

But one of the more amazing “gifts” to come out of the abundance of materials on the internet is the growth of democratization of the historical process. As more items that were previously inaccessible to the general public are digitized and compiled in databases like the Internet Archive, more grassroots historians are popping up each day. They read, research and comment on what they see and learn, adding a unique perspective and voice to literature on a variety of historical topics. They now have not only the ability to share their interpretations of historical events, but now they also have the primary and secondary sources to support it.

Don’t get me wrong, even the internet hasn’t been able to overcome the entrenched intellectual hierarchy which privileges the contributions made to the field by professional scholars in universities and other elite institutions [3]. I’m not even arguing that it necessarily should. After all, academic historians have spent years training in proper research methodology and historiography. But neither should the amateur historian be dismissed out of hand. As someone who is untrained in this type of research, I believe that there is a lot that historians can learn from these amateurs about how those outside of the ivory tower interact with history, what they need and want from it, and perhaps where academic research needs to go in order to satisfy the interests of the wider public.

And that’s where I believe public historians come in. I think the task now falls to us to mediate between the two groups (the public and the academics) in order to ensure that the contributions of both sides are acknowledged, valued and made accessible to the other.

Sources:

[1] Turkel, William J. Digital History Hacks: Methodology for the Infinite Archive (2005-08). [Weblog.]

[2] Cohen, Daniel and Roy Rosenzweig, “Web of lies? Historical knowledge on the Internet,” First Monday, Volume 10, Number 12 - 5 December 2005, http://firstmonday.org/htbin/cgiwrap/bin/ojs/index.php/fm/article/view/1299/1219.

[3] Cohen, Daniel and Roy Rosenzweig, “Exploring the History Web,” Digital History: A Guide to Gathering, Preserving and Presenting to Past on the Web, http://chnm.gmu.edu/digitalhistory/exploring/1.php.

Thursday, December 17, 2009

My Inner Feminist Hates Jane Austen

I know that as historians we are not supposed to judge the past by the standards of the present, but in this instance, I just can’t help myself; the piece of my personality that’s a little bit feminist hates Jane Austen novels. I think her male characters could benefit from being socked in the jaw a time or two and I dislike her female characters enough to think that they might actually deserve the idiots they end up married to.

This week’s experiment in data mining reminded me of a paper I had to write for my first year English Literature course on the theme of love and marriage in Jane Austen’s Pride and Prejudice. After reading the novel, I ended up going through the book page-by-page and physically highlighting all of the sections pertaining to either love or marriage before I began to formulate my arguments and write my essay. Needless to say, it was an extremely tedious process and one I’m not eager to repeat.

But now it seems there’s a simpler way to gather the same information using a free, Canadian-made digital data mining tool. Using a plain text copy of the novel available on the Project Gutenberg website, I copied the URL into TAPoR (Text Analysis Portal for Research)’s Word List and Concordance Tools, and ran three different word/pattern searches for “love,” “marriage” and “money.” I was interested in seeing how many instances of each word were present in the novel, if the words were used in relation to each other, where in the novel the majority of these three words were used, and if that would help me to come to any new conclusions about the themes of the novel.

The first thing I found out from using TAPoR’s Word List Tool was that the word “love” was used 91 times, “marriage” 66 times and “money” only 29 times in Pride and Prejudice. Considering that we know money was a major factor in the marriages of the period in which Jane Austen was writing, these results were somewhat surprising. I then used the Concordance Tool to tell me where and in what context each word was used. “Love” was used most often in the first third of the novel, with a small spike again towards the very end. “Marriage” spikes mildly towards the end of the first third of the book but is practically off the charts in the last quarter of the book when Austen is working to tie all of the relationships between the characters into neat little bows. And “money”, of course, is mentioned most often just as the discussions about marriage start to heat up.

My favourite aspect of the Concordance Tool is that although you can set the program to only return the specific words you are looking for, it also lets one look at their search results in terms of the context of the words, lines, sentences or paragraphs surrounding the specified words. Essentially the tool allows a historian to look at a text both quantitatively and qualitatively.

The one downfall, as Sara mentioned in her blog post about TAPoR, is that the returned search results lack page references. I understand that this is due to the fact that one can currently only use a plain text copy of a text with the TAPoR tools, but it does seem like an aspect of the program that could benefit from some improvements. This improvement would be especially beneficial for those in an academic environment who must be careful to adhere to copyright protections and site all of their sources in footnotes/endnotes and bibliographies.