Saturday, November 6, 2010
Wk 10: Muddiest Points
Since there are so many challenges to efforts to preserve digital materials (cost, maintenance, labor demands, technology demands), especially for long-term preservation needs, what is the future of digital archiving? i.e. do you think organizations and institutions will begin to abandon the idea of digital preservation due to daunting costs and challenges, or this the future for digital archiving healthy?
Wk 10 Readings
Challenges of Digital Preservation
I plan to take Digital Preservation next term, and these readings have just answered a nagging questions I have had about how the topic of Digital Preservation could fill an entire term. It turns out that digital preservation is much more difficult and tenuous than I imagined.
Begin with Technology Watch Report (2004): in which Brian F. Lavoie takes us through the creation of formal standards for long-term storage of digital data. Beginning with CCSDS's project for digital data generated from space missions, and ending with the impressive, but still rather tenuous Open Archival Information System (OAIS ) standard. Note the six OAIS criteria on page 3. On page 16 of the Technology Watch Report, Lavoie points out that the meaning of OAIS-compliant is "necessarily vague, but that "rigorous interpretation of OAIS compliance will likely emerge... In further readings, the looming problems of cost and human labor associated with digital preservation are keeping this topic on shaky ground.
In Research Challengs in Digital Archiving and Long-term Preservation, Margaret Hedstrom points out that "digital collections are vast, heterogeneous, and growing at a rate that outpaces our ability to manage and preserve them " (1) and that "digital objects require constant maintenance and elaborate life-support systems to remain viable" (2). Security against attacks and system failures pose further challenges to manage. As Lavoie and Hedstrom both point out, "there are no formal economic or business models for digital preservation activities" (2). On this issue of cost, in order to move forward in the world of digital preservation, Hedstrom explains that "there is a critical need to develop tools that automatically supply core metadata, extract metadat from resources at ingest, and restructure and manage metadata over time" (3). On one hand, the field of digital preservation seems almost too daunting; however, the silver lining could be in the obvious need for many jobs in this area as long as organizations can afford to pursue the goal of preservation this type of material.
In Actualized Preservation Threats: Practical Lessons from Chronicling America, Justin Littman gives a great account of some major challenges in digitization projects, as he identifies four preservation threats encountered on the NDNP project digitizing historically significant newspapers. These four categories are: media failure, hardware failure, software failures, and most of all - operator errors, which stands to reason. As Littman points out, hopefully we can all learn from these studies and work toward creating more stable digitization environments.
I plan to take Digital Preservation next term, and these readings have just answered a nagging questions I have had about how the topic of Digital Preservation could fill an entire term. It turns out that digital preservation is much more difficult and tenuous than I imagined.
Begin with Technology Watch Report (2004): in which Brian F. Lavoie takes us through the creation of formal standards for long-term storage of digital data. Beginning with CCSDS's project for digital data generated from space missions, and ending with the impressive, but still rather tenuous Open Archival Information System (OAIS ) standard. Note the six OAIS criteria on page 3. On page 16 of the Technology Watch Report, Lavoie points out that the meaning of OAIS-compliant is "necessarily vague, but that "rigorous interpretation of OAIS compliance will likely emerge... In further readings, the looming problems of cost and human labor associated with digital preservation are keeping this topic on shaky ground.
In Research Challengs in Digital Archiving and Long-term Preservation, Margaret Hedstrom points out that "digital collections are vast, heterogeneous, and growing at a rate that outpaces our ability to manage and preserve them " (1) and that "digital objects require constant maintenance and elaborate life-support systems to remain viable" (2). Security against attacks and system failures pose further challenges to manage. As Lavoie and Hedstrom both point out, "there are no formal economic or business models for digital preservation activities" (2). On this issue of cost, in order to move forward in the world of digital preservation, Hedstrom explains that "there is a critical need to develop tools that automatically supply core metadata, extract metadat from resources at ingest, and restructure and manage metadata over time" (3). On one hand, the field of digital preservation seems almost too daunting; however, the silver lining could be in the obvious need for many jobs in this area as long as organizations can afford to pursue the goal of preservation this type of material.
In Actualized Preservation Threats: Practical Lessons from Chronicling America, Justin Littman gives a great account of some major challenges in digitization projects, as he identifies four preservation threats encountered on the NDNP project digitizing historically significant newspapers. These four categories are: media failure, hardware failure, software failures, and most of all - operator errors, which stands to reason. As Littman points out, hopefully we can all learn from these studies and work toward creating more stable digitization environments.
Saturday, October 30, 2010
Wk 9 Readings
Week 9 Readings: Originally Scheduled for 11/2, but switched to 10/26
All About Searching and Who is In-charge of the Index!
All About Searching and Who is In-charge of the Index!
Readings this week focused on searching and covered a wide range of interesting topics from the history of Z39.50 (a protocol that specifies data structures in the internet environ) to challenges faced when it comes to figuring out what users want and need when it comes to searching for information (Federated Searching: Put it in its Place).
In the article, Definitions and Origins of OAI-PMH, we are shown an example of how easy it is to run into a problem with tagging that results in an inability for the user to search properly. When users searched the “History of Women in America” using the search term “women” they did not get any hits because metadata records were compromised (pg. 6).
When it comes to users and what they want/need, Todd Miller makes a great point in his article, Federated Searching: Put It in Its Place when he explains that “librarians like to search, but users like to find” and that Federated Search and the Library Catalog should have a relationship. The question is how to relax security to make the “information banks” compatible with Federated Searching, but we are still left with question marks about how to makes this happen.
Norbert Lossau makes it clear in Search Engine Technology and Digital Libraries that he favors library collaborations (between libraries and also partnerships with others) to make information available and searchable, and that this would be preferable to leaving scholarly searching collaborations in the hands of larger entities such as Google (2,11).
Finally, The Truth About Federated Searching keeps us hones as it debunks some widely myths about federated searching. We are left with the realization that
- authentication makes it impossible for federated search engines to search all dbs.
- true de-dupe is impossible
- relevancy ranking is not totally relevant.
- federated searching is best consumed as a service not as software
- results are not necessarily better with a federated search engine over a native database search.
These insights boil down to the fact that it is very difficult to get all aspects of information searching in the hands of one central trustworthy organization. Tagging for good metadata search results is daunting, but an even larger problem seems to be in identifying who should build partnerships to create this giant index – should it be library systems partnered with each other, partnered with superpowers like Google and Amazon, or partnered with other organizations. Then what… very complex indeed but very interesting at the same time.
Saturday, October 9, 2010
Wk 7 Questions and Muddiest Points
When constructing a digital library, is one of the goals to incorporate crawler algorithms, indexers, and inquiry handlers? I assume even a tech savvy digital librarian would need outside experts to do this.
Wk 7 Readings
Wk 7 Readings: Access in Digital Libraries - 10/19
The readings offered very interesting details regarding how massive amounts of information are made "miraculously" searchable through the use of crawler algorithms, indexers, and inquiry handlers. Hawkings clear and concise accounts of the challenges of developing crawlers and indexing strategies gave me an understanding of just how complicated it is to develop an intelligent search engine given the volume and unwieldiness of webpages. The equally daunting task of indexing a seeming infinite number of terms due to the influx of many languages used on the web. Hawkings makes it clear that the world of search engine development is not for the novice.
Henzinger's piece is more scientific, and for me somewhat difficult to follow, but definitely pulls together many of the concepts brought forth by Hawkings with regard to seeking improvements to crawlers and search engines. I can see through this article that a lot of research goes into developing search engine stragegies to reduce server overload and improve efficiency. Effective search engines really are amazing.
The readings offered very interesting details regarding how massive amounts of information are made "miraculously" searchable through the use of crawler algorithms, indexers, and inquiry handlers. Hawkings clear and concise accounts of the challenges of developing crawlers and indexing strategies gave me an understanding of just how complicated it is to develop an intelligent search engine given the volume and unwieldiness of webpages. The equally daunting task of indexing a seeming infinite number of terms due to the influx of many languages used on the web. Hawkings makes it clear that the world of search engine development is not for the novice.
Henzinger's piece is more scientific, and for me somewhat difficult to follow, but definitely pulls together many of the concepts brought forth by Hawkings with regard to seeking improvements to crawlers and search engines. I can see through this article that a lot of research goes into developing search engine stragegies to reduce server overload and improve efficiency. Effective search engines really are amazing.
Tuesday, October 5, 2010
Assignment #2 - Flickr URL
The URL for my Flickr collection is: http://www.flickr.com/photos/54577108@N06/
Subscribe to:
Posts (Atom)