Monday, September 22, 2008
Extracurricular activity
On a related note, I am going to be facilitating an extracurricular activity on Thursday afternoons at YouthBuild Charter School. We really don't know what form this will take -- It will depend on what the students are interested in -- but it's our vision to give them a way to put their year-culminating portfolio projects on the web, help them build simple websites to promote their businesses or passions, and reinforce computer skills. I am going to be creating a wiki as an example, and then linking it to and from this blog, so hopefully that will mean I will create more meaningful posts for this site as well -- Lead by example, right?
New Special Collections website
Well, I am still here! We recently launched a
new Special Collections website. So I may be not writing, but at least I am cranking! We have gotten through our first 60 batch of Owen Wister letters, and are working on creating our Digital Wisteriana / La Salliana concept. We've been spending a lot of time updating Wikipedia with information on the Wisters. Given that the Wisters' estate was on the La Salle campus, we are really poised to be an authority on the family. Many Google searches take you to another website that I am now in charge of -- Belfield and Wakefield: A Link to La Salle's Past. I recently took over this website from Dr. Jim Butler, who had made it with an honors English class in something like 1994. I think its going to be a wonderful opportunity to incorporate digital collections and things like internet based geneologies and such and breathe some life into a website that hasn't been updated in a while. I have changed the style sheet on this website to make it conform to the La Salle "brand" more and started incorporating some images from the Wister Special Collection, so please look for this to develop more over time.
new Special Collections website. So I may be not writing, but at least I am cranking! We have gotten through our first 60 batch of Owen Wister letters, and are working on creating our Digital Wisteriana / La Salliana concept. We've been spending a lot of time updating Wikipedia with information on the Wisters. Given that the Wisters' estate was on the La Salle campus, we are really poised to be an authority on the family. Many Google searches take you to another website that I am now in charge of -- Belfield and Wakefield: A Link to La Salle's Past. I recently took over this website from Dr. Jim Butler, who had made it with an honors English class in something like 1994. I think its going to be a wonderful opportunity to incorporate digital collections and things like internet based geneologies and such and breathe some life into a website that hasn't been updated in a while. I have changed the style sheet on this website to make it conform to the La Salle "brand" more and started incorporating some images from the Wister Special Collection, so please look for this to develop more over time.
Friday, July 18, 2008
Developments
We are cranking! It's amazing what a new budget year can do to a project. We have set up our digitization lab and are beginning to scan. The first project I have decided to tackle is the Wister family papers -- We are beginning with the letters of Owen Wister, author of the Virginian. I am busy developing my cataloging database, which we will be attempting to use METS and MODS, so my head is all over the LC website today. I will add some pictures soon.
Friday, June 27, 2008
Greenstone integration
When my puppy was just a few months old, we got a silly little card from our vet saying that "everyday is a new adventure." In spite of how corny this sounds, such is the same for a librarian trying to plan their first digital library project.
My project, I feel, has really taken on the shape and direction that I was initially lacking ever since our Director of Academic Computing suggested that whatever I choose should dovetail into Sakai. Whether we use Sakai or not is almost moot, at this point -- the real idea here is that our systems should be wide open and able to "talk" to our other systems. AC is working on that on their end; likewise, we should too.
So... I just learned yesterday about Greenstone's ability to integrate with DSpace. Apparently, you can develop in Greenstone and then transfer your collections and metadata to DSpace rather easily, and vice versa. You can use Greenstone as the frontend for DSpace, or you can scrap it entirely.
We are still looking for our "quick and dirty" UI to mount our collections soon -- This is looking like a highly likely route.
Further reading:
Computer Science Thesis from Virginia Tech, outlining an open XML schema for DSpace
My project, I feel, has really taken on the shape and direction that I was initially lacking ever since our Director of Academic Computing suggested that whatever I choose should dovetail into Sakai. Whether we use Sakai or not is almost moot, at this point -- the real idea here is that our systems should be wide open and able to "talk" to our other systems. AC is working on that on their end; likewise, we should too.
So... I just learned yesterday about Greenstone's ability to integrate with DSpace. Apparently, you can develop in Greenstone and then transfer your collections and metadata to DSpace rather easily, and vice versa. You can use Greenstone as the frontend for DSpace, or you can scrap it entirely.
We are still looking for our "quick and dirty" UI to mount our collections soon -- This is looking like a highly likely route.
Further reading:
Computer Science Thesis from Virginia Tech, outlining an open XML schema for DSpace
Wednesday, June 4, 2008
New PACSCL website
I am fortunate to be La Salle's web liasion to the Philadelphia Area Consortium of Special Collection Libraries. They have just launched a new dynamic website for member libraries to announce collections, exhibitions, lectures, etc. It is set up as a content management system, for member libraries to manage their own posts.
Check it out:
www.pacscl.org
Check it out:
www.pacscl.org
And the prize goes to... DSpace!
At long last, I have settled on developing DSpace for La Salle. Everytime that I talk repositories to faculty, they want to start throwing teaching materials into it. I found an interesting paper that talks about the possible integrations of DSpace with Sakai, an alternative to Blackboard which I was asked to find out how my systems may integrate with. Also, I saw a posting yesterday that Fedora and DSpace have recently decided to try to identify some areas to work together. We still don't feel confident that we could develop Fedora here without a programmer, but we feel good about DSpace being within our skillsets. I still have a crush on Fedora, and would welcome the intersection of the two systems to allow me more opportunity to play with that system as well.
I had originally not considered DSpace, as I always thought that it was just for ETDs. People are using it differently than that, and as I started demoing CONTENTdm and asking users about their experiences, I started getting pointed to DSpace more and more, given our desire to manage diverse scholarly materials and didactic resources. The dealsealer was seeing how Australian National University migrated from CONTENTdm into DSpace and came up with something what I would be looking to create for our Special Collections: http://anulib.anu.edu.au/subjects/ap/digilib/.
We're still early enough in our digitization game here that I don't want to end up locking us into a proprietary database without getting a greater sense of the digital directions of the campus. The openness of the code and flexibility of metadata input/output will set us up for optimum interoperability with future systems. Also, DSpace runs on Dublin Core, which is far from intimidating, but also allows input of METS, which I think should allow for a fun challenge.
Now the preservation question... OCLC's digital archive service is still looking pretty good, as they are reasonably priced compared to other services. I also may be looking at using the money that I save on software acquisition to mount an argument for Sun Microsystems' Honeycomb or something along those lines. It is too onorous to throw manual digital preservation administration and data backup on top of a librarian's other duties -- Therefore, all redundancies and reporting will be automated and/or outsourced so that it is done efficiently and effectively.
I had originally not considered DSpace, as I always thought that it was just for ETDs. People are using it differently than that, and as I started demoing CONTENTdm and asking users about their experiences, I started getting pointed to DSpace more and more, given our desire to manage diverse scholarly materials and didactic resources. The dealsealer was seeing how Australian National University migrated from CONTENTdm into DSpace and came up with something what I would be looking to create for our Special Collections: http://anulib.anu.edu.au/subjects/ap/digilib/.
We're still early enough in our digitization game here that I don't want to end up locking us into a proprietary database without getting a greater sense of the digital directions of the campus. The openness of the code and flexibility of metadata input/output will set us up for optimum interoperability with future systems. Also, DSpace runs on Dublin Core, which is far from intimidating, but also allows input of METS, which I think should allow for a fun challenge.
Now the preservation question... OCLC's digital archive service is still looking pretty good, as they are reasonably priced compared to other services. I also may be looking at using the money that I save on software acquisition to mount an argument for Sun Microsystems' Honeycomb or something along those lines. It is too onorous to throw manual digital preservation administration and data backup on top of a librarian's other duties -- Therefore, all redundancies and reporting will be automated and/or outsourced so that it is done efficiently and effectively.
Monday, May 19, 2008
DAM list
So, I keep using this blog more for the list of quick links at the right to cool digital libraries and less like a "web log." The bad librarian in me doesn't tag these entries anyway. I do think there is a shortage of comprehensive lists of systems, and I will do my best to keep this one up over time.
Wednesday, April 23, 2008
Hiatus over, summer's coming...
So, I have successfully lasted in my new position for about 9 months now, and can see the end of the school year. At some point around February, my schedule started to get crazier, and I knew that is a sign of my successes. Here's brief summary of our activities for the last 2 months.
DAM software -- Jury is still out, but we're getting close to a decision on how to proceed. After much deliberation, open source is out, as we just don't have the infrastructure to support it. CONTENTdm is still in the running. I am going to be looking at 3 more products -- VTLS's Vital, which runs on FEDORA, ArchivalWare and Olive. Hopefully we will have 1 -2 possible scenarios by next week.
OCLC's Dark Archive is another service that we are looking at, which would run complementary to the asset management system. I am not sure how widely publicized this is, as there are only currently about 30 customers for it. But OCLC will can be used as an outsourcing agent to manage the big wieldy masters, through a combo of tape and hard disk. They keep one copy of everything online on disk, and then keep 3-5 copies rotating through on tape. I feel really good about our possible adoption of this service.
I also successfully implemented our first distance library instruction service -- One of our prof's wanted me to go out to Buck's County to talk about the AV department, so I worked with our Academic Computing department to use recently acquired video conferencing equipment. It was a feel good experience for all, as the class was on new technology in the classroom, and we were able to lead by example. Also AC was excited to have an opportunity to use the equipment in a non-traditional environment, meaning not in one of the smart classrooms. I feel like in time library instruction will have to have this service readily available to accomodate the needs of further distance education iniatives.
I've been busy gallavanting around, too -- Attended the Computers in Libraries conference, Grant Writing at Palinet and a Web 2.0 / Scholarly Communication Symposium at Drexel. Comparable to drinking from a hose, I am fully information saturated. I will be excited to announce our new initiatives as they are fully flushed out in the next couple of weeks, though.
DAM software -- Jury is still out, but we're getting close to a decision on how to proceed. After much deliberation, open source is out, as we just don't have the infrastructure to support it. CONTENTdm is still in the running. I am going to be looking at 3 more products -- VTLS's Vital, which runs on FEDORA, ArchivalWare and Olive. Hopefully we will have 1 -2 possible scenarios by next week.
OCLC's Dark Archive is another service that we are looking at, which would run complementary to the asset management system. I am not sure how widely publicized this is, as there are only currently about 30 customers for it. But OCLC will can be used as an outsourcing agent to manage the big wieldy masters, through a combo of tape and hard disk. They keep one copy of everything online on disk, and then keep 3-5 copies rotating through on tape. I feel really good about our possible adoption of this service.
I also successfully implemented our first distance library instruction service -- One of our prof's wanted me to go out to Buck's County to talk about the AV department, so I worked with our Academic Computing department to use recently acquired video conferencing equipment. It was a feel good experience for all, as the class was on new technology in the classroom, and we were able to lead by example. Also AC was excited to have an opportunity to use the equipment in a non-traditional environment, meaning not in one of the smart classrooms. I feel like in time library instruction will have to have this service readily available to accomodate the needs of further distance education iniatives.
I've been busy gallavanting around, too -- Attended the Computers in Libraries conference, Grant Writing at Palinet and a Web 2.0 / Scholarly Communication Symposium at Drexel. Comparable to drinking from a hose, I am fully information saturated. I will be excited to announce our new initiatives as they are fully flushed out in the next couple of weeks, though.
Tuesday, February 5, 2008
Photomerge in CS3
We have only very small scanners at La Salle. I have a nice new Epson Expression 10000XL on order, but in the meantime I am using circulation's little 8 1/2 x 11 scanner that they use for e-reserves. Its a little cumbersome and far from ideal, but it allows me the opportunity to get to know the Photomerge function in Photoshop CS3 -- which is a certainly handy enough tool to share.
So here were my raw scans of an object that is approx. 20 in X 16 in.
So here were my raw scans of an object that is approx. 20 in X 16 in.
I opened all these up in Photoshop. All I did was go to File -- Automate -- Photomerge, and clicked add all open files, click OK -- And viola! Like a dream, it matches up all the scans perfectly to create a cohesive image, and even comspenates for seam lines, edges, and variations in scan quality and color.
I over cleaned it up a little bit in Photoshop, so this is not an archival quality object, but it gives me a "real" enough object to work with through our demos, and to see how MARC records will relate to it.
Pretty handy, eh? Created a good looking digital object in about 15 min, from scanning to finish.
Wednesday, January 16, 2008
Academy of Natural Sciences Digital Collections
I suppose this is also also a repository spotlight, but it is also a plug for a friend's institution. A friend of mine works as the cataloger at the Academy of Natural Science's library. She forwarded me the link to their digital collections today, and I really love the design of the site. They don't necessarily have a lot of their digitization work up there, but what they do have up there is colorful, bright and well-layed out. I would love to be able to click on a picture to get the full information, but it does a wonderful job of placing images in context and explaining what they are. Its quality of content over quantity, and I think in my heart that's what I'm feeling with all my introspection on design. I would go back to the site simply because it has beautiful vibrant images, and I can recall what the site made me feel after I leave it.
They, from what I know, don't have an internet access system, such as ContentDM or DigiTool or the other ones that I'm looking at. They rely on just solid design that integrates tightly and seemlessly into the Academy's website. The other sites that have content in an out-of-the-box looking DAM just don't give you the feeling of a) the institution's values; b) the richness of the materials; c) the method and rationale for digitization; d) connection to the rest of the institution.
Thanks, ANS -- I think you got it goin' on.
http://www.ansp.org/museum/digital_collections/
They, from what I know, don't have an internet access system, such as ContentDM or DigiTool or the other ones that I'm looking at. They rely on just solid design that integrates tightly and seemlessly into the Academy's website. The other sites that have content in an out-of-the-box looking DAM just don't give you the feeling of a) the institution's values; b) the richness of the materials; c) the method and rationale for digitization; d) connection to the rest of the institution.
Thanks, ANS -- I think you got it goin' on.
http://www.ansp.org/museum/digital_collections/
Tuesday, January 15, 2008
Repository spotlight: Joan Flasch Artist's Book Collection
http://digital-libraries.saic.edu/cdm4/index_jfabc.php?CISOROOT=/jfabc
I just got back to the office from ALA's midwinter meeting, and will be sharing my notes about the conference over the course of the next few weeks.
I did want to make a note of the School of the Art Institute's Joan Flasch Artist's Book Collection, as design-wise it's one of the first ContentDM sites that I have seen that looks like it has transcended the look and structure that ContentDM gives you right out of the box. I thought this was notable, given my last posting. I like the interface -- Simple, straightforward, unintimidating and unique.
I got the impression that they are using ContentDM because it is available, rather than having chosen that for themselves. They are using simple Dublin Core for their metadata, as that is what the software supports. Their ILS is Voyager and they use OCLC Connexion to push a MARC record into ContentDM. The woman presenting this collection ended with plea: "Ask OCLC to use METS and MODS *please!*"
I just got back to the office from ALA's midwinter meeting, and will be sharing my notes about the conference over the course of the next few weeks.
I did want to make a note of the School of the Art Institute's Joan Flasch Artist's Book Collection, as design-wise it's one of the first ContentDM sites that I have seen that looks like it has transcended the look and structure that ContentDM gives you right out of the box. I thought this was notable, given my last posting. I like the interface -- Simple, straightforward, unintimidating and unique.
I got the impression that they are using ContentDM because it is available, rather than having chosen that for themselves. They are using simple Dublin Core for their metadata, as that is what the software supports. Their ILS is Voyager and they use OCLC Connexion to push a MARC record into ContentDM. The woman presenting this collection ended with plea: "Ask OCLC to use METS and MODS *please!*"
Monday, January 7, 2008
Starting off the new year...
So, we are pretty much at the same place as we were in the middle of December. I got myself lost in a bunch of Cascading Style Sheets and PhotoShop tutorials and have recently reemerged with nothing really to show for my time. I do feel very confident in my CSS skills, which I believe will be of great help in the long run.
What I have noticed is that I am fixated on design over content right now. I don't know if that is where I really should be, but I do think there is merit to fretting over how something looks. I am a fully aesthetic person, and I feel like I don't go back to a digital library site unless it looks good to me. So many sites I have checked out may have great content, but if the design looks like a DAM system right out of the box, I turn off to it.
Allowing the content only to fuel a site is good for researchers, yes, as (hopefully) they can get on and find what they need in the most efficient manner. La Salle's Connelly Library, it has been noted by my director, is not a research library, nor will it most likely ever be. The Special Collections are strong and relatively comprehensive -- thought my director has also pointed out that he's not sure the research value in our bible collection, nor is our Viet Nam collection a "traditional" collection, as it is based somewhat in popular culture.
What I am getting at here is that I am rather unconvinced that the strength of our content alone would allow for us to slack off on design. I want my DAM to look awesome, so that repeat visitors will be coming back for both the content and the design of the site. Whichever DAM I choose will integrate itself back and forth between our library website and its special collections pages. I want our DAM to look first like La Salle University's digital collections, rather than firstly like just another university digital collections site.
What I have noticed is that I am fixated on design over content right now. I don't know if that is where I really should be, but I do think there is merit to fretting over how something looks. I am a fully aesthetic person, and I feel like I don't go back to a digital library site unless it looks good to me. So many sites I have checked out may have great content, but if the design looks like a DAM system right out of the box, I turn off to it.
Allowing the content only to fuel a site is good for researchers, yes, as (hopefully) they can get on and find what they need in the most efficient manner. La Salle's Connelly Library, it has been noted by my director, is not a research library, nor will it most likely ever be. The Special Collections are strong and relatively comprehensive -- thought my director has also pointed out that he's not sure the research value in our bible collection, nor is our Viet Nam collection a "traditional" collection, as it is based somewhat in popular culture.
What I am getting at here is that I am rather unconvinced that the strength of our content alone would allow for us to slack off on design. I want my DAM to look awesome, so that repeat visitors will be coming back for both the content and the design of the site. Whichever DAM I choose will integrate itself back and forth between our library website and its special collections pages. I want our DAM to look first like La Salle University's digital collections, rather than firstly like just another university digital collections site.
Monday, December 3, 2007
Where we are at...
I am putting off thinking about what my access tool will be until the ALA Mid-Winter Meeting, which is here in Philadelphia, and will allow for me to be totally overwhelmed by the vendor exhibits. My director and I have talked briefly, however, and I am on track to acquire my software sometime between April and June.
In the meantime, I get to play with Special Collections. I'm really listening to Dr. Skinner when she said that you need to know your collection before you choose the tool, so I'm revamping the Special Collections website and photographing the collections to learn what we have, meaning both physical objects and their bibliographic control. What diverse, fun stuff!
Here's a couple of highlights:

This is an illuminated bible from the 14th Century.

This is an artist's book of Ho Chi Minh, handprinted on rice paper in a handmade case.

And then there are kitchen matches from the premier of Rambo 3 in France.
This quickly shows the range of materials I will be working with. We have both the written word, preserved carefully in our Bible collection, and then there is the Vietnam War collection, intentionally a collection of artistic expressions including both fine art and ephemeral novelties. My challenge here is to come up a methodology for exposing these materials in a thought-out, cohesive manner. Hopefully by the time I'm ready to choose my software, I'll be ready to tackle the digital needs of the collections.
In the meantime, I get to play with Special Collections. I'm really listening to Dr. Skinner when she said that you need to know your collection before you choose the tool, so I'm revamping the Special Collections website and photographing the collections to learn what we have, meaning both physical objects and their bibliographic control. What diverse, fun stuff!
Here's a couple of highlights:
This is an illuminated bible from the 14th Century.
This is an artist's book of Ho Chi Minh, handprinted on rice paper in a handmade case.
And then there are kitchen matches from the premier of Rambo 3 in France.
This quickly shows the range of materials I will be working with. We have both the written word, preserved carefully in our Bible collection, and then there is the Vietnam War collection, intentionally a collection of artistic expressions including both fine art and ephemeral novelties. My challenge here is to come up a methodology for exposing these materials in a thought-out, cohesive manner. Hopefully by the time I'm ready to choose my software, I'll be ready to tackle the digital needs of the collections.
Thursday, November 29, 2007
Stewardship of Digital Assets @ Palinet (Post 2)
Alright -- Onto Day 2 of my notes from the Stewardship of Digital Assets workshop!
Dr. Katherine Skinner from Emory University and the MetaArchive spoke for 3/4 of the day. Her interest in digital preservation has risen from her doctoral research on emergent fields. While digital preservation is a few steps beyond just stumbling in the dark at this point, it is still an emergent field and we should "get used to the discomfort" of what goes along with being on the bleeding edge.
In her definition, Skinner describes digital preservation as the "management and maintenance of digital information over a long period of time." How long, she says, is unknown at this point, but its sure is longer than many of our access systems will be around, which is currently running at about 3 years per system. This is why adopting recognized standards are important -- for interoperability -- between systems and between collaborating institutions.
Skinner's largest project of the last few years has been in developing the MetaArchive, which was in turn developed out of the Lots of Copies Keeps Stuff Safe (LOCKSS) framework created by Stanford. The idea behind LOCKSS and the MetaArchive is that a minimum of 6 copies of information is stored across a large geographic space (could be even across several continents). Automated checks continually verify the accuracy and completeness of data. If anything happens to a single file in one of the repositories then the other 5 check to make sure their data is complete, and replace the bad or corrupt file in the 6th location. This idea of spreading information across a great geographic area could help restore information in the case of a major disaster. An example that Skinner gave was NPR content that was destroyed by Hurricane Katrina was able to be restored from the mirrored content kept at Emory University.
The biggest "AHA!" that came out of the day for me was that if you choose your access tool too early in the digital library design process, then your collection ends up becoming bound by the tool. By examining and fully understanding your collection, you can choose a tool that allows for curation of the materials that complements the care received by the analog materials. Process for selection should be identifying materials, seeing what information is available (ie. cataloging record) and then choose the metadata. Only when this process is complete should the tool be selected. I will have more thoughts on how this will change my approach to developing my repository in a later post.
The presentation was great, and the instructors knowledgable and friendly. They realize fully that as the field is new, no one really truly knows what they are doing for the long haul, and the more that we can help each other out, the more successful the digital preservation program will be.
Dr. Katherine Skinner from Emory University and the MetaArchive spoke for 3/4 of the day. Her interest in digital preservation has risen from her doctoral research on emergent fields. While digital preservation is a few steps beyond just stumbling in the dark at this point, it is still an emergent field and we should "get used to the discomfort" of what goes along with being on the bleeding edge.
In her definition, Skinner describes digital preservation as the "management and maintenance of digital information over a long period of time." How long, she says, is unknown at this point, but its sure is longer than many of our access systems will be around, which is currently running at about 3 years per system. This is why adopting recognized standards are important -- for interoperability -- between systems and between collaborating institutions.
Skinner's largest project of the last few years has been in developing the MetaArchive, which was in turn developed out of the Lots of Copies Keeps Stuff Safe (LOCKSS) framework created by Stanford. The idea behind LOCKSS and the MetaArchive is that a minimum of 6 copies of information is stored across a large geographic space (could be even across several continents). Automated checks continually verify the accuracy and completeness of data. If anything happens to a single file in one of the repositories then the other 5 check to make sure their data is complete, and replace the bad or corrupt file in the 6th location. This idea of spreading information across a great geographic area could help restore information in the case of a major disaster. An example that Skinner gave was NPR content that was destroyed by Hurricane Katrina was able to be restored from the mirrored content kept at Emory University.
The biggest "AHA!" that came out of the day for me was that if you choose your access tool too early in the digital library design process, then your collection ends up becoming bound by the tool. By examining and fully understanding your collection, you can choose a tool that allows for curation of the materials that complements the care received by the analog materials. Process for selection should be identifying materials, seeing what information is available (ie. cataloging record) and then choose the metadata. Only when this process is complete should the tool be selected. I will have more thoughts on how this will change my approach to developing my repository in a later post.
The presentation was great, and the instructors knowledgable and friendly. They realize fully that as the field is new, no one really truly knows what they are doing for the long haul, and the more that we can help each other out, the more successful the digital preservation program will be.
The internet is running out of space!
"Consumer demand for bandwidth could see the internet running out of capacity as early as 2010, a new study warns."
http://news.bbc.co.uk/2/hi/technology/7103426.stm
I can't stop thinking about this study that reports that the internet could run out of bandwidth by 2010, and wonder how we -- internet users, content creators, and/or librarians -- could help solve this problem. Digital information is "invisible," or so it seems... You upload to the internet and its physically away from you -- Yet accessible as need be. But how much stuff on the internet is junk? I know that I personally have all the images and pages from my old website still hanging out on an FTP site somewhere -- And those images were probably unnecessarily large, as that was before I know anything about correct imaging for projects and webpages. If everyone just cleaned out their Flickr accounts, or deleted old webpages, could we then get another year or two out of the internet at its current capacity? By realizing that data truly are physical objects, perhaps we feel a greater sense of responsibility to the upkeep and care.
Its physics, its kind of string theory, I know, but as I sit here right now I am surrounded by the internet. I can't see it or feel it, but its here, physical, voluminous and completely disorganized. A Tech Director at a former job of mine always wanted to take a filing cabinet, cram if full of paper in dissarray, with materials falling out of it and jammed into the drawers. "This is what the server actually looks like," would be the message. I am sure that the same analogy would apply to the internet.
http://news.bbc.co.uk/2/hi/technology/7103426.stm
I can't stop thinking about this study that reports that the internet could run out of bandwidth by 2010, and wonder how we -- internet users, content creators, and/or librarians -- could help solve this problem. Digital information is "invisible," or so it seems... You upload to the internet and its physically away from you -- Yet accessible as need be. But how much stuff on the internet is junk? I know that I personally have all the images and pages from my old website still hanging out on an FTP site somewhere -- And those images were probably unnecessarily large, as that was before I know anything about correct imaging for projects and webpages. If everyone just cleaned out their Flickr accounts, or deleted old webpages, could we then get another year or two out of the internet at its current capacity? By realizing that data truly are physical objects, perhaps we feel a greater sense of responsibility to the upkeep and care.
Its physics, its kind of string theory, I know, but as I sit here right now I am surrounded by the internet. I can't see it or feel it, but its here, physical, voluminous and completely disorganized. A Tech Director at a former job of mine always wanted to take a filing cabinet, cram if full of paper in dissarray, with materials falling out of it and jammed into the drawers. "This is what the server actually looks like," would be the message. I am sure that the same analogy would apply to the internet.
Labels:
BBC,
digital information,
digital preservation,
internet,
report,
storage
Tuesday, November 27, 2007
Stewardship of Digital Assets @ Palinet
On November 14 and 15th I had the opportunity to attend the first in a series of workshops on digital preservation. Developed by the North East Document Conservation Center, "Stewardship of Digital Assets" took digital projects to the next level -- No longer we were just talking about just digitizing materials, but we are beginning to look at sustaining the collections.
It's taken me about 2 weeks to get around to writing about this workshop, and I don't think that I would have been able to tackle any sooner. What a dense presentation! The faculty's experience was diverse and well reaching -- from working with the National Archives (NARA), the Research Libraries Group (RLG), Lots of Copies Keeps Stuff Safe (LOCKSS) framework developed at Stanford, the powerhouse Online Computer Learning Center (OCLC), the Pennsylvania Library Network (Palinet), Institute of Museum and Library Science (IMLS), and National Information Standards Organization (NISO), and that's just to name a few of the distinguished organizations that these four people worked for, developing standards and managing digital content. Liz Bischoff, Tom Clareson, Robin Dale, and Dr. Katherine Skinner were kind enough to share with the forty attendees the results of the rich work that they have been doing in encouragement of collaboration.
Overall, the librarians and archivists in attendence are newbies to the field, as 1/3 of participants have yet to begin digitizing anything. That is a welcome number to hear -- as it means my institution is not falling behind of the herd. What I have found happened to many early digitization projects, anyway, is that objects were not created at a high enough resolution to warrant digital preservation, and some projects may have to rescan to render objects to meet new standards. The prevailing wind for digital projects these days is to scan BIG (and I really mean BIG -- as big as your institution can afford to maintain) ie. full size @ 300 - 600 dpi. Save that BIG scan as an unweildy tiff, and then make smaller derivative jpegs from that file for use copies.
But I digress -- This workshop was not aimed at the creation of these files -- Moreover it addressed what you do with the files now that you have them. Digitizing costs lots of money -- Improper care of files can lead to obselence or corrpuption of materials, rendering all your hard work wasted.
The workshop began with a lesson in assessment, simply meaning that if you can't articulate your institution's needs, then you can't apply standards. Not everything can be preserved all the time, and what becomes preserved should be a conscious choice, not a passive decision out of laziness. Whether it be analog or digital, a preservation program costs money to upkeep, and lots of it -- Why save junk? If the file is in a lesser format at alower quality level then the institution should make less of a commitment to preservation.
A fundamental shift is happening in the digitization world -- no longer are digital projects finite, but morever regarded as a cradle to cradle process. To ensure the integrity and authenticity of a digital document over time, digital objects needs to have a sense of curation present -- one that guarantees coceptually that the information comes out the same way that it went in to the system. This can be accomplished through preservation metadata that tracks the lifecycle of the object.
Reshaping digital projects to digital programs clarifies who is in charge of what tasks, and where and how the information is stored. Many organizations contain disparate silos of information across campus. Let IT run the servers and back everything up. Let the digital repository manage access. That way the organization knows clearly where all the information is, how it is being managed, and by who. Informed individuals make educated decisions, and can also identify potential risks. For example, in just the last few years reccommendations have moved away from storage on CD/DVD to spinning disk, hard drive and tape. Why? CDs aren't scalable, are hard to manage, and have been found to fail -- 15% of all information at Emory University failed when checked on CD -- That may be 15% of ALL digital information. If you can get boxes of CDs out of peoples individual offices and centralized in one preservation office, risk management becomes all the easier.
Grants at this point will not give monies for long term management of objects -- But the language of many grants includes a commitment to the long-term access and preservation of the objects. As institutions complete the digitization project and the grant money runs out, institutions may find the best way to manage preservation is to repurposing time and leveraging the skills of existing staff. The teams that are created through this process should be evenly weighted between curatorial and tech staff -- Curators choose content, techies digitize.
Preservation must be stabilized -- Choose a recognized standard and keep it there, don't unnecessarily migrate information, secure funding and document EVERYTHING.
Yikes, that's already a lot of notes -- I think this will be best broken up into a series of posts.
It's taken me about 2 weeks to get around to writing about this workshop, and I don't think that I would have been able to tackle any sooner. What a dense presentation! The faculty's experience was diverse and well reaching -- from working with the National Archives (NARA), the Research Libraries Group (RLG), Lots of Copies Keeps Stuff Safe (LOCKSS) framework developed at Stanford, the powerhouse Online Computer Learning Center (OCLC), the Pennsylvania Library Network (Palinet), Institute of Museum and Library Science (IMLS), and National Information Standards Organization (NISO), and that's just to name a few of the distinguished organizations that these four people worked for, developing standards and managing digital content. Liz Bischoff, Tom Clareson, Robin Dale, and Dr. Katherine Skinner were kind enough to share with the forty attendees the results of the rich work that they have been doing in encouragement of collaboration.
Overall, the librarians and archivists in attendence are newbies to the field, as 1/3 of participants have yet to begin digitizing anything. That is a welcome number to hear -- as it means my institution is not falling behind of the herd. What I have found happened to many early digitization projects, anyway, is that objects were not created at a high enough resolution to warrant digital preservation, and some projects may have to rescan to render objects to meet new standards. The prevailing wind for digital projects these days is to scan BIG (and I really mean BIG -- as big as your institution can afford to maintain) ie. full size @ 300 - 600 dpi. Save that BIG scan as an unweildy tiff, and then make smaller derivative jpegs from that file for use copies.
But I digress -- This workshop was not aimed at the creation of these files -- Moreover it addressed what you do with the files now that you have them. Digitizing costs lots of money -- Improper care of files can lead to obselence or corrpuption of materials, rendering all your hard work wasted.
The workshop began with a lesson in assessment, simply meaning that if you can't articulate your institution's needs, then you can't apply standards. Not everything can be preserved all the time, and what becomes preserved should be a conscious choice, not a passive decision out of laziness. Whether it be analog or digital, a preservation program costs money to upkeep, and lots of it -- Why save junk? If the file is in a lesser format at alower quality level then the institution should make less of a commitment to preservation.
A fundamental shift is happening in the digitization world -- no longer are digital projects finite, but morever regarded as a cradle to cradle process. To ensure the integrity and authenticity of a digital document over time, digital objects needs to have a sense of curation present -- one that guarantees coceptually that the information comes out the same way that it went in to the system. This can be accomplished through preservation metadata that tracks the lifecycle of the object.
Reshaping digital projects to digital programs clarifies who is in charge of what tasks, and where and how the information is stored. Many organizations contain disparate silos of information across campus. Let IT run the servers and back everything up. Let the digital repository manage access. That way the organization knows clearly where all the information is, how it is being managed, and by who. Informed individuals make educated decisions, and can also identify potential risks. For example, in just the last few years reccommendations have moved away from storage on CD/DVD to spinning disk, hard drive and tape. Why? CDs aren't scalable, are hard to manage, and have been found to fail -- 15% of all information at Emory University failed when checked on CD -- That may be 15% of ALL digital information. If you can get boxes of CDs out of peoples individual offices and centralized in one preservation office, risk management becomes all the easier.
Grants at this point will not give monies for long term management of objects -- But the language of many grants includes a commitment to the long-term access and preservation of the objects. As institutions complete the digitization project and the grant money runs out, institutions may find the best way to manage preservation is to repurposing time and leveraging the skills of existing staff. The teams that are created through this process should be evenly weighted between curatorial and tech staff -- Curators choose content, techies digitize.
Preservation must be stabilized -- Choose a recognized standard and keep it there, don't unnecessarily migrate information, secure funding and document EVERYTHING.
Yikes, that's already a lot of notes -- I think this will be best broken up into a series of posts.
Friday, November 16, 2007
Sarah Thomas at PMA
On Wednesday evening, I attended the Sarah Thomas library lecture at the Philadelphia Museum of Art, "First Sell the First Folio." The title of the lecture was in reference to how, many years ago, the Bodleian Library at Oxford deaccessioned a first folio of Shakespeare's plays, not seeing the potential value of the 1st edition. Thomas used this analogy to make a leap to why the physical existance of libraries is still relevant, and how dismissing libraries as obselete in the digital age can seem as erroneous as a decision to get rid of a first edition of a potential future rare title.
Thomas made the argument that we can't digitize everything, especially if its analog existence doesn't have strict bibliographic control already. Even if it was digitized, there would be no access information. Digitization also means we would need to make choices about what to keep and allow access to -- and if we destroy materials after they are digitized we may be losing the future's equivalent to the First Folio. Likewise, if libraries only exist digitally, then our choices about what to digitize are as weighted as what to purchase, preserve or destroy -- if patrons cannot access the materials it is the same as choosing them to not be in the collection.
She referred to the traditional round reading room as the "center of the universe where one would consult the oracle." As libraries have been replaced as information hub by the internet, the library must shift from being a box of books to a suite of services. To remain relevant, the library must connect the information it stores to what is being created outside the library.
A nice change from a text-based lecture, Thomas infused her lecture with photographs of library architecture and rare books. Her words were infused with a sweet reverence for what she refered to as the "power of the artifact." The lecture was ended by bringing Anne d'Harnoncourt, director of the Philadelphia Museum of Art, to tears by lauding the newly redesigned library in the newly opened Perelman Building of the PMA. A touching and soothing experience, in praise of libraries of the past, and in argument, gentle and hopeful, for the libraries of the future.
Thomas made the argument that we can't digitize everything, especially if its analog existence doesn't have strict bibliographic control already. Even if it was digitized, there would be no access information. Digitization also means we would need to make choices about what to keep and allow access to -- and if we destroy materials after they are digitized we may be losing the future's equivalent to the First Folio. Likewise, if libraries only exist digitally, then our choices about what to digitize are as weighted as what to purchase, preserve or destroy -- if patrons cannot access the materials it is the same as choosing them to not be in the collection.
She referred to the traditional round reading room as the "center of the universe where one would consult the oracle." As libraries have been replaced as information hub by the internet, the library must shift from being a box of books to a suite of services. To remain relevant, the library must connect the information it stores to what is being created outside the library.
A nice change from a text-based lecture, Thomas infused her lecture with photographs of library architecture and rare books. Her words were infused with a sweet reverence for what she refered to as the "power of the artifact." The lecture was ended by bringing Anne d'Harnoncourt, director of the Philadelphia Museum of Art, to tears by lauding the newly redesigned library in the newly opened Perelman Building of the PMA. A touching and soothing experience, in praise of libraries of the past, and in argument, gentle and hopeful, for the libraries of the future.
Labels:
architecture of libraries,
lecture,
Oxford,
PMA,
Sarah Thomas
Monday, October 29, 2007
Non-use of ContentDM and other sites
One software that we are looking at is ContentDM, which is a great software for accessioning and displaying digital assets. It is clear, easy to use, and can pretty much work right out of the box. I have yet to come across any librarian who really dislikes the software. The one comment that I have heard, however, is that no one really uses the site. So why go through the effort of digitization if no one is really using the software? Digitization is anything but cheap -- Between licenses and staff, a digitization project can easily cost tens of thousands of dollars.
Cornell, too, published a report evaluating the non-use of D-Space at their university. The whole report can be found at http://www.dlib.org/dlib/march07/davis/03davis.html. Part of the reason for non-use is that each discipline already has mechanisms for publication of materials in place and that their DSpace is additional, operating outside the sphere of traditional avenues.
Librarians are dreaming up these wonderful digital repositories, and companies are creating awesome software packages to host them, but is it all just Library Science laboratory work? How can digital initiatives be made more relevant and integrative into the academic sphere? Blackboard and other class content management and presentation software systems are thriving, because students and faculty have to go to them for class materials. It almost makes sense to grow a repository as attached to Blackboard, as people are already there.
I see now why ArtStor has taken off so well in the academic art realm -- Faculty can direct their students to the repository, add to the population of the image collection, sort and create their presentations and pretty much base classes out of there. I don't know if ContentDM can offer similar flexibility, or if the collections that we are speaking of digitizing have as much relevance to the coursework as would images of art for art history classes.
But it could.
And with careful planning, maybe it will?
Cornell, too, published a report evaluating the non-use of D-Space at their university. The whole report can be found at http://www.dlib.org/dlib/march07/davis/03davis.html. Part of the reason for non-use is that each discipline already has mechanisms for publication of materials in place and that their DSpace is additional, operating outside the sphere of traditional avenues.
Librarians are dreaming up these wonderful digital repositories, and companies are creating awesome software packages to host them, but is it all just Library Science laboratory work? How can digital initiatives be made more relevant and integrative into the academic sphere? Blackboard and other class content management and presentation software systems are thriving, because students and faculty have to go to them for class materials. It almost makes sense to grow a repository as attached to Blackboard, as people are already there.
I see now why ArtStor has taken off so well in the academic art realm -- Faculty can direct their students to the repository, add to the population of the image collection, sort and create their presentations and pretty much base classes out of there. I don't know if ContentDM can offer similar flexibility, or if the collections that we are speaking of digitizing have as much relevance to the coursework as would images of art for art history classes.
But it could.
And with careful planning, maybe it will?
Wednesday, October 10, 2007
Indiana University, My Heroes!
This summer I attended the Visual Resources Assoc.'s Summer Educational Institute for Visual Resources and Image Management (or something along those lines)...
That's how I found out about the Digital Library Program at IU. I just wanted to give them some props, as they do a great job of sharing information about how/why they create their collections to look and function as they do. For example, the Cushman Photo collection (which is in the links of this page) includes a history of the project, including proposals, tech specs and rationales. I think that's very awesome -- helping others help themselves through knowledge sharing. More info can be found at: http://www.dlib.indiana.edu/
That's how I found out about the Digital Library Program at IU. I just wanted to give them some props, as they do a great job of sharing information about how/why they create their collections to look and function as they do. For example, the Cushman Photo collection (which is in the links of this page) includes a history of the project, including proposals, tech specs and rationales. I think that's very awesome -- helping others help themselves through knowledge sharing. More info can be found at: http://www.dlib.indiana.edu/
A Decision, I believe...
So Fedora may have gotten the best of me. I don't think that I want to pursue an open source solution me, by my lonesome, as my one mind blowing project. I really wanted to, in a way, to make myself feel smart -- Now I'm feeling like it would be an overwhelming decision. So... We may go with ContentDM for images.
My hesitations with ContentDM or any proprietary software:
1) Lack of ability to customize and to integrate into the library's website
2) Cost of updates / Inability to stay up with updates, if cost prohibitive
3) Creating a "cookie cutter" respository
I will be learning more about ContentDM next month at Palinet. My hypothesis is:
ContentDM will prove to be much more intuitive and easier to work with out of the box than an open source system. ContentDM will be less customizable and harder to integrate into the library's website, but may be worth the sacrifice, due to the lack of labor that the program will require to set up. ContentDM is making strides in abililty to handle text-based assets, but will continue to be outshined by DSpace or other institutional academic respositories.
I want to keep my hand in open source work to some degree so maybe we will look at implementing DSpace or EPrints for an academic respository of research and writing. The open source community is hard at work deploying cutting edge technologies and programming to make collections more accessible - I think it would be a good experience to have to do some of the original programming myself, and not rely completely on tech support and help lines to get things done...
My hesitations with ContentDM or any proprietary software:
1) Lack of ability to customize and to integrate into the library's website
2) Cost of updates / Inability to stay up with updates, if cost prohibitive
3) Creating a "cookie cutter" respository
I will be learning more about ContentDM next month at Palinet. My hypothesis is:
ContentDM will prove to be much more intuitive and easier to work with out of the box than an open source system. ContentDM will be less customizable and harder to integrate into the library's website, but may be worth the sacrifice, due to the lack of labor that the program will require to set up. ContentDM is making strides in abililty to handle text-based assets, but will continue to be outshined by DSpace or other institutional academic respositories.
I want to keep my hand in open source work to some degree so maybe we will look at implementing DSpace or EPrints for an academic respository of research and writing. The open source community is hard at work deploying cutting edge technologies and programming to make collections more accessible - I think it would be a good experience to have to do some of the original programming myself, and not rely completely on tech support and help lines to get things done...
Subscribe to:
Posts (Atom)