<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Bplale on Journal of Digital Humanities</title><link>https://journalofdigitalhumanities.org/author/bplale/</link><description>Recent content in Bplale on Journal of Digital Humanities</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sat, 01 Jun 2013 00:00:00 +0000</lastBuildDate><atom:link href="https://journalofdigitalhumanities.org/author/bplale/index.xml" rel="self" type="application/rss+xml"/><item><title>Architecture to Enable Large-Scale Computational Analysis of Millions of Volumes</title><link>https://journalofdigitalhumanities.org/2-3/architecture-to-enable-large-scale-computational-analysis-of-millions-of-volumes/</link><pubDate>Sat, 01 Jun 2013 00:00:00 +0000</pubDate><guid>https://journalofdigitalhumanities.org/2-3/architecture-to-enable-large-scale-computational-analysis-of-millions-of-volumes/</guid><description>&lt;h3 id="poster"&gt;Poster&lt;/h3&gt;
&lt;iframe class="gde-frame" scrolling="no" src="https://docs.google.com/viewer?url=http%3A%2F%2Fjournalofdigitalhumanities.org%2Fwp-content%2Fuploads%2F2013%2F11%2FLarge-Scale-Computational-Analysis.pdf&amp;amp;hl=en_US&amp;amp;embedded=true" style="width:100%; height:500px; border: none;"&gt;&lt;/iframe&gt;
&lt;p&gt;&lt;a href="https://journalofdigitalhumanities.org/wp-content/uploads/2013/11/Large-Scale-Computational-Analysis.pdf"&gt;Download (PDF, 1.5MB)&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="abstract"&gt;Abstract&lt;/h3&gt;
&lt;p&gt;The HathiTrust Research Center (HTRC) is a collaborative research center that provides Digital Humanities researchers access to not only millions of volumes from the HathiTrust (HT) digital library, but also cutting-edge software tools and cyber infrastructure to perform advanced computational analysis over the corpus at an unprecedented scale.&lt;/p&gt;
&lt;p&gt;The corpus at the HTRC currently consists of over 3 million public domain volumes and anticipates access to an additional 6 million in-copyright volumes. In their raw form at the HathiTrust, these volumes are stored as files on special hardware using an internal Pairtree structure. The internal HathiTrust structure is optimal for its primary function of the digital page image delivery to digital library patrons for viewing; however, it does not support well the large-scale computational analysis which is the primary function of the HTRC. Navigating the Pairtree and uncompressing the text data would encounter major performance and scalability issues. While researchers from other scientific communities have been addressing aspects of the “Big Data” problem with success, the large corpus that HTRC hosts to support computational analysis presents a unique setting in that it consists of a massive number of small text-based files, whereas most solutions from the scientific communities are tailored towards large files and non-text-based content. In this poster, we will present the approach the HTRC takes to solve this problem — the HTRC keeps the Pairtree only for the purpose of synchronization with the HT, and processes and pushes the volume data from the local Pairtree to a NoSQL storage cluster using Apache Cassandra hosted on conventional hardware during the ingest process. In order to balance the data store and ingest workload, the developers at the HTRC and the HT also devised a very simple yet effective way to parallelize the rsync of the single source Pairtree at the HT on all Cassandra nodes by starting rsync at lower branches instead of at the root.&lt;/p&gt;</description></item></channel></rss>