<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Cblack on Journal of Digital Humanities</title><link>https://journalofdigitalhumanities.org/author/cblack/</link><description>Recent content in Cblack on Journal of Digital Humanities</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Thu, 01 Dec 2011 00:00:00 +0000</lastBuildDate><atom:link href="https://journalofdigitalhumanities.org/author/cblack/index.xml" rel="self" type="application/rss+xml"/><item><title>Clustering with Compression for the Historian</title><link>https://journalofdigitalhumanities.org/1-1/clustering-with-compression-for-the-historian-by-chad-black/</link><pubDate>Thu, 01 Dec 2011 00:00:00 +0000</pubDate><guid>https://journalofdigitalhumanities.org/1-1/clustering-with-compression-for-the-historian-by-chad-black/</guid><description>&lt;h3 id="introduction"&gt;INTRODUCTION&lt;/h3&gt;
&lt;p&gt;I mentioned &lt;a href="https://parezcoydigo.wordpress.com/2011/09/23/an-algorithmic-approach-to-legal-culture-in-the-early-modern-spanish-empire/" title="Chad Black, 'An algorithmic approach to legal culture in the early modern Spanish empire'"&gt;in my blog&lt;/a&gt; that I’m playing around with a variety of clustering techniques to identify patterns in legal records from the early modern Spanish Empire. In this post, I will discuss the first of my training experiments using Normalized Compression Distance (NCD). I’ll look at what NCD is, some potential problems with the method, and then the results from using NCD to analyze the Criminales Series descriptions of the Archivo Nacional del Ecuador’s (&lt;a href="http://www.ane.gob.ec/" title="The National Archive of Ecuador Homepage"&gt;ANE&lt;/a&gt;) Series Guide. For what it’s worth, this is a very easy and approachable method for measuring similarity between documents and requires almost no programming chops. So, it’s perfect for me!&lt;/p&gt;</description></item></channel></rss>