<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Dblei on Journal of Digital Humanities</title><link>https://journalofdigitalhumanities.org/author/dblei/</link><description>Recent content in Dblei on Journal of Digital Humanities</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sat, 01 Dec 2012 00:00:00 +0000</lastBuildDate><atom:link href="https://journalofdigitalhumanities.org/author/dblei/index.xml" rel="self" type="application/rss+xml"/><item><title>Topic Modeling and Digital Humanities</title><link>https://journalofdigitalhumanities.org/2-1/topic-modeling-and-digital-humanities-by-david-m-blei/</link><pubDate>Sat, 01 Dec 2012 00:00:00 +0000</pubDate><guid>https://journalofdigitalhumanities.org/2-1/topic-modeling-and-digital-humanities-by-david-m-blei/</guid><description>&lt;h3 id="introduction"&gt;Introduction&lt;/h3&gt;
&lt;p&gt;Topic modeling provides a suite of algorithms to discover hidden thematic structure in large collections of texts. The results of topic modeling algorithms can be used to summarize, visualize, explore, and theorize about a corpus.&lt;/p&gt;
&lt;p&gt;A topic model takes a collection of texts as input. It discovers a set of “topics” — recurring themes that are discussed in the collection — and the degree to which each document exhibits those topics. Figure 1 illustrates topics found by running a topic model on 1.8 million articles from the &lt;em&gt;New York Times&lt;/em&gt;. The model gives us a framework in which to explore and analyze the texts, but we did not need to decide on the topics in advance or painstakingly code each document according to them. The model algorithmically finds a way of representing documents that is useful for navigating and understanding the collection.&lt;/p&gt;</description></item></channel></rss>