Midterm Data Analysis

I choose to use the New Testament King James Version xml file for my data for the midterm project. Using this data, I decided that I wanted to analyze what are unanimously considered to be the seven Authentic Letters of the Apostle Paul that are compiled in the New Testament. I wanted to see simply which words are used most in order to get a good understanding of what topics are most prevelant in Paul’s confirmed writings. I wanted to get a selection out of the whole New Testament because I wanted to have a more isolated voice out of the whole corpus in order to see if there is anything especially noticeable about these works in particular. I thought it would be interesting to look at the work of just one author, at least one confirmed author. These works in the order that they are represented on the voyant site are 2 Corinthians, 1 Thessalonians, Philemon, Romans, Galatians, 1 Corinthians, and Philippians. I don’t know why they appeared this way, because they are in chronological order in the html file that I uploaded, as far as I know. There has definitely been some user error on my part.

Because I wanted to look only at a selection of the books of the New Testament, I had to do some data cleaning. I knew because my data consisted of text which was conviently divied up into differents, I decided to my data cleaning manually. If remember correctly, the first thing that I did was transfer the xml file into an html file, because I feel much more comfortable with html, as I have done some coding with it. I put the xml code for the new testament file into notepad and saved it as an html file. At this point, I wasn’t quite sure what exactly that I wanted to do with the new testament text, so I went ahead and made a csv out of the file by putting it into Google sheets. I then explored different options that were on the midterm schedule, such as Flourish, RAWGraphs.io, and Gephi. Nothing really made sense to me to visualize my data through this different applications, because they were much arreigned to deal with numerical data, as opposed to textual data. I looked at Open Refine as well, but that didn’t lead anywhere interesting for me. It was around this point that I thought of looking at the seven authentic letters of Paul because I thought of one of friends has a particular interest in them. I then took the html file that I made and selected the seven books out of the file, and combined them into another notepad file, and saved them as an html file. I think that I might have messed up with my copying, because one of the book tags was missing a bracket. I was thinking about separating the books into the individual contents of the chapters to see the most frequent words in each chapter, but I couldn’t figure out a way that wouldn’t be manually, and I didn’t want to put that much work into something that could be difficult to find anything out from. I eventually thought of using Voyant, but it wasn’t working the way I wanted it to, with each book being deliniated into separate categories. I asked Austin for help with making the books separate documents, and fixed the html tag error. The process of making lots of different csv, tsv, and html files eventually resulted in the voyant link that is here: https://voyant-tools.org/?corpus=96f54db1728e4ffe26891bcfaf2753b8

Here is my html file that I entered in Voyant: file:///C:/Users/griff/Downloads/AuthenticPaul.html

WordPress hasn’t allowed me to me to the Voyant website that I made, which honestly doesn’t surprise me.

From this data, the most interesting thing to me are the most common words. Non-prepostional words not included, standout words include, in order, God, Christ, Lord, Man, Jesus, Law, Spirit, Faith, Brethern, Glory, Flesh, Gospel, Grace, and Righteousness. It may be pretty obvious, considering these texts are included in the Bible, but God, Jesus, and names associated with them are the most prevelant words. Following those are Law, Spirit, and Faith. One interesting thing to look at are the distinctive words, because they show that around three quarters of the mentions of the word Law occur in two books, Romans and Galatians. This shows that the Jewish Law, and how it applies to followers of Jesus is mostly covered in those two books. Other words, like Spirit, Faith, Grace, Righteousness are spread throughout the corpus, indicating that these concepts are at the core to Paul’s teachings. The fact that God is the most prevelant word by a wide margin shows how much these works care about God, even if it tracks with the common assumption. This project relates to Digital Humanities in the sense that it provides a simplistic summary for the topics of Paul’s Authentic letters.

css.php