VeRBATIM
Description
VeRBATIM (Versatile Resource to Bridge Ancient Texts Interpretation and Metadata) is a tool which allows the tokenization of a huge amount of data from the EDCS Database (Clauss-Slaby).
From a source input text which may contain up to thousands of inscriptions in the EDCS format, the program is capable of isolating each individual word and associating it with the metadata of the text whence it is sourced in a .csv output file.
The application is particularly useful for quantitative and qualitative analyses of epigraphic texts.
User's Guide
Copy a given dataset of texts (even up to hundreds of thousands) from EDCS and paste in a .txt file. Please use the Italian version of EDCS, as the β version of VeRBATIM is trained to recognize the metadata header lines in the Italian language.

Copy the EDCS results starting from the row beginning with “pubblicazione”, below the line “risultati trovati: xxxx” (i.e., not starting from the very beginning of the screen).
Run the program (β version) and select the data source, i.e., EDCS.
Upload or Drag and Drop the .txt file.
Press the button “Analyze” to process your file. Please wait for notification of completion.
You can export the .csv output file and select different options for displaying data (they can also be combined all together):
Omit rows: If selected, rows containing non-linguistically relevant data (e.g., empty rows, rows containing only /, [, or ]) will be excluded. At the end of the operation, the total number of the rows omitted will be displayed.
Original format: If selected, the token will be displayed according to the format of the data source, i.e., EDCS.
Clean token (1): If selected, the token will be cleaned according to diplomatic criteria, i.e., by eliminating the characters which were added in the source edition (i.e., characters between round brackets) or by adding the characters which were deleted (i.e., characters between curly brackets).
Clean token (2): If selected, Clean token (1) will be cleaned by eliminating characters which were supplemented by the editor (i.e., characters between square brackets).
Normalized token: If selected, the token will be displayed by normalizing divergent spellings and supplementing abbreviations or textual lacunas.
Remarks: If selected, linguistic and philological remarks based on the regular expressions of the data source will be displayed.
Omit supplemented rows: If selected, rows containing words completely supplemented by the editor will be deleted.
Token-ID: If selected, each token will be associated with a univocal ID (resulting from the combination of the EDCS-ID and the number associated with the position of the token within the texts).
Text and Context: If selected, two columns will appear: the first with the full text wherein the token is attested, the second with the two words preceding and following the selected token.
Other metadata: If selected, columns with extra-linguistic variables will appear, i.e., publication, publication ID, dating, geographic provenance, text type, writing material, additional remarks.
To import the .csv file into Excel, please select the UTF-8 encoding of characters and the comma as delimiter mark.
Test the β Version
Please download the application to test the β version and send a feedback to serena.barchiATunitus.it to report problems and help improving the program.
The application is available for the moment only for macOS (at least 11.0 version).

It is possible that the system will resist running the application due to security restrictions. In that case, just right-click on the icon and select “open”; alternatively (not recommended), you can change the security options of the system itself.
download VeRBATIM (send request access)
sample .txt file 1 Aemilia / Regio VIII (about 5,000 inscriptions)
sample .txt file 2 Africa proconsularis (about 33,000 inscriptions)