Questa è una vecchia versione del documento!
VeRBATIM
Description
VeRBATIM (Versatile Resource to Bridge Ancient Texts Interpretation and Metadata) is a tool which allows the tokenization of a huge amount of data from the EDCS Database (Clauss-Slaby).
From a source input text which may contain up to thousands of inscriptions in the EDCS format, the program is capable of isolating each individual word and associating it with the metadata of the text whence it is sourced in a .csv output file.
The application is particularly useful for quantitative and qualitative analyses of epigraphic texts.
User's Guide
Copy a given dataset of texts (even up to hundreds of thounsands) from EDCS and paste in a .txt file. Please use the Italian version of EDCS, as the β version of VeRBATIM is trained to recognize the metadata header lines in the Italian language.

Copy the EDCS results starting from the row beginning with “pubblicazione”, below the line “risultati trovati: xxxx” (i.e., not starting from the very beginning of the screen).
Run the program (β version) and select the data source, i.e., EDCS.
Upload or Drag and Drop the .txt file.
Press the button “Analyze” to process your file. Please wait for notification of completion.
You can export the .csv output file and select different options for displaying data (they can also be combined all together):
Omit rows: If selected, rows containing non-linguistically relevant data (e.g., empty rows, rows containing only /, [, or ]) will be excluded. At the end of the operation, the total number of the rows omitted will be displayed.
Original format: If selected, the token will be displayed according to the format of the data source, i.e., EDCS.
Clean token (1): If selected, the token will be cleaned according to diplomatic criteria, i.e., by eliminating the characters which were added in the source edition (i.e., characters between round brackets) or by adding the characters which were deleted (i.e., characters between curly brackets).
Clean token (2): If selected, Clean token (1) will be cleaned by eliminating characters which were supplemented by the editor (i.e., characters between square brackets).
Normalized token: If selected, the token will be displayed by normalizing divergent spellings and supplementing abbreviations or textual lacunas.
Remarks: If selected, linguistic and philological remarks based on the regular expressions of the data source will be displayed.
Omit supplemented rows: If selected, rows containing words completely supplemented by the editor will be deleted.
Token-ID: If selected, each token will be associated with a univocal ID (resulting from the combination of the EDCS-ID and the number associated with the position of the token within the texts).
Text and Context: If selected, two columns will appear: the first with the full text wherein the token is attested, the second with the two words preceding and following the selected token.
Other metadata: If selected, columns with extra-linguistic variables will appear, i.e., publication, publication ID, dating, geographic provenance, text type, writing material, additional remarks.
To import the .csv file into Excel, please select the UTF-8 encoding of characters and the comma as delimiter mark.
Future developments
In the final version, the dating column will feature a normalized and coherent format that will allow for a better cross-referencing of the data. In addition, geographic coordinates will be associated with each column containing the geographic provenance, thus allowing a fine-grained mapping of the entries. There will also be an additional export option that will allow the researcher to export only those tokens displaying a given string of characters or just a given letter.
The combination with automatic lemmatization softwares is being worked on.
The developer is implementing the program to read the data format of EDR (Epigraphic Database Roma), papyri.info and PHI Greek Inscriptions. It will also be possible to implement the program for other databases that may be requested.
Test the β Version
Please download the application to test the β version and send a feedback to serena.barchiATunitus.it to report bugs or glitches and help improving the program.
The application is available for the moment only for macOS (at least 11.0 version).

It is possible that the system will resist running the application due to security restrictions. In that case, just right-click on the icon and select “open”; alternatively, you can change the security options of the system itself.
verbatim_macos.app.zip