A comparative study of likelihood ratio based forensic text comparison procedures: Multivariate kernel density with lexical features vs. word N-grams vs. character N-grams

Shunichi Ishihara*

*Corresponding author for this work

    Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

    2 Citations (Scopus)

    Abstract

    This is a comparative study to empirically investigate the performances of three different procedures for calculating authorship attribution likelihood ratios (LR). The procedures to be compared are: 1) a procedure based on multivariate kernel density (MVKD) with lexical features; 2) a procedure based on word N-grams; and 3) a procedure based on character N-grams. Furthermore, the best-performing LRs of these three procedures are fused into combined single LRs using a logistic-regression fusion, in order to investigate the extent of the improvement/deterioration that the fusion brings about. This study uses chatlog messages, which were presented as evidence to prosecute paedophiles, for testing. The numbers of word tokens used to model the authorship attribution of each message group are 500 and 1000 words. This was done to examine the effect of sample size on the performance of a system. The performance of a system is assessed with regard to its validity (= accuracy) and reliability (= precision) using the log-likelihood-ratio cost (Cllr) and 95% credible intervals (CI), respectively. While describing the different characteristics of these three procedures in their outcomes, this study demonstrates that the MVKD procedure was the best-performing procedure out of the three in terms of Cllr. This study also demonstrates that a logistic-regression fusion is useful for combining the LRs obtained from the three procedures in question, resulting in a good improvement in performance.

    Original languageEnglish
    Title of host publicationProceedings - 5th Cybercrime and Trustworthy Computing Conference, CTC 2014
    PublisherInstitute of Electrical and Electronics Engineers Inc.
    Pages1-11
    Number of pages11
    ISBN (Electronic)9781479988259
    DOIs
    Publication statusPublished - 15 Apr 2015
    Event5th Cybercrime and Trustworthy Computing Conference, CTC 2014 - Aukland, New Zealand
    Duration: 24 Nov 201425 Nov 2014

    Publication series

    NameProceedings - 5th Cybercrime and Trustworthy Computing Conference, CTC 2014

    Conference

    Conference5th Cybercrime and Trustworthy Computing Conference, CTC 2014
    Country/TerritoryNew Zealand
    CityAukland
    Period24/11/1425/11/14

    Fingerprint

    Dive into the research topics of 'A comparative study of likelihood ratio based forensic text comparison procedures: Multivariate kernel density with lexical features vs. word N-grams vs. character N-grams'. Together they form a unique fingerprint.

    Cite this