Large-Corpus AI Text Detection Using Token Probability Distributions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for detecting AI-generated text in large document corpora are unstable and biased, particularly affecting non-native English speakers, and lack accurate estimation of AI-generated text presence and extent.
Innovation Solution
A method utilizing maximum likelihood estimation of text distributions between AI-generated and human-written document corpora, employing token statistics and neural networks to generate detection flags for AI-generated text, providing more accurate and stable detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing detection methods are used, then detection can be performed, but detection accuracy and stability deteriorate, particularly for non-native English speakers
Solution Approach 1:
The patent changes the detection parameters from traditional n-gram analysis and perplexity scores to token-level probability distributions and likelihood ratios. By transforming the detection metric into a probability distribution over tokens and using likelihood ratios to compare AI-generated versus human-written text, the system achieves more accurate and stable detection results, particularly for non-native English speakers.
Solution Approach 2:
The patent replaces traditional mechanical detection mechanisms (n-gram matching, perplexity calculation) with a probabilistic model based on token-level probability distributions. This substitution allows for more nuanced detection that accounts for linguistic nuances and contextual variations, thereby improving both accuracy and stability across diverse writing styles and languages.
2Productivity
If large language models generate documents, then document production speed increases, but distinguishing machine-generated from human-written text becomes more difficult
Solution Approach 1:
The patent introduces an intermediary detection system that mediates between the AI-generated text and the detection process. By using token-level probability distributions and likelihood ratios as intermediate representations, the system can accurately distinguish AI-generated from human-written text even when the generated text closely mimics human writing styles, thus maintaining high productivity while enabling reliable detection.
Solution Approach 2:
The patent adds a new dimension to text analysis by moving from sentence-level or paragraph-level analysis to token-level probability distribution analysis. This dimensional change enables the system to detect subtle statistical patterns in token generation that are invisible to traditional detection methods, thereby maintaining high document production speeds while improving detection capability.
3Productivity
If traditional detection methods are applied to large document corpora, then processing can be performed, but detection results become biased and unstable
Solution Approach 1:
The patent segments the detection task into token-level analyses rather than processing entire documents as a single unit. By analyzing token probability distributions individually and aggregating results through likelihood ratios, the system maintains the ability to process large corpora efficiently while ensuring consistent and unbiased detection results, as each token's contribution is independently evaluated.
Data Source
AI summary
Systems and methods for detecting artificial intelligence generated text in large document corpora. An AI-generated document corpus and a human-written document corpus can be generated using identified prompts. Text distributions of human-written text and AI-generated text from the corpus of AI generated documents and the corpus of human-written documents can be estimated using token statistics. A detection distribution of AI generated documents from a target corpus can be estimated with maximum likelihood estimation using the text distributions. Detection flags for the AI generated documents from the target corpus can be generated based on the detection distribution.


