Candidate Word Filtering for AI-Generated Text Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
AI models generate persuasive but false or outdated text data, leading to decreased accuracy and performance in machine learning models due to training on such data, which introduces bias and inaccuracies.
Innovation Solution
A method and system for generating candidate words based on character codes and predefined criteria to detect and identify text data generated by AI models, including a first set of candidate words formed from character code combinations and a second set applied with likelihood criteria to identify AI-generated text.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If AI models generate text data in large volumes to meet demand, then productivity and accessibility of AI-generated content improve, but the quantity of false and outdated text data increases, leading to decreased accuracy and performance in machine learning models
Solution Approach 1:
The system performs preliminary detection and classification of AI-generated text data before it enters the machine learning training pipeline. By identifying false and outdated text data in advance and filtering it out, the system prevents contaminated data from affecting model accuracy, thus resolving the contradiction between high-volume data generation and maintained model reliability
Solution Approach 2:
The patent introduces an intermediary detection system that acts as a filter between AI text generation and machine learning training. This intermediary layer classifies text data as AI-generated, false, or outdated, and selectively filters out problematic data while allowing high-quality data to pass through, thereby maintaining both high productivity and reliability
2Productivity
If machine learning models train on large volumes of AI-generated text data to improve learning capacity, then productivity of data processing improves, but false and outdated data introduces bias and decreases model performance
Solution Approach 1:
The system performs preliminary detection and classification of AI-generated text data before it enters the machine learning training pipeline. By identifying false and outdated text data in advance and filtering it out, the system prevents contaminated data from affecting model accuracy, thus resolving the contradiction between high-volume data generation and maintained model reliability
Solution Approach 2:
The patent introduces an intermediary detection system that acts as a filter between AI text generation and machine learning training. This intermediary layer classifies text data as AI-generated, false, or outdated, and selectively filters out problematic data while allowing high-quality data to pass through, thereby maintaining both high productivity and reliability
3Ease of operation
If AI models generate persuasive text data to satisfy user demand, then ease of operation and user interaction improve, but the text data becomes false or outdated, introducing harmful factors into the system
Solution Approach 1:
The system detects harmful factors (false and outdated text data) and converts this information into beneficial classification labels. By identifying AI-generated text data and marking it as such, the system enables users to understand the nature of the content, transforming the harmful false data into useful metadata for informed decision-making and selective data usage
Data Source
AI summary
Generation of candidate words for detection of text data generated by an artificial intelligence (AI) model includes obtaining a plurality of character codes associated with a plurality of characters. The plurality of characters is associated with a plurality of words. Based on the plurality of character codes, a first set of candidate words is generated. Each candidate word of the first set of candidate words comprises a combination of at least two character codes of the plurality of character codes. Further, based on an application of a set of predefined criteria on the first set of candidate words, a second set of candidate words is generated. The set of predefined criteria is associated with a likelihood of generation of each of the first set of candidate words by the AI model. The second set of candidate words is output for detecting the text data generated by the AI model.


