Data Processing for Argot Risk Identification via Corpus Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current risk prevention and control systems are inefficient and inaccurate in identifying argots with hidden meanings, leading to low data processing efficiency and accuracy, which allows malicious third parties to bypass security measures.
Innovation Solution
A data processing method that involves obtaining a target object, identifying matching argots, and using a pre-constructed corpus database to determine risk based on similarity with risk corpora, including risk words associated with the argot, to improve identification accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual determination method is used to identify argots based on context, then identification accuracy can be improved, but data processing efficiency deteriorates due to large data amount
Solution Approach 1:
The system pre-constructs a corpus database containing risk corpora and risk words before actual risk identification. This preliminary preparation allows the system to quickly retrieve and compare against pre-analyzed risk patterns during runtime, achieving both high accuracy and efficiency without manual intervention during processing
Solution Approach 2:
The system creates a corpus database that copies and stores pre-analyzed risk patterns, risk words, and contextual information. Instead of manually analyzing each new input, the system compares against these pre-copied risk patterns, maintaining accuracy while dramatically improving processing speed
2Productivity
If simple word matching is used to identify argots, then data processing efficiency is improved, but identification accuracy deteriorates because argots have high similarity to risk-free words
Solution Approach 1:
The system enhances basic word matching by incorporating local contextual quality checks. It examines surrounding words, sentence structure, and semantic context to determine whether a matched argot appears in a risky or safe context, thereby improving accuracy without sacrificing the efficiency of automated processing
Solution Approach 2:
The corpus database acts as an intermediary between simple word matching and complex manual analysis. It provides pre-analyzed risk contexts and associated risk words that mediate the matching process, enabling automated systems to achieve accuracy previously requiring human judgment
3Device complexity
If no pre-constructed corpus database is used, then device complexity is reduced, but risk prevention and control accuracy deteriorates
Solution Approach 1:
The corpus database is constructed in advance with pre-analyzed risk patterns, risk words, and contextual information. This preliminary action creates a ready-to-use reference system that improves risk identification accuracy without adding complexity to the runtime processing system
Solution Approach 2:
The corpus database is constructed in advance with pre-analyzed risk patterns, risk words, and contextual information. This preliminary action creates a ready-to-use reference system that improves risk identification accuracy without adding complexity to the runtime processing system
Data Source
AI summary
Embodiments of this specification provide a data processing method, apparatus, and device. The method includes: obtaining a to-be-identified target object; if the target object includes a word matching a first argot, obtaining a target corpus corresponding to the target object from a corpus included in a pre-constructed corpus database, where the pre-constructed corpus database includes a first corpus, the first corpus is a risk corpus constructed based on a second argot and a target risk corpus, and the target risk corpus includes a risk word that has a preset association relationship with the second argot; and determining, based on a similarity between the target object and the target corpus and a risk label of the target corpus, whether the target object has a risk.


