Data Processing for Argot Risk Identification via Corpus Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current risk prevention and control systems are inefficient and inaccurate in identifying argots with hidden meanings, leading to low data processing efficiency and accuracy, which allows malicious third parties to bypass security measures.

Innovation Solution

A data processing method that involves obtaining a target object, identifying matching argots, and using a pre-constructed corpus database to determine risk based on similarity with risk corpora, including risk words associated with the argot, to improve identification accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual determination method is used to identify argots based on context, then identification accuracy can be improved, but data processing efficiency deteriorates due to large data amount

Engineering Contradiction:
Improveidentification accuracyVSAvoiddata processing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system pre-constructs a corpus database containing risk corpora and risk words before actual risk identification. This preliminary preparation allows the system to quickly retrieve and compare against pre-analyzed risk patterns during runtime, achieving both high accuracy and efficiency without manual intervention during processing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a corpus database that copies and stores pre-analyzed risk patterns, risk words, and contextual information. Instead of manually analyzing each new input, the system compares against these pre-copied risk patterns, maintaining accuracy while dramatically improving processing speed

Inventive Principle:
Principle #26Copying

2Productivity

If simple word matching is used to identify argots, then data processing efficiency is improved, but identification accuracy deteriorates because argots have high similarity to risk-free words

Engineering Contradiction:
Improvedata processing efficiencyVSAvoididentification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system enhances basic word matching by incorporating local contextual quality checks. It examines surrounding words, sentence structure, and semantic context to determine whether a matched argot appears in a risky or safe context, thereby improving accuracy without sacrificing the efficiency of automated processing

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The corpus database acts as an intermediary between simple word matching and complex manual analysis. It provides pre-analyzed risk contexts and associated risk words that mediate the matching process, enabling automated systems to achieve accuracy previously requiring human judgment

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If no pre-constructed corpus database is used, then device complexity is reduced, but risk prevention and control accuracy deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidrisk identification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The corpus database is constructed in advance with pre-analyzed risk patterns, risk words, and contextual information. This preliminary action creates a ready-to-use reference system that improves risk identification accuracy without adding complexity to the runtime processing system

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The corpus database is constructed in advance with pre-analyzed risk patterns, risk words, and contextual information. This preliminary action creates a ready-to-use reference system that improves risk identification accuracy without adding complexity to the runtime processing system

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250328635A1Data processing method, apparatus and device
Publication Date: 2025.10.23 ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
  • US20250328635A1 patent drawing
  • US20250328635A1 patent drawing
  • US20250328635A1 patent drawing

AI summary

Embodiments of this specification provide a data processing method, apparatus, and device. The method includes: obtaining a to-be-identified target object; if the target object includes a word matching a first argot, obtaining a target corpus corresponding to the target object from a corpus included in a pre-constructed corpus database, where the pre-constructed corpus database includes a first corpus, the first corpus is a risk corpus constructed based on a second argot and a target risk corpus, and the target risk corpus includes a risk word that has a preset association relationship with the second argot; and determining, based on a similarity between the target object and the target corpus and a risk label of the target corpus, whether the target object has a risk.