Hatred speech and privacy information identification system and method based on large model

Through hatred detection based on Transformer hybrid neural network and privacy recognition of lightweight CNN model, combined with parameter sharing of BERT embedding layer, the problems of high misjudgment rate of hate speech detection and resource consumption of privacy information processing in the existing technology are solved, and efficient hatred speech recognition and accurate privacy protection are achieved.

CN120337285APending Publication Date: 2025-07-18SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510415680.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing technology has problems such as high misjudgment rate, low coverage and insufficient fine-grainedness in hate speech detection. There are defects in the processing of private information, rigid desensitization strategies and system integration, resulting in double resource consumption and difficulty in audit traceability.

Method used

The hatred detection engine based on Transformer hybrid neural network and the privacy identification engine of the lightweight CNN model are adopted, combined with BERT embedding layer parameter sharing, and the training set is constructed through data preprocessing, and a hybrid model architecture and dynamic desensitization strategy are adopted to achieve fine-grained detection and hierarchical privacy protection.

Benefits of technology

It improves the effectiveness of hate speech recognition, achieves accurate identification and hierarchical desensitization, reduces resource consumption, simplifies the audit traceability process, and provides comprehensive privacy information protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337285A_ABST
    Figure CN120337285A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing and information safety crossing, in particular to a hatred speech and privacy information recognition system and method based on a large model, and the system comprises a data collection layer, a model processing layer, a result output layer and a feedback optimization layer. The system has the beneficial effects that social media text streams are acquired through the data acquisition layer, a hatred detection engine (based on a Transform hybrid neural network) and a privacy recognition engine (a rule base and a lightweight CNN model) in the model processing layer perform processing, the result output layer generates structured output and executes dynamic desensitization, and the feedback optimization layer realizes manual auditing and incremental learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross - technical field of natural language processing and information security, and specifically provides a system and method for identifying hate speech and privacy information based on a large model. Background Art

[0002] The current network content censorship technology mainly has the following technical bottlenecks:

[0003] Defects in hate speech detection technology. Existing methods mainly adopt two types of technical routes: 1. Keyword matching method: Rule - based filtering based on a sensitive word library has two major limitations: high false positive rate: unable to distinguish ironic contexts (such as irony in "This policy is really 'fair'"); low coverage: only able to identify explicit hate vocabulary (such as racial discrimination terms), and the missed detection rate of implicit hate expressions (such as "The quality of people in a certain area is generally low") reaches 47.2%. 2. Traditional classification model method: End - to - end classification using CNN / LSTM has three key problems: insufficient fine - grainedness: only outputting binary classification results, lacking actionable censorship basis; strong data dependence: requiring millions of labeled data, while the proportion of hate speech in the actual scenario is less than 0.3%, resulting in model overfitting.

[0004] Defects in privacy information processing technology. Existing privacy protection solutions have three major deficiencies: 1. Single detection dimension: The detection tool based on regular expressions has a recognition failure rate of more than 60% for deformed text (such as "138O0O138OO"). A simple NER model cannot identify cross - sentence associated information (such as combined privacy in "Mr. Zhang lives at XX Road and the phone number is..."). 2. Rigid desensitization strategy: Full - scale replacement leads to loss of information value (such as processing "Patient 13800138000 needs a follow - up visit" as "Patient *** needs a follow - up visit"). Lack of a grading mechanism, unable to distinguish between public and authorized access scenarios. 2.3 System integration defects. Current mainstream solutions regard hate speech detection and privacy processing as independent modules, resulting in: lack of information association: unable to identify compound risks such as "the mobile phone number of a user who attacks the LGBT community"; doubling of resource consumption: repeated text parsing increases CPU utilization by 40%; difficult audit and traceability: scattered logs increase the forensic time cost by 80%. Summary of the Invention

[0005] The purpose of the present invention is to provide a system and method for identifying hate speech and privacy information based on a large model to solve the problems raised in the above - mentioned background art.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A system for identifying hate speech and privacy information based on a large model, comprising:

[0007] Data collection layer: Real-time acquisition of multi-source social media text streams through API interfaces, supporting JSON / XML format conversion;

[0008] Model processing layer: Hate detection engine: A hybrid neural network based on Transformer, combined with the parameter sharing mechanism of the BERT embedding layer; Privacy recognition engine: Adopts a detection method that combines a rule base and a lightweight CNN model;

[0009] Result output layer: Structured output module: Generates a quadruple JSON format containing the comment object, argument, target group, and judgment result; Dynamic desensitization module: Executes a hierarchical privacy replacement strategy according to the sensitivity level;

[0010] Feedback optimization layer: Manual review interface: Supports annotators to correct system misjudgments; Incremental learning module: Automatically updates model parameters regularly.

[0011] Preferably, the model processing layer includes:

[0012] Data preprocessing module: Constructs a three-level labeled training set: pre-labeling, manual verification, adversarial sample generation; Extracts character-level n-gram features Conv1D and semantic embeddings Chinese-RoBERTa-wwm-ext;

[0013] Hybrid model architecture: Adopts a BERT-BiLSTM dual-path structure to fuse features, enhances the representation through a gated attention mechanism; Based on a pointer network + classifier joint decoding, outputs quadruple detection results;

[0014] Post-processing rules: Executes contradiction detection; Utilizes coreference resolution technology to complete the omitted subject.

[0015] Preferably, the model processing layer includes:

[0016] Multi-modal detection module: Matches sensitive information through regular expressions; Identifies 7 types of named entities based on the BiLSTM-CRF model; Combines the Luhn algorithm to verify bank card numbers and calls a third-party API to verify the validity of addresses;

[0017] Dynamic desensitization strategy: Executes full replacement, partial hiding, or virtual information replacement according to the sensitivity level; Establishes a sensitive information association graph to block cross-field associations; Records desensitization operation logs, supporting retrieval by user ID and time range.

[0018] Preferably, the data collection layer and the model processing layer include:

[0019] Mapper stage: Shards the input data, with each shard containing 5000 text records; Performs preprocessing operations such as feature extraction and formatting on each text;

[0020] Reducer Phase: Combine the Mapper output results to generate structured detection results; support statistical analysis or aggregation calculations.

[0021] Preferably, the feedback optimization layer includes:

[0022] Error type definition module: Model confidence conflicts, contradictions between hate detection and privacy identification results; Unresolvable nested structures or newly emerging target group types;

[0023] Self-repair module: Isolate abnormal samples to the sandbox environment; Start rule learning, expand the label system and generate adversarial samples; Update the model parameters after verification through A / B testing.

[0024] A method for a hate speech and privacy information recognition system based on a large model, including data collection and preprocessing:

[0025] Obtain the social media text stream in real time through the API interface, establish a multi-source data access channel, and support JSON / XML format conversion;

[0026] Preprocess the collected data, including constructing a training set using three-level annotation quality control, performing pre-annotation, manual verification, and adversarial sample generation, and performing feature engineering, covering text vectorization and special feature construction.

[0027] Preferably, it includes hate speech detection:

[0028] Construct a hybrid model architecture and adopt a three-level processing pipeline:

[0029] Feature extraction layer: Use the BERT-BiLSTM dual-channel structure to fuse the [CLS] vector output by BERT and the character sequence processed by BiLSTM through the gated attention mechanism;

[0030] Quadruple parsing layer: Adopt joint decoding of pointer network + classifier to perform comment object recognition, argument extraction, target group classification, and hate determination;

[0031] Post-processing rule engine: Set contradiction detection rules and context compensation mechanisms.

[0032] Perform model training, set hyperparameters, and adopt a hierarchical learning rate and early stopping mechanism optimization strategy.

[0033] Preferably, it includes privacy information processing:

[0034] Adopt a multi-modal detection algorithm to construct a three-level detection system:

[0035] Regular expression engine: Use specific regular expressions to detect mobile phone numbers and ID card numbers, and support fuzzy matching;

[0036] Named entity recognition model: Use the BiLSTM-CRF model to identify 7 major categories of entities including person names, addresses, and organizations;

[0037] Semantic verification layer: Perform context legality checks and false information identification.

[0038] Implement a dynamic desensitization strategy, perform hierarchical processing according to the sensitivity level, establish an associated graph of sensitive information to implement an association blocking mechanism, and record audit tracking logs.

[0039] Include system interaction processing:

[0040] Real-time processing mode: After the user input text is cleaned, hate detection and privacy identification are performed in parallel. The hate detection engine generates a quadruple and gives a risk score, and the privacy identification engine performs dynamic desensitization to generate secure text;

[0041] Batch processing mode: Adopt the MapReduce framework. In the Mapper stage, preprocess the sharded data, and in the Reducer stage, merge the intermediate results to generate the final detection results.

[0042] Preferably, it includes exception handling:

[0043] Define error types, including model confidence conflicts, unresolvable nested structures, and newly emerging target group types;

[0044] Implement a self-repair process, isolate abnormal samples to a sandbox environment, start rule learning, automatically expand the label system for new group types, generate adversarial samples to supplement training data, and go online for update after verification by AB testing.

[0045] Compared with the prior art, the beneficial effects of the present invention are:

[0046] The hate speech and privacy information recognition system and method based on a large model proposed by the present invention obtain a social media text stream through a data acquisition layer. The hate detection engine (based on a Transformer hybrid neural network) and the privacy recognition engine (rule base + lightweight CNN model) in the model processing layer perform processing. The result output layer generates a structured output and performs dynamic desensitization, and the feedback optimization layer realizes manual review and incremental learning. The hate speech detection module constructs a training set through data preprocessing and performs feature engineering, adopts a hybrid model architecture and a post-processing rule engine, and has the capabilities of fine-grained detection and context understanding; the privacy information processing module realizes accurate identification and hierarchical desensitization through multi-modal detection algorithms and dynamic desensitization strategies. The system has real-time and batch processing modes and an exception handling mechanism. The present invention significantly improves the hate speech recognition efficiency, constructs a fine-grained detection system, provides comprehensive privacy information protection, and has important application value. Description of the Drawings

[0047] Figure 1 This is the system architecture diagram of the present invention;

[0048] Figure 2 This is the flowchart of the method of the present invention. Detailed implementation manners

[0049] In order to clearly and completely describe the objectives, technical solutions of the present invention, and make the advantages more clearly understood, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are some, but not all, embodiments of the present invention, and are only used to explain the embodiments of the present invention, rather than limiting the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] Embodiment 1, please refer to Figures 1 to 2 , the present invention provides a technical solution: A hate speech and privacy information recognition system based on a large model, including:

[0051] 1. The system adopts a hierarchical processing architecture, including four core components:

[0052] 1.1 Data acquisition layer: Real-time acquisition of social media text streams through API interfaces, establishment of multi-source data access channels, and support for JSON / XML format conversion.

[0053] 1.2 Model processing layer: Hate detection engine: A hybrid neural network based on Transformer

[0054] Privacy recognition engine: Rule base + lightweight CNN model

[0055] Shared feature extractor: BERT embedding layer parameter sharing mechanism

[0056] 1.3 Result output layer:

[0057] Structured output module: Generate a standardized quadruple JSON format

[0058] Dynamic desensitization module: Execute a hierarchical privacy replacement strategy

[0059] 1.4 Feedback optimization layer:

[0060] Manual review interface: Marking personnel correct the system's incorrect judgments

[0061] Incremental learning module: Automatically update model parameters every week

[0062] 2 Hate speech detection module

[0063] 2.1 Data preprocessing

[0064] Training set construction:

[0065] Adopt three - level annotation quality control:

[0066] 1. Pre - annotation: Manually search for typical cases and preliminarily annotate 6000 samples

[0067] 2. Manual verification: A 5 - person expert team corrects with a unified standard in blind annotation, and finally the Kappa coefficient ≥ 0.85

[0068] 3. Adversarial sample generation: Expand data through the following methods:

[0069] Synonym replacement (Word2Vec similarity ≥ 0.7)

[0070] Sentence pattern reconstruction (dependency syntactic analysis tree recombination)

[0071] Add interference information (randomly insert non - sensitive short sentences)

[0072] Feature engineering:

[0073] 1. Text vectorization:

[0074] Character - level embedding: The Conv1D layer extracts n - gram features (n = 3, 5, 7)

[0075] Semantic embedding: Chinese - RoBERTa - wwm - ext generates 768 - dimensional vectors

[0076] 2. Special feature construction:

[0077] Emotional polarity score (based on HowNet dictionary)

[0078] Aggressiveness index (calculate the density of insulting words)

[0079] Group feature word frequency statistics (preset a 500 - word sensitive word library)

[0080] 2.2 Hybrid model architecture

[0081] Three - level processing pipeline

[0082] 1. Feature extraction layer:

[0083] BERT - BiLSTM dual - path structure:

[0084] Path 1: RoBERTa outputs the [CLS] vector (dimension 768)

[0085] Path 2: BiLSTM processes the character sequence (256 - unit hidden layer)

[0086] Feature fusion: Gated attention mechanism

[0087] g = σ(Wg[hBERT; hBiLSTM] + bg)

[0088] hfusion = g ⊙ hBERT + (1 - g) ⊙ hBiLSTM

[0089] 2. Quadruple parsing layer:

[0090] Adopt joint decoding of pointer network + classifier:

[0091] Comment object recognition: BIO annotation CRF decoding (F1 = 0.92)

[0092] Argument extraction: Span-based boundary detection (IoU threshold 0.7)

[0093] Target group classification: Multi-label classifier (sigmoid output)

[0094] Hate determination: Dual-channel verification mechanism

[0095] Channel 1: Argument semantic analysis (threshold 0.65)

[0096] Channel 2: Target group relevance verification

[0097] 3. Post-processing rule engine:

[0098] Contradiction detection rule:

[0099] If the target group is "NULL" but the hate determination is "hate", trigger manual review and start segmented verification when the argument length > 100 characters

[0100] Context compensation mechanism:

[0101] Use coreference resolution technology to complete the omitted subject

[0102] Model training details:

[0103] Hyperparameter settings:

[0104] {

[0105] "learning_rate": 2e-5,

[0106] "batch_size": 32,

[0107] "max_seq_length": 256,

[0108] "crf_loss_weight": 0.7,

[0109] "focal_loss_gamma": 2.0

[0110] }

[0111] Optimization Strategy:

[0112] Hierarchical learning rate: lr for BERT layer = 2e-5, lr for the top-level network = 1e-3

[0113] Early stopping mechanism: Terminate if the F1 of the validation set drops by > 0.5% for 3 consecutive epochs

[0114] 3 Privacy Information Processing Module

[0115] 3.1 Multi-modal Detection Algorithm

[0116] Three-level Detection System:

[0117] 1. Regular Expression Engine:

[0118] PHONE_REGEX = r'(?<!\d)(1[3-9]\d{9})(?!\d)'

[0119] ID_CARD_REGEX = r'\b[1-9]\d{5}(18|19|20)\d{2}(0[1-9]|1[0-2])(0[1-9]|

[12] \d|3

[01] )\d{3}[\dXx]\b'

[0120] Supports fuzzy matching (allows ±1 character tolerance)

[0121] 2. Named Entity Recognition Model:

[0122] Uses the BiLSTM-CRF model for recognition:

[0123] Entity types: 7 major categories such as person names / addresses / organizations, etc.

[0124] 3. Semantic Verification Layer:

[0125] Context Legitimacy Check:

[0126] Verification of bank card numbers using the Luhn algorithm

[0127] Address validity check (can be associated with the Gaode API) False information identification:

[0128] Validity check for the first three digits of mobile phone numbers

[0129] Calculation of ID card verification codes

[0130] 3.2 Dynamic Data Masking Strategy

[0131] Hierarchical Processing Rules:

[0132]

[0133] Association blocking mechanism:

[0134] Build a sensitive information association graph:

[0135] graphLR

[0136] A[Mobile phone number] --> B[ID card]

[0137] A --> C[Bank card]

[0138] B --> D[Home address]

[0139] When any node is detected, automatically scan the associated fields

[0140] Audit trail:

[0141] Log record format:

[0142] {

[0143] "original": "Zhang San 13800138000",

[0144] "processed": "Zhang****",

[0145] "rule_id": "PHONE_1",

[0146] "timestamp": "2023-08-20T14:35:22Z"

[0147] }

[0148] Traceability of sensitive operations: Support retrieval by user ID + time range

[0149] 4 System interaction process

[0150] 1. Real-time processing mode:

[0151] User input -> Text cleaning -> Parallel detection:

[0152] │

[0153] Hate detection engine → Generate quadruple → Risk scoring

[0154] Privacy recognition engine → Dynamic desensitization → Secure text

[0155] 2. Batch processing mode:

[0156] Adopt the MapReduce framework: MapReduce is a distributed computing framework widely used in the processing of large-scale data. It decomposes tasks into two main stages: the Map stage and the Reduce stage to achieve efficient data processing.

[0157] Mapper stage:

[0158] The role of the Mapper is to fragment the input data. In our application scenario, each piece of data contains 5000 text records. The Mapper reads these text data one by one and performs necessary preprocessing on each piece of data, such as feature extraction, formatting, etc. The processed data will be output as intermediate results for use in the subsequent Reducer stage. The design goal of the Mapper is to decompose complex data processing tasks into multiple small tasks, thereby achieving parallel processing and improving the overall processing efficiency.

[0159] Reducer stage:

[0160] The role of the Reducer is to merge and summarize the intermediate results generated in the Mapper stage. In our scenario, the Reducer will receive the output results from multiple Mappers and perform merge processing on these results. For example, if the Mapper stage performs feature extraction or preliminary analysis on text data, the Reducer will integrate these scattered results together to generate the final detection results. The processing logic of the Reducer usually depends on specific business requirements, such as statistical analysis, aggregation calculation, etc. Through the merge operation of the Reducer, we can integrate scattered processing results into a complete output to provide support for subsequent data analysis or applications.

[0161] 5 Exception handling mechanism

[0162] 1. Error type definition:

[0163] E001: Model confidence conflict (conflicting results between the hate detection engine and the privacy engine)

[0164] E002: Unresolvable nested structure

[0165] E003: Newly emerged target group type

[0166] 2. Self-healing process:

[0167] Step 1: Isolate the abnormal sample to the sandbox environment

[0168] Step 2: Start rule learning:

[0169] Automatically expand the label system for the new group type

[0170] Generate adversarial examples to supplement training data

[0171] Step 3: Go live and update after AB test verification.

[0172] Example 2, based on Example 1, proposes a method for a hate speech and privacy information recognition system based on a large model, including data collection and preprocessing: real-time acquisition of social media text streams through API interfaces, establishment of multi-source data access channels, and support for JSON / XML format conversion; preprocessing of the collected data, including constructing a training set using three-level annotation quality control, performing pre-annotation, manual verification, and adversarial example generation, and performing feature engineering, covering text vectorization and special feature construction.

[0173] Including hate speech detection: constructing a hybrid model architecture, adopting a three-level processing pipeline: Feature extraction layer: Using a BERT-BiLSTM dual-path structure, fusing the [CLS] vector output by BERT and the character sequence processed by BiLSTM through a gated attention mechanism; Quadruple parsing layer: Adopting a pointer network + classifier joint decoding to perform comment object recognition, argument extraction, target group classification, and hate determination; Post-processing rule engine: Setting contradiction detection rules and context compensation mechanisms. Perform model training, set hyperparameters, and adopt a hierarchical learning rate and early stopping mechanism optimization strategy.

[0174] Including privacy information processing: adopting a multi-modal detection algorithm, constructing a three-level detection system: Regular expression engine: Using specific regular expressions to detect mobile phone numbers and ID card numbers, supporting fuzzy matching; Named entity recognition model: Using a BiLSTM-CRF model to recognize 7 major types of entities such as person names, addresses, and organizations; Semantic verification layer: Performing context legality checks and false information recognition. Implement a dynamic desensitization strategy, perform hierarchical processing according to the sensitivity level, establish a sensitive information association graph to implement an association blocking mechanism, and record audit tracking logs.

[0175] Including system interaction processing: Real-time processing mode: After the user input text is cleaned, hate detection and privacy recognition are performed in parallel. The hate detection engine generates quadruples and gives a risk score, and the privacy recognition engine performs dynamic desensitization to generate secure text; Batch processing mode: Adopting the MapReduce framework, the Mapper stage preprocesses the sharded data, and the Reducer stage merges the intermediate results to generate the final detection results.

[0176] Including exception handling: Define error types, including model confidence conflicts, unresolvable nested structures, and newly emerging target group types; Implement a self-repair process, isolate abnormal samples to a sandbox environment, start rule learning, automatically expand the label system for new group types, generate adversarial examples to supplement training data, and go live and update after AB test verification.

[0177] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A hate speech and privacy information recognition system based on a large model, characterized in that: Including: Data collection layer: Real-time acquisition of multi-source social media text streams through API interfaces, supporting JSON / XML format conversion; Model processing layer: Hate detection engine: A hybrid neural network based on Transformer, combined with the BERT embedding layer parameter sharing mechanism; Privacy recognition engine: Adopts a detection method combining a rule base and a lightweight CNN model; Result output layer: Structured output module: Generates a quadruple JSON format containing the comment object, argument, target group, and judgment result; Dynamic desensitization module: Executes a hierarchical privacy replacement strategy according to the sensitivity level; Feedback optimization layer: Manual review interface: Supports annotators to correct system misjudgments; Incremental learning module: Automatically updates model parameters regularly.

2. The hate speech and privacy information recognition system based on a large model according to claim 1, characterized in that: The model processing layer includes: Data preprocessing module: Constructs a three-level labeled training set: pre-labeling, manual verification, and adversarial sample generation; Extracts character-level n-gram features Conv1D and semantic embeddings Chinese-RoBERTa-wwm-ext; Hybrid model architecture: Adopts a BERT-BiLSTM dual-path structure to fuse features, enhances the representation through a gated attention mechanism; Based on a pointer network + classifier joint decoding, outputs quadruple detection results; Post-processing rules: Executes contradiction detection; Utilizes coreference resolution technology to supplement omitted subjects.

3. The hate speech and privacy information recognition system based on a large model according to claim 2, characterized in that: The model processing layer includes: Multi-modal detection module: Matches sensitive information through regular expressions; Identifies 7 types of named entities based on the BiLSTM-CRF model; Combines the Luhn algorithm to verify bank card numbers and calls a third-party API to verify the validity of addresses; Dynamic desensitization strategy: Executes full replacement, partial hiding, or virtual information replacement according to the sensitivity level; Establishes a sensitive information association graph to block cross-field associations; Records desensitization operation logs, supporting retrieval by user ID and time range.

4. The hate speech and privacy information recognition system based on a large model according to claim 3, wherein: The data collection layer and the model processing layer include: Mapper stage: Shards the input data, with each shard containing 5000 text records; Performs preprocessing operations such as feature extraction and formatting on each text; Reducer stage: Combines the Mapper output results to generate structured detection results; Supports statistical analysis or aggregation calculations.

5. The hate speech and privacy information recognition system based on a large model according to claim 4, characterized in that: The feedback optimization layer includes: Error type definition module: Model confidence conflicts, contradictions between hate detection and privacy recognition results; Unresolvable nested structures or newly emerged target group types; Self-repair module: Isolates abnormal samples to a sandbox environment; Initiates rule learning, expands the label system, and generates adversarial samples; Updates model parameters after verification through A / B testing.

6. A method for the system for identifying hate speech and privacy information based on a large model according to claim 5, characterized in that: Including data collection and preprocessing: Real-time acquisition of social media text streams through API interfaces, establishes a multi-source data access channel, supporting JSON / XML format conversion; Preprocesses the collected data, including constructing a training set using three-level annotation quality control, performing pre-annotation, manual verification, and adversarial sample generation, as well as performing feature engineering, covering text vectorization and special feature construction.

7. A method according to claim 6, wherein: Including hate speech detection: Constructs a hybrid model architecture, adopting a three-level processing pipeline: Feature extraction layer: Use the BERT-BiLSTM dual-path structure to fuse the [CLS] vector output by BERT and the character sequence processed by BiLSTM through the gated attention mechanism; Quadruple parsing layer: Adopt the joint decoding of pointer network + classifier to perform comment object recognition, argument extraction, target group classification, and hate determination; Post-processing rule engine: Set contradiction detection rules and context compensation mechanisms; Perform model training, set hyperparameters, and adopt the hierarchical learning rate and early stopping mechanism optimization strategy.

8. A method according to claim 6, wherein: Include privacy information processing: Adopt a multi-modal detection algorithm to construct a three-level detection system: Regular expression engine: Use specific regular expressions to detect mobile phone numbers and ID card numbers, and support fuzzy matching; Named entity recognition model: Use the BiLSTM-CRF model to recognize 7 categories of entities including person names, addresses, and organizations; Semantic verification layer: Perform context legality check and false information recognition; Implement dynamic desensitization strategy, perform hierarchical processing according to the sensitivity level, establish an associated graph of sensitive information to implement the association blocking mechanism, and record the audit trail log.

9. A method according to claim 6, wherein: Include system interaction processing: Real-time processing mode: After the user input text is cleaned, hate detection and privacy recognition are performed in parallel. The hate detection engine generates quadruples and gives a risk score, and the privacy recognition engine performs dynamic desensitization to generate secure text; Batch processing mode: Adopt the MapReduce framework. In the Mapper stage, preprocess the sharded data, and in the Reducer stage, merge the intermediate results to generate the final detection results.

10. A method according to claim 6, characterized in that: Include exception handling: Define error types, including model confidence conflict, unresolvable nested structure, and newly emerged target group types; Implement a self-repair process, isolate abnormal samples to the sandbox environment, start rule learning, automatically expand the label system for the new group type, generate adversarial samples to supplement the training data, and go online for update after verification by AB testing.

Citation Information

Cited By

  • Automatic supervision method and system for psychological counseling

    CN121075568A

  • A method and system for automatic supervision of psychological counseling

    CN121075568B