Method and system for accurately detecting sensitive content of reimbursement information based on multilayer semantic coding and triple learning

By using multi-layer semantic coding and triple-group learning methods in the detection of reimbursement information sensitive content, the problems of insufficient semantic understanding and poor generalization of the model in the prior art are solved, and efficient and accurate detection of sensitive content is achieved.

CN120030144APending Publication Date: 2025-05-23COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411989396.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient semantic understanding depth, limited feature representation ability and insufficient generalization of the model in the detection of reimbursement information sensitive content, resulting in low detection accuracy and efficiency.

Method used

The detection method based on multi-layer semantic coding and triple learning is adopted, and multi-layer semantic coding is performed through BERT-BiLSTM combined with multi-head attention mechanism deep learning architecture, and the feature mapping model is trained using triple loss function to construct semantic mapping relationships to accurately identify sensitive content.

Benefits of technology

It realizes a deep semantic understanding of complex texts, improves the accuracy of feature representation and the utilization of multi-level semantic information, enhances the generalization ability and recognition efficiency of the model, and significantly improves the accuracy and efficiency of the detection of sensitive content of reimbursement information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030144A_ABST
    Figure CN120030144A_ABST
Patent Text Reader

Abstract

The invention provides a reimbursement information sensitive content accurate detection method and system based on multilayer semantic coding and triple learning, and belongs to the field of natural language processing. Through systematic analysis of mistakenly screened cases in real reimbursement data, a set of multilayer semantic coding framework is designed, and the accuracy of sensitive content recognition is remarkably improved in combination with a triple learning method. Compared with the prior art, the method has the main innovation points that three types of typical pseudo-sensitive content scenes are summarized through real data analysis; according to the method, a triple structure is innovatively constructed for the reimbursement text containing the sensitive words, accurate judgment on whether the text containing the sensitive words contains the sensitive content or not is achieved through a multi-layer semantic coding and deep learning method, and the real sensitive content and the false sensitive content are effectively distinguished. According to the method, the accuracy of detecting the sensitive content of the reimbursement information is remarkably improved, and the method has important practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical field to which the present invention belongs is the intersection of natural language processing and financial compliance management. Specifically, it integrates the semantic understanding technology in natural language processing and the demand for sensitive content detection in financial information management, aiming to improve the compliance and efficiency of reimbursement information processing through innovative technical means. Background Art

[0002] As the scale of enterprises expands and compliance requirements become increasingly stringent, compliance review of reimbursement information has become an important part of enterprise risk management. By screening sensitive content in the bill details and reimbursement reasons in the reimbursement information, not only can potential risks be discovered in time at the early stage of filling in and pre-warning can be achieved, but also statistical analysis of historical data can provide strong support for post-risk tracing and management.

[0003] At present, the detection of sensitive content in reimbursement information has mainly gone through the following development stages: Rule-based detection method: It mainly relies on the preset sensitive word library for word matching and screening, which is simple to implement and highly efficient. Detection method that introduces semantic analysis: It uses shallow natural language processing technology to perform preliminary semantic analysis on the text, which improves the detection accuracy to a certain extent. Machine learning-assisted detection method: It combines statistical learning models to model text features and enhances the system's adaptability. However, the existing technical solutions still have the following problems and shortcomings:

[0004] Insufficient depth of semantic understanding: Existing systems find it difficult to accurately distinguish the semantic changes of the same sensitive word in different contexts, and have limited ability to understand texts containing complex syntactic structures such as modifiers and qualifiers. At the same time, they lack effective modeling methods when dealing with contextual semantic dependencies.

[0005] Limited feature representation capabilities: Due to the rough division of the feature space, it is difficult for the system to effectively distinguish between similar texts that also contain sensitive words. The single dimension of feature extraction leads to the failure to fully utilize the multi-level semantic information of the text.

[0006] Insufficient model generalization: The system performs unstably when faced with new reimbursement scenarios. In particular, when encountering sensitive content that does not appear in the training data, the recognition ability is significantly reduced, exposing the problem of poor scenario adaptability.

[0007] These problems in the existing technical solutions seriously affect the accuracy and efficiency of reimbursement information screening. A large number of false screening results not only increase the workload of manual review, but also reduce the overall effectiveness of corporate compliance management. Therefore, it is of great practical significance to develop an intelligent detection method that can accurately understand text semantics and accurately identify real sensitive content. Summary of the invention

[0008] In view of the problems existing in the prior art, the purpose of the present invention is to provide a method and system for accurately detecting sensitive content of reimbursement information based on multi-layer semantic coding and triple learning.

[0009] The technical solution adopted by the present invention is as follows:

[0010] A method for accurately detecting sensitive content of reimbursement information based on multi-layer semantic coding and triple learning, comprising the following steps:

[0011] Use historical reimbursement data to build a training sample set containing typical misscreening situations, and establish a "anchor text-similar text-comparison text" triple for each sensitive word to form a sensitive content text pool;

[0012] The deep learning architecture of BERT-BiLSTM combined with multi-head attention mechanism is used for multi-layer semantic encoding to convert the "anchor text-similar text-comparison text" triple into a high-dimensional feature vector;

[0013] A semantic mapping relationship is constructed in the feature space based on the high-dimensional feature vector of the "anchor text-similar text-comparison text" triple, and the feature mapping model is trained through the triple loss function;

[0014] For the reimbursement text to be judged, a triplet of "anchor text-sensitive sample-safe sample" is constructed. The trained feature mapping model is used to convert the text in the triplet of "anchor text-sensitive sample-safe sample" into a high-dimensional feature vector. By calculating the similarity between the feature vectors of the anchor text and the sensitive sample, as well as between the anchor text and the safe sample, it is determined whether the reimbursement text to be judged contains sensitive content.

[0015] Furthermore, the training sample set containing typical misscreening situations includes word meaning ambiguity samples, modifier misjudgment samples and sentiment context misjudgment samples, and the samples are divided into sensitive texts and safe texts through annotation.

[0016] Furthermore, the "anchor text-similar text-comparison text" triplet includes two types of triplet training sets: taking sensitive text as anchor text, combining sensitive text and safe text containing the same sensitive words, to form a triplet of (sensitive anchor text, similar sensitive text, comparison safety text); taking safe text as anchor text, combining safe text and sensitive text containing the same sensitive words, to form a triplet of (safe anchor text, similar safe text, comparison sensitive text).

[0017] Furthermore, the deep learning architecture using BERT-BiLSTM combined with a multi-head attention mechanism for multi-layer semantic encoding includes:

[0018] Initially encode the text through the BERT pre-trained model to obtain the basic word vector representation;

[0019] The basic word vector representation is input into the BiLSTM network, and the forward and backward recurrent neural network structures are used to capture the long-range semantic dependencies in the text;

[0020] A multi-head attention mechanism is introduced to deeply process the output of BiLSTM, identify key semantic information and assign higher weights;

[0021] The outputs of multiple attention heads are concatenated and linearly transformed to generate a comprehensive text semantic feature vector representation.

[0022] Furthermore, the training of the feature mapping model by the triplet loss function includes:

[0023] Map the text in the triple to the corresponding feature vector;

[0024] For each type of triplet, the triplet loss function minimizes the distance between the anchor text and similar text, while maximizing the distance between the anchor text and the contrasting text;

[0025] The back-propagation algorithm is used to continuously optimize the model parameters, form a clear semantic distribution structure in the feature space, and train a high-quality feature mapping model.

[0026] Furthermore, the reimbursement text to be judged is constructed into a triple of "anchor text-sensitive sample-safe sample", including:

[0027] Identify sensitive words in the reimbursement text to be judged through a sensitive word library;

[0028] The reimbursement text to be judged is used as the anchor text, and sensitive samples and safety samples related to sensitive words are retrieved from the sample library to form a triplet of "anchor text-sensitive sample-safe sample".

[0029] Furthermore, the similarity is cosine similarity.

[0030] A system for accurately detecting sensitive content of reimbursement information based on multi-layer semantic coding and triple learning, comprising:

[0031] The sensitive content text pool construction module is responsible for using historical reimbursement data to construct a training sample set containing typical misscreening situations, and establishes a "anchor text-similar text-comparison text" triple for each sensitive word to form a sensitive content text pool;

[0032] The multi-layer semantic encoding module is responsible for using the deep learning architecture of BERT-BiLSTM combined with the multi-head attention mechanism to perform multi-layer semantic encoding and convert the "anchor text-similar text-comparison text" triple into a high-dimensional feature vector;

[0033] The triple learning feature space mapping module is responsible for constructing semantic mapping relationships in the feature space based on the high-dimensional feature vectors of the "anchor text-similar text-comparison text" triples, and training the feature mapping model through the triple loss function;

[0034] The new data sensitive content identification module is responsible for constructing the "anchor text-sensitive sample-safe sample" triplet for the reimbursement text to be judged, and using the trained feature mapping model to convert the text in the "anchor text-sensitive sample-safe sample" triplet into a high-dimensional feature vector. By calculating the similarity between the feature vectors of the anchor text and the sensitive sample and between the anchor text and the safe sample, it is determined whether the reimbursement text to be judged contains sensitive content.

[0035] The beneficial effects of the present invention are as follows:

[0036] (1) Technical advantages: The multi-layer feature extraction architecture based on deep semantic understanding accurately captures the contextual associations and semantic dependencies of the text through a multi-head attention mechanism and a bidirectional recurrent network, effectively solving the limitations of traditional rule-based methods in dealing with complex contexts such as word ambiguity and modifier interference.

[0037] (2) Recognition efficiency and accuracy: Innovatively adopt a "coarse screening + fine screening" double-layer discrimination strategy, combining traditional sensitive word matching with semantic similarity calculation of deep learning. By optimizing the triple training set, a clear semantic distribution structure is constructed in the feature space to provide a reliable measurement basis for new data discrimination. Among them, coarse screening refers to the use of word segmentation processing and regular expression matching of predefined sensitive word libraries in the sensitive content text pool construction module and the new data recognition module of the present invention to preliminarily screen texts containing sensitive words; the fine screening process refers to further judging whether the text actually contains sensitive content through deep semantic analysis and feature space measurement.

[0038] (3) Practicality and scalability: The system supports the update of sensitive word libraries and can adapt to the recognition needs of new sensitive content. The discrimination mechanism based on multi-dimensional similarity calculation ensures the stable performance and flexible expansion of the system in different scenarios.

[0039] (4) In practical applications, the present invention can improve the efficiency of corporate compliance management through accurate identification of sensitive content and effectively reduce the workload of manual review. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flow chart of the steps of the method of the present invention.

[0041] Figure 2 It is a schematic diagram of the training process and new data processing process of the method of the present invention. DETAILED DESCRIPTION

[0042] The present invention is further described in detail below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention. The present invention is suitable for use in situations where sensitive content screening is required in the field of financial reimbursement. The specific implementation methods of the present invention are further described in detail below.

[0043] The present invention mainly includes the following contents:

[0044] (1) Deep understanding of text semantics: The present invention adopts the BERT-BiLSTM-multi-head attention network architecture to realize multi-level semantic feature extraction of reimbursement text. This architecture captures basic semantic representation through a pre-trained language model, combines a bidirectional recurrent neural network to process sequence dependencies, and uses an attention mechanism to highlight key semantic information, thereby building a complete semantic representation system.

[0045] (2) Accurate sensitivity identification: This paper designs a feature space mapping mechanism based on the triple learning framework and optimizes the vector representation of text features through a metric learning method. This mechanism uses loss function constraints to make semantically similar sensitive text pairs form a tight cluster structure in the feature space, while text pairs with significant semantic differences maintain a sufficient distance boundary, thereby achieving accurate quantification of sensitivity.

[0046] (3) Model generalization and adaptation: We build a diverse sensitive text training set that covers three typical scenarios: word meaning ambiguity, misjudgment of modifiers, and misjudgment of sentiment context. The discrimination process always calculates similarity based on the triple structure. Even when faced with new sensitive words or unseen expressions, we only need to construct them into corresponding triple structures to maintain consistent recognition results, ensuring the stable performance of the model in a dynamically changing business environment.

[0047] The core architecture of the present invention consists of three key components: a sensitive content text pool construction module, a multi-layer semantic encoding module, and a triple learning feature space mapping module. The sensitive content text pool construction module is responsible for organizing and analyzing historical reimbursement data, constructing a training sample set containing typical misscreening situations, and establishing a standardized triple structure for each sensitive word. The multi-layer semantic encoding module adopts a deep learning architecture of BERT-BiLSTM combined with a multi-head attention mechanism to convert triple text into a high-dimensional feature vector. The triple learning feature space mapping module constructs a semantic mapping relationship in the feature space based on these feature vectors, and ensures that text feature vectors with similar semantics are clustered through the optimization of the loss function, while text feature vectors with significant semantic differences maintain a sufficient distance. In practical applications, when the system receives a new text to be judged, it first identifies and extracts the sensitive words therein, retrieves sensitive samples and safety samples related to these sensitive words from the sample library, and constructs the triples required for evaluation (anchor text, sensitive sample, safety sample). Subsequently, the triples are converted into feature vectors through a multi-layer semantic encoding module and mapped to the feature space. The cosine similarity between the feature vector of the anchor text to be identified and the feature vectors of sensitive samples and safe samples is calculated and compared to finally determine whether the text to be identified contains sensitive content.

[0048] The technical solution adopted by the present invention is as follows Figure 1 As shown, it includes the following modules:

[0049] (1) Sensitive content text pool construction module: This module is mainly responsible for the collection, cleaning and organization of historical reimbursement data. This module focuses on constructing a sample set covering three typical misscreening situations: word meaning ambiguity, modifier misjudgment and emotional context misjudgment, and constructs a triple structure of "anchor text-similar text-comparison text" for each sensitive word.

[0050] (2) Multi-layer semantic encoding module: It adopts a deep learning architecture, relies on the BERT pre-training model to extract basic semantic features, combines the BiLSTM network to capture the long-distance dependencies of text, and introduces a multi-head attention mechanism to highlight key semantic information. Through the feature fusion layer, multi-level semantic features are integrated to achieve efficient conversion of text to feature vectors, providing a solid vector representation for subsequent feature space mapping.

[0051] (3) Triplet learning discriminant module: This module uses the triplet loss function to adaptively learn the feature space mapping relationship. The module minimizes the distance between the anchor text and similar texts while maximizing the distance between the anchor text and the contrasting text, so that the model can automatically form a reasonable semantic distribution structure in the feature space.

[0052] (4) New data sensitive content identification module: This module is responsible for processing new reimbursement texts and first performs preprocessing. Then, it uses a multi-layer semantic encoding module to extract text feature vectors, calculates and compares the cosine similarity between the feature vectors of the text to be identified and the sensitive samples and safe samples, and finally determines whether the text to be identified contains sensitive content.

[0053] The above four modules constitute the accurate detection system of sensitive content of reimbursement information based on multi-layer semantic coding and triple learning of the present invention. The detailed technical process of the present invention is carried out in sequence according to the four core modules to realize the intelligent processing of the whole process from data preparation to final judgment.

[0054] (1) In the sensitive content text pool construction module, such as Figure 2 As shown in module (a), the historical reimbursement data is segmented, and the text containing sensitive words is preliminarily screened by matching the predefined sensitive word library with regular expressions. On this basis, the screened text is deeply analyzed to construct a sample set containing three typical misscreening scenarios: word meaning ambiguity, modifier misjudgment, and emotional context misjudgment. Among them, word meaning ambiguity means that some sensitive words have multiple meanings, and when they express non-sensitive meanings in the reimbursement text, they are mistakenly identified as sensitive content; modifier misjudgment means that sensitive words are used as modifiers or descriptive words in the text, but their core reimbursement content does not actually involve sensitive information; emotional context misjudgment means that some sensitive words constitute sensitive content only in specific emotional contexts, but not in other emotional contexts. Through professional manual annotation, the samples are divided into sensitive samples (texts that do contain sensitive content) and safe samples (texts that only contain sensitive words but do not constitute sensitive content). On this basis, a triple structure of (anchor text, similar text, comparison text) is constructed, which contains two types of triple training sets: using sensitive text as anchor text, combining sensitive text and safe text containing the same sensitive words, to form a triple of (sensitive anchor text, similar sensitive text, comparison safe text); using safe text as anchor text, combining safe text and sensitive text containing the same sensitive words, to form a triple of (safe anchor text, similar safe text, comparison sensitive text). This two-way triple construction method lays the data foundation for subsequent feature learning.

[0055] (2) In the multi-layer semantic encoding module, such as Figure 2As shown in module (b), the system uses the BERT-BiLSTM-multi-head attention deep learning architecture to achieve multi-level feature extraction of text. First, the text is initially encoded through the BERT pre-trained model to obtain the basic word vector representation. These word vectors are input into the BiLSTM network layer, and the forward and backward recurrent neural network structures are used to effectively capture the long-distance semantic dependencies in the text. Subsequently, the multi-head attention mechanism is introduced to deeply process the output of BiLSTM, identify key semantic information and assign higher weights. Finally, the outputs of multiple attention heads are spliced ​​and linearly transformed to generate a comprehensive text semantic feature vector representation.

[0056] (3) In the triple learning module, Figure 2 As shown in module (c), the system maps the text in the triplet to the corresponding feature vector and optimizes the feature space distribution structure through the triplet loss function. The triplet structure is (anchor text, similar text, comparison text). The loss function calculates the distance between the anchor text and the similar text and the distance between the anchor text and the comparison text to ensure that the distance between semantically similar text pairs (such as sensitive anchor text and similar sensitive text, or safe anchor text and similar safe text) in the feature space is significantly smaller than the distance between semantically different comparison texts. The system uses the back-propagation algorithm to continuously optimize the model parameters, and finally forms a clear semantic distribution structure in the feature space. The system trains a high-quality feature mapping model to provide a reliable measurement basis for new sample discrimination.

[0057] The specific calculation formula of the triplet loss function is as follows:

[0058] L=max(0,D(a,p)-D(a,n)+margin)

[0059] Among them, L represents the loss value; D(a,p) represents the square of the Euclidean distance between the anchor sample (anchor) and the positive sample (positive), where the positive sample is the similar text in the triple structure; D(a,n) represents the square of the Euclidean distance between the anchor sample (anchor) and the negative sample (negative), where the negative sample is the contrasting text in the triple structure. Margin is a boundary value used to control the minimum interval between the positive and negative samples. A too small margin value will lead to insufficient distinction between positive and negative samples. A too large margin may cause training difficulties and the model is difficult to converge. Therefore, it is necessary to determine a suitable threshold through experiments to balance the model's discriminative ability and training stability. The optimization goal of the triple loss function is to minimize the distance D(a,p) between the anchor sample and the positive sample, and maximize the distance D(a,n) between the anchor sample and the negative sample, ensuring that the distance difference between the positive and negative samples is greater than the margin. In this way, discriminative feature representations can be effectively learned.

[0060] (4) In the new data sensitive content identification module, such as Figure 2 As shown in module (d), the system performs intelligent analysis and processing on the newly received text to be judged. The system first performs word segmentation on the text, identifies sensitive words in the text by matching the predefined sensitive word library with regular expressions, and preliminarily screens out texts containing sensitive words. Then, the screened text to be judged is used as anchor text, and relevant sensitive samples and security samples are retrieved from the sample library to construct evaluation triples. Then, the trained feature mapping model is called to convert the text in the triplet into a high-dimensional feature vector. Finally, the system calculates the cosine similarity between the new sample and the sensitive sample and the security sample, and determines whether the text to be detected contains sensitive content.

[0061] An embodiment of the present invention provides a method for accurately screening sensitive content of reimbursement information based on semantics, and the specific implementation includes the following steps:

[0062] (1) Construction of sensitive content text pool, mainly completing data preprocessing and training sample construction:

[0063] ① Data preprocessing: Perform text segmentation on historical reimbursement data using mainstream Chinese word segmentation tools such as Jieba Word Segmentation; perform text cleaning operations, including removing special characters, unifying uppercase and lowercase, and normalizing punctuation marks; standardize the formats of numbers and dates; and remove stop words and meaningless strings.

[0064] ②Sensitive sample screening: perform text matching based on a predefined sensitive word library; use regular expressions and other technologies to locate text fragments containing sensitive words; extract sensitive word context information and build an initial candidate set.

[0065] ③ Manual labeling and classification: Professionals are organized to label the candidate set and classify samples containing sensitive words into two categories: safe samples and sensitive samples. Among them, safe samples refer to texts that contain sensitive words but do not constitute sensitive content. Such samples mainly involve three types of misscreening: ambiguous word meanings, misjudgment of modifiers, and misjudgment of emotional contexts; sensitive samples refer to texts that do contain sensitive content.

[0066] ④ Construction of triple training set: Two types of triples are constructed with sensitive words as the center. When sensitive text is selected as anchor text, sensitive text and safe text containing the same sensitive word are combined; when safe text is selected as anchor text, safe text and sensitive text containing the same sensitive word are combined. The triple structures of (sensitive anchor text, similar sensitive text, comparative safe text) and (safe anchor text, similar safe text, comparative sensitive text) are formed respectively.

[0067] (2) Multi-layer semantic encoding module: This module uses a deep learning architecture for feature extraction:

[0068] ① BERT pre-training layer: The pre-trained Chinese BERT model is used as the basic encoder. The input text is input into the model after WordPiece segmentation, and the deep semantic features of the text are extracted using BERT's multi-layer Transformer structure.

[0069] ②BiLSTM processing layer: Design a bidirectional LSTM network structure, the forward LSTM captures the semantic dependency from left to right, the backward LSTM captures the semantic dependency from right to left, and the bidirectional features are concatenated to obtain enhanced semantic representation.

[0070] ③Multi-head attention layer: introduces a multi-head self-attention mechanism to capture the multi-dimensional feature associations of the text through parallel attention calculation heads. Each attention head focuses on different semantic features of the text by learning different query-key-value transformation matrices. Finally, the output features of multiple attention heads are weighted and fused to obtain feature representations that fully consider contextual relationships.

[0071] ④Fully connected layer: The fused features are subjected to nonlinear transformation and dimensionality reduction processing through a multi-layer fully connected network, and finally a text feature vector of fixed dimension is output to provide a standardized feature representation for subsequent similarity calculation.

[0072] (3) Triplet learning feature space mapping module: This module optimizes the semantic distribution of feature space:

[0073] ① Feature vector mapping: Map the text in the triplet to a feature space of uniform dimension through feature encoding to ensure the integrity and representation consistency of the original semantic information.

[0074] ② Loss function design: Construct a triplet loss function. For each type of triplet, minimize the distance between the anchor text and similar text, maximize the distance between the anchor text and the comparison text, and introduce boundary constraints to ensure sufficient category discrimination.

[0075] ③Distance metric selection: Cosine similarity is used as the distance metric between feature vectors to effectively measure the similarity of text semantics.

[0076] (4) New data sensitive content identification module:

[0077] ① Sensitive word identification: Accurately match the text to be judged based on the preset sensitive word library.

[0078] ②Related sample retrieval: Retrieve relevant sensitive samples and security samples in the sample library based on the identified sensitive words, and construct triples for evaluation.

[0079] ③Similarity evaluation: Use multi-layer semantic coding to extract the semantic features of the text, convert the text in the triple into a feature vector through the trained feature mapping method, and calculate the cosine similarity between the text to be judged and the sensitive sample and the safe sample. If the similarity between the text to be judged and the sensitive sample is greater than the similarity with the safe sample, it is determined that the text contains sensitive content; otherwise, it is determined that it does not contain sensitive content.

[0080] It should be understood that the methods and systems disclosed in the above embodiments provided by the present invention can be implemented in other ways. For example, the above module division can have other division methods in actual implementation, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0081] Each module in the present invention can be implemented in the form of a software functional unit, which can be stored in a computer-readable storage medium, including several instructions for enabling a computer device to perform some or all of the steps of the method of the present invention. For example, one embodiment of the present invention provides a computer device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention. For example, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, CD, etc.), the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the various steps of the method of the present invention are implemented.

[0082] Although the specific embodiments of the present invention are disclosed for the purpose of illustration, the purpose is to help understand the content of the present invention and implement it accordingly, those skilled in the art will understand that various substitutions, changes and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiment, and the scope of the present invention is subject to the scope defined in the claims.

Claims

1. A method for accurately detecting sensitive content of reimbursement information based on multi-layer semantic coding and triple learning, characterized in that: The following steps are involved: Use historical reimbursement data to build a training sample set containing typical misscreening situations, and establish a "anchor text-similar text-comparison text" triple for each sensitive word to form a sensitive content text pool; The deep learning architecture of BERT-BiLSTM combined with multi-head attention mechanism is used for multi-layer semantic encoding to convert the "anchor text-similar text-comparison text" triple into a high-dimensional feature vector; Based on the high-dimensional feature vector of the "anchor text-similar text-comparison text" triple, a semantic mapping relationship is constructed in the feature space, and the feature mapping model is trained through the triple loss function; The "anchor text-sensitive sample-safe sample" triplet is constructed for the reimbursement text to be judged. The trained feature mapping model is used to convert the text in the "anchor text-sensitive sample-safe sample" triplet into a high-dimensional feature vector. By calculating the similarity between the feature vectors of the anchor text and the sensitive sample, and between the anchor text and the safe sample, it is determined whether the reimbursement text to be judged contains sensitive content.

2. The method according to claim 1, characterized in that The training sample set containing typical misscreening situations includes samples of ambiguous word meanings, samples of misjudged modifiers, and samples of misjudged emotional contexts, and the samples are divided into sensitive texts and safe texts through annotation.

3. The method according to claim 1, characterized in that The "anchor text-similar text-comparison text" triples include two types of triple training sets: using sensitive text as anchor text, combining sensitive text and safe text containing the same sensitive words to form a triple of (sensitive anchor text, similar sensitive text, comparison safe text); using safe text as anchor text, combining safe text and sensitive text containing the same sensitive words to form a triple of (safe anchor text, similar safe text, comparison sensitive text).

4. The method according to claim 1, characterized in that: The deep learning architecture using BERT-BiLSTM combined with a multi-head attention mechanism for multi-layer semantic encoding includes: Initially encode the text through the BERT pre-trained model to obtain the basic word vector representation; The basic word vector representation is input into the BiLSTM network, and the forward and backward recurrent neural network structures are used to capture the long-range semantic dependencies in the text; A multi-head attention mechanism is introduced to deeply process the output of BiLSTM, identify key semantic information and assign higher weights; The outputs of multiple attention heads are concatenated and linearly transformed to generate a comprehensive text semantic feature vector representation.

5. The method according to claim 1, characterized in that The method of training the feature mapping model by using a triplet loss function includes: Map the text in the triple to the corresponding feature vector; For each type of triplet, the distance between the anchor text and the similar text is minimized through the triplet loss function, while the distance between the anchor text and the contrast text is maximized; The back-propagation algorithm is used to continuously optimize the model parameters, form a clear semantic distribution structure in the feature space, and train a high-quality feature mapping model.

6. The method according to claim 1, characterized in that The reimbursement text to be judged is constructed into a triple of "anchor text-sensitive sample-safe sample", including: Identify sensitive words in the reimbursement text to be judged through a sensitive word library; The reimbursement text to be identified is used as the anchor text, and sensitive samples and safety samples related to sensitive words are retrieved from the sample library to form a triplet of "anchor text-sensitive sample-safe sample".

7. The method according to claim 1, characterized in that The similarity is cosine similarity.

8. A system for accurately detecting sensitive content in reimbursement information based on multi-layer semantic coding and triple learning, characterized in that: include: The sensitive content text pool construction module is responsible for using historical reimbursement data to construct a training sample set containing typical misscreening situations, and establishes a "anchor text-similar text-comparison text" triple for each sensitive word to form a sensitive content text pool; The multi-layer semantic encoding module is responsible for using the deep learning architecture of BERT-BiLSTM combined with the multi-head attention mechanism to perform multi-layer semantic encoding and convert the "anchor text-similar text-comparison text" triple into a high-dimensional feature vector; The triple learning feature space mapping module is responsible for building semantic mapping relationships in the feature space based on the high-dimensional feature vectors of the "anchor text-similar text-comparison text" triples, and training the feature mapping model through the triple loss function; The new data sensitive content identification module is responsible for constructing the "anchor text-sensitive sample-safe sample" triplet for the reimbursement text to be judged, and using the trained feature mapping model to convert the text in the "anchor text-sensitive sample-safe sample" triplet into a high-dimensional feature vector. By calculating the similarity between the feature vectors of the anchor text and the sensitive sample and between the anchor text and the safe sample, it is determined whether the reimbursement text to be judged contains sensitive content.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Efficient sensitive data detection method

    CN121561540A

  • An efficient sensitive data detection method

    CN121561540B