Sensitive information identification method and device, equipment, storage medium and program product
By combining the technology of the large language model and the Milvus vector library, the regular expression and vector database are used for preliminary screening and similarity search, and context analysis is combined with the large language model, the false alarms, missed alarms and misjudgment problems of sensitive information recognition in the existing technology are solved, and the recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510840024.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Due to the limitations of the rules, existing sensitive information recognition methods are prone to false alarms, misreports, and misjudgment, resulting in low accuracy of the identification results.
The pre-trained large language model is used to combine search enhancement generation (RAG) technology and Milvus vector library, and preliminary screening is performed through regular expressions, similarity search is performed using vector database, context analysis is performed with large language models, and sensitive information recognition reports are generated.
It improves the accuracy of sensitive information recognition, avoids false alarms, misreports and misjudgments, and enhances the ability to identify sensitive information in natural language.
Smart Images

Figure CN120337938A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information recognition technology, and particularly to a method, device, equipment, storage medium and program product for sensitive information recognition. Background Art
[0002] In today's digital age, data security has become a key concern for enterprises and organizations. The importance of sensitive data recognition lies in protecting the security of sensitive information, especially during the process of handling sensitive information. By identifying sensitive data, enterprises can not only effectively control the accessibility, utilization rights, and protection measures of their sensitive information, but also enhance customer trust and satisfaction.
[0003] Currently, there are mainly three ways to identify sensitive information: the first is based on sensitive feature rules, but its flexibility is insufficient, which may affect the accuracy of the recognition result; the second is by calculating the similarity of sensitive data, but it is affected by the threshold setting and may cause false alarms or missed detections; the third is text entity recognition based on the Natural Language Processing (NLP) model, but the model performance is affected by the training data and is also prone to misjudgment.
[0004] Therefore, due to the limitations of the rules, the current methods for sensitive information recognition are prone to false alarms, missed detections, misjudgments, etc., resulting in a relatively low accuracy of the recognition result. Summary of the Invention
[0005] The main purpose of this application is to provide a method, device, equipment, storage medium and program product for sensitive information recognition, aiming to solve the technical problem that the current methods for sensitive information recognition are prone to false alarms, missed detections, misjudgments, etc. due to the limitations of the rules, resulting in a relatively low accuracy of the recognition result.
[0006] To achieve the above purpose, this application proposes a method for sensitive information recognition, and the method includes: Conduct a preliminary screening on the text to be detected to obtain sensitive marked text and unmarked text; Perform a similarity retrieval on the unmarked text based on a preset vector database to obtain multiple candidate sensitive data that are semantically similar to the unmarked text; Fuse the multiple candidate sensitive data and the text to be detected to obtain a fused text; Perform a context analysis on the fused text through a pre-trained large language model to obtain sensitive confidence data; Generate a sensitive information recognition report according to the sensitive marked text and the sensitive confidence data.
[0007] In one embodiment, the step of preliminarily screening the text to be detected to obtain sensitive marked text and unmarked text includes: Perform text cleaning on the text to be detected to obtain the cleaned text to be detected; Perform word segmentation on the cleaned text to be detected according to semantic features to obtain segmented text; Perform preliminary screening on the segmented text through a preset regular expression pattern library to obtain sensitive marked text, and the regular expression pattern library is constructed based on regular expressions of known sensitive information types.
[0008] In one embodiment, the step of performing similarity retrieval on the unmarked text based on a preset vector database to obtain multiple candidate sensitive data semantically similar to the unmarked text includes: Convert the unmarked text into a low-dimensional dense vector through a pre-trained large language model; Extract features from the unmarked text to obtain key feature information; Based on the low-dimensional dense vector and the key feature information, perform retrieval on the unmarked text in the preset vector database through a similarity algorithm to obtain multiple candidate sensitive data semantically similar to the unmarked text.
[0009] In one embodiment, the step of fusing the multiple candidate sensitive data and the text to be detected to obtain a fused text includes: Perform correlation ranking on the multiple candidate sensitive data to obtain ranked candidate data; Based on the ranked candidate data and the text to be detected, perform context construction through a pre-trained large language model to obtain a fused text.
[0010] In one embodiment, before the step of preliminarily screening the text to be detected to obtain sensitive marked text and unmarked text, it further includes: Perform sensitive information annotation on the collected data text set to obtain an annotated sensitive data set; Construct an initial model through a Transformer architecture; Divide the annotated sensitive data set into a training set, a validation set, and a test set; Train the initial model according to the training set, the validation set, and the test set to obtain a training result; When the training result reaches a preset stop condition, obtain a pre-trained large language model.
[0011] In one embodiment, after the step of generating a sensitive information recognition report according to the sensitive marked text and the sensitive confidence data, it further includes: Manually sample the sensitive information identification report to obtain misjudged texts and missed-judged texts; Analyze the misjudged texts and the missed-judged texts to determine error factors; When the error factor is a regular expression error, correct the regular expression pattern library; When the error factor is a large language model error, perform fine-tuning training on the pre-trained large language model.
[0012] In addition, to achieve the above object, the present application also proposes a sensitive information identification device, the device includes: A regular expression module for preliminarily screening the text to be detected to obtain sensitive marked texts and unmarked texts; A similarity retrieval module for performing similarity retrieval on the unmarked texts based on a preset vector database to obtain multiple candidate sensitive data semantically similar to the unmarked texts; A text fusion module for fusing the multiple candidate sensitive data and the text to be detected to obtain a fused text; A model analysis module for performing context analysis on the fused text through a pre-trained large language model to obtain sensitive confidence data; A report generation module for generating a sensitive information identification report according to the sensitive marked texts and the sensitive confidence data.
[0013] In addition, to achieve the above object, the present application also proposes a sensitive information identification device, the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program is configured to implement the steps of the sensitive information identification method as described above.
[0014] In addition, to achieve the above object, the present application also proposes a storage medium, the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the sensitive information identification method as described above.
[0015] In addition, to achieve the above object, the present application also provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the sensitive information identification method as described above.
[0016] One or more technical solutions proposed in this application have at least the following technical effects: The sensitive information recognition method of this application includes: preliminarily screening the text to be detected to obtain sensitive marked text and unmarked text; performing similarity retrieval on the unmarked text based on a preset vector database to obtain multiple candidate sensitive data semantically similar to the unmarked text; fusing the multiple candidate sensitive data and the text to be detected to obtain a fused text; performing context analysis on the fused text through a pre-trained large language model to obtain sensitive confidence data; generating a sensitive information recognition report according to the sensitive marked text and the sensitive confidence data.
[0017] Since this application performs retrieval by fusing a vector database, situations such as false positives, false negatives, and misjudgments are avoided; and it combines a large language model with advanced language understanding capabilities for analysis to obtain sensitive confidence data, which helps to accurately identify and process sensitive information in natural language. Through the layer-by-layer recognition of the vector database and the large language model, the accuracy of sensitive information recognition is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the sensitive information recognition method of this application; Figure 2 It is a flowchart of sensitive information recognition and extraction provided for Embodiment 1 of this application; Figure 3 It is a schematic flowchart provided for Embodiment 2 of the sensitive information recognition method of this application; Figure 4 It is a process diagram of generating a sensitive data set provided for Embodiment 2 of this application; Figure 5 It is a schematic diagram of realizing sensitive data recognition by the Llama model + RAG technology + Milvus vector library provided for Embodiment 2 of this application; Figure 6 It is a flowchart of training and application of a large language model provided for Embodiment 2 of this application; Figure 7 It is a schematic diagram of the module structure of the sensitive information recognition device for the embodiments of this application; Figure 8 This is a schematic diagram of the device structure of the hardware operating environment involved in the sensitive information recognition method in the embodiments of the present application.
[0021] The realization of the purpose of the present application, functional features and advantages will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments
[0022] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0023] In order to better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific embodiments.
[0024] It should be noted that currently, sensitive information recognition is roughly divided into three methods: The first is to recognize based on sensitive feature rules, and different feature rules are written according to different types of data. For example, the feature rule for the ID number is: area address code (the first 6 digits) + date of birth (the 7th - 14th digits) + sequence code (the 15th - 17th digits) + check code (the 18th digit). Another example is that the feature rule for the mobile phone number is: network identification number (the first 3 digits), area code (the 4th - 7th digits), user number (the 8th - 11th digits).
[0025] The second is to recognize based on text similarity, calculate the similarity between the target text and the sensitive data, and obtain a similarity value. If the obtained similarity value is within the preset threshold range, the target text is determined as sensitive data.
[0026] The third is the text entity recognition method based on natural language processing. Collect text big data, preprocess the text big data to obtain a text standard data set; according to the text standard data set, and based on the pre-trained language sub-model, recurrent neural network, graph neural network and attention mechanism of NLP technology, establish a text entity recognition model; obtain the text data to be recognized, and input the text data to be recognized into the text entity recognition model for text entity recognition to obtain the text entity recognition result.
[0027] However, the problems with the above methods are as follows: For the first method, there are certain rule limitations. The rules may not cover all types of sensitive information, especially newly emerged or uncommon sensitive information types. At the same time, there is insufficient flexibility. The rules may be too strict and difficult to adapt to different text styles and expressions, which may lead to a decrease in the accuracy of the recognition results.
[0028] In the second method, if the similarity threshold is set improperly, a large number of false positives or false negatives may occur; moreover, the effectiveness of similarity comparison depends on an effective feature extraction method. If the feature extraction is inaccurate or incomplete, the result of similarity comparison will also be affected.
[0029] In the third method, the performance of the NLP model depends to a great extent on the quality and diversity of the training data. If the training data is insufficient, biased, or the model fails to be updated in a timely manner, the model may have difficulty accurately identifying all types of sensitive information; and there are limitations in context understanding. Although progress has been made in context understanding, it may still be unable to fully understand complex language usage scenarios, resulting in misjudgments.
[0030] To solve the above problems, this application proposes a sensitive information recognition method that uses a pre-trained large language model + Retrieval-Augmented Generation (RAG) technology + Milvus (vector database management system) vector library, and combines regular expressions to achieve sensitive information recognition. The pre-trained large language model provides advanced language understanding capabilities, which helps to accurately identify and process sensitive information in natural language; the RAG technology realizes real-time and timely updates of knowledge through vector library retrieval without the need to retrain the large model. At the same time, combining regular expressions can effectively improve the accuracy of the system in recognizing sensitive information.
[0031] It should be noted that the execution subject of this embodiment can be a computing service device with functions such as text pre-screening, similarity retrieval, and context analysis, such as a personal computer, server, etc., or an electronic device capable of implementing the above functions, a sensitive information recognition device (abbreviated as recognition device) that executes the sensitive information recognition method of this application, etc. This embodiment does not limit this. Hereinafter, taking the recognition device as an example, this embodiment and the following embodiments will be described.
[0032] Based on this, Embodiment 1 of this application is proposed. Embodiment 1 of this application provides a sensitive information recognition method, referring to Figure 1 , Figure 1 which is the flowchart provided for Embodiment 1 of the sensitive information recognition method of this application.
[0033] In this embodiment, the sensitive information recognition method includes steps S10 to S50: Step S10: Perform a preliminary screening on the text to be detected to obtain sensitive marked text and unmarked text.
[0034] It should be noted that in the preliminary screening process, it can be implemented through a regular expression pattern library. The regular expression pattern library can be a resource library that centrally stores various regular expression patterns related to sensitive words. The regular expression pattern library can be constructed based on known types of sensitive information. After defining these sensitive information, corresponding regular expressions can be constructed according to the characteristics of this information. By setting a series of keywords related to sensitive information, the parts of the text that may contain sensitive information can be quickly located in the text.
[0035] Furthermore, the regular expression pattern library can be updated regularly, and manually or automatically expanded according to newly emerging forms of expression of sensitive information to optimize the regular expressions.
[0036] It can be understood that the text to be detected can be a text waiting to be identified for sensitive information. There may be sensitive information in the text, and sensitive information identification is required.
[0037] It should be understood that the sensitive marked text can be the text that has been initially identified through the regular expression pattern library and marked with sensitive content. The unmarked text can be the text in which no sensitive information has been detected through the regular expression pattern library.
[0038] In a specific implementation, after receiving the text to be detected, the identification device can preliminarily screen the text to be detected through the regular expression pattern library. If there are sensitive words related to various regular expressions in the text to be detected, the sensitive words are marked to obtain the sensitive marked text. The remaining text to be detected then enters the subsequent sensitive information identification as the unmarked text.
[0039] In a feasible implementation manner, step S10 of this embodiment may include the steps of: performing text cleaning on the text to be detected to obtain the cleaned text to be detected; performing word segmentation on the cleaned text to be detected according to semantic features to obtain the segmented text; and performing preliminary screening on the segmented text through a preset regular expression pattern library to obtain the sensitive marked text, where the regular expression pattern library is constructed based on regular expressions of known types of sensitive information.
[0040] Exemplarily, text cleaning may include: removing the HyperText Markup Language (HTML) tags, special symbols (such as redundant punctuation, garbled characters, etc.) from the detected text, uniformly converting full-width characters to half-width, identifying and removing invisible characters, control characters, etc. The obtained cleaned text to be detected can avoid interfering with subsequent analysis, ensure the text format is standardized, and facilitate subsequent processing.
[0041] It should be noted that the semantic features can be different semantic information contained in each word in different languages. For Chinese semantic features, the Chinese text can be segmented according to vocabulary and semantic units to obtain segmented text; for English semantic features, the word can be segmented as a basic unit to obtain segmented text; for other languages, word segmentation can also be performed according to different semantic features, which is not limited in this embodiment.
[0042] For example, the Chinese sentence “I love my hometown” can be segmented into “I”, “love”, “my” and “hometown”; the English sentence “I love my hometown” can be segmented into “I”, “love”, “my” and “hometown”.
[0043] The segmented text can be stored in a list format, making it easier for subsequent modules to process word by word or unit by unit.
[0044] In this embodiment, after receiving the text to be detected, the recognition device first performs a text cleaning step on the input text to be detected to obtain the cleaned text to be detected with a standardized text format. Then, word segmentation is performed, and the Chinese text is segmented according to vocabulary and semantic units; for English text, the words are segmented as basic units to obtain the segmented text. Finally, the segmented text is traversed, and each word or text unit is matched one by one with the regular expression pattern library. Once a match is found, the text fragment is immediately marked and its sensitive type is recorded. For the text fragment that is successfully matched, a sensitive marked text containing metadata such as its location information and sensitive type is generated for preliminary warning. The sensitive marked text enters the subsequent sensitive information identification. Therefore, through regular expressions, a variety of sensitive information patterns can be flexibly defined and identified.
[0045] Step S20: performing similarity search on the unlabeled text based on a preset vector database to obtain a plurality of candidate sensitive data having semantic similarity to the unlabeled text.
[0046] It should be noted that the vector database (such as the Milvus vector database) can be a database used for sensitive information similarity retrieval, in which a massive set of sensitive information knowledge graph vectors is pre-stored.
[0047] Among them, the knowledge graph vector set may include historical sensitive information texts, relevant legal provisions, industry norms and other knowledge texts saved in vector form, covering a wide range of sensitive information semantic scenarios.
[0048] It is understandable that the candidate sensitive data can be sensitive content data with semantics similar to the unlabeled text retrieved from the vector database. If they are semantically similar, the corresponding vector distance in the vector space may be relatively close, and the two can be determined as candidate sensitive data with semantics similar.
[0049] In a specific implementation, after the initial screening in step S10, the recognition device can retrieve the unlabeled text in the massive sensitive information knowledge graph vector set of the vector database to obtain multiple candidate sensitive data that is semantically similar to the unlabeled text.
[0050] In a feasible implementation manner, step S20 of this embodiment may include the steps of: converting the unlabeled text into a low-dimensional dense vector through a pre-trained large language model; extracting feature information of the unlabeled text to obtain key feature information; and based on the low-dimensional dense vector and the key feature information, retrieving the unlabeled text in a preset vector database through a similarity algorithm to obtain multiple candidate sensitive data that is semantically similar to the unlabeled text.
[0051] It should be noted that the large language model (Large Language Model Meta AI, Llama) can be a natural language processing model that can perform various tasks such as text generation, question answering, and machine translation on text. The low-dimensional dense vector can be a vector representation obtained by converting the unlabeled text into a low-dimensional dense vector, which helps to reduce the computational complexity and improve the processing efficiency, and can better capture the key features of sensitive information. The key feature information can be the feature information representing the key content of the unlabeled text.
[0052] Exemplarily, for the convenience of understanding the above implementation process, an example is given below. For example, after being processed by a pre-trained large language model, a text "The weather is nice today" can obtain a vector with 768 dimensions, and this vector contains the semantic features of the text. The keyword "weather" in the above text can be extracted as the key feature information.
[0053] It can be understood that the similarity algorithm can be a method for measuring the similarity between the unlabeled text and the data in the vector database, such as the Euclidean distance algorithm, the cosine similarity algorithm, etc. This embodiment does not limit this.
[0054] In this embodiment, after obtaining the above unlabeled text, the recognition device can use a pre-trained large language model (such as the embedding layer of the Llama model) to convert the unlabeled text into a low-dimensional dense vector representation. At the same time, key feature information such as keywords and named entities in the unlabeled text is extracted as auxiliary retrieval information. Then, the generated low-dimensional dense vector and the key feature information are passed into the above vector database. The vector database performs a fast retrieval in the pre-stored massive sensitive information knowledge graph vector set according to the vector similarity algorithm. The retrieval result returns several candidate sensitive information text fragments that are semantically similar to the input unlabeled text and their corresponding vectors (i.e., multiple candidate sensitive data). Thus, the similarity recognition of sensitive information can be improved through the vector database.
[0055] Step S30: Integrate the multiple candidate sensitive data with the text to be detected to obtain an integrated text.
[0056] It should be noted that the integrated text can be a text generated by contextually associating and integrating candidate sensitive data with the text to be detected. This process is a Retrieval-Augmented Generation (RAG) process, which can improve the accuracy and quality of the generated content.
[0057] Step S40: Perform context analysis on the multiple candidate sensitive data and the text to be detected through a pre-trained large language model to obtain sensitive confidence data.
[0058] It should be noted that the sensitive confidence data can be a judgment result of sensitive information indicating the confidence level that the text belongs to sensitive information in the form of probability output through the above-mentioned pre-trained large language model (such as Llama model) for context recognition.
[0059] In a specific implementation, after obtaining the above multiple candidate sensitive data, the recognition device can integrate the multiple candidate sensitive data with the text to be detected to obtain an integrated text; then perform context analysis on the integrated text through a pre-trained large language model to obtain sensitive confidence data indicating the confidence level that the text belongs to sensitive information in the form of probability.
[0060] In a feasible implementation manner, step S30 of this embodiment may include the steps of: sorting the multiple candidate sensitive data according to relevance to obtain the sorted candidate data; based on the sorted candidate data and the text to be detected, construct context through a pre-trained large language model to obtain an integrated text.
[0061] It can be understood that the relevance sorting can determine the degree of tight association between multiple candidate sensitive data and sensitive features according to specific criteria. For example, if the privacy leakage risk is used as the criterion, data directly related to personal core privacy can be regarded as highly relevant sensitive data and ranked at the top; while some data that is relatively easy to obtain and has a slightly lower privacy risk has a lower relevance and is ranked at the back.
[0062] The pre-trained large language model, through a multi-layer neural network structure, weighs factors such as semantic logic, sentiment tendency, and potential intention in the text, and can output 0.8, indicating that the text has an 80% possibility of containing sensitive information, along with detailed sensitive types, sensitive content fragments, etc., to determine the sensitivity level.
[0063] In this embodiment, the recognition device can sort and organize multiple candidate sensitive data retrieved from the vector database according to relevance to obtain the sorted candidate data. Then, in combination with the original input text to be detected, the generation ability of the pre-trained large language model is utilized to construct a new context. Through the above RAG process, a fused text can be obtained. Finally, the pre-trained large language model receives the constructed fused text and, based on its powerful language understanding ability, deeply analyzes the sensitive tendency of the overall text, outputs the final judgment result of sensitive information, represents the confidence level that the text belongs to sensitive information in the form of probability, and obtains sensitive confidence data. Thus, by combining the RAG technology with the pre-trained large language model, the accuracy of sensitive information recognition is greatly improved.
[0064] Step S50: Generate a sensitive information recognition report according to the sensitive marked text and the sensitive confidence data.
[0065] In a specific implementation, finally, the recognition device aggregates and integrates the sensitive marked text in the initial screening stage of the above regular expression pattern library and the above sensitive confidence data to form a complete sensitive information recognition report.
[0066] Further, after step S50 of this embodiment, the following steps may also be included: manually sampling the sensitive information recognition report to obtain misjudged texts and missed judged texts; analyzing the misjudged texts and the missed judged texts to determine the error factors; when the error factor is a regular expression error, correcting the regular expression pattern library; when the error factor is a large language model error, performing fine-tuning training on the pre-trained large language model.
[0067] It should be noted that due to certain errors and limitations of the large model, especially in the initial stage of insufficient training, the recognition results of the large model can be manually reviewed and corrected. Adopting this advanced mode with the large model's automatic data recognition as the main and manual review as the supplement helps to accelerate the data annotation work process and significantly reduce the cost of sensitive information recognition. And the differential data found during the manual review process can be used to optimize the training of the large model, making the recognition ability of the large model more and more accurate and having the characteristic of continuous optimization.
[0068] In this embodiment, after obtaining the above-mentioned sensitive information recognition report, it can also be randomly inspected by an artificial review team. If misjudged texts are found (such as normal text being recognized as sensitive information) or missed judgments (such as implicit sensitive expressions not being recognized), they are promptly recorded and fed back to the above-mentioned model training and optimization process. For misjudged texts, analyze the reasons. If it is a misjudgment by the regular expression, it may be that the regular expression is too broad, then adjust the regular expression pattern accordingly; if it is a misjudgment by the model, it may be that the pre-trained large language model has an incorrect understanding of the special context, and the pre-trained large language model can be fine-tuned. Thus, through the above feedback, the model and regular expression are continuously improved to improve the accuracy of sensitive information recognition.
[0069] Exemplarily, to help understand the implementation process of the sensitive information recognition method in this embodiment, please refer to Figure 2 , Figure 2 which is the flowchart of sensitive information recognition and extraction provided in Embodiment 1 of this application. Specifically: 1. Text cleaning and attachment parsing. The recognition device first parses the input text to be detected, and passes the parsed text through the text cleaning step to ensure that the text format is standardized for subsequent processing.
[0070] 2. Word segmentation processing. Use a professional word segmentation tool to segment Chinese text according to words and semantic units; for English text, segment it with words as the basic unit. The segmented text can be stored in a list form for subsequent modules to process word by word or unit by unit.
[0071] 3. Initial screening of the regular expression pattern library. Traverse the preprocessed text list, match each word or text unit with the regular expression pattern library one by one to determine whether it contains sensitive data. Once a matching item is found, immediately mark the text segment, generate sensitive marked text, and issue a preliminary warning.
[0072] 4. RAG+Milvus knowledge enhanced retrieval. For the unmarked text that has not been marked by the regular expression, use the pre-trained language model to convert the unmarked text into a low-dimensional dense vector, and at the same time extract key feature information as auxiliary retrieval information. Pass the generated vector and key feature information into the vector database (such as the Milvus vector database). And use the generation ability of the pre-trained large language model (hereinafter illustrated by the Llama model) to construct a new context and generate a fused text.
[0073] 5. Deep Judgment of Llama Model. The pre-trained large language model receives the constructed integrated text. Based on its powerful language understanding ability, it deeply analyzes the overall sensitive tendency of the text to determine whether it contains sensitive information. If it does not contain sensitive information, the constructed integrated text is discarded, and the deep judgment process of the Llama model ends; if it contains sensitive information, the confidence level that the text belongs to sensitive information is expressed in the form of probability, and the final sensitive information judgment result is output for model warning. Finally, it is aggregated and integrated with the results of the regular expression preliminary screening stage to generate a complete sensitive information identification report.
[0074] 6. Manual Review and Correction. Finally, the manual review team can conduct spot checks on the sensitive information identification report output by the pre-trained large language model to determine whether correction is needed. If no correction is needed, the sensitive information identification report can be directly output. If misjudgment or missed judgment is found, it is recorded in a timely manner and feedback is sent to the model training and optimization link. For missed judgment situations, the missed judgment text content can be marked, including the complete expression of the text, the context in which it appears, and the output batch information, etc.; then the reason for the missed judgment is analyzed, for example, it is judged whether the loophole is caused by the fact that the sensitive word library does not contain specific words or phrases in the text or whether there is a deviation in the model's understanding of certain semantics; if the sensitive word library is incomplete, the sensitive word library is updated in a timely manner to supplement the relevant sensitive words or phrases in the missed judgment text; if it is a problem with the model's own semantic understanding, more similar missed judgment text examples are collected to strengthen the training of the model on the missed judgment text. For misjudged text, the reason is analyzed. If it is an expression problem, the regular expression is adjusted; if it is a model problem, the pre-trained large language model is fine-tuned.
[0075] In the technical solution provided in this embodiment, after receiving the text to be detected, the recognition device can perform a preliminary screening of the text to be detected through the regular expression pattern library. If there are sensitive words related to various regular expressions in the text to be detected, the sensitive words are marked to obtain sensitive marked text. The remaining text to be detected enters the subsequent sensitive information recognition as unmarked text. Then, the vector database can be used to search the massive sensitive information knowledge graph vector set of the unmarked text in the vector database to obtain multiple candidate sensitive data with similar semantics to the unmarked text. Then, the multiple candidate sensitive data and the text to be detected are fused to obtain a fused text; then, the fused text is contextually analyzed through a pre-trained large language model to obtain sensitive confidence data that represents the confidence that the text belongs to sensitive information in a probabilistic form. Finally, the recognition device summarizes and integrates the sensitive marked text in the initial screening stage of the regular expression pattern library with the above-mentioned sensitive confidence data to form a complete sensitive information recognition report. Since this embodiment integrates the vector database for retrieval, it avoids false positives, missed negatives, and misjudgments; and combined with a large language model with advanced language understanding capabilities for analysis, sensitive confidence data is obtained, which helps to accurately identify and process sensitive information in natural language. Through the layered recognition of the vector database and the large language model, the accuracy of sensitive information recognition is improved.
[0076] Based on the above-mentioned embodiment 1 of the present application, the embodiment 2 of the present application is proposed. In the second embodiment of the present application, the same or similar contents as those of the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated later. On this basis, please refer to Figure 3 , Figure 3 A flowchart diagram is provided for Example 2 of the sensitive information identification method of this application.
[0077] Before step S10 of this example, the sensitive information identification method further includes steps S01 to S05: Step S01: label the collected data text set with sensitive information to obtain a labeled sensitive data set.
[0078] Exemplary, reference Figure 4 , Figure 4 A process diagram for generating a sensitive data set provided in Example 2 of the present application. First, a large amount of text information can be collected. This data can come from various sources, such as social media, news websites, forums, chat records, etc. After collecting the data, data preprocessing is performed, which may include deleting duplicate data, repairing erroneous data, deleting irrelevant data, etc. Then all sensitive information is found in the text and the data is annotated. Human participation is required in the initial stages of this process; a labeling team can be established to do this work, and over time, machine learning algorithms can be used to automate part or all of the labeling process. Finally, a data review is performed to check the accuracy of the annotations and fix any errors.
[0079] Step S02: Construct an initial model through the Transformer architecture.
[0080] It should be noted that the Transformer architecture can be a deep learning architecture for building large language models, which is mainly constructed based on the attention mechanism.
[0081] Specifically, the process of model construction and initialization can include: First, preprocess the above-mentioned labeled sensitive dataset again, such as data cleaning, deduplication, denoising, and data standardization. Remove unnecessary data, repair missing values and errors in the dataset, handle abnormal data and noise, convert the data into a unified format and unit, ensure data quality, avoid interfering with the model, and improve the efficiency of model training. During the data collation process, in order to enable the model to better learn and understand the data, labels and annotations can be added to the data. This can be done through manual annotation or by using automatic annotation techniques to automatically add labels to the data through machine learning algorithms. To facilitate model training and evaluation, the dataset can be further divided into three datasets: a training set, a validation set, and a test set. Cross-validation is used to evaluate the performance of the model, and stratified sampling is used to ensure that the data of each category is representative in the three test sets, avoiding data bias.
[0082] Then comes the construction of the initial model. The large language model can be constructed using variants of the Transformer architecture, such as GPT-3 or similar structures. The process of constructing the model mainly includes the following steps: Model design: Select a suitable model structure according to the task requirements, including components such as an embedding layer, an encoding layer, and a decoding layer. Parameter initialization: Randomly initialize the model parameters (such as weights, biases, etc.) or use pre-trained weights for initialization. Model implementation: Write code using a deep learning framework (such as PyTorch, TensorFlow, etc.) to implement the model structure and define the forward propagation and backward propagation processes.
[0083] Step S03: Divide the labeled sensitive dataset into a training set, a validation set, and a test set; Step S04: Train the initial model according to the training set, the validation set, and the test set to obtain a training result; Step S05: When the training result reaches a preset stop condition, obtain a pre-trained large language model.
[0084] It should be noted that the preset stop condition can be the stop condition for training the model, such as reaching the maximum number of iterations, the validation loss no longer decreasing, etc. This embodiment does not limit this.
[0085] Specifically, before model training, relevant training parameters and hyperparameters can be configured to control the training process. For example, learning rate, batch size, optimizer, number of iterations, hardware resources, etc. After completing data preparation and model construction, model training begins. First, load the data, load the training data into memory, and divide it according to the batch size. Then perform forward propagation, input the batch data into the model, and calculate the model output and loss function value. Next, perform backward propagation, calculate the gradient according to the loss function value, and use the optimizer to update the model parameters. Finally, after each iteration cycle, use the validation set to evaluate the model and save the best model weights. At the same time, adjust hyperparameters such as the learning rate according to the validation results, and continue to train the model until the stop condition is met.
[0086] After the model training is completed, the test set can be used to evaluate the model, and the model can be optimized according to the evaluation results. Among them, appropriate evaluation metrics (such as accuracy, F1 value, perplexity, etc.) can be selected to evaluate the model. Include error analysis, feature importance analysis, etc. to find out the problems and improvement directions of the model. Optimize the model according to the evaluation results and analysis results, including adjusting the model structure, adding regularization terms, adopting more advanced training strategies, etc. Finally, deploy the optimized model to the actual application scenario, and conduct continuous monitoring and maintenance to ensure the stability and reliability of the model performance.
[0087] Among them, regarding the above description of the amount of training data, the pre-trained large language model can be trained using up to 1.5 trillion Tokens of data in the pre-training stage. In the fine-tuning stage, although such a large amount of data is not required, it is still important to select an appropriate amount of data. The fine-tuning dataset should be representative enough of the target task so that the model can learn relevant features. The specific amount of data depends on the complexity of the task and the capacity of the model, but generally speaking, a few thousand to tens of thousands of samples may be sufficient for effective fine-tuning.
[0088] Furthermore, for ease of understanding, the following uses the Llama model and the Milvus vector database as examples to illustrate the sensitive information identification process of the pre-trained large language model, but does not limit this solution. In the sensitive information identification process of the large language model, it can also be implemented by combining the RAG technology + Milvus vector database, refer to Figure 5 , Figure 5Schematic diagram of implementing sensitive data recognition using the Llama model + RAG technology + Milvus vector library provided in the second embodiment of this application. After document parsing, the document is segmented to obtain tokenized text. Then the text is embedded into the vector library for retrieval to obtain corresponding text fragments, which will undergo format conversion and be restored to natural language text. These texts are sorted by relevance and, in the form of additional information, follow the original text to be recognized seamlessly and are fed into the input port of the Llama model, integrated into the model's understanding process, and finally the recognition result is output.
[0089] Exemplarily, to help understand the implementation process of this embodiment in combination with the first embodiment above, please refer to Figure 6 , Figure 6 Flowchart of large language model training and application provided in the second embodiment of this application. Specifically: Step A1: First, collect a large number of data text sets, perform preprocessing operations such as data cleaning and data segmentation on the text, clarify the classification and level of sensitive information, and obtain a labeled sensitive data set.
[0090] Step A2: Then use the labeled sensitive data set to train the Llama model, and adjust the parameters and structure of the model to obtain a pre-trained large language model.
[0091] Step A3: Then comes the use of the model. Based on the machine learning of the Llama model and the Milvus search enhancement of the RAG technology, various types of regular and simple algorithms are integrated at the same time to accurately identify sensitive information.
[0092] In the sensitive information recognition process, the Llama model acts as the core semantic understanding engine. To adapt to this task, when the text to be recognized enters the system, first use the retrieval of RAG to quickly find relevant historical sensitive information cases or auxiliary knowledge texts in the external knowledge base (integrated with the Milvus vector library) based on the key semantic features of the text. The retrieved text will be concatenated with the original input text and fed into the Llama model together as a new input. This approach enables the Llama model to not only rely on its own large model capabilities but also refer to the knowledge of past similar cases to enhance the judgment accuracy.
[0093] When RAG interacts with the Milvus vector database, multi-layer filtering logic is added. In the initial rough screening stage, based on the vector distance threshold, vectors that are too far away and have extremely low relevance are quickly excluded; in the subsequent fine screening process, using the sensitive information category labels, vectors that are more precisely matched under the same category are secondarily screened, reducing the interference of irrelevant information and accelerating the retrieval process. For the integration details of the Milvus vector database, in the Milvus vector database, the embedded vectors are organized according to the classification system of sensitive information. For example, vectors corresponding to privacy-sensitive information are grouped into one group, and politically sensitive information is grouped into another group, and each group is further subdivided into multiple subcategories. When each vector is inserted, it is attached with the detailed annotation of the corresponding sensitive information category for convenient subsequent precise retrieval.
[0094] When the Milvus vector database completes the retrieval, the text fragments corresponding to the output embedded vectors will undergo format conversion to restore them to natural language text. These texts are sorted by relevance and, in the form of additional information, follow the original text to be recognized seamlessly and are fed into the input port of the Llama model to be integrated into the model's understanding process. At the same time, to improve the retrieval performance of the Milvus vector database, vector quantization technology can also be used to compress the vector storage space, reducing memory overhead and accelerating the retrieval speed. And the parameters can be adjusted through the Hierarchical Navigable Small World (HNSW, an algorithm for performing similarity search tasks in large-scale datasets that can efficiently find data objects similar to the query object) algorithm, appropriately increasing the number of nearest neighbor connections and reducing the graph construction cost to ensure that in a large-scale sensitive information vector dataset, highly relevant vectors can be quickly located.
[0095] Summarize the above three recognition results and use the characteristics of data structures (such as sets or hash tables) or database queries (such as DISTINCT) to remove duplicates in the merged list. Define the sorting rules according to actual needs, which may include sorting by the type of sensitive information recognized, the frequency of occurrence, the position in the text, etc., and output the sorted recognition results in an appropriate format (such as text files, CSV files, database records, etc.). According to requirements, it can also include detailed information about the recognized sensitive information, such as type, position, context, etc.
[0096] Step A4: Finally, establish an artificial review mechanism to review important decisions. At the same time, regularly collect new sensitive information samples, including sensitive expressions in emerging Internet terms, sensitive topics derived from current affairs hotspots, etc., to expand the knowledge graph of the vector database and update the regular expression pattern library. Use the new samples to perform incremental training on the Llama model to make it continuously adapt to the dynamic changes of sensitive information and improve the subsequent recognition accuracy. At the same time, deploy the optimized model, rules, etc. back to the sensitive information recognition system to achieve closed-loop optimization.
[0097] In this embodiment, the pre-trained large language model after training provides advanced language understanding capabilities, which helps to accurately identify and process sensitive information in natural language; the RAG technology realizes real-time and timely update of knowledge through vector library retrieval without retraining the large model. At the same time, the integration of regular expressions can effectively improve the accuracy of the system in identifying sensitive information.
[0098] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the sensitive information recognition method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0099] This application also provides a sensitive information recognition device. Please refer to Figure 7 , Figure 7 which is the module structure diagram of the sensitive information recognition device in the embodiment of this application; the sensitive information recognition device includes: A regular expression module 701, configured to perform preliminary screening on the text to be detected to obtain sensitive marked text and unmarked text; A similarity retrieval module 702, configured to perform similarity retrieval on the unmarked text based on a preset vector database to obtain multiple candidate sensitive data semantically similar to the unmarked text; A text fusion module 703, configured to fuse the multiple candidate sensitive data and the text to be detected to obtain a fused text; A model analysis module 704, configured to perform context analysis on the fused text through a pre-trained large language model to obtain sensitive confidence data; A report generation module 705, configured to generate a sensitive information recognition report according to the sensitive marked text and the sensitive confidence data.
[0100] As an implementation manner, the regular expression module 701 is further configured to perform text cleaning on the text to be detected to obtain the cleaned text to be detected; perform word segmentation processing on the cleaned text to be detected according to semantic features to obtain segmented text; perform preliminary screening on the segmented text through a preset regular expression pattern library to obtain sensitive marked text, and the regular expression pattern library is constructed based on regular expressions of known sensitive information types.
[0101] As an implementation manner, the similarity retrieval module 702 is further configured to convert the unmarked text into a low-dimensional dense vector through a pre-trained large language model; perform feature extraction on the unmarked text to obtain key feature information; perform retrieval on the unmarked text in the preset vector database through a similarity algorithm based on the low-dimensional dense vector and the key feature information to obtain multiple candidate sensitive data semantically similar to the unmarked text.
[0102] As an implementation, the text fusion module 703 is further configured to perform correlation sorting on the multiple candidate sensitive data to obtain the sorted candidate data; and based on the sorted candidate data and the text to be detected, construct a context through a pre-trained large language model to obtain a fused text.
[0103] As an implementation, the model analysis module 704 is further configured to perform sensitive information annotation on the collected data text set to obtain an annotated sensitive data set; construct an initial model through the Transformer architecture; divide the annotated sensitive data set into a training set, a validation set, and a test set; train the initial model according to the training set, the validation set, and the test set to obtain a training result; and when the training result reaches a preset stop condition, obtain a pre-trained large language model.
[0104] As an implementation, the model analysis module 704 is further configured to perform manual sampling inspection on the sensitive information recognition report to obtain misjudged texts and missed judged texts; analyze the misjudged texts and the missed judged texts to determine error factors; when the error factor is a regular expression error, correct the regular expression pattern library; and when the error factor is a large language model error, perform fine-tuning training on the pre-trained large language model.
[0105] Other embodiments or specific implementation manners of the sensitive information recognition device of the present application may refer to the above method embodiments, and will not be elaborated here.
[0106] The sensitive information recognition device provided by the present application adopts the sensitive information recognition method in the above embodiments, and can solve the technical problem that the current sensitive information recognition method is prone to false alarms, missed reports, misjudgments, etc. due to rule limitations, resulting in low accuracy of the recognition result. Compared with the prior art, the beneficial effects of the sensitive information recognition device provided by the present application are the same as those of the sensitive information recognition method provided by the above embodiments, and other technical features in the sensitive information recognition device are the same as those disclosed in the above embodiment method, and will not be elaborated here.
[0107] The present application provides a sensitive information recognition device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the sensitive information recognition method in the first embodiment above.
[0108] Next, refer to Figure 8 , Figure 8The figure is a schematic diagram of the device structure of the hardware operating environment involved in the sensitive information recognition method in the embodiments of the present application, which shows the schematic diagram of the structure of the sensitive information recognition device suitable for implementing the embodiments of the present application. The sensitive information recognition device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The shown sensitive information recognition device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0109] As Figure 8 shown, the sensitive information recognition device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the sensitive information recognition device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the sensitive information recognition device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a sensitive information recognition device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.
[0110] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0111] The sensitive information recognition device provided by the present application adopts the sensitive information recognition method in the above-mentioned embodiment, and can solve the technical problem that the current sensitive information recognition method is prone to false alarms, missed alarms, misjudgments, etc. due to the limitations of rules, resulting in low accuracy of recognition results. Compared with the prior art, the beneficial effects of the sensitive information recognition device provided by the present application are the same as those of the sensitive information recognition method provided by the above-mentioned embodiment, and other technical features in the sensitive information recognition device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0112] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0113] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0114] The present application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the sensitive information recognition method in the above-mentioned embodiment.
[0115] The computer-readable storage medium provided by this application can, for example, be a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0116] The above computer-readable storage medium can be included in a sensitive information identification device; or it can exist separately without being assembled into a sensitive information identification device.
[0117] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by a sensitive information identification device, the sensitive information identification device is caused to: preliminarily screen the text to be detected to obtain sensitive marked text and unmarked text; perform a similarity search on the unmarked text based on a preset vector database to obtain multiple candidate sensitive data that are semantically similar to the unmarked text; fuse the multiple candidate sensitive data and the text to be detected to obtain a fused text; perform context analysis on the fused text through a pre-trained large language model to obtain sensitive confidence data; and generate a sensitive information identification report based on the sensitive marked text and the sensitive confidence data.
[0118] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by connecting through an Internet service provider via the Internet).
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0120] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0121] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned sensitive information recognition method, and can solve the technical problems that the current methods of sensitive information recognition are prone to false positives, false negatives, misjudgments, etc. due to the limitations of the rules, resulting in low accuracy of the recognition results. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the sensitive information recognition method provided in the above embodiments, and will not be elaborated here.
[0122] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the sensitive information recognition method as described above.
[0123] The computer program product provided by the present application can solve the technical problem that the current method of sensitive information recognition is prone to false alarms, missed alarms, misjudgments, etc. due to the limitations of rules, resulting in a low accuracy of the recognition result. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the sensitive information recognition method provided by the above embodiments, and will not be elaborated herein.
[0124] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A sensitive information recognition method, characterized in that, The method includes: Performing preliminary screening on the text to be detected to obtain sensitive marked text and unmarked text; Performing similarity retrieval on the unmarked text based on a preset vector database to obtain multiple candidate sensitive data that are semantically similar to the unmarked text; Fusing the multiple candidate sensitive data and the text to be detected to obtain a fused text; Performing context analysis on the fused text through a pre-trained large language model to obtain sensitive confidence data; Generating a sensitive information recognition report according to the sensitive marked text and the sensitive confidence data.
2. The method according to claim 1, wherein The step of performing preliminary screening on the text to be detected to obtain sensitive marked text and unmarked text includes: Performing text cleaning on the text to be detected to obtain the cleaned text to be detected; Performing word segmentation on the cleaned text to be detected according to semantic features to obtain a segmented text; Performing preliminary screening on the segmented text through a preset regular expression pattern library to obtain sensitive marked text, and the regular expression pattern library is constructed based on regular expressions of known sensitive information types.
3. The method according to claim 1, wherein The step of performing similarity retrieval on the unmarked text based on a preset vector database to obtain multiple candidate sensitive data that are semantically similar to the unmarked text includes: Converting the unmarked text into a low-dimensional dense vector through a pre-trained large language model; Performing feature extraction on the unmarked text to obtain key feature information; Based on the low-dimensional dense vector and the key feature information, performing retrieval on the unmarked text in the preset vector database through a similarity algorithm to obtain multiple candidate sensitive data that are semantically similar to the unmarked text.
4. The method according to claim 1, wherein The step of fusing the multiple candidate sensitive data and the text to be detected to obtain a fused text includes: Performing correlation sorting on the multiple candidate sensitive data to obtain sorted candidate data; Based on the sorted candidate data and the text to be detected, performing context construction through a pre-trained large language model to obtain a fused text.
5. The method according to any one of claims 1 to 4, characterized in that, Before the step of performing preliminary screening on the text to be detected to obtain sensitive marked text and unmarked text, it further includes: Performing sensitive information annotation on the collected data text set to obtain an annotated sensitive data set; Constructing an initial model through a Transformer architecture; Dividing the annotated sensitive data set into a training set, a validation set, and a test set; Training the initial model according to the training set, the validation set, and the test set to obtain a training result; When the training result reaches a preset stop condition, obtaining a pre-trained large language model.
6. The method according to claim 2, wherein After the step of generating a sensitive information recognition report according to the sensitive marked text and the sensitive confidence data, it further includes: Performing manual sampling inspection on the sensitive information recognition report to obtain misjudged text and missed judged text; Analyzing the misjudged text and the missed judged text to determine error factors; When the error factor is a regular expression error, correcting the regular expression pattern library; When the error factor is a large language model error, performing fine-tuning training on the pre-trained large language model.
7. A sensitive information recognition device, characterized in that The device includes: A regular expression module for preliminarily screening the text to be detected to obtain sensitive marked text and unmarked text; A similarity retrieval module for performing similarity retrieval on the unmarked text based on a preset vector database to obtain multiple candidate sensitive data semantically similar to the unmarked text; A text fusion module for fusing the multiple candidate sensitive data and the text to be detected to obtain a fused text; A model analysis module for performing context analysis on the fused text through a pre-trained large language model to obtain sensitive confidence data; A report generation module for generating a sensitive information identification report according to the sensitive marked text and the sensitive confidence data.
8. A sensitive information recognition device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the sensitive information identification method according to any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the sensitive information identification method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the sensitive information identification method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text data sensitivity-related detection method and device, equipment and medium
CN116681083A
Sensitive information identification method, system and device, storage medium and program product
CN119513231A
Self-adaptive sensitive information intelligent identification method and device, equipment, storage medium and product
CN119599130A
Sensitive text classification method and apparatus, computer device, and storage medium
WO2025124024A1
Cited By
Voice interaction method, server and medium
CN120913560A
Government affair service-oriented master-slave multi-agent collaboration method and system
CN121092699A
Multi-expert sensitive information identification method and system based on large model, product and medium
CN121210661A
A method, system, product, and medium for identifying sensitive information based on a large model and multiple experts.
CN121210661B