Sensitive information processing method and system based on pre-training model and hybrid model architecture

Through sensitive information processing methods based on pre-trained models and hybrid model architectures, the problems of high data labeling cost, poor model adaptability and excessive forgetting in the prior art are solved, and efficient identification and precise forgetting of sensitive information in unstructured text are achieved.

CN120162835APending Publication Date: 2025-06-17ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510238647.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When detecting and processing sensitive information in unstructured texts, the prior art has problems such as high data labeling cost, poor adaptability of the model to different texts, excessive forgetting and difficulty in distinguishing sensitive information from non-sensitive information.

Method used

Using sensitive information processing methods based on pretrained models and hybrid model architectures, an unstructured sensitive information text recognition model is constructed, preprocessed and divided through data sets, and sequence annotation is used using CRF decoder, combining MemFlex technology to achieve accurate identification and forgetting of sensitive entities.

Benefits of technology

It significantly reduces the cost of manual labeling, improves the model's adaptability to different unstructured texts, avoids excessive forgetting, and ensures the effectiveness of data security and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162835A_ABST
    Figure CN120162835A_ABST
Patent Text Reader

Abstract

The invention discloses a sensitive information processing method and system based on a pre-training model and hybrid model architecture, and relates to the technical field of data security and privacy protection. Comprising the following steps: S1, constructing a data set; s2, preprocessing the data set; s3, data division; s4, constructing and training a model; s5, positioning sensitive information; and S6, forgetting sensitive information. In terms of model performance, the unstructured sensitive information text recognition model adopts vocabulary-level and character-level marking processing and feature enhancement, so that the recognition capability of sensitive information is remarkably enhanced, and meanwhile, the adaptability of the model to different unstructured texts is improved; in a data security and privacy protection level, a sensitive entity erasure area in a text is analyzed and determined based on a gradient information key area, accurate forgetting of sensitive information is realized, excessive forgetting is avoided, and data security compliance is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security and privacy protection, and particularly to a sensitive information processing method and system based on a pre-trained model and a hybrid model architecture. Background Art

[0002] With the acceleration of the digitalization process, the application of artificial intelligence models faces huge challenges in data security and privacy protection. On the one hand, sensitive information such as personal identity information (PII), financial data, health records, etc. widely exists in unstructured text. Once leaked, it will bring serious consequences to enterprises and individuals. On the other hand, large language models (LLMs) inevitably retain sensitive data such as privacy information and copyright materials during the training process, which raises concerns about the security and integrity of the models.

[0003] In existing sensitive information detection technologies, traditional rule-based methods, such as regular expressions and content fingerprinting, although having relatively high precision, require continuous maintenance and are difficult to handle the diversity and complexity of data. Machine learning methods, although having a certain degree of flexibility, require a large amount of labeled data during the training process, which is difficult to obtain in the field of sensitive information because the sensitivity of the data limits the sharing and annotation of the data. In addition, existing machine learning models often generate a large number of false positives when processing sensitive information in unstructured text, or are too sensitive to specific types of data, such as emails, account numbers, or social security numbers, etc., and cannot comprehensively cover various types of sensitive information.

[0004] At the same time, existing knowledge forgetting methods also face many problems. For example, the ambiguity of the forgetting boundary makes it difficult for the model to accurately locate when forgetting specific knowledge, often resulting in over-forgetting, that is, when removing sensitive information, some non-sensitive and crucial knowledge for the model function is also deleted. This over-forgetting not only affects the performance of the model but also may cause the model to degrade in its ability to handle tasks related to the forgotten knowledge. In addition, existing forgetting methods are insufficient in distinguishing between knowledge to be forgotten and knowledge to be retained, and cannot effectively identify which knowledge belongs to sensitive information and which knowledge belongs to the public domain or information beneficial to the model function. For example, when processing information related to public figures, the model may not be able to accurately distinguish their private life details (such as home address, contact information, etc.) and information contributing to their public image (such as career achievements, works, etc.). This ambiguity makes it difficult for knowledge forgetting technology to achieve an ideal privacy protection effect in practical applications and also limits its wide application in the field of data privacy protection. In summary, how to ensure the performance of the model and the integrity of knowledge while protecting user privacy is an urgent problem to be solved in current sensitive information detection and knowledge forgetting technologies.

[0005] Therefore, it is an urgent problem for those skilled in the art to propose a sensitive information processing method and system based on a pre-trained model and a hybrid model architecture to solve the difficulties existing in the prior art. Summary of the Invention

[0006] In view of this, the present invention provides a sensitive information processing method and system based on a pre-trained model and a hybrid model architecture, which can effectively detect and process sensitive information in unstructured text in a data-constrained environment.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A sensitive information processing method based on a pre-trained model and a hybrid model architecture includes the following steps:

[0009] S1. Dataset construction: Collect personal original sensitive entity sentences and convert them into template-form sentences, and use forged sensitive entity sentences to replace the original sensitive entity sentences to obtain a dataset;

[0010] S2. Dataset preprocessing: Preprocess the obtained dataset and output the preprocessed dataset;

[0011] S3. Data partitioning: Partition the preprocessed dataset into a training set and a test set;

[0012] S4. Model construction and training: Construct an unstructured sensitive information text recognition model, input the training set into the constructed unstructured sensitive information text recognition model for training, and obtain a trained unstructured sensitive information text recognition model;

[0013] S5. Locate sensitive information: Input the test set into the trained unstructured sensitive information text recognition model to obtain an output result, and use a CRF decoder to perform sequence labeling on the output result to realize the recognition of sensitive entities in the text;

[0014] S6. Forget sensitive information: Integrate the MemFlex technology, determine the erasure area of sensitive entities in the text based on the analysis of key regions of gradient information, and realize the forgetting of sensitive entity information.

[0015] Optionally, the specific content of collecting personal original sensitive entity sentences and converting them into template-form sentences, and using forged sensitive entity sentences to replace the original sensitive entity sentences to obtain a dataset in S1 is:

[0016] The set of personal original sensitive entity sentences collected includes:

[0017] OS = {PII Name , PII Organization , PII Location,…,Others}

[0018] Among them, OS is the set of original sentences, and PII Name is the personal name in sensitive information, and PII Organization is the organizational information in sensitive information, and PII Location is the address information in sensitive information, and Others is the part of the sentence excluding PII;

[0019] Using placeholders to replace the PII in the original sensitive entity sentences, the resulting template-form sentences include:

[0020] TS = { <nam> , <org> , <loc>,…,Others}

[0021] Among them, TS is the set of template-form sentences, <nam>Placeholder for the personal name in sensitive information, <org>Placeholder for organizational information in sensitive information, <loc>A placeholder for address information in sensitive information, and Others is the part of the sentence excluding PII;

[0022] Replace the original sensitive entity sentence with a forged sensitive entity sentence, that is, modify the placeholder into incorrect PII to generate a new sentence, so as to obtain a dataset including:

[0023] GS = {FPII Name , FPII Organization , FPII Location , …, Others}

[0024] Among them, GS is the set of generated sentences, and FPII Name is the incorrect personal name, and FPII Organization is the incorrect organizational information, and FPII Location is the incorrect address information.

[0025] Optionally, the unstructured sensitive information text recognition model constructed in S4 includes a pre-trained model processing unit, a feature enhancement unit, and a CRF decoder connected in sequence;

[0026] The pre-trained model processing unit is used to process unstructured text;

[0027] The feature enhancement unit is used to implement character-level feature enhancement and word-level feature enhancement;

[0028] The CRF decoder is used to classify sensitive entities.

[0029] Optionally, the specific content of using the CRF decoder to perform sequence annotation on the output result to realize the recognition of sensitive entities in the text in S5 is:

[0030]

[0031] Among them, P(Y|X) is the probability of the output label sequence Y = {y1, y2, …, y n} under the condition of the given input sequence X = {x1, x2, …, x n}, w is the weight vector of the feature function, Z(X) is the normalization factor, f(y t , y t-1 , X, t) is the feature function describing the dependence relationship between the label and the input feature and between labels, and n is the total number of elements constituting the input sequence X.

[0032] Optionally, the specific content of integrating the MemFlex technology in S6 to determine the erasure area of sensitive entities in the text based on the key area analysis of gradient information and realize the forgetting of sensitive entity information is:

[0033] Determine the sensitive entity erasure region θ in the text loc which is

[0034]

[0035] where μ and σ are selected thresholds, and G ULi is the unlearned gradient related to the i-th parameter, i is the i-th parameter, is the average value to obtain a stable erasure gradient matrix, k is the k-th batch, and g k is the gradient information obtained by the k-th pass through backpropagation, N is the total number of samples or batches, θ is the parameters of the model, μ is the threshold of cosine similarity, and σ is the threshold of gradient magnitude.

[0036] A sensitive information processing system based on a pre-trained model and a hybrid model architecture, applying a sensitive information processing method based on a pre-trained model and a hybrid model architecture according to any one of the above, includes: a dataset construction module, a dataset preprocessing module, a data partitioning module, a model construction and training module, a sensitive information localization module, and a sensitive information forgetting module;

[0037] The dataset construction module, connected to the input end of the dataset preprocessing module, is used to collect personal original sensitive entity sentences and convert them into template-form sentences, and replace the original sensitive entity sentences with forged sensitive entity sentences to obtain a dataset;

[0038] The dataset preprocessing module, connected to the input end of the data partitioning module, is used to preprocess the obtained dataset and output the preprocessed dataset;

[0039] The data partitioning module, connected to the input end of the model construction and training module, is used to partition the preprocessed dataset into a training set and a test set;

[0040] The model construction and training module, connected to the input end of the sensitive information localization module, is used to construct an unstructured sensitive information text recognition model, input the training set into the constructed unstructured sensitive information text recognition model for training, and obtain a trained unstructured sensitive information text recognition model;

[0041] The sensitive information localization module, connected to the input end of the sensitive information forgetting module, is used to input the test set into the trained unstructured sensitive information text recognition model to obtain an output result, and use a CRF decoder to perform sequence labeling on the output result to achieve the recognition of sensitive entities in the text;

[0042] The sensitive information forgetting module is used to integrate the MemFlex technology, determine the sensitive entity erasure region in the text based on the analysis of key regions of gradient information, and achieve the forgetting of sensitive entity information.

[0043] As can be seen from the above technical solutions, compared with the prior art, the present invention provides a sensitive information processing method and system based on a pre-trained model and a hybrid model architecture, which has the following beneficial effects:

[0044] (1) In terms of data annotation, by increasing the number of sentences and the placeholder replacement method, the manual annotation cost is significantly reduced, and the data preparation efficiency is improved;

[0045] (2) In terms of model performance, the unstructured sensitive information text recognition model uses lexical-level and character-level tokenization processing and feature enhancement, which significantly enhances the recognition ability of sensitive information and improves the adaptability of the model to different unstructured texts;

[0046] (3) At the level of data security and privacy protection, based on the analysis of key regions of gradient information, the sensitive entity erasure region in the text is determined to achieve precise forgetting of sensitive information, avoid excessive forgetting, and ensure data security compliance;

[0047] (4) The present invention also reduces the consumption of training resources, reduces the occupation of human, time, computing and storage resources, improves the training efficiency and effect, and has important application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0049] Figure 1 It is a flowchart of a sensitive information processing method based on a pre-trained model and a hybrid model architecture provided by the present invention;

[0050] Figure 2 It is a block diagram of the structure of a sensitive information processing system based on a pre-trained model and a hybrid model architecture provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0052] Refer to Figure 1 As shown in the figure, the present invention discloses a sensitive information processing method based on a pre-trained model and a hybrid model architecture, including the following steps:

[0053] S1. Dataset construction: Collect personal original sensitive entity sentences and convert them into template-form sentences, and use forged sensitive entity sentences to replace the original sensitive entity sentences to obtain a dataset;

[0054] S2. Dataset preprocessing: Preprocess the obtained dataset and output the preprocessed dataset;

[0055] S3. Data partitioning: Partition the preprocessed dataset into a training set and a test set;

[0056] S4. Model construction and training: Construct an unstructured sensitive information text recognition model, input the training set into the constructed unstructured sensitive information text recognition model for training, and obtain a trained unstructured sensitive information text recognition model;

[0057] S5. Locate sensitive information: Input the test set into the trained unstructured sensitive information text recognition model to obtain an output result, and use a CRF decoder to perform sequence annotation on the output result to realize the recognition of sensitive entities in the text;

[0058] S6. Forget sensitive information: Integrate the MemFlex technology, determine the sensitive entity erasure area in the text based on the gradient information key area analysis, and realize the forgetting of sensitive entity information.

[0059] Further, the specific content of collecting personal original sensitive entity sentences and converting them into template-form sentences, and using forged sensitive entity sentences to replace the original sensitive entity sentences to obtain a dataset in S1 is as follows:

[0060] The set of personal original sensitive entity sentences collected includes:

[0061] OS = {PII Name , PII Organization , PII Location , …, Others}

[0062] Among them, OS is the set of original sentences, PII Name is the personal name in the sensitive information, PII Organization is the organization information in the sensitive information, PII Location is the address information in the sensitive information, and Others is the part of the sentence excluding PII;

[0063] Use placeholders to replace the PII in the original sensitive entity sentences, and the template-form sentences obtained include:

[0064] TS = { <nam> , <org> , <loc>,…,Others}

[0065] Among them, TS is a set of template-form sentences, <nam>Placeholder for the personal name in sensitive information, <org>Placeholder for organizational information in sensitive information <loc>Placeholder for address information in sensitive information, Others is the part of the sentence excluding PII;

[0066] Replace the original sensitive entity sentence with a forged sensitive entity sentence, that is, modify the placeholder into incorrect PII to generate a new sentence, so as to obtain a dataset including:

[0067] GS = {FPII Name , FPII Organization , FPII Location , …, Others}

[0068] Among them, GS is the generated sentence set, FPII Name is the incorrect personal name, FPII Organization is the incorrect organization information, FPII Location is the incorrect address information.

[0069] Furthermore, the unstructured sensitive information text recognition model constructed in S4 includes a pre-trained model processing unit, a feature enhancement unit, and a CRF decoder connected in sequence;

[0070] The pre-trained model processing unit is used to process unstructured text;

[0071] The feature enhancement unit is used to achieve character-level feature enhancement and word-level feature enhancement;

[0072] The CRF decoder is used to classify sensitive entities.

[0073] Furthermore, the specific content of using the CRF decoder to perform sequence labeling on the output result to realize the recognition of sensitive entities in the text in S5 is:

[0074]

[0075] Among them, P(Y|X) is the probability of the output label sequence Y = {y1, y2, …, y n} under the condition of the given input sequence X = {x1, x2, …, x n}, w is the weight vector of the feature function, Z(X) is the normalization factor, f(y t , y t-1 , X, t) is the feature function describing the dependence relationship between the label and the input feature and between labels, and n is the total number of elements constituting the input sequence X.

[0076] Furthermore, the specific content of integrating the MemFlex technology in S6 to determine the sensitive entity erasure area in the text based on the key area analysis of gradient information and realize the forgetting of sensitive entity information is:

[0077] Determine the sensitive entity erasure region θ in the text loc is:

[0078]

[0079] where μ and σ are selected thresholds, and G ULi is the unlearned gradient related to the i-th parameter, i is the i-th parameter, is the average value to obtain a stable erasure gradient matrix, k is the k-th batch, and g k is the gradient information obtained through backpropagation for the k-th time, N is the total number of samples or batch numbers, θ is the parameters of the model, μ is the threshold of cosine similarity, and σ is the threshold of gradient magnitude.

[0080] In a specific embodiment, it includes the following:

[0081] S1: In the dataset construction stage, collect multi-modal data from public sources such as social media platforms, forums, and news comments. Use web crawler technology to collect data from public sources such as social media platforms, online forums, and news comment sections. Convert the original sensitive entity sentences containing personal identity information and other specific entities into template forms. Subsequently, replace them with forged sensitive entity sentences to construct a dataset. During this process, rely on a pre-trained language model based on the Transformer architecture to amplify the number of sentences, thereby effectively reducing the workload of manual annotation.

[0082] For example, the original sentence is "Bank <org>will notify Borrower <per>when it debitsBorrower <per>"account.”, generate a new sentence "Wells Fargo" by templatizing and replacing the fake PII <org>will notify Jack Milton <per>when it debits Jack Milton <per>"account."

[0083] Specifically, the collected set of original sensitive entity sentences of individuals includes:

[0084] OS = {PII Name , PII Organization , PII Location , …, Others}

[0085] Among them, OS is the set of original sentences, PII Name is the personal name in sensitive information, PII Organization is the organizational information in sensitive information, PII Location is the address information in sensitive information, and Others is the part of the sentence excluding PII;

[0086] Replace the PII in the original sensitive entity sentence with placeholders to obtain the template form sentences including:

[0087] TS = { <nam> , <org> , <loc>,…,Others}

[0088] Among them, TS is a set of template-form sentences, <nam>Placeholder for personal names in sensitive information, <org>Placeholder for the organizational information in sensitive information, <loc>A placeholder for address information in sensitive information, and Others is the part of the sentence excluding PII;

[0089] Use forged sensitive entity sentences to replace the original sensitive entity sentences, that is, modify the placeholder into incorrect PII to generate new sentences, so as to obtain a dataset including:

[0090] GS = {FPII Name , FPII Organization , FPII Location , …, Others}

[0091] Among them, GS is the set of generated sentences, and FPII Name is an incorrect personal name, and FPII Organization is incorrect organization information, and FPII Location is incorrect address information.

[0092] S2: Preprocess the obtained dataset and output the preprocessed dataset;

[0093] S3: Divide the preprocessed dataset into a training set and a test set;

[0094] S4: Build an unstructured sensitive information text recognition model, input the training set into the built unstructured sensitive information text recognition model for training, and obtain a trained unstructured sensitive information text recognition model;

[0095] Specifically, the unstructured sensitive information text recognition model passes through a pre-trained model processing unit for processing unstructured text; through a feature enhancement unit for implementing character-level feature enhancement and vocabulary-level feature enhancement; and then through a CRF decoder unit for classifying sensitive entities. The data enhancement result is to splice the character-level and vocabulary-level features. Among them, the BERT model is used to perform context semantic modeling on the word-level embeddings of the input sequence. The CharacterBERT model extracts the representation of words through character-level information, making up for the deficiency of BERT in fine-grained information modeling within words. Through vocabulary-level and character-level feature enhancement means, the accurate recognition ability of the model for sensitive entities is significantly improved.

[0096] The unstructured sensitive information text recognition model uses a pre-trained model based on the Transformer architecture, which combines character-level and word-level features. For character-level embedding, only the character-level embedding layer architecture is used. Each token is segmented into a character sequence of length "N" and input into multiple one-dimensional convolutional neural networks with different filter sizes. The model is initialized with open-source pre-trained weights. Finally, the final concatenated representation is input into a conditional random field (CRF) decoder, which outputs each entity type, including other entities represented by "O". The CRF decoder includes the probability of outputting a label sequence (such as entity labels) given an input sequence (such as words or feature vectors), which is used to capture the dependencies between adjacent labels and locate sensitive information. The weight vector of the feature function, the feature function, and the normalization factor work together to achieve accurate classification of sensitive entities.

[0097] S5: Input the test set into the trained unstructured sensitive information text recognition model to obtain the output results, and use the CRF decoder to perform sequence annotation on the output results to achieve the recognition of sensitive entities in the text;

[0098] Using the conditional random field (CRF) as the decoder, given an input sequence (such as words or feature vectors)) X = {x1, x2, …, x n}, the probability P(Y|X) of outputting a label sequence (such as entity labels) Y = {y1, y2, …, y n} is used to capture the dependencies between adjacent labels and locate sensitive information.

[0099] Annotate the sequence output by the model to accurately identify the sensitive entities f(y t , y t-1 , X, t).

[0100]

[0101] Among them, Z(X) is the normalization factor, and n is the total number of elements that make up the input sequence X.

[0102] S6: Integrate the MemFlex technology, determine the sensitive entity erasure area in the text based on the critical region analysis of gradient information, and achieve the forgetting of sensitive entity information;

[0103] Determine the sensitive entity erasure area θ loc as:

[0104]

[0105] Among them, μ and σ are the selected thresholds, G ULi is the unlearned gradient related to the i-th parameter, and i is the i-th parameter, To obtain a stable erasure gradient matrix for the average value, k is the k-th batch, and g k is the gradient information obtained by backpropagation for the k-th time. N is the total number of samples or batches, θ is the various parameters of the model, μ is the threshold of the cosine similarity, and σ is the threshold of the gradient magnitude.

[0106] and Figure 1 corresponding to the method described above, an embodiment of the present invention further provides a sensitive information processing system based on a pre-trained model and a hybrid model architecture, and its structural schematic diagram is as Figure 2 shown, including: a dataset construction module, a dataset preprocessing module, a data partitioning module, a model construction and training module, a sensitive information localization module, and a sensitive information forgetting module;

[0107] The dataset construction module, connected to the input end of the dataset preprocessing module, is used to collect personal original sensitive entity sentences and convert them into template-form sentences, and replace the original sensitive entity sentences with forged sensitive entity sentences to obtain a dataset;

[0108] The dataset preprocessing module, connected to the input end of the data partitioning module, is used to preprocess the obtained dataset and output the preprocessed dataset;

[0109] The data partitioning module, connected to the input end of the model construction and training module, is used to partition the preprocessed dataset into a training set and a test set;

[0110] The model construction and training module, connected to the input end of the sensitive information localization module, is used to construct an unstructured sensitive information text recognition model, input the training set into the constructed unstructured sensitive information text recognition model for training, and obtain a trained unstructured sensitive information text recognition model;

[0111] The sensitive information localization module, connected to the input end of the sensitive information forgetting module, is used to input the test set into the trained unstructured sensitive information text recognition model to obtain an output result, and use a CRF decoder to perform sequence annotation on the output result to realize the recognition of sensitive entities in the text;

[0112] The sensitive information forgetting module is used to integrate the MemFlex technology, determine the erasure area of sensitive entities in the text based on the key area analysis of gradient information, and realize the forgetting of sensitive entity information.

[0113] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0114] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.< / loc> < / org> < / nam> < / loc> < / org> < / nam> < / per> < / per> < / org> < / per> < / per> < / org> < / loc> < / org> < / nam> < / loc> < / org> < / nam> < / loc> < / org> < / nam> < / loc> < / org> < / nam>

Claims

1. A sensitive information processing method based on a pre-trained model and a hybrid model architecture, characterized in that: The following steps are involved: S1. Dataset construction: Collect personal original sensitive entity sentences and convert them into template sentences, and use forged sensitive entity sentences to replace the original sensitive entity sentences to obtain the dataset; S2. Dataset preprocessing: preprocess the acquired data set and output the preprocessed data set; S3. Data division: Divide the preprocessed data set into training set and test set; S4. Model construction and training: construct an unstructured sensitive information text recognition model, input the training set into the constructed unstructured sensitive information text recognition model for training, and obtain a trained unstructured sensitive information text recognition model; S5. Locating sensitive information: Input the test set into the trained unstructured sensitive information text recognition model to obtain the output results, and use the CRF decoder to sequence the output results to achieve sensitive entity recognition in the text; S6. Forget sensitive information: Integrate MemFlex technology to determine the sensitive entity erasure area in the text based on gradient information key area analysis, and realize the forgetting of sensitive entity information.

2. According to claim 1, a sensitive information processing method based on a pre-trained model and a hybrid model architecture is characterized in that: In S1, personal original sensitive entity sentences are collected and converted into template sentences. The original sensitive entity sentences are replaced with forged sensitive entity sentences. The specific content of the data set is as follows: The collected personal original sensitive entity sentences include: OS={PII Name ,PII Organization ,PII Location ,···,Others} Among them, OS is the original sentence set, PII Name The name of a person in sensitive information, PII Organization Organizational information in sensitive information, PII Location is the address information in the sensitive information, and Others is the part of the sentence without PII; Use placeholders to replace the PII in the original sensitive entity sentence, and the resulting template sentence includes: TS={ <nam> , <org> , <loc> ,···,Others}< / loc> < / org> < / nam> Among them, TS is a set of template sentences. <nam>A placeholder for a personal name in sensitive information. <org>It is a placeholder for organization information in sensitive information. <loc> It is a placeholder for address information in sensitive information, and Others is the part of the sentence without PII;< / loc> < / org> < / nam> The original sensitive entity sentences are replaced with forged sensitive entity sentences, that is, the placeholders are modified into incorrect PII to generate new sentences, so that the dataset includes: GS={FPII Name ,FPII Organization ,FPII Location ,···,Others} Among them, GS is the generated sentence set, FPII Name For wrong personal name, FPII Organization For incorrect organizational information, FPII Location Incorrect address information.

3. According to claim 1, a sensitive information processing method based on a pre-trained model and a hybrid model architecture is characterized in that: The unstructured sensitive information text recognition model constructed in S4 includes a pre-trained model processing unit, a feature enhancement unit, and a CRF decoder connected in sequence; Pre-trained model processing unit for processing unstructured text; A feature enhancement unit, used to achieve character-level feature enhancement and vocabulary-level feature enhancement; CRF decoder for classifying sensitive entities.

4. A sensitive information processing method based on a pre-trained model and a hybrid model architecture according to claim 1 or 3, characterized in that: In S5, the CRF decoder is used to sequence the output results to realize the sensitive entity recognition in the text. The specific content is: Where P(Y|X) is the given input sequence X = {x1,x2,···,x n }, the output label sequence Y = {y1,y2,···,y n }, w is the weight vector of the feature function, Z(X) is the normalization factor, f(y t ,y t-1 ,X,t) is the feature function that describes the dependency between labels and input features and labels, and n is the total number of elements that constitute the input sequence X.

5. The sensitive information processing method based on a pre-trained model and a hybrid model architecture according to claim 1, characterized in that: MemFlex technology is integrated in S6. Based on the key area analysis of gradient information, sensitive entity erasure areas in the text are determined. The specific contents of forgetting sensitive entity information are as follows: Determine the sensitive entity erasure area in the text θ loc for: Among them, μ and σ are the selected thresholds, G ULi is the unlearned gradient associated with the i-th parameter, i is the 9th parameter, is the average value to obtain a stable erased gradient matrix, k is the kth batch, g k is the gradient information obtained through back propagation for the kth time, N is the total number of samples or batches, θ is the parameters of the model, μ is the threshold of cosine similarity, and σ is the threshold of gradient size.

6. A sensitive information processing system based on a pre-trained model and a hybrid model architecture, characterized in that: A sensitive information processing method based on a pre-trained model and a hybrid model architecture according to any one of claims 1 to 5 is applied, comprising: a data set construction module, a data set pre-processing module, a data partitioning module, a model construction and training module, a sensitive information positioning module, and a sensitive information forgetting module; The data set construction module is connected to the input end of the data set preprocessing module, and is used to collect personal original sensitive entity sentences and convert them into template sentences, and replace the original sensitive entity sentences with forged sensitive entity sentences to obtain the data set; A data set preprocessing module, connected to the input end of the data partitioning module, is used to preprocess the acquired data set and output the preprocessed data set; A data partitioning module, connected to the input end of the model building and training module, is used to partition the preprocessed data set into a training set and a test set; A model building and training module is connected to the input end of the sensitive information positioning module, and is used to build an unstructured sensitive information text recognition model, input a training set into the built unstructured sensitive information text recognition model for training, and obtain a trained unstructured sensitive information text recognition model; The sensitive information location module is connected to the input end of the forget sensitive information module to input the test set into the trained unstructured sensitive information text recognition model to obtain the output result, and the CRF decoder is used to sequence the output result to realize sensitive entity recognition in the text; The module for forgetting sensitive information is used to integrate MemFlex technology, determine the sensitive entity erasure area in the text based on gradient information key area analysis, and realize the forgetting of sensitive entity information.