Network security sensitive information desensitization method and system for data leakage

By extracting and analyzing multi-dimensional contextual features of data, and combining pre-trained models to identify sensitive information and generate desensitization schemes, this approach solves the problem that semantic and temporal associations are not considered in existing technologies, and achieves efficient desensitization of sensitive information while preserving semantic integrity.

CN120893070APending Publication Date: 2025-11-04HEDUN DIGITAL (SHANGHAI) INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510999883.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing methods for de-identifying sensitive information fail to effectively consider the semantic and temporal relationships between data units, resulting in the loss of the original semantic integrity of the de-identified data and posing security risks.

Method used

By extracting multi-dimensional contextual features from the original dataset, including semantic and temporal correlation features, a pre-trained sensitive information identification model is invoked for localization processing, generating a desensitization processing scheme, and the effectiveness is verified.

Benefits of technology

Ensure the accuracy and effectiveness of data anonymization, preserve the semantic integrity of the data, and avoid damaging the semantic structure of the data after anonymization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120893070A_ABST
    Figure CN120893070A_ABST
Patent Text Reader

Abstract

The invention provides a network security sensitive information desensitization method and system for data leakage, and the method comprises the steps: obtaining a to-be-processed original data set, extracting the multi-dimensional context features, including a semantic association feature and a time sequence association feature, of the to-be-processed original data set; and calling a pre-trained sensitive information identification model to carry out sensitive information positioning processing on the multi-dimensional context features, generating a sensitive information identifier set, generating a corresponding desensitization processing scheme according to the sensitive information identifier set, including a replacement rule and a mask strategy, executing an effect verification operation on the desensitization processing scheme, and obtaining a desensitization result of the desensitization processing scheme. And a desensitization result verification report is generated, and the retention degree of the desensitized data set on the semantic integrity of the original data is evaluated, so that sensitive information can be comprehensively and accurately identified, accurate desensitization is realized, the semantic integrity of the data is retained, and the network security protection level is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information security, in particular to a network security sensitive information desensitization method and system for data leakage. BACKGROUND

[0002] In today's digital age, data leakage incidents occur frequently, causing huge losses to enterprises and individuals. In the field of network security, sensitive information desensitization is one of the key technologies to prevent data leakage. However, existing sensitive information desensitization methods often only focus on the surface form of data, such as simple character replacement or masking, while ignoring the semantic association and temporal association between data units. This desensitization method is easy to cause the desensitized data to lose the original semantic integrity, and even may produce new security risks. For example, in the medical, financial, e-commerce and other fields, the sensitive information in the data is often closely related to the context, and simple desensitization processing may destroy the semantic structure of the data, making the desensitized data unable to be used for subsequent analysis and processing. Therefore, a sensitive information desensitization method that can comprehensively consider the semantic association and temporal association of data units is needed to improve the desensitization effect and preserve the semantic integrity of the data. SUMMARY

[0003] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide a network security sensitive information desensitization method for data leakage, which comprises:

[0004] Obtaining a set of original data to be processed, the set of original data comprising a plurality of data units, each data unit consisting of text content and metadata information;

[0005] Extracting multi-dimensional context features of the set of original data, the multi-dimensional context features comprising semantic association features of data units and temporal association features between data units;

[0006] Calling a pre-trained sensitive information recognition model to perform sensitive information positioning processing on the multi-dimensional context features, generating a set of sensitive information identifiers of the set of original data, the set of sensitive information identifiers being used to indicate specific positions and types of data units that need to be desensitized;

[0007] Generating a corresponding desensitization processing scheme according to the set of sensitive information identifiers, the desensitization processing scheme comprising replacement rules and masking strategies for different types of sensitive information;

[0008] Performing an effect verification operation on the desensitization processing scheme, generating a desensitization result verification report, the desensitization result verification report being used to evaluate the preservation degree of the semantic integrity of the original data set after desensitization.

[0009] In still another aspect, the embodiment of the present application also provides a network security sensitive information desensitization system for data leakage, comprising a processor, a machine readable storage medium, the machine readable storage medium and the processor are connected, the machine readable storage medium is used for storing programs, instructions or codes, and the processor is used for executing the programs, instructions or codes in the machine readable storage medium to realize the above method.

[0010] Based on the above aspects, the embodiment of the present application can comprehensively and accurately understand the internal relationship between data units by extracting multi-dimensional context features of the original data set, including semantic association features of data units and timing association features between data units, calling a pre-trained sensitive information identification model to perform sensitive information positioning processing on the multi-dimensional context features, accurately identifying specific positions and types that need to be desensitized, generating a corresponding desensitization processing scheme according to the sensitive information identification set, the desensitization processing scheme containing replacement rules and mask strategies for different types of sensitive information, ensuring the effectiveness and pertinence of desensitization processing, performing effect verification operation on the desensitization processing scheme, generating a desensitization result verification report, and evaluating the retention degree of the original data semantic integrity of the desensitized data set, thereby ensuring that the desensitization processing will not damage the original semantic structure of the data, thereby significantly improving the accuracy and effectiveness of sensitive information desensitization while preserving the semantic integrity of the data. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is the execution flow diagram of the network security sensitive information desensitization method for data leakage provided by the embodiment of the present application.

[0012] Figure 2 is the schematic diagram of exemplary hardware and software components of the network security sensitive information desensitization system for data leakage provided by the embodiment of the present application. DETAILED DESCRIPTION

[0013] The present application will be described in detail below with reference to the accompanying drawings, Figure 1 is the flow diagram of the network security sensitive information desensitization method for data leakage provided by an embodiment of the present application, and the network security sensitive information desensitization method for data leakage will be described in detail below.

[0014] Step S110: obtaining an original data set to be processed, the original data set containing a plurality of data units, each data unit consisting of text content and metadata information.

[0015] In today's digital age, the generation and storage of data are showing an explosive growth trend. In order to protect the security of data and avoid the leakage of sensitive information, it is necessary to first obtain the original data set to be processed. The source of the original data set is diverse, which can come from the internal database system of the enterprise, which may store customer transaction records, personal information, business data, etc.; it can also come from the data crawled by web crawlers from the Internet, such as news articles, social media posts, etc.; it can also come from the data collected by Internet of Things devices, such as environmental data collected by sensors, device operating status data, etc.

[0016] For example, a large e-commerce enterprise has a database storing a large amount of user transaction data. These user transaction data contain numerous data units, each consisting of text content and metadata information. The text content can be the text information input by the user when filling in the delivery address, product evaluation, message feedback, etc., such as "Please deliver the goods to XX community XX building XX number, contact phone XXXXXXXX", "This mobile phone has excellent shooting effect and high cost performance" etc. Metadata information includes data generation time, data source device, user IP address, operation type, etc.

[0017] Step S120: Extracting multi-dimensional context features of the original data set, the multi-dimensional context features including semantic association features of data units and time sequence association features between data units.

[0018] The extraction of multi-dimensional context features is crucial for accurately identifying sensitive information in the original data set. The semantic association features of data units can reflect the semantic relationship between the internal vocabulary of data units, while the time sequence association features between data units can reflect the association degree of different data units in the time dimension and source dimension. By comprehensively considering the semantic association features of data units and the time sequence association features between data units, the context information of data can be more comprehensively understood, thereby improving the accuracy of sensitive information identification.

[0019] Step S121: Performing word segmentation processing on each data unit in the original data set to obtain a word segmentation result set containing basic vocabulary units and combined phrase units.

[0020] The purpose of word segmentation processing is to divide continuous text content into meaningful vocabulary units. For each data unit in the original data set, word segmentation processing is needed to obtain a word segmentation result set containing basic vocabulary units and combined phrase units.

[0021] Step S1211: Performing punctuation symbol segmentation processing on the text content of the data unit to generate a text segmentation set composed of sentence units.

[0022] Before performing the word segmentation processing, first, the punctuation symbol segmentation processing needs to be performed on the text content of the data unit. The punctuation symbol is an important separator in the text, which can divide the text content into individual sentence units. For example, in Chinese text, full stop, question mark, exclamation mark, etc. usually indicate the end of a sentence. Taking an evaluation of an e-commerce user "The shooting effect of this mobile phone is very good! The battery life is also good, but it is a bit expensive." as an example, through punctuation symbol segmentation processing, it can be divided into two sentence units: "The shooting effect of this mobile phone is very good!" and "The battery life is also good, but it is a bit expensive.".

[0023] Step S1212: Perform single-character segmentation processing on each sentence unit to generate an initial word segmentation candidate set containing continuous single-character sequences.

[0024] After obtaining the text segmentation set, single-character segmentation processing needs to be performed on each sentence unit. Single-character segmentation is to take each character in the sentence unit as an independent element to generate an initial word segmentation candidate set containing continuous single-character sequences. For example, for the sentence unit "The shooting effect of this mobile phone is very good", the initial word segmentation candidate set obtained after single-character segmentation is "this", "mobile phone", "shooting effect", "very good".

[0025] Step S1213: Perform vocabulary matching processing on the initial word segmentation candidate set based on the pre-constructed domain dictionary to identify the basic vocabulary units that meet the dictionary entries.

[0026] The pre-constructed domain dictionary is a vocabulary set constructed according to the vocabulary characteristics and semantic rules of a specific domain. After obtaining the initial word segmentation candidate set, vocabulary matching processing needs to be performed based on the pre-constructed domain dictionary. By matching the continuous single-character sequences in the initial word segmentation candidate set with the entries in the domain dictionary, the basic vocabulary units that meet the dictionary entries are identified. For example, in the dictionary of the e-commerce domain, it may contain vocabulary entries such as "mobile phone", "shooting effect", etc. For the initial word segmentation candidate set "this", "mobile phone", "shooting effect", "very good", through vocabulary matching processing, "mobile phone" and "shooting effect" can be identified as basic vocabulary units.

[0027] Step S1214: Perform adjacent word combination analysis processing on the single-character sequences that do not match the basic vocabulary units to extract double-word or multi-word combinations with semantic coherence as combined phrase units.

[0028] After performing lexical matching processing, there may be some single - character sequences that do not match the basic lexical units. For these single - character sequences, adjacent character combination analysis needs to be performed. By combining adjacent single characters, analyze whether they have semantic coherence. If they have semantic coherence, extract them as combined phrase units. For example, for the single - character sequence "这" and "款" that do not match the basic lexical units, after analysis, it is found that "这款" has semantic coherence and can be used as a combined phrase unit.

[0029] Step S1215: Arrange and combine the basic lexical units and the combined phrase units in the order of their appearance in the sentence unit to generate a set of word - segmentation results containing position index information.

[0030] After identifying the basic lexical units and combined phrase units, they need to be arranged and combined in the order of their appearance in the sentence unit. At the same time, for the convenience of subsequent processing, position index information also needs to be added to each lexical unit and combined phrase unit. For example, for the sentence unit "这款手机的拍照效果很好", the set of word - segmentation results obtained after word - segmentation processing may be ("这款", position index 1 - 2), ("手机", position index 3 - 4), ("的", position index 5), ("拍照效果", position index 6 - 9), ("很好", position index 10 - 11).

[0031] Step S122: Call a pre - trained context encoder to perform semantic encoding processing on the set of word - segmentation results, generating a semantic vector representation for each data unit, and the semantic vector representation is used to reflect the semantic association relationship between the internal words of the data unit.

[0032] The pre - trained context encoder is a model trained with a large amount of text data. It can encode the lexical units and combined phrase units in the set of word - segmentation results into semantic vectors. Through semantic encoding processing, the text information can be converted into a vector form that can be understood and processed by a computer, thus better reflecting the semantic association relationship between the internal words of the data unit.

[0033] When calling the pre-trained context encoder, the segmented result set needs to be input into the encoder first. The context encoder will generate a corresponding semantic vector representation for each data unit according to the semantic information of the lexical units and combined phrase units. For example, for the segmented result set ("this", position index 1-2), ("mobile phone", position index 3-4), ("of", position index 5), ("photographing effect", position index 6-9), ("very good", position index 10-11), the context encoder will encode it into a multi-dimensional semantic vector. Each dimension in the semantic vector represents certain semantic information, and through the distance and similarity between vectors, the semantic correlation between different data units can be measured.

[0034] Step S123: Analyze the metadata information of the data units in the original data set, extract the timestamp label and source identification information of the data units.

[0035] Metadata information plays an important role in understanding the background and context of data units. In the original data set, each data unit contains certain metadata information, among which the timestamp label and source identification information are very important information.

[0036] The timestamp label records the specific time when the data unit is generated, which can help understand the timeliness and sequence of the data. The source identification information indicates the source of the data unit, such as which device, system or user the data comes from. By analyzing the metadata information of the data units, these important information can be extracted. For example, in the original data set of an e-commerce platform, for each user's transaction record, the metadata information may include the transaction time (timestamp label) and the user's device ID (source identification information).

[0037] Step S124: Based on the timestamp label, construct a time sequence arrangement sequence of the data units, and calculate the interval duration parameter of adjacent data units in the time sequence arrangement sequence.

[0038] After extracting the timestamp label of the data unit, the time sequence arrangement sequence of the data unit can be constructed based on these timestamp labels. The time sequence arrangement sequence is a sequence obtained by sorting the data units according to the time sequence of the data units. By constructing the time sequence arrangement sequence, the time relationship between the data units can be clearly understood.

[0039] After the time sequence sequence is constructed, the interval duration parameter of adjacent data units in the time sequence sequence needs to be calculated. The interval duration parameter can reflect the time interval between adjacent data units, which is of great significance for analyzing the time sequence correlation characteristics between data units. For example, in the transaction records of an e-commerce platform, by calculating the time interval between two adjacent transactions, the user's shopping frequency and behavior pattern can be understood.

[0040] Step S125: generating a time sequence correlation characteristic between data units according to the interval duration parameter and the source identification information, the time sequence correlation characteristic being used to reflect the correlation degree of different data units in the time dimension and the source dimension.

[0041] The time sequence correlation characteristic between data units can comprehensively consider the information in the time dimension and the source dimension. By combining the interval duration parameter and the source identification information, a time sequence correlation characteristic reflecting the correlation degree of different data units in these two dimensions can be generated.

[0042] For example, if the source identification information of two data units is the same and the interval duration parameter is short, it can be considered that the two data units have a strong correlation degree in the time dimension and the source dimension. On the contrary, if the source identification information of two data units is different and the interval duration parameter is long, the correlation degree between them may be weak. By generating the above time sequence correlation characteristic, the relationship between data units can be better understood.

[0043] Step S126: performing feature fusion processing on the semantic vector representation and the time sequence correlation characteristic to obtain a multi-dimensional context feature with unified dimensional representation.

[0044] After obtaining the semantic vector representation of the data unit and the time sequence correlation characteristic between the data units, they need to be processed by feature fusion. Feature fusion can integrate different types of feature information together to form a multi-dimensional context feature with unified dimensional representation.

[0045] There are many methods for feature fusion, for example, the semantic vector representation and the time sequence correlation characteristic can be simply connected by splicing. In this embodiment, the semantic vector representation and the time sequence correlation characteristic are fused by splicing. For example, assuming that the semantic vector representation is an n-dimensional vector and the time sequence correlation characteristic is an m-dimensional vector, then the fused multi-dimensional context feature is an (n+m)-dimensional vector. Through the above feature fusion processing, a more comprehensive and richer multi-dimensional context feature can be obtained.

[0046] Step S130: calling a pre-trained sensitive information recognition model to perform sensitive information positioning processing on the multi-dimensional context features, to generate a sensitive information identification set of the original data set, which is used to indicate the specific location and type of data units that need to be desensitized.

[0047] The pre-trained sensitive information recognition model is an artificial intelligence model trained by a large amount of labeled data, which can analyze and process the input multi-dimensional context features, identify the sensitive information therein, and determine the specific location and type of the sensitive information.

[0048] Step S131: inputting the multi-dimensional context features into the feature input layer of the sensitive information recognition model to perform feature dimension alignment processing, to generate a standardized feature vector meeting the model input requirements.

[0049] Before inputting the multi-dimensional context features into the sensitive information recognition model, it needs to be processed for feature dimension alignment. Feature dimension alignment is to ensure that the input feature vector has a uniform dimension, meeting the input requirements of the model. Different models may have different requirements for the dimension of the input features, so the multi-dimensional context features need to be processed accordingly.

[0050] For example, the feature input layer of the sensitive information recognition model may require the input feature vector to be a fixed-dimensional vector. If the dimension of the multi-dimensional context features is inconsistent with the required dimension of the model, dimension adjustment is needed. The feature vector can be processed by padding or truncation to make its dimension meet the requirements of the model. After feature dimension alignment processing, the generated standardized feature vector can be better processed by the model.

[0051] Step S132: performing deep semantic analysis processing on the standardized feature vector by the semantic understanding module of the sensitive information recognition model, to extract potential sensitive segments containing personal identity information, financial account information and private communication information in the data units.

[0052] The semantic understanding module of the sensitive information recognition model can perform deep semantic analysis processing on the standardized feature vector to identify potential sensitive segments therein. The potential sensitive segments may contain sensitive content such as personal identity information, financial account information and private communication information.

[0053] Step S1321: inputting the standardized feature vector into the word embedding layer of the semantic understanding module to generate word vector representation of each word unit.

[0054] The word embedding layer of the semantic understanding module can convert the lexical units in the standardized feature vector into word vector representations. The word vector representations can map the semantic information of the words into a low-dimensional vector space, enabling the computer to better understand the semantic relationships between the words.

[0055] For example, for each lexical unit in the standardized feature vector, the word embedding layer generates a corresponding word vector according to a pre-trained word vector model. These word vectors can reflect the semantic features of the words, and the degree of semantic association between the words can be measured by calculating the similarity between the word vectors.

[0056] Step S1322: The word vector representations are processed by the bidirectional long short-term memory network layer of the semantic understanding module to model the temporal dependencies, generating context word vectors containing the semantic associations of the preceding and following texts.

[0057] The bidirectional long short-term memory network (Bi-LSTM) layer can model the temporal dependencies of the word vector representations, taking into account the semantic information of the preceding and following texts of the lexical units in the sentence. Through the bidirectional long short-term memory network layer, context word vectors containing the semantic associations of the preceding and following texts can be generated.

[0058] Step S13221: The word vector representations are input into the forward propagation sublayer of the bidirectional long short-term memory network layer in the order of the appearance of the lexical units in the sentence, generating a forward hidden state sequence reflecting the semantic information of the preceding text.

[0059] The bidirectional long short-term memory network layer includes a forward propagation sublayer and a backward propagation sublayer. In the forward propagation sublayer, the word vector representations are input in the order of the appearance of the lexical units in the sentence, and through the calculation and transmission of the network, a forward hidden state sequence reflecting the semantic information of the preceding text can be generated. For example, for the sentence "This mobile phone has a good camera effect", the word vector representations of the lexical units "This", "mobile phone", "and" can be processed in sequence in the forward propagation sublayer to generate the corresponding forward hidden state sequence.

[0060] Step S13222: The word vector representations are input into the backward propagation sublayer of the bidirectional long short-term memory network layer in the reverse order of the appearance of the lexical units in the sentence, generating a backward hidden state sequence reflecting the semantic information of the following text.

[0061] In the backward propagation sublayer, the word vector representations are input in the reverse order of the appearance of the lexical units in the sentence, and through the calculation and transmission of the network, a backward hidden state sequence reflecting the semantic information of the following text can be generated. For example, for the sentence "This mobile phone has a good camera effect", the word vector representations of the lexical units "good", "effect", "camera" can be processed in sequence in the backward propagation sublayer to generate the corresponding backward hidden state sequence.

[0062] Step S13223: performing element-wise addition processing on the forward hidden state sequence and the backward hidden state sequence to generate a combined hidden state sequence that fuses the context semantic information.

[0063] After obtaining the forward hidden state sequence and the backward hidden state sequence, element-wise addition processing needs to be performed on them. By element-wise addition, the semantic information of the context can be fused to generate a combined hidden state sequence. The combined hidden state sequence can comprehensively consider the context semantic information of the lexical unit.

[0064] Step S13224: inputting the combined hidden state sequence into the output layer of the bidirectional long short-term memory network layer to perform linear transformation processing to generate a context word vector sequence consistent with the word vector representation dimension.

[0065] The output layer of the bidirectional long short-term memory network layer can perform linear transformation processing on the combined hidden state sequence. Through linear transformation, the combined hidden state sequence can be converted into a context word vector sequence consistent with the word vector representation dimension, so as to ensure that the context word vector sequence has the same dimension as the original word vector representation.

[0066] Step S13225: extracting the context word vector corresponding to each lexical unit from the context word vector sequence as the final output.

[0067] After obtaining the context word vector sequence, the context word vector corresponding to each lexical unit needs to be extracted from it as the final output. These context word vectors contain the context semantic information of the lexical unit, which can more accurately reflect the semantic role and association relationship of the lexical unit in the sentence.

[0068] Step S1323: calling the attention mechanism layer of the semantic understanding module to perform importance weight allocation processing on the context word vector to generate an attention weight value reflecting the semantic contribution degree of the lexical unit in the sentence.

[0069] The attention mechanism layer can perform importance weight allocation processing on the context word vector. By calculating the importance weight value of each context word vector, the semantic contribution degree of the lexical unit in the sentence can be reflected.

[0070] For example, in a sentence, some words may be more important to express the core semantics of the sentence, while some words may be relatively secondary. The attention mechanism layer can allocate a corresponding attention weight value to each context word vector according to the context information and semantic relationship of the lexical unit. The greater the weight value, the higher the semantic contribution degree of the lexical unit in the sentence.

[0071] Step S1324: Perform weighted aggregation processing on the context word vectors based on the attention weight values, to generate a sentence-level semantic representation vector.

[0072] After obtaining the attention weight values, it is necessary to perform weighted aggregation processing on the context word vectors based on these weight values. Weighted aggregation can weight and sum each context word vector according to its importance, to generate a sentence-level semantic representation vector.

[0073] For example, assuming that the sequence of context word vectors is a set containing multiple context word vectors, and the attention weight values are a set of weights corresponding to the context word vectors one by one. By multiplying each context word vector by the corresponding attention weight value and then adding them up, a sentence-level semantic representation vector can be obtained. This semantic representation vector can more comprehensively reflect the semantic information of the sentence.

[0074] Step S1325: Perform similarity matching processing between the semantic representation vector and the preset sensitive information feature template, and extract the combination of words whose matching degree exceeds the preset standard as a potential sensitive fragment. The preset sensitive information feature template is pre-constructed according to common sensitive information types and semantic features. These templates can cover typical semantic patterns of various sensitive information such as personal identity information, financial account information, and private communication information. When performing similarity matching processing, the sentence-level semantic representation vector is compared with each sensitive information feature template.

[0075] The comparison process is based on the similarity measure between vectors. This similarity measure can reflect the closeness of the semantic representation vector and the sensitive information feature template in the semantic space. When the similarity matching degree exceeds the preset standard, it means that the corresponding combination of words may contain sensitive information, which is extracted as a potential sensitive fragment.

[0076] For example, for a sentence describing a user's transaction record, if the matching degree exceeds the preset standard after the semantic representation vector is matched with the sensitive information feature template of financial account information, then the combination of words related to account number, transaction amount, etc. in the sentence will be identified as a potential sensitive fragment.

[0077] Step S133: Use the position positioning module of the sensitive information recognition model to perform coordinate labeling processing on the start position and end position of the potential sensitive fragment in the data unit, to generate a set of sensitive fragment position coordinates.

[0078] The position positioning module of the sensitive information recognition model is responsible for determining the specific position of the potential sensitive fragment in the data unit, and it can accurately label the start position and end position of the potential sensitive fragment in the original text with coordinates.

[0079] When performing coordinate labeling, the position positioning module refers to the position index information in the word segmentation result set. Through these index information, the starting and ending positions of the potential sensitive fragments in the data unit text content can be accurately found. For example, for a potential sensitive fragment "user's ID number XXXXXXX", the position positioning module will determine the starting and ending positions of "ID number" and the specific number after it in the text according to the word segmentation result set, and record these position information in the form of coordinates.

[0080] The position coordinates of all potential sensitive fragments are summarized to generate a sensitive fragment position coordinate set. The sensitive fragment position coordinate set contains the specific position information of each potential sensitive fragment in the data unit.

[0081] Step S134: calling the type classification module of the sensitive information recognition model to perform type discrimination processing on the potential sensitive fragments, and determining the sensitive information type corresponding to each sensitive fragment.

[0082] The role of the type classification module of the sensitive information recognition model is to accurately discriminate the types of potential sensitive fragments. It classifies potential sensitive fragments into different sensitive information types, such as personal identity information, financial account information, private communication information, etc., according to their semantic features and context information.

[0083] The type classification module can combine preset sensitive information type rules and semantic patterns when performing discrimination. For example, if the potential sensitive fragment contains information such as name and ID number, the type classification module will determine it as a personal identity information type; if it contains bank card number and transaction amount, it will be determined as a financial account information type.

[0084] During the discrimination process, the type classification module will also consider the context information of the potential sensitive fragment. Sometimes, a word or phrase may have different sensitive information types in different contexts. For example, the word "number" can be determined as a private communication information type in the context of "mobile phone number", but as a personal identity information type in the context of "ID number".

[0085] Through the processing of the type classification module, each potential sensitive fragment is accurately determined to correspond to a sensitive information type.

[0086] Step S135: associating and mapping the sensitive fragment position coordinate set and the corresponding sensitive information type, generating a sensitive information identification set containing position coordinates and type identifiers.

[0087] After obtaining the sensitive segment position coordinate set and the sensitive information type corresponding to each potential sensitive segment, they need to be associated and mapped. The purpose of association and mapping is to correspond the position coordinates and the sensitive information type one by one to form a complete sensitive information identification set.

[0088] The specific association and mapping process is that for each position coordinate in the sensitive segment position coordinate set, the corresponding potential sensitive segment is found, and the sensitive information type of the potential sensitive segment is associated with it. For example, for the position coordinate (start position 10, end position 20) corresponding to the potential sensitive segment "ID number XXXXXXX", the sensitive information type "personal information" is associated with the position coordinate.

[0089] Through the above association and mapping process, the generated sensitive information identification set contains the position coordinates of each sensitive segment and the corresponding type identifier. The sensitive information identification set indicates the specific position in the data unit that needs to be desensitized and the corresponding sensitive information type.

[0090] Step S140: generating a corresponding desensitization processing scheme according to the sensitive information identification set, the desensitization processing scheme containing replacement rules and masking strategies for different types of sensitive information.

[0091] In this embodiment, according to the different sensitive information types and the corresponding position information in the sensitive information identification set, the corresponding replacement rules and masking strategies need to be formulated to ensure that the sensitive information is effectively protected.

[0092] Step S141: parsing the sensitive information type in the sensitive information identification set, and extracting the classification identifier of the personal information type, the financial account information type and the private communication information type.

[0093] First, the sensitive information identification set needs to be parsed to extract the classification identifier of different sensitive information types. The personal information type may include name, ID number, date of birth, etc.; the financial account information type may include bank card number, account balance, transaction record, etc.; and the private communication information type may include mobile phone number, email address, etc.

[0094] In the parsing process, different types of sensitive information can be classified according to the type identifier recorded in the sensitive information identification set. For example, for a sensitive information identification set containing multiple potential sensitive segments, the segments belonging to the personal information type can be filtered out, and their corresponding classification identifiers are extracted. Similarly, the financial account information type and the private communication information type are processed.

[0095] Step S142: For each classification identifier, match the corresponding basic replacement rule in the preset de-sensitization rule library, which contains character replacement patterns and length preservation requirements.

[0096] The preset de-sensitization rule library is pre-constructed and contains basic replacement rules for different types of sensitive information. For each extracted classification identifier, the corresponding basic replacement rule needs to be found in the de-sensitization rule library.

[0097] The character replacement pattern in the basic replacement rule specifies how to replace the characters in the sensitive information. For example, for an ID number, a possible character replacement pattern is to replace part of the digits with a specific character, such as “*”. The length preservation requirement specifies the length of the sensitive information that needs to be preserved during the replacement process. For example, for a mobile phone number, it may be required to preserve the first three and last four digits, and replace the middle digits.

[0098] For example, for the name in the personal identity information type, the basic replacement rule may be to replace the last character of the name with “certain”; for the bank card number in the financial account information type, the basic replacement rule may be to replace some of the middle digits with “*”, while preserving the first few and last few digits of the card number.

[0099] Step S143: Analyze the sensitive segment position coordinates in the set of sensitive information identifiers to determine the context information of the sensitive segment in the data unit.

[0100] The context information of the sensitive segment in the data unit has significant value for generating appropriate de-sensitization processing schemes. By analyzing the sensitive segment position coordinates in the set of sensitive information identifiers, the text content around the sensitive segment can be determined, thus understanding its context.

[0101] For example, for a potential sensitive segment “user's ID number XXXXXXX”, the specific position in the data unit text can be found through the position coordinates, and then the text content before and after it can be analyzed. If the text before is “in order to handle business, please provide”, and the text after is “in order to perform identity verification”, then the semantic and role of the sensitive segment in the above context can be understood.

[0102] The context information can include the grammatical structure, semantic logic, and association relationship with other words of the sentence where the sensitive segment is located. These information are helpful to judge how to maintain the semantic coherence and rationality of the sentence during the de-sensitization process.

[0103] Step S144: According to the context information, adaptively adjust the basic replacement rule to generate a customized replacement rule that conforms to the semantic coherence of the context.

[0104] Step S1441: Extract the leading and trailing words of the sensitive fragment in the data unit, and construct a context window containing the leading and trailing words.

[0105] In order to more accurately adjust the basic replacement rule according to the context information, first, the leading and trailing words of the sensitive fragment in the data unit are extracted, and a context window is constructed. The context window contains the words within a certain range before and after the sensitive fragment, which can more comprehensively reflect the context environment of the sensitive fragment.

[0106] For example, for the sensitive fragment "identity card number XXXXXXX", the leading word "user's" and the trailing word "for identity verification" are extracted, and a context window containing these leading and trailing words is constructed. The context window can help analyze the semantic role of the sensitive fragment in the sentence and the relationship with other words.

[0107] Step S1442: Analyze the semantic categories of the words in the context window, and determine the grammatical structure type of the sentence in which the sensitive fragment is located.

[0108] By analyzing the semantic categories of the words in the context window, the semantic properties of these words, such as nouns, verbs, adjectives, etc., can be understood. By analyzing the semantic categories of the words, the grammatical structure type of the sentence in which the sensitive fragment is located can be further determined.

[0109] For example, in the context window "user's identity card number for identity verification", "user's" is a limiting word of the possession relationship, "identity card number" is a noun, "for" is a preposition, and "identity verification" is a noun phrase. By analyzing the semantic categories of these words, it can be determined that the grammatical structure type of the sentence is subject-predicate-object structure, and "identity card number" is the subject in the sentence.

[0110] Step S1443: Determine the grammatical function of the sensitive fragment in the sentence according to the grammatical structure type, the grammatical function including subject, object or modifier.

[0111] According to the determined grammatical structure type, the grammatical function of the sensitive fragment in the sentence can be determined. The sensitive fragment may appear in the sentence as a subject, an object or a modifier.

[0112] For example, in the above example, "identity card number" is the subject of the sentence, and it assumes the grammatical function of expressing the main information. If the sentence is "the staff checked the user's identity card number", then "identity card number" is the object, which is the object of the action "check". If the sentence is "the file with the identity card number", then "identity card number" is a modifier, which modifies "file".

[0113] Step S1444: Adjust the character replacement pattern in the base replacement rule for sensitive fragments as subjects, preserving lexical attribute features matching subject grammatical functions.

[0114] When sensitive fragments are subjects, the character replacement pattern in the base replacement rule needs to be adjusted to preserve lexical attribute features matching subject grammatical functions. Subjects are usually the core part of a sentence expressing the main information, so their semantic integrity and rationality should be maintained as much as possible during replacement.

[0115] For example, for a name as a subject, the base replacement rule may originally replace the last character with "X". But if the name as a subject assumes the function of emphasizing a specific identity in a specific context, the replacement pattern may need to be adjusted to replace only part of the character slightly to preserve the recognizability of the name and its matching with the grammatical function of the subject.

[0116] Step S1445: Adjust the length preservation requirement in the base replacement rule for sensitive fragments as objects, ensuring that the replaced word is reasonable in combination with the predicate verb.

[0117] When sensitive fragments are objects, the length preservation requirement in the base replacement rule needs to be adjusted. The object is the object of the action, and the replaced word needs to be semantically and grammatically reasonable in combination with the predicate verb.

[0118] For example, for a bank card number as an object, the base replacement rule may stipulate replacing several middle digits with "*". But if the predicate verb is "check", in order to ensure the rationality of the "check" action, the length preservation requirement may need to be adjusted to preserve more digits so that effective checking operations can still be performed after replacement.

[0119] Step S1446: Adjust the replacement symbol type in the base replacement rule for sensitive fragments as modifiers, maintaining the semantic modification relationship between the modifier and the center word.

[0120] When sensitive fragments are modifiers, the replacement symbol type in the base replacement rule needs to be adjusted. The role of the modifier is to modify and limit the center word, so the semantic modification relationship between the modifier and the center word should be maintained during replacement.

[0121] For example, for a mobile phone number as a modifier, if the base replacement rule is to replace the middle digits with "*", but in a specific context, the semantic relationship between the modifier and the center word requires that the replacement symbol not affect the modification effect too much, the replacement symbol type may need to be adjusted to choose a more moderate replacement method, such as replacing some digits with "#".

[0122] Step S1447: Perform rule integration processing on the adjusted character replacement pattern, length preservation requirement, and replacement symbol type to generate a customized replacement rule.

[0123] After adapting the basic replacement rule, the adjusted character replacement pattern, length preservation requirement, and replacement symbol type need to be processed for rule integration. By integrating these adjusted rules, a customized replacement rule that meets the semantic coherence of the context is generated.

[0124] The customized replacement rule can better adapt to the characteristics of sensitive fragments in specific contexts, ensuring that sensitive information is protected while maintaining the semantic integrity and reasonableness of the data unit text during the desensitization process.

[0125] Step S145: For sensitive fragments containing consecutive numbers or letters, generate a fixed-length mask-based masking strategy containing a mask symbol type and a mask coverage ratio.

[0126] For sensitive fragments containing consecutive numbers or letters, a fixed-length mask-based masking strategy needs to be generated. The mask symbol type in the masking strategy specifies what symbol to use to replace the characters in the sensitive information, such as "*" and "#". The mask coverage ratio specifies the length ratio of sensitive information that needs to be masked.

[0127] For example, for a bank card number containing consecutive numbers, the masking strategy may specify using "*" as the mask symbol, and the mask coverage ratio is 80%, i.e., 80% of the numbers in the bank card number are masked.

[0128] When generating the masking strategy, the type and context of the sensitive information need to be considered. For some sensitive information that needs to partially retain recognizability, such as a mobile phone number, the mask coverage ratio may need to be adjusted to retain some numbers for easy identification.

[0129] Step S146: Perform combination and packaging processing on the customized replacement rule and the masking strategy to generate a desensitization processing scheme containing rule application order and strategy effectiveness conditions.

[0130] After obtaining the customized replacement rule and the masking strategy, they need to be combined and packaged. During the combination and packaging process, the rule application order and the strategy effectiveness conditions need to be determined.

[0131] The rule application order specifies whether to use the customized replacement rule first or the masking strategy first during the desensitization process, or how to alternate the use of the two. The strategy effectiveness condition specifies under what circumstances the corresponding rule and strategy should be applied. For example, for some specific types of sensitive information, the masking strategy should only be applied when the set context conditions are met.

[0132] By combining the packaging process, the generated desensitization processing scheme contains detailed rule and policy information.

[0133] Step S150: Perform an effect verification operation on the desensitization processing scheme to generate a desensitization result verification report, which is used to evaluate the degree of preservation of the semantic integrity of the original data set after desensitization.

[0134] In order to ensure the effectiveness and rationality of the desensitization processing scheme, it is necessary to perform an effect verification operation and generate a desensitization result verification report.

[0135] Step S151: Apply the desensitization processing scheme to the original data set to perform simulated desensitization processing and generate a simulated desensitized data set.

[0136] Apply the generated desensitization processing scheme to the original data set to perform simulated desensitization processing. In the simulated desensitization processing, according to the rules and policies in the desensitization processing scheme, the sensitive information in the original data set is replaced and masked.

[0137] For example, for a sentence containing an ID number in the original data set, according to the customized replacement rules and masking policies in the desensitization processing scheme, the ID number is replaced and masked to generate a simulated desensitized sentence. After all the original data units are processed in the above manner, the simulated desensitized data set is obtained.

[0138] Step S152: Call the pre-trained semantic similarity evaluation model to perform semantic similarity calculation processing on the original data set and the simulated desensitized data set to generate a semantic similarity score.

[0139] Step S1521: Perform data alignment processing on the original data set and the simulated desensitized data set to ensure that each data unit has a one-to-one correspondence in the two sets.

[0140] Before performing semantic similarity calculation, data alignment processing is required on the original data set and the simulated desensitized data set. The purpose of data alignment is to ensure that the data units in the original data set and the simulated desensitized data set can be one-to-one corresponding, so as to accurately compare their semantics.

[0141] For example, for each data unit in the original data set, find the corresponding desensitized data unit in the simulated desensitized data set. Data alignment can be performed through some unique identification information of the data unit, such as timestamp, data source, etc.

[0142] Step S1522: Extract the text content of each corresponding data unit, input the two input channels of the semantic similarity evaluation model respectively, perform feature encoding processing, and generate the original data feature vector and the desensitized data feature vector.

[0143] For the aligned data unit, extract its text content. Input the text content of the original data unit into one input channel of the semantic similarity evaluation model, and input the text content of the corresponding simulated desensitized data unit into the other input channel.

[0144] The semantic similarity evaluation model will perform feature encoding processing on the input text content. Through feature encoding, the text content is converted into the form of a feature vector. The text content of the original data unit is encoded into an original data feature vector, and the text content of the simulated desensitized data unit is encoded into a desensitized data feature vector.

[0145] Step S1523: Calculate the cosine similarity value of the original data feature vector and the desensitized data feature vector through the cosine similarity calculation layer of the semantic similarity evaluation model.

[0146] The cosine similarity calculation layer of the semantic similarity evaluation model can calculate the cosine similarity value of the original data feature vector and the desensitized data feature vector. Cosine similarity is a commonly used method for measuring the similarity between vectors, which reflects the similarity between two vectors by calculating the cosine value of the included angle between them.

[0147] The closer the cosine similarity value is to 1, the more similar the two vectors are, that is, the closer the semantics of the original data unit and the simulated desensitized data unit are; the closer the cosine similarity value is to 0, the less similar the two vectors are, and the greater the semantic difference.

[0148] Step S1524: Standardize the cosine similarity value to generate a corresponding standardized similarity index.

[0149] In order to facilitate comparison and analysis, the calculated cosine similarity value needs to be standardized. Standardization can convert the cosine similarity value to a unified range and generate a corresponding standardized similarity index.

[0150] The method of standardization can be selected according to specific needs and model settings. For example, the cosine similarity value can be mapped to the range of 0 to 100 to generate a standardized similarity index, which can more intuitively represent the semantic similarity between the original data unit and the simulated desensitized data unit.

[0151] Step S1525: Perform arithmetic average processing on the standardized similarity indexes of all corresponding data units to generate a semantic similarity score reflecting the overall semantic similarity degree.

[0152] The standardized similarity indicators of all corresponding data units are arithmetically averaged to obtain a semantic similarity score reflecting the overall semantic similarity degree. The semantic similarity score can comprehensively evaluate the semantic similarity degree between the simulated de-identification data set and the original data set.

[0153] If the semantic similarity score is high, it means that the de-identification process has better preserved the semantic integrity of the original data while protecting sensitive information; if the score is low, the de-identification processing scheme needs to be further adjusted to improve the semantic preservation degree.

[0154] Step S153: Extract the residual information of the de-identification processing position in the simulated de-identified data set, analyze whether the residual information contains clues that can infer the original sensitive information, and generate a residual information risk assessment result.

[0155] In the simulated de-identified data set, the residual information of the de-identification processing position needs to be extracted. Residual information refers to the part of sensitive information or information related to sensitive information that remains after de-identification processing.

[0156] The residual information is analyzed to see if it contains clues that can infer the original sensitive information. This process needs to be considered from multiple angles such as semantics, logic, and context. For example, when de-identifying an ID number, if only part of the number is replaced, and the remaining number combination may be associated with information such as a specific region or birth year, there is a risk of inferring the original sensitive information from these residual information.

[0157] The first step of analysis is to perform semantic analysis on the residual information. Identify the key elements in the residual information and determine whether they have a specific semantic direction. For example, the residual letters or numbers may correspond to certain industry codes, region identifiers, etc. Then, in combination with the context information, view the association between the residual information and the surrounding text. If the context can provide more background information, the residual information is more likely to be used to infer the original sensitive information. For example, in a text about employee information, the residual information after de-identification may be associated with the department and position of the employee, thereby increasing the risk of information leakage.

[0158] The logical relationship between residual information also needs to be considered. Some residual information may not have obvious directionality when viewed individually, but when combined, they may form a logical chain, revealing the original sensitive information. For example, the residual date and part of the number may jointly point to a specific transaction record.

[0159] According to the above analysis, the residual information risk assessment result is generated. The assessment result can be divided into different levels, such as high risk, medium risk and low risk. High risk means that the residual information is extremely likely to be used to infer the original sensitive information, and the desensitization processing scheme needs to be re-adjusted; medium risk indicates that there is a certain inference possibility, and the scheme needs to be further reviewed and optimized; low risk means that the residual information basically will not lead to the leakage of the original sensitive information, but still needs to be vigilant.

[0160] Step S154: Count the number of positions in the simulated desensitized data set where the sentence structure is broken or the logic is not coherent due to desensitization processing, and generate a semantic coherence damage assessment result.

[0161] During the simulation of desensitization processing, the sentence structure may be broken or the logic may be not coherent due to the replacement or masking of sensitive information. In order to evaluate this situation, the number of positions in the simulated desensitized data set where such problems occur needs to be counted.

[0162] First, each data unit in the simulated desensitized data set is analyzed sentence by sentence. Check the grammatical structure of the sentence to see if there are problems such as incomplete components, improper collocation, etc. caused by desensitization processing. For example, when replacing sensitive information, if the key noun in the sentence is replaced, the subject or object of the sentence may be missing, causing the sentence structure to be broken.

[0163] At the same time, the logical relationship between sentences is analyzed. Logical incoherence may manifest as unclear causality, unreasonable transitions, etc. For example, in a discussion, because the desensitization processing changes the key information, the logical deduction between the previous and subsequent sentences cannot be established.

[0164] When counting the number of positions, each position where the problem occurs needs to be accurately recorded. It can be counted by sentence or accurately to the specific word position. According to the statistical result, the semantic coherence damage assessment result is generated. The assessment result can reflect the degree of influence of desensitization processing on the semantic coherence of the data unit. If the number of positions is large, it means that the semantic coherence damage is large, and the desensitization processing scheme needs to be adjusted to ensure that the data can still maintain reasonable semantic expression after desensitization.

[0165] Step S155: Comprehensive analysis and processing of the semantic similarity score, residual information risk assessment result and semantic coherence damage assessment result to generate a desensitization result verification report containing assessment index values and risk level descriptions.

[0166] Comprehensive analysis and processing is to integrate the semantic similarity score, residual information risk assessment result and semantic coherence damage assessment result, and this process needs to consider the mutual relationship and weight of each assessment index.

[0167] Firstly, according to the semantic similarity score, it can be understood that the semantic similarity degree of the desensitized data set and the original data set. The higher the semantic similarity score, the better the semantic preservation. The residual information risk assessment result reflects the possibility of original sensitive information leakage caused by residual information after desensitization processing. The higher the risk level, the more likely there is a security risk in the desensitization scheme. The semantic coherence damage assessment result reflects the influence of desensitization processing on data semantic coherence. The more the position quantity, the more serious the semantic coherence damage.

[0168] In the comprehensive analysis, each evaluation index can be assigned a corresponding weight. The allocation of weights is determined according to the specific application scene and security requirements. For example, in the scene with high security requirements, the weight of the residual information risk assessment result can be relatively high; while in the scene with high semantic integrity requirements, the weights of the semantic similarity score and the semantic coherence damage assessment result can be appropriately increased.

[0169] According to the weights and the values of each evaluation index, a comprehensive calculation is performed. The specific method of calculation can be selected according to the actual situation, such as weighted summation. Through comprehensive calculation, a comprehensive evaluation index value is obtained.

[0170] According to the comprehensive evaluation index value, combined with the pre-set risk level division standard, the final risk level is determined. The risk level can be divided into multiple levels, such as excellent, good, medium, poor, and bad. Each risk level corresponds to different risk descriptions, such as excellent, which means that the desensitization processing well preserves the semantic integrity and coherence of the original data while protecting the sensitive information; poor means that the desensitization processing has serious problems and needs to be comprehensively modified and optimized.

[0171] The comprehensive evaluation index value and the risk level description are arranged into a desensitization result verification report. The report can also include detailed analysis and explanation of each evaluation index, as well as improvement suggestions for the current desensitization processing scheme, which can provide important reference for subsequent desensitization processing and ensure that the data can meet the security requirements after desensitization and maintain high usability.

[0172] Figure 2 The figure shows an exemplary hardware and software components of the network security sensitive information desensitization system 100 for data leakage provided by some embodiments of the present application, which can implement the idea of the present application. For example, the processor 120 can be used in the network security sensitive information desensitization system 100 for data leakage, and used to execute the functions in the present application.

[0173] The network security sensitive information desensitization system 100 for data leakage can be a general server or a special-purpose server, both of which can be used to implement the network security sensitive information desensitization method for data leakage of the present application. The present application only shows one server, but for the sake of convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0174] For example, the network security sensitive information desensitization system 100 for data leakage can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, a ROM, or a RAM, or any combination thereof. Exemplarily, the network security sensitive information desensitization system 100 for data leakage can also include program instructions stored in a ROM, a RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The network security sensitive information desensitization system 100 for data leakage also includes an I / O interface 150 between the computer and other input / output devices.

[0175] For the sake of illustration, only one processor is described in the network security sensitive information desensitization system 100 for data leakage. However, it should be noted that the network security sensitive information desensitization system 100 for data leakage in the present application can also include multiple processors, so the steps performed by one processor described in the present application can also be jointly performed or separately performed by multiple processors. For example, if the processor of the network security sensitive information desensitization system 100 for data leakage performs steps A and B, it should be understood that steps A and B can also be jointly performed by two different processors or separately performed in one processor. For example, a first processor performs step A, a second processor performs step B, or the first processor and the second processor jointly perform steps A and B.

[0176] In addition, the present application also provides a readable storage medium, in which computer executable instructions are pre-set, and when a processor executes the computer executable instructions, the network security sensitive information desensitization method for data leakage is implemented.

[0177] It should be noted that, in order to simplify the description of the present application and to help understand one or more embodiments of the present application, in the foregoing description of embodiments of the present application, various features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. A method for de-identifying sensitive network security information in response to data leakage, characterized in that, The method includes: Obtain the raw data set to be processed, which contains multiple data units, each of which consists of text content and metadata information; Extract multi-dimensional contextual features from the original dataset, wherein the multi-dimensional contextual features include semantic association features of data units and temporal association features between data units; A pre-trained sensitive information identification model is invoked to perform sensitive information localization processing on the multi-dimensional context features, generating a sensitive information identifier set for the original data set. The sensitive information identifier set is used to indicate the specific location and type of data unit that needs to be desensitized. Based on the set of sensitive information identifiers, a corresponding desensitization processing scheme is generated. The desensitization processing scheme includes replacement rules and masking strategies for different types of sensitive information. Perform an effectiveness verification operation on the aforementioned desensitization processing scheme to generate a desensitization result verification report. The desensitization result verification report is used to evaluate the degree to which the desensitized data set retains the semantic integrity of the original data.

2. The method for desensitizing network security sensitive information in response to data leakage as described in claim 1, characterized in that, The extraction of multi-dimensional contextual features from the original dataset, wherein the multi-dimensional contextual features include semantic association features of data units and temporal association features between data units, including: Each data unit in the original dataset is segmented into words to obtain a segmentation result set containing basic vocabulary units and combined phrase units; The pre-trained context encoder is invoked to perform semantic encoding processing on the word segmentation result set, generating a semantic vector representation for each data unit. The semantic vector representation is used to reflect the semantic relationship between words within the data unit. Analyze the metadata information of the data units in the original dataset, and extract the timestamp and source identifier information of the data units; Based on the timestamp, a temporal sequence of data units is constructed, and the interval duration parameter between adjacent data units in the temporal sequence is calculated. Based on the interval duration parameter and the source identification information, a temporal correlation feature is generated between data units. The temporal correlation feature is used to reflect the degree of correlation between different data units in the time dimension and the source dimension. The semantic vector representation and the temporal correlation features are fused to obtain multi-dimensional context features with a unified dimensional representation.

3. The method for desensitizing network security sensitive information in response to data leakage as described in claim 2, characterized in that, The step of performing word segmentation on each data unit in the original dataset yields a word segmentation result set containing basic vocabulary units and combined phrase units, including: The text content of the data unit is segmented using punctuation marks to generate a text segment set composed of sentence units; Each sentence unit is segmented into individual characters to generate an initial candidate set of word segments containing a continuous sequence of individual characters; The initial word segmentation candidate set is subjected to lexical matching processing based on a pre-built domain dictionary to identify basic lexical units that match the dictionary entries; For single-character sequences that do not match basic vocabulary units, adjacent character combination analysis is performed to extract two-character or multi-character combinations with semantic coherence as combined phrase units; The basic vocabulary units and the combined phrase units are arranged and combined according to their order of appearance in the sentence units to generate a set of word segmentation results containing position index information.

4. The method for desensitizing network security sensitive information in response to data leakage as described in claim 1, characterized in that, The process of calling a pre-trained sensitive information recognition model to perform sensitive information localization processing on the multi-dimensional context features, generating a set of sensitive information identifiers for the original dataset, including: The multi-dimensional context features are input into the feature input layer of the sensitive information recognition model, and feature dimension alignment processing is performed to generate a standardized feature vector that meets the model input requirements. The semantic understanding module of the sensitive information identification model performs deep semantic parsing on the standardized feature vector to extract potential sensitive fragments containing personal identity information, financial account information and privacy communication information in the data unit; The location module of the sensitive information identification model is used to perform coordinate labeling on the start and end positions of the potential sensitive segments in the data unit, generating a set of sensitive segment position coordinates. The type classification module of the sensitive information identification model is invoked to perform type discrimination processing on the potential sensitive segments, and to determine the sensitive information type corresponding to each sensitive segment; The set of sensitive segment location coordinates and the corresponding sensitive information types are associated and mapped to generate a set of sensitive information identifiers containing location coordinates and type identifiers.

5. The method for desensitizing network security sensitive information in response to data leakage as described in claim 4, characterized in that, The semantic understanding module of the sensitive information identification model performs deep semantic parsing on the standardized feature vector to extract potentially sensitive fragments containing personal identity information, financial account information, and private communication information from the data unit, including: The standardized feature vectors are input into the word embedding layer of the semantic understanding module to generate word vector representations for each lexical unit; The word vector representation is processed by the bidirectional long short-term memory network layer of the semantic understanding module to perform temporal dependency modeling, thereby generating context word vectors that contain semantic relationships between the preceding and following text. The attention mechanism layer of the semantic understanding module is invoked to perform importance weight allocation on the context word vectors, generating attention weight values ​​that reflect the semantic contribution of lexical units in the sentence; The context word vectors are weighted and aggregated based on the attention weight values ​​to generate sentence-level semantic representation vectors. The semantic representation vector is used to perform similarity matching with the preset sensitive information feature template, and the combination of word units with a matching degree exceeding the preset standard is extracted as potential sensitive segments. The step of performing temporal dependency modeling on the word vector representation through the bidirectional long short-term memory network layer of the semantic understanding module to generate contextual word vectors containing semantic relationships between the preceding and following text includes: The word vector representations are input into the forward propagation sublayer of the bidirectional long short-term memory network layer in the order in which the vocabulary units appear in the sentence, to generate a forward hidden state sequence that reflects the semantic information of the preceding text. The word vector representations are input into the backward propagation sublayer of the bidirectional long short-term memory network layer in reverse order of the lexical units in the sentence to generate a backward hidden state sequence that reflects the semantic information of the following text. The forward hidden state sequence and the backward hidden state sequence are added element by element to generate a combined hidden state sequence that integrates the semantic information of the preceding and following contexts. The combined hidden state sequence is input into the output layer of the bidirectional long short-term memory network and linearly transformed to generate a context word vector sequence with the same dimension as the word vector representation. The context word vector corresponding to each lexical unit is extracted from the context word vector sequence as the final output.

6. The method for desensitizing network security sensitive information in response to data leakage as described in claim 1, characterized in that, The step involves generating a corresponding de-identification scheme based on the set of sensitive information identifiers. The de-identification scheme includes replacement rules and masking strategies for different types of sensitive information, including: The sensitive information types in the set of sensitive information identifiers are analyzed, and classification identifiers for personal identity information types, financial account information types, and privacy communication information types are extracted. For each category identifier, a basic replacement rule from a preset desensitization rule library is matched. The basic replacement rule includes character replacement patterns and length retention requirements. Analyze the location coordinates of sensitive segments in the set of sensitive information identifiers to determine the contextual information of the sensitive segments in the data unit; The basic replacement rules are adaptively adjusted based on the contextual information to generate customized replacement rules that conform to the semantic coherence of the context. For sensitive segments containing consecutive numbers or letters, a masking strategy based on a fixed-length mask is generated, wherein the masking strategy includes the mask symbol type and the mask coverage ratio; The customized replacement rules and the masking strategy are combined and encapsulated to generate a de-identification processing scheme that includes the rule application order and the policy effective conditions.

7. The method for desensitizing network security sensitive information in response to data leakage as described in claim 6, characterized in that, The step of adaptively adjusting the basic replacement rules based on the contextual information to generate customized replacement rules that conform to the semantic coherence of the context includes: Extract the preceding and following words of the sensitive segment in the data unit, and construct a context window containing the preceding and following words; Analyze the semantic categories of words in the context window to determine the grammatical structure type of the sentence containing the sensitive segment; The grammatical function of the sensitive segment in the sentence is determined based on the grammatical structure type, and the grammatical function includes subject, object or modifier; For sensitive segments that serve as the subject, the character replacement pattern in the basic replacement rules is adjusted to retain the lexical attribute features that match the grammatical function of the subject; For sensitive segments used as objects, the length retention requirements in the basic replacement rules have been adjusted to ensure the appropriate collocation of the replaced words with the predicate verb; For sensitive segments that function as modifiers, the type of replacement symbol in the basic replacement rules is adjusted to maintain the semantic modification relationship between the modifier and the headword; The adjusted character replacement pattern, length retention requirements, and replacement symbol types are integrated into a set of rules to generate customized replacement rules.

8. The method for desensitizing network security sensitive information in response to data leakage as described in claim 1, characterized in that, The process of verifying the effectiveness of the de-identification scheme generates a de-identification result verification report. This report assesses the degree to which the de-identified dataset retains the semantic integrity of the original data, including: The original dataset is subjected to simulated desensitization processing using the aforementioned desensitization processing scheme to generate a simulated desensitized dataset. A pre-trained semantic similarity evaluation model is invoked to perform semantic similarity calculation on the original dataset and the simulated de-identified dataset, generating a semantic similarity score. Extract residual information from the desensitized data set after simulated desensitization, analyze whether the residual information contains clues that can infer the original sensitive information, and generate residual information risk assessment results; The number of positions in the simulated desensitized data set where sentence structure breaks or logical incoherence are caused by desensitization processing is counted, and a semantic coherence impairment assessment result is generated. The semantic similarity score, residual information risk assessment result, and semantic coherence impairment assessment result are comprehensively analyzed and processed to generate a desensitization result verification report containing assessment index values ​​and risk level descriptions.

9. The method for desensitizing network security sensitive information in response to data leakage as described in claim 8, characterized in that, The pre-trained semantic similarity evaluation model is invoked to perform semantic similarity calculation on the original dataset and the simulated de-identified dataset, generating a semantic similarity score, including: The original data set and the simulated desensitized data set are aligned to ensure that each data unit has a one-to-one correspondence in the two sets. Extract the text content of each corresponding data unit, input it into the two input channels of the semantic similarity evaluation model, perform feature encoding processing, and generate the original data feature vector and the desensitized data feature vector; The cosine similarity value between the original data feature vector and the desensitized data feature vector is calculated through the cosine similarity calculation layer of the semantic similarity evaluation model. The cosine similarity values ​​are standardized to generate corresponding standardized similarity indices; The standardized similarity indices of all corresponding data units are averaged to generate a semantic similarity score that reflects the overall semantic similarity.

10. A network security sensitive information desensitization system for data leakage, characterized in that, The device includes a processor and a memory, the memory being connected to the processor. The memory is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the memory to implement the network security sensitive information desensitization method for data leakage as described in any one of claims 1-9.

Citation Information

Cited By

  • Method and system for safely storing population information of digital country

    CN121211509A

  • Clinical data management system, clinical data privacy protection method, equipment and medium

    CN121393705A