Detection method, system and equipment for structured sensitive data and medium
By performing token sequence processing and adaptive masking of attention weights on structured data, the problem of model performance degradation caused by label conflicts is solved, and the accuracy and robustness of sensitive data detection are improved.
Patent Information
- Application Number
- CN202511277839.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing sensitive data detection methods suffer from label conflict problems in structured data scenarios, which leads to decreased model detection performance. In particular, they are prone to misjudgment and attention drift in minority class detection tasks, making it difficult to identify low-quality samples, thus affecting detection accuracy and robustness.
By preprocessing structured data into token sequences, extracting contextual features, calculating attention weights and generating anomaly-aware masks, adaptively masking and normalizing, identifying and masking abnormal attention weights, and improving the robustness and accuracy of the model.
It achieves accurate identification and effective shielding of label conflicting samples, improves the detection accuracy and generalization ability of the model in complex data scenarios, and enhances the robustness and adaptability of the model.
Smart Images

Figure CN120805199A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data security, in particular to a feature anomaly perception and adaptive shielding technology for structured sensitive data detection, which is used to improve the accuracy of sensitive data identification and the robustness of the model in the presence of label conflict samples. BACKGROUND
[0002] With the acceleration of digitalization, data has become a national strategic resource and a key production factor. As a fundamental security issue in the digital age, data security is directly related to national security, enterprise compliance operation, and personal privacy protection, especially in high-risk industries such as finance, healthcare, and government affairs, where its protection capability is increasingly important. In recent years, frequent sensitive data leakage incidents have exposed the weak links in the detection and protection of sensitive data in information systems, and more effective detection and prevention methods are urgently needed.
[0003] Under this background, sensitive information detection in structured data has become one of the core technologies to ensure data compliance and security. Structured data widely exists in database systems and is stored in table form, covering various high-sensitive attributes such as ID numbers, phone numbers, addresses, and bank card numbers. Current mainstream detection methods can be roughly divided into two categories: one is the traditional rule-based method, such as regular expressions, keyword matching, and column name recognition, which has certain accuracy but lacks generalization ability for unknown types and context changes; the other is the automatic detection method based on deep learning, such as LSTM and BERT models, which have stronger semantic understanding and context modeling capabilities and have gradually become the research mainstream.
[0004] However, existing sensitive data detection methods still face key challenges in actual structured data scenarios, the most typical problem being the "label conflict problem" (LCP). When structured data is concatenated into a text sequence, there may be multiple sensitive types mixed in a sample, leading to ambiguous label semantics and causing the model to misjudge or attention drift during training. Such "perception shift" samples significantly interfere with the model learning process, reducing detection accuracy and generalization ability.
[0005] In addition, current methods generally treat all training samples "equally", making it difficult to identify "low-quality samples" in the training set that have annotation errors, semantic noise, or context confusion, especially in minority class detection tasks. Such samples are extremely likely to cause training process deviation and amplify incorrect learning signals, thus making the model unstable in actual deployment.
[0006] Therefore, there is an urgent need for a new technical solution for structured sensitive data detection, which can automatically identify and perceive abnormal tokens or abnormal samples during model training, and perform targeted processing and shielding, thereby reducing semantic interference, improving the adaptability of the model to complex samples, improving the accuracy and robustness of detection, and assisting the reliable landing of sensitive data identification technology in high-risk application scenarios. SUMMARY
[0007] Based on this, the present application provides a structured sensitive data detection method, system, device and medium to solve the defect of model detection performance decline caused by label conflict in the prior art.
[0008] In a first aspect, the present application provides a structured sensitive data detection method, comprising the following steps: Obtaining structured data and preprocessing into token sequences; Context feature extraction is performed on the token sequence to obtain a context feature vector for each token; Based on the context feature vector, the attention weight of each token is calculated to form an initial attention weight distribution; Abnormal perception is performed on the initial attention weight distribution to identify and quantify abnormal attention weights, and a final attention weight mask is generated; The final attention weight mask is applied to shield and normalize the initial attention weight distribution to obtain an adjusted attention weight; Based on the adjusted attention weight and the context feature vector, a context representation vector of the sample is calculated; According to the context representation vector, the structured data is classified and predicted for sensitive categories.
[0009] As an optional implementation of the first aspect of the present application, the step of performing abnormal perception on the initial attention weight distribution specifically includes: within a batch of model training, for each token position, the statistical features of the attention weights of all samples at the position in the batch are counted, the statistical features include the first quartile and the third quartile, and the interquartile range is calculated based on the first quartile and the third quartile; according to the attention weight, the first quartile, the third quartile and the interquartile range, the deviation degree of the attention weight of each sample at each token position is calculated; the deviation degree is nonlinearly transformed to obtain a final deviation degree which is not sensitive to extreme values and contains deviation direction information.
[0010] As an optional implementation of the first aspect of the application, the step of the nonlinear transformation is: calculating a direction vector of the deviation degree; and the final deviation degree is obtained by multiplying the direction vector and the logarithmic value of the absolute value of the deviation degree plus 1, and the specific calculation formula is: wherein, is the final deviation degree, D is the deviation degree, sign() is a sign function, B is the number of training samples per batch, and L is the maximum length of tokens.
[0011] As an optional implementation of the first aspect of the application, the step of generating the final attention weight mask includes generating a first layer mask and a second layer mask; the first layer mask is a global token-level mask, which is generated by: calculating the variance of the final deviation degree at each token position within the batch; comparing the variance with a preset variance threshold, and marking the token position as to-be-processed in the first layer mask when the variance is greater than the variance threshold.
[0012] As an optional implementation of the first aspect of the application, the second layer mask is a sample-level mask, which is generated by: for the token positions marked as to-be-processed in the first layer mask, calculating the mean and standard deviation of the final deviation degree within the batch; setting dynamic upper and lower boundary thresholds based on the mean and standard deviation, wherein the upper boundary threshold is the sum of the mean and a first coefficient multiplied by the standard deviation, and the lower boundary threshold is the difference between the mean and a second coefficient multiplied by the standard deviation; for each sample, if the final deviation degree at the to-be-processed token position is greater than the upper boundary threshold or less than the lower boundary threshold, the corresponding position in the second layer mask is marked; and the final attention weight mask is obtained by performing a logical AND operation on the first layer mask and the second layer mask.
[0013] As an optional implementation of the first aspect of the application, the step of applying the final attention weight mask to mask and normalize the initial attention weight distribution specifically includes: according to the final attention weight mask, setting the corresponding value of the identified abnormal attention weight in the initial attention weight distribution to zero to form a masked attention weight; calculating the sum of the masked attention weights of each sample; if the weight sum of a certain sample is zero, replacing the masked attention weight of the sample with a uniform distribution to prevent division by zero error in subsequent normalization operation; and normalizing the attention weight after masking and zero-sum processing to obtain the adjusted attention weight.
[0014] As an optional implementation of the first aspect of the application, the step of obtaining structured data and preprocessing into a token sequence specifically comprises: reading data from the structured table by column, splicing multiple cell data in the same column into a field set; connecting the data in the field set using a preset delimiter to form a string sample, and if some column data is missing, using a preset placeholder to make up; performing character-level segmentation on the string sample to obtain an initial token sequence; and truncating or padding the initial token sequence to a preset maximum sequence length to obtain the token sequence.
[0015] In a second aspect, the embodiments of the application provide a structured sensitive data detection system, comprising: a preprocessing module configured to obtain structured data and preprocess into a token sequence; a feature extraction module configured to extract context features of the token sequence to obtain a context feature vector of each token; an adaptive feature masking and context representation construction module configured to calculate attention weights of each token based on the context feature vector to form an initial attention weight distribution, perform anomaly perception on the initial attention weight distribution to identify and quantify abnormal attention weights, and generate a final attention weight mask; apply the final attention weight mask to mask and normalize the initial attention weight distribution to obtain adjusted attention weights; and calculate a context representation vector of a sample based on the adjusted attention weights and the context feature vector; a classification prediction module configured to perform sensitive category classification prediction on the structured data according to the context representation vector.
[0016] In a third aspect, the embodiments of the application provide an electronic device, which comprises a processor, a memory, and a program or instructions stored on the memory and executable on the processor, and the program or instructions are executed by the processor to implement the steps of the method according to the first aspect.
[0017] In a fourth aspect, the embodiments of the application provide a readable storage medium, which stores a program or instructions, and the program or instructions are executed by a processor to implement the steps of the method according to the first aspect.
[0018] Compared with the prior art, the application has the following beneficial effects: 1. Precise anomaly perception: by introducing a deviation degree modeling based on quartile range, the abnormal degree of attention weight can be accurately quantified, and attention drift caused by label conflict can be effectively identified.
[0019] 2. Hierarchical adaptive masking: An innovative two-layer masking mechanism that realizes adaptive masking from coarse-grained (identify problematic tokens) to fine-grained (locate abnormal samples), which is more flexible and robust than single-threshold methods.
[0020] 3. Enhanced model robustness: By masking abnormal attention and designing a robust normalization strategy, the interference of noisy samples on model training is effectively suppressed, significantly improving the detection accuracy and generalization ability of the model in complex and noisy data scenarios.
[0021] 4. Good integration and practicality: This technology is completely realized inside the model through statistical analysis, without additional manual labeling or complex network structure modification, easy to integrate into existing deep learning detection frameworks, with high practical application value. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is a flowchart of a structured sensitive data detection method according to an embodiment of the present application; Figure 2 is a technical roadmap of a structured sensitive data detection method according to an embodiment of the present application; Figure 3 is a structural schematic diagram of a structured sensitive data detection system according to an embodiment of the present application.
[0023] The following specific embodiments will further illustrate the present application in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0025] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise explicitly specified.
[0026] Embodiment 1 Please refer to Figure 1 andFigure 2 FIG. 1 and FIG. 2 are respectively a flowchart and a technical roadmap of a structured sensitive data detection method provided by an embodiment of the present application. The method can include the following steps: S1: Obtain structured data and preprocess into token sequences; The goal of this step is to convert structured table data into a text sequence format suitable for model input, and construct training samples. Specifically, the following operations are included: Step 1.1 Column sample splicing construction: Read data from the original table and organize samples by column. Each sample is composed of multiple (e.g., 5) cells in the same column spliced together.
[0027] For example, a column contains the following five rows of data: Beijing ID 7458 Chaoyang District No. 12 Postal code 100022.
[0028] Step 1.2 String sequence construction: Use the English comma “,” as a delimiter to splice the above fields and construct a string sample (text_sample) where, represents the cell (Cell) content in the structured data table, i.e., the specific data at position i, and n represents the number of cells participating in splicing, i.e., the sample string is composed of n cells combined in order.
[0029] For example, text_sample = "Beijing, ID 7458, Chaoyang District, No. 12, Postal code 100022".
[0030] If a column is missing, use "0" to fill it in to ensure that the number of spliced columns is consistent.
[0031] Step 1.3 Character-level tokenization processing: Use a character-level tokenizer Split the spliced string by character to get a tokenization vector where, represents each word after tokenizer processing, and the tokenization process is as follows: where, is the number of tokens in the sequence, and here the maximum token sequence length is set to For example: tk = ['Beijing', 'ID', '7458', 'Chaoyang', 'District', '1', '2', 'No.', '1', '0', '0', '0', '2', '2'].
[0032] S2: Context feature extraction is performed on the token sequence to obtain a context feature vector of each token; The goal of this step is to convert the token sequence generated by splicing into a vector representation that can reflect its context semantics. The specific operation is as follows: Step 2.1 Index mapping and padding processing: The obtained word segmentation vector , each word is mapped according to a fixed vocabulary , and the index mapping operation is performed , the index position of the mapped character is given for the character existing in the vocabulary, and for the character <unk>Index position is recorded as 1, and an index mapping vector is generated The index mapping processing procedure is shown as follows: wherein represents the i-th token after tokenization, represents the index position of the token in the vocabulary, represents the index position value of the i-th token.
[0033] An index mapping vector is obtained After that, for the index vector that does not satisfy the maximum length , the vector needs to be filled with 0 representing the padding character <pad>The index of the token sequence is filled in the processing process as follows: For example, the sample token sequence is: "North" → 12, "Capital" → 15, "ID" → 28, "7458" → 3, "morning" → 46,..., "2" → 9 The mapped index sequence is obtained as follows: [12, 15, 28, 3, 46, 52, 36, 3, 19, 31, 44, 3, 7, 9, 60, 3, 23, 45, 7, 8, 8, 8, 9, 9] Since the model is set to have a maximum token sequence length of 128 and the current input length is 24, the end needs to be filled in: The filled index sequence = [12, 15, 28, 3,..., 9, 9, 0, 0,..., 0] (total 128 bits) where 0 represents <pad>Position.
[0034] Step 2.2 Word Embedding Mapping: For the padded vector , we perform embedding operation , the main process is to look up the parameter matrix by padding the vector to obtain the embedding vector , and finally obtain the embedding matrix , the embedding process is represented as follows: For example: "North" → [0.12, -0.03,..., 0.98] "Beijing" → [0.22, 0.01,..., 0.91] ", " → [0.07, -0.02,..., 0.61] "1" → [0.09, -0.08,..., 0.83] "0" → [0.15, -0.05,..., 0.76] ... After embedding, we get an embedding matrix of size: (128 × 128), representing a total of 128 characters (including padding), each character is represented as a 128-dimensional vector.
[0035] Step 2.3 Context Feature Extraction (LSTM Layer): This step sends the entire token sequence into the LSTM layer for processing. That is, after obtaining the embedding matrix , the hidden layer (hidden_output) matrix is calculated through the LSTM layer, and the generated hidden layer matrix is represented as follows: Through the above steps, the original character sequence is efficiently converted into a semantic representation with context awareness, providing rich input features for subsequent attention modeling, anomaly detection and classification.
[0036] For example, if the LSTM output dimension is set to 256, the output shape is: (128 × 256), representing each token being encoded as a 256-dimensional context vector.
[0037] ("North") -> [0.18, 0.35,..., 0.44] ("Beijing") -> [0.21, 0.39,..., 0.53] ("ID") -> [0.27, 0.41,..., 0.67] ... ( <pad>) -> [0.00, 0.00, …, 0.00] Finally, the context feature tensor of the sample is obtained: hidden_output= ∈R 128×256 .
[0038] S3: Calculate the attention weight of each token based on the context feature vector to form an initial attention weight distribution; perform anomaly perception on the initial attention weight distribution, identify and quantify the abnormal attention weight, and generate a final attention weight mask; apply the final attention weight mask to shield and normalize the initial attention weight distribution to obtain the adjusted attention weight; calculate the context representation vector of the sample based on the adjusted attention weight and the context feature vector; The core of this step is to identify abnormal points in the attention distribution and automatically construct a mask to shield interfering tokens. Specifically, the following steps are included: Step 3.1: Calculate token attention weight: Map the context vector of each token to a scalar score through a linear layer.
[0039] Normalize using the Softmax operation to obtain the attention weight of each token (soft_attn).
[0040] For example, for a sample with a token sequence length of 128, its partial attention weights are as follows: soft_attn = [0.00149, 0.00221, 0.00175, …, 0.00092, 0.00325, 0.00118] Where the attention weight of the 10th token is: soft_attn
[10] = 0.00325 Step 3.2: Anomaly perception of token attention weight distribution During training, each token of all samples in each batch forms a sample space, and we need to calculate the , and features of each token; The quantile(A, p, dim=0) function calculates the p-th quantile of matrix A in the specified dimension. dim represents the dimension when calculating quantiles, dim=0 represents calculating in the 0th dimension, that is, calculating the statistics of attention weights of the same token position in different samples; the statistical method is represented as follows: For each token of all samples in a batch, we calculate the deviation degree according to the statistical features The deviation degree calculation formula is defined as follows: Wherein represents the deviation degree of the jth feature of the ith sample, represents the value of the ith sample in the jth feature dimension, represents the interquartile range of the jth feature dimension.
[0041] According to the analysis of LCP Optimization Analysis, for LCP samples, the abnormal weight values are distributed upwards and downwards, resulting in a large variance, while the abnormal weight values of normal samples are almost distributed upwards, so we introduce a direction vector to measure the deviation degree of abnormal samples when calculating the deviation degree, and the direction vector is defined as follows: Considering the calculated deviation degree , the direction vector and the need to prevent extreme deviation degree values from dominating subsequent statistical calculations, we define the following nonlinear replacement to calculate a more representative deviation degree , which is represented as follows: Wherein, is the final deviation degree, D is the deviation degree, sign( ) is the sign function, B is the number of training samples per batch, which is set to a constant value of 80, L is the maximum length of tokens, which is set to a constant value of 128.
[0042] For example, the relevant statistics calculated for the 10th token are as follows: Deviation degree D
[10] = 1.208964 E (average of the sample deviation degree) = 0.1545735 std (standard deviation of the sample deviation degree) = 0.349563 Step 3.3 First layer mask generation First, we calculate the variance features of the deviation degree samples and mean characteristics The statistical method is represented as follows: Where mean() represents the operation of calculating the average value according to the specified dimension.
[0043] According to the weight analysis result, we found that the variance of the deviation of normal samples and LCP samples is very different, and the variance of LCP samples is generally higher than that of normal samples, so we set the hyperparameter variance threshold to automatically generate a global mask for each token in each batch , and the first layer mask generation is represented as follows: For example, the first layer mask is = [1,1,0,1,0,...,1] Step 3.4 Second layer mask generation In the second layer mask, we conduct more fine-grained sample-level processing on the deviation weight of the token filtered out by the first layer mask, and we filter out the samples whose weight deviation of each token in each sample is higher than and lower than and set them to 1 to generate the second layer mask , and the second layer mask generation is represented as follows: Where the boundary values and are defined as follows: Finally, the generated first layer mask and the second layer mask are subjected to operation (AND operation in logical operation) to obtain the final mask , and the generation of the final mask is represented as follows: For example, with the current sample, the token whose deviation is greater than E + 0.5 × std ≈ 0.329 will be marked as abnormal. If the 10th deviation is 1.208964, it meets the condition:
[10] = 1 The masks of other positions are as follows (partially): = [[0, 0,..., 1,..., 0], [0, 1,..., 0,..., 1], … [1, 1,..., 0,..., 0]] ∈ R B×L The final mask is: = & = [[0, 0,..., 1,..., 0], [0, 1,..., 0,..., 1], … [1, 1,..., 0,..., 0]] ∈R B×L .
[0044] Step 3.5 Attention Masking and Normalization: After obtaining the final mask , we perform the following mask dropout operation on the weight matrix : 1. When the current bit is 1, we discard the current weight; 2. When the current bit is 0, we keep the current weight.
[0045] The overall mask dropout operation is represented as follows: For example, the original attention vector is: [0.01, 0.02, 0.05, 0.08, 0.01,..., 0.00] The mask is: [0, 0, 1, 0, 1,..., 0] The reserved result after the mask is: [0.01, 0.02, 0.00, 0.08, 0.00,..., 0.00] For the generated mask dropout matrix , we first calculate the token weight sum for each sample, and the calculation process is represented as follows: For the zero-sum case (i.e., the sum of all token values is 0), we set the overall distribution to a uniform distribution, which prevents the occurrence of division by zero anomalies. For normal cases, the original weight value is retained, and the processing process is represented as follows: In this way, a modified mask drop matrix is obtained, and the sum of the weights of each sample token is recalculated on this matrix to ensure that the sum of the weights will not be 0. For example, if the weights of all tokens of a sample are 0 after masking, that is: [0.0, 0.0, ..., 0.0] (128 dimensions in total) Now replace it with a uniform distribution: [1 / 128, 1 / 128, ..., 1 / 128].
[0046] At this time, normalization operation is performed to ensure that the overall weight sum is 1 and the overall expectation remains unchanged, generating the final weight matrix , the normalization operation is expressed as follows: For example, the weights retained after dropping outliers are: [0.01, 0.02, 0.00, 0.08, 0.00, ..., 0.04] The sum of the weights is 0.15, which after normalization becomes: [0.0667, 0.1333, 0.00, 0.5333, 0.00, ..., 0.2667] (round to four decimal places).
[0047] Use the new weight matrix And the output matrix of the LSTM layer Perform matrix multiplication to calculate the context vector , the process is as follows: For example: context = [0.2532, 0.7649, 0.9984, ..., 1.4335] (256 dimensions in total) For example, the token weight of a sample before masking is: [0.05, 0.1, 0.2, 0.15, 0.05, ..., 0.0] After masking and normalization, it becomes: [0.0, 0.15, 0.3, 0.25, 0.0, ..., 0.0].
[0048] The corresponding context vector is updated to focus more on the direction of the key token, and the final classification result is more discriminative.
[0049] S4: Perform sensitive category classification prediction on the structured data according to the context representation vector.
[0050] Get the context vector After that, through a fully connected layer, the number of categories to be classified is , as the final output of the entire model , and is a parameter, and the process is expressed as follows: Through all the above processes, the text sequence vector is successfully converted into a dimension with the number of categories. The output vector , for the predicted category, directly pass the vector through a layer, the maximum value is the predicted category, where is the natural base, The prediction process for any category is expressed as follows: in represents the predicted category label of the i-th sample, Indicates taking the category index that maximizes the value in the brackets. represents the kth element of the vector, that is, the prediction score of the i-th sample in the k-th category. Indicates the exponential operation of the prediction score in the softmax function.
[0051] That is, the context vector is sent to the fully connected layer and Softmax for classification, and the sensitive category (such as phone number, email, address, etc.) to which it belongs is output.
[0052] For example: =[0.76, 1.23, 1.77, …, 2.11] =[0.09, 0.17, 0.20, …, 0.26] The largest value is selected as the output, and the output is mapped to the corresponding label. For example, if 0.26 is selected, the output classification category is Address.
[0053] The above is the complete implementation process of the method of the present invention. By introducing the anomaly perception mechanism of attention distribution and the adaptive masking strategy, it is possible to effectively suppress the interference of label conflicting samples in structured data, thereby improving the accuracy and robustness of sensitive data detection.
[0054] Example 2 See also Figure 3 , as shown in the structural schematic diagram of a structured sensitive data detection system proposed in the second embodiment of the present application, the system comprises the following key modules: A preprocessing module 100 is configured to obtain structured data and preprocess the structured data into a token sequence. A feature extraction module 200 is configured to perform context feature extraction on the token sequence to obtain a context feature vector of each token. An adaptive feature masking and context representation construction module 300 is configured to calculate an attention weight of each token based on the context feature vector to form an initial attention weight distribution, perform abnormality perception on the initial attention weight distribution to identify and quantify an abnormal attention weight, generate a final attention weight mask, mask and normalize the initial attention weight distribution by using the final attention weight mask to obtain an adjusted attention weight, and calculate a context representation vector of a sample based on the adjusted attention weight and the context feature vector. A classification prediction module 400 is configured to perform sensitive category classification prediction on the structured data according to the context representation vector.
[0055] The structured sensitive data detection system in the embodiments of the present application can be a device, a component in a terminal, an integrated circuit, or a chip. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a wearable device, an Ultra-mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a Personal Computer (PC), and the like, which are not limited in the embodiments of the present application.
[0056] The structured sensitive data detection system in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an IOS operating system, or other possible operating systems, which are not limited in the embodiments of the present application.
[0057] The structured sensitive data detection system provided in the embodiments of the present application can implement the method provided in the method embodiments. Figure 1 The processes implemented by the structured sensitive data detection method in the method embodiments are not repeated here to avoid repetition.
[0058] Optionally, the embodiments of the present application further provide an electronic device, comprising a processor, a memory, a program or instructions stored in the memory and executable in the processor, which, when executed by the processor, implement each process of the above-mentioned embodiment of the method for detecting structured sensitive data and achieve the same technical effects. To avoid repetition, details are not described herein.
[0059] The embodiments of the present application further provide a readable storage medium, which stores a program or instructions, which, when executed by a processor, implement each process of the above-mentioned embodiment of the method for detecting structured sensitive data and achieve the same technical effects. To avoid repetition, details are not described herein.
[0060] The processor is the processor of the electronic device in the above-mentioned embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, etc.
[0061] It should be noted that, in this document, the terms "comprise", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of another identical element in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to the order of performing functions as shown or discussed, but can also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to certain examples can be combined in other examples.
[0062] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and the necessary general hardware platform, of course, also can be through hardware, but many cases the former is the better embodiment. Based on such understanding, the technical solutions of the present application essentially or say the part of the contribution to the prior art can be embodied in the form of software products, the computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including a number of instructions to make a terminal (may be a mobile phone, computer, server, air conditioner, or network equipment, etc.) executes the method described in various embodiments of the present application.
[0063] The embodiments of the present application are described above in conjunction with the drawings, but the present application is not limited to the above-mentioned specific embodiments, the above-mentioned specific embodiments are only illustrative, but not limited, those skilled in the art can make many forms without departing from the purpose of the present application and the scope of the claims under the inspiration of the present application, all belong to the protection of the present application.< / pad> < / pad> < / pad> < / unk>
Claims
1. A method for detecting structured sensitive data, characterized in that: The following steps are involved: Get structured data and preprocess it into token sequences; Extracting context features from the token sequence to obtain a context feature vector for each token; Calculate the attention weight of each token based on the context feature vector to form an initial attention weight distribution; Performing abnormality perception on the initial attention weight distribution, identifying and quantifying abnormal attention weights, and generating a final attention weight mask; Applying the final attention weight mask to mask and normalize the initial attention weight distribution to obtain an adjusted attention weight; Calculating a context representation vector of the sample based on the adjusted attention weight and the context feature vector; A sensitive category classification prediction is performed on the structured data according to the context representation vector.
2. The method for detecting structured sensitive data according to claim 1, characterized in that: The step of performing abnormality perception on the initial attention weight distribution specifically includes: Within the batch of model training, for each token position, statistical features of the attention weights of all samples in the batch at that position are counted. The statistical features include the first quartile and the third quartile, and the interquartile range is calculated based on the first quartile and the third quartile. Calculate the deviation of the attention weight of each sample at each token position based on the attention weight, the first quartile, the third quartile, and the interquartile range; A nonlinear transformation is performed on the deviation to obtain a final deviation that is insensitive to extreme values and contains deviation direction information.
3. The method for detecting structured sensitive data according to claim 2, wherein: The steps of the nonlinear transformation are: Calculating a direction vector of the deviation; The final deviation is obtained by multiplying the direction vector by the logarithm of the absolute value of the deviation plus 1. The specific calculation formula is: in, is the final deviation, D is the deviation, sign() is the sign function, B is the number of training samples per batch, and L is the maximum length of token.
4. A method for detecting structured sensitive data according to claim 2 or 3, characterized in that: The step of generating the final attention weight mask includes generating a first layer mask and a second layer mask; The first-level mask is a global token-level mask, which is generated as follows: Calculate the variance of the final deviation at each token position within the batch; The variance is compared with a preset variance threshold, and when the variance is greater than the variance threshold, the token position is marked in the first layer mask as to be processed.
5. The method for detecting structured sensitive data according to claim 4, characterized in that: The second-level mask is a sample-level mask, which is generated as follows: For the token positions marked as to-be-processed by the first-layer mask, calculate the mean and standard deviation of their final deviation within the batch; Setting dynamic upper and lower thresholds based on the mean and standard deviation, wherein the upper threshold is the sum of the mean and a first coefficient multiplied by the standard deviation, and the lower threshold is the difference between the mean and a second coefficient multiplied by the standard deviation; For each sample, if its final deviation at the position of the token to be processed is greater than the upper boundary threshold or less than the lower boundary threshold, then mark the corresponding position of the second layer mask; The final attention weight mask is obtained by performing a logical AND operation on the first layer mask and the second layer mask.
6. The method for detecting structured sensitive data according to claim 1, characterized in that: The step of applying the final attention weight mask to mask and normalize the initial attention weight distribution specifically comprises: According to the final attention weight mask, the corresponding values of the identified abnormal attention weights in the initial attention weight distribution are set to zero to form masked attention weights; Calculate the sum of the attention weights after masking each sample; If the sum of the weights of a sample is zero, the masked attention weight of the sample is replaced with a uniform distribution to prevent division by zero errors in subsequent normalization operations; The masked and zero-summed attention weights are normalized to obtain the adjusted attention weights.
7. The method for detecting structured sensitive data according to claim 1, wherein: The step of obtaining structured data and preprocessing it into a token sequence specifically includes: Read data from a structured table by column, and concatenate multiple cell data in the same column into a field set; Use the preset delimiter to connect the data in the field set to form a string sample. If a column of data is missing, use the preset placeholder to fill it. Perform character-level segmentation on the string sample to obtain an initial token sequence; The initial token sequence is truncated or padded to a preset maximum sequence length to obtain the token sequence.
8. A detection system for structured sensitive data, characterized in that: include: The preprocessing module is used to obtain structured data and preprocess it into token sequences; A feature extraction module is used to extract context features from the token sequence to obtain a context feature vector for each token; An adaptive feature masking and context representation building module is used to calculate the attention weight of each token based on the context feature vector to form an initial attention weight distribution; perform abnormality perception on the initial attention weight distribution, identify and quantify abnormal attention weights, and generate a final attention weight mask; Applying the final attention weight mask to mask and normalize the initial attention weight distribution to obtain an adjusted attention weight; Calculating a context representation vector of the sample based on the adjusted attention weight and the context feature vector; A classification prediction module is used to perform sensitive category classification prediction on the structured data based on the context representation vector.
9. An electronic device, characterized in that: The system comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for detecting structured sensitive data as described in any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the method for detecting structured sensitive data as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Text similarity recognition method, system and equipment and storage medium
CN120012777A
Chinese named entity recognition system, method and equipment based on multi-scale features and medium
CN120409476A
Table-level data classification and grading method and system based on big data
CN120561755A
Automatically labeling data using natural language processing
EP4040330A1
Cited By
Alarm information processing method, system and device and computer storage medium
CN121193586A