Data processing method and device, equipment and storage medium
By using a self-attention transformer model to extract features and calculate attention for various types of sensitive data, the problem of low recognition efficiency and low accuracy in existing technologies is solved, and efficient and accurate recognition and sensitivity level determination of resources such as text, images, and audio are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 中移信息技术有限公司
- Filing Date
- 2022-12-26
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies for identifying sensitive data are relatively limited and cannot effectively identify multiple types of sensitive data, resulting in low identification efficiency and low accuracy.
The self-attention transformer model is used to extract features from the resources to be identified. By calculating the attention-hidden features of the resources, it can identify various types of sensitive data, including text, images, and audio. The relationship between resources is considered during the identification process.
It improves the efficiency of sensitive data identification, avoids errors in identifying relationships when identifying different types of resources separately, and can accurately identify and determine the sensitivity level of sensitive data, facilitating subsequent processing.
Smart Images

Figure CN115827870B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and specifically relates to a data processing method, apparatus, electronic device and storage medium. Background Technology
[0002] With the rapid development of information technology, the amount of data generated by people in production and daily life is growing exponentially. How to identify sensitive data in the massive amount of data and protect it has become an urgent issue.
[0003] While text recognition technology can identify sensitive text data in massive datasets, it cannot identify non-text sensitive data such as images and audio. Thus, the method for identifying sensitive data is relatively limited and cannot identify multiple types of sensitive data. Summary of the Invention
[0004] This application provides a data processing method, apparatus, device, and storage medium that can solve the problem that the existing technology has a relatively single method for identifying sensitive data, which makes it impossible to identify multiple types of sensitive data.
[0005] In a first aspect, embodiments of this application provide a data processing method, which may include:
[0006] Obtain the resource to be identified and its resource information. The resource to be identified includes N types of resources. The resource information includes the type identifier and position vector of each type of resource in the N types of resources, where N is an integer greater than 1.
[0007] Input the resource to be identified and the resource information into the sensitive data identification model. The sensitive data identification model extracts features from the resource to be identified to obtain hidden features of N types of resources.
[0008] Based on the hidden features of any two types of resources out of N types of resources, calculate the attention hidden features of each type of resource in any two types. The attention hidden features are used to characterize the attention distribution of the hidden features of one type of resource to the hidden features of the other type of resource in any two types of resources.
[0009] Based on the attention-hidden features of N types of resources, the identification results of the resources to be identified are output from the sensitive data identification model.
[0010] Secondly, embodiments of this application provide a data processing apparatus, which may include:
[0011] The acquisition module is used to acquire the resource to be identified and its resource information. The resource to be identified includes N types of resources, and the resource information includes the type identifier and position vector of each type of resource in the N types of resources, where N is an integer greater than 1.
[0012] The processing module is used to input the resource to be identified and resource information into the sensitive data identification model, and to extract features of the resource to be identified through the sensitive data identification model to obtain hidden features of N types of resources;
[0013] The calculation module is used to calculate the attention hidden features of each of the two types of resources in any two of the N types of resources, based on the hidden features of any two types of resources. The attention hidden features are used to characterize the attention distribution of the hidden features of one type of resource to the hidden features of the other type of resource in any two types of resources.
[0014] The output module is used to output the recognition result of the resource to be identified from the sensitive data recognition model based on the attention-hidden features of N types of resources.
[0015] Thirdly, embodiments of this application provide a computing device, which includes: a processor and a memory storing computer program instructions;
[0016] When the processor executes computer program instructions, it implements the data processing method as described in the first aspect.
[0017] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the data processing method as described in the first aspect.
[0018] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the data processing method as shown in the first aspect.
[0019] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the data processing method as described in the first aspect.
[0020] The data processing method, apparatus, device, and storage medium of this application embodiment acquire resources to be identified, including N types of resources, and resource information of the resources to be identified. The resource information includes the type identifier and position vector of each type of resource in the N types of resources, where N is an integer greater than 1. Then, the resources to be identified and the resource information are input into a sensitive data identification model. The sensitive data identification model extracts features from the resources to be identified to obtain hidden features of the N types of resources. Then, based on the hidden features of any two types of resources in the N types of resources, attention hidden features of each type of resource in any two types are calculated. The attention hidden features are used to characterize the attention distribution of the hidden features of one type of resource to the hidden features of the other type of resource in any two types of resources. Based on the attention hidden features of the N types of resources, the identification result of the resources to be identified is output from the sensitive data identification model. In this way, both text-sensitive data and non-text-sensitive data such as images and audio can be identified in resources. There is no need to use different identification methods for resources of different types. This improves the efficiency of sensitive data identification and avoids the problem of low identification accuracy caused by identifying the relationships between different types of resources without any connection. In addition, the sensitivity level of sensitive data in the resource to be identified can be determined by the above method, which makes it convenient for users to process the resource to be identified based on the sensitivity level. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a data processing method provided in an embodiment of this application;
[0023] Figure 2 One of the schematic diagrams of the structure of the initial sensitive data identification model for training a data processing method provided in this application embodiment;
[0024] Figure 3 A second schematic diagram of the structure of a data processing method for training an initial sensitive data identification model provided in an embodiment of this application;
[0025] Figure 4 A third schematic diagram of the structure of the initial sensitive data identification model for training a data processing method provided in this application embodiment;
[0026] Figure 5Fourth schematic diagram of the structure of the initial sensitive data identification model for training a data processing method provided in this application embodiment;
[0027] Figure 6 This is a schematic diagram of the structure of a data processing apparatus provided in one embodiment of this application;
[0028] Figure 7 This is a schematic diagram of the structure of a data processing device provided in one embodiment of this application. Detailed Implementation
[0029] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0030] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0031] In recent years, many users and merchants have suffered heavy losses due to the leakage of sensitive data (such as users' personal information and merchants' confidential information). Therefore, data security is becoming increasingly important. To protect data security, we must first understand which data should be protected most. In the era of big data, the amount of data is enormous. To select the more sensitive data for protection, we need to quickly and accurately identify and classify sensitive data from the massive amount of data. This facilitates effective subsequent processing of sensitive data. Therefore, the identification of sensitive data is of great significance.
[0032] In related technologies, sensitive data can be identified in the following ways: for example, obtaining feature parameters for locating the target data to be identified and a regular expression for identifying sensitive data in the target data; obtaining the target object containing the target data based on the feature parameters; and then identifying the target data within the target object line by line according to the regular expression to determine whether the target object contains sensitive data. Although the above method can identify sensitive data in massive amounts of data, due to the large number of data types in massive amounts of data, such as text sensitive data and non-text sensitive data such as images, audio, and audio-visual (i.e., voice + video can be considered as multiple images, text, and audio), and the huge differences in feature parameters between different data types, the above method cannot comprehensively identify all types of sensitive data. As a result, the accuracy of the identification results is low, making it difficult to guarantee the protection of the various types of sensitive data required by users. Alternatively, feature engineering can be performed on the raw data. Full Chinese datasets, non-full Chinese datasets, and image datasets can be input into corresponding classification models for training, resulting in different classification models. These models are then input into a specified classification model based on the type of sensitive data to be identified, yielding classification labels. This allows for comprehensive identification of different types of data. However, while different classification models are used for identification, this method uses two different models with different classification methods to identify both text-sensitive and non-text-sensitive data. Therefore, different identification methods are required to traverse massive amounts of data separately. This not only affects the efficiency of sensitive data identification but also, because the classification models identify massive amounts of data separately without considering the relationships between different types of resources, it can cause identification anomalies, low accuracy, and even errors if different types of resources such as text, images, and audio / video are present simultaneously.
[0033] Therefore, the data processing method provided in this application embodiment uses a self-attention transformer model to identify sensitive information in various types of resources, including images, text, audio, and audio-visual materials. It can also predict text that will be obscured or replaced.
[0034] Based on this, embodiments of this application provide a data processing method, apparatus, device, and storage medium. The following will describe in conjunction with the appendix... Figures 1 to 6 This application describes in detail the data processing methods, apparatus, servers, and storage media of the embodiments thereof. It should be noted that these embodiments are not intended to limit the scope of this application.
[0035] The following is combined Figure 1 The data processing method provided in the embodiments of this application will be described in detail.
[0036] Figure 1This is a flowchart of a data processing method provided in an embodiment of this application.
[0037] like Figure 1 As shown, this data processing method can be applied to, for example... Figure 1 The data processing architecture shown may include the following steps in its specific data processing method:
[0038] Step 110: Obtain the resource to be identified and its resource information. The resource to be identified includes N types of resources. The resource information includes the type identifier and position vector of each type of resource in the N types of resources. The position vector is used to represent the position of each type of resource in the resource to be identified. N is an integer greater than 1. Step 120: Input the resource to be identified and its resource information into the sensitive data identification model. The sensitive data identification model extracts features from the resource to be identified to obtain the hidden features of the N types of resources. Step 130: Based on the hidden features of any two types of resources in the N types of resources, calculate the attention hidden features of each type of resource in any two types of resources. The attention hidden features are used to characterize the attention distribution of the hidden features of one type of resource to the hidden features of the other type of resource in any two types of resources. Step 140: Based on the attention hidden features of the N types of resources, output the identification result of the resource to be identified from the sensitive data identification model. The identification result includes the sensitivity level of each type of resource.
[0039] Therefore, the above method can identify both text-sensitive data and non-text-sensitive data such as images and audio in resources. It eliminates the need to use different identification methods for different types of resources, increases the types of resources that can be identified, improves the efficiency of sensitive data identification, and avoids the problem of low accuracy caused by identifying the relationships between different types of resources without any connection. In addition, the above method can also determine the sensitivity level of sensitive data in the resource to be identified, making it convenient for users to process the resource to be identified based on the sensitivity level.
[0040] The above steps are explained in detail below:
[0041] First, regarding step 110, the N types of resources can include text, images, audio, and audio / video.
[0042] For example, if N is 2, the resource to be identified can be text and image. For example, 139AAAAAA is the ID card number of user A, and an image of user A's ID card number is attached. The resource information of text and image can include the type identifier of each type of resource, such as the type identifier-1 of text and the type identifier-2 of image, i.e., 1, 2 - text and image. The position vector is used to represent the position vector of the text in the first half of the resource to be identified, and to represent the position vector of the image in the second half of the resource to be identified.
[0043] Similarly, if N is 3, the resource to be identified can be text, image, and audio. For example, 139AAAAAA is user A's ID card number, accompanied by an image of user A's ID card number and an audio file of a commitment letter being read aloud. The resource information for text, image, and audio can include the type identifier of each of the three types of resources, such as the type identifier-1 for text, the type identifier-2 for image, and the type identifier-3 for audio, i.e., 1, 2, 3 - text, image, and audio. The position vector is used to represent the position vector of the text in the first half of the resource to be identified, the position vector of the image in the middle of the resource to be identified, and the position vector of the audio file in the latter half of the resource to be identified.
[0044] Here, prior to step 120, the data processing method may also include a process of training a sensitive data identification model. Based on this, in one or more possible embodiments, the data processing method may also include steps 1501 to 1509, as detailed below.
[0045] Step 1501: Obtain the sample set. The sample set includes sample training resources and sample resource information of the sample training resources. The sample training resources include M types of sample resources. The sample resource information includes the sample type identifier and sample position vector of each type of sample resource in the M types. The sample position vector is used to represent the position of each type of sample resource in the sample training resources. M can be an integer greater than 1 or a positive integer.
[0046] For example, resources containing sensitive data that are actually used are acquired and identified as sample training resources, wherein the sample training resources may include any one or any combination of the following: text, images, audio, audio and video.
[0047] Furthermore, if M is 1, then the sample training resource is a plain text resource, such as 139BBBBBB, which is user B's mobile phone number. The sample type identifier of the sample resource is plain text-1, the sample training sensitivity level is level three (confidential), and the sample position vector is used to represent the position vector of the text at the beginning of the resource to be identified. Similarly, if the sample training resource is a pure image resource, such as image 1, the sample type identifier of the sample resource is pure image-2, the sample training sensitivity level is level two (confidential), and the sample position vector is used to represent sensitive data, such as the sensitive image region in image 1, which is located in the upper left corner of image 1.
[0048] Furthermore, if M is 2, the sample training resources can be text and image resources, such as 139BBBBBB, which is user B's mobile phone number, and is accompanied by an image of user A's ID card number. The sample type identifier of the sample resources is 1, 2 - text and image, the sample training sensitivity level is level three (secret), and the sample position vector is used to represent the position vector of the text in the first half of the sample training resource and the position vector of the image in the second half of the resource to be identified.
[0049] It should be noted that, in one example, the aforementioned sample resource information may further include a sample sensitive data location vector. This vector represents the position of the pre-defined sensitive data in each type of sample resource within the sample training resource. This allows for subsequent verification based on the sample sensitive data location vector to determine whether the position of the sample sensitive data output by the initial sensitive data recognition model in the sample training resource is the same as the sample sensitive data location vector. If they are the same, it indicates that the initial sensitive data recognition model training is complete. Conversely, if they are different, the difference between the sample sensitive data location vector and the position of the sample sensitive data output by the initial sensitive data recognition model in the sample training resource can be used to determine the adjustment gradient. This gradient can then be used to continue training the initial sensitive data recognition model using gradient descent. In another example, the sample type identifier can identify whether the sample resource is plain text, plain image, or a combination of text and images, and can also identify the sensitivity type of the sample sensitive data contained within it. In yet another example, the sample training sensitivity level can include Level 1 (Top Secret), Level 2 (Confidential), Level 3 (Secret), or Level 4 (Ordinary).
[0050] Step 1502: Input the sample set into the initial sensitive data recognition model, and extract features from the sample training resources through the encoder of the initial sensitive data recognition model to obtain the hidden features of the sample resources of M types.
[0051] For example, such as Figure 2As shown, taking M=2 and sample text and sample images as examples, the initial sensitive data recognition model is the content of the sample set. This initial sensitive data recognition model can be a Transformer. The content of the sample set can be passed to the encoder in the Transformer to encode the content of the sample set through the editor, extract the features of each type of sample resource, and output the sample hidden features of two types of sample resources, such as the sample hidden features of text (sample HT) and the sample hidden features of images (sample HV), so that sample HT and sample HV can be used as the input of the same decoder.
[0052] Furthermore, before passing the content of the sample set to the encoder in the Transformer, word embedding can be used to map the words in the text resource space of the sample training resources to multidimensional vectors in another space. Additionally, linear projection of flattened patches can be used to map each image patch in the image to a one-dimensional vector, transforming the input image format [H, W, C] into a standard Transformer encoder input format token (vector) sequence, i.e., a two-dimensional matrix [num_token, token_dim]. In this way, the data processed by word embedding and linear projection of flattened patches can be input to the Transformer encoder, enabling the calculation of the text-corresponding sample HT and the image-corresponding sample HV based on the Transformer encoder.
[0053] It should be noted that if the sample training resources include audio, the audio can be converted into frequency domain audio images and time domain audio images, and the audio file can be converted into an image for processing using the Linear Projection of Flattened Patches described above; or, the audio can be converted into audio text for processing using the Word Embedding described above.
[0054] Step 1503: Based on the sample hidden features of any two types of sample resources among the sample hidden features of M types of sample resources, calculate the sample attention hidden features of each type of sample resource among the two types. The sample attention hidden features are used to characterize the sample attention distribution of the sample hidden features of one type of sample resource to the sample hidden features of the other type of sample resource among the two types of sample resources.
[0055] For example, such as Figure 3 As shown, in the same Transformer decoder, the inputs are samples HT and HV. Based on sample HT, a sample text query vector (QT), a sample text key vector (KT), and a sample text value vector (VT) are generated. Similarly, based on sample HV, a sample image query vector (QV), a sample text key vector (KV), and a sample text value vector (VV) are generated. Based on this, samples QT and QV are swapped to obtain samples HT (QT, KT, QV) and HV (QT, KV, VV). Multi-head attention is then performed on samples (QT, KT, QV) and (QT, KV, VV) respectively. Following this, the results of the multi-head attention calculation are sequentially processed with residual connections and normalization (Add & Norm), and then fed forward neural network processing (Feed). (Forward), then perform residual connection and normalization (Add & Norm). Finally, extract the sample attention hidden features, such as sample HT←V (image to text sample attention hidden features) corresponding to sample HT, and sample HV←T (text to image sample attention hidden features) corresponding to sample HV. Then, add and merge sample HT←V and sample HV←T, and process them differently depending on whether the sample training resources in the sample set include anomalous resources. If the sample training resources do not include anomalous resources, the result of adding and merging sample HT←V and sample HV←T is used as the input to pooling and fully connected layers to train the pooling and fully connected layers in the initial sensitive data recognition model. Conversely, if the sample training resources include anomalous resources, the result of adding and merging sample HT←V and sample HV←T is used as the input to the Multilayer Perceptron (MLP) to train the MLP in the initial sensitive data recognition model.
[0056] Step 1504: Based on the sample attention hiding features of M types of sample resources, output the sample recognition results of the sample training resources from the initial sensitive data recognition model. The sample recognition results include the sample sensitive data of each type of sample resource in any two types and the sample sensitivity level corresponding to the sample sensitive data.
[0057] For example, the result of adding and merging samples HT←V and HV←T is used to train the layer between neurons in the pooling and fully connected (FC) layers. Since all neurons in the pooling and fully connected layers are weighted connections, training the layer between neurons in the pooling and fully connected (FC) layers in this embodiment can be understood as training the weights of all neurons in the pooling and fully connected layers. Thus, after the pooling and fully connected layers, the sample recognition result of the text and image sample training resources is obtained, i.e., which are the sample-sensitive data of the sample resources and the probability of the sample sensitivity level of the sample-sensitive data, such as "Top Secret," "Confidential," "Secret," and "Ordinary."
[0058] Step 1505: Compare the sample sensitive data with the sample training sensitive data corresponding to the sample training resources to obtain the first contrast gradient; and compare the sample sensitivity level with the sample training sensitivity level corresponding to the sample training resources to obtain the second contrast gradient.
[0059] For example, if the probability of "Top Secret" is the highest among the four probabilities corresponding to the four sample sensitivity levels output in step 1504 above, such as "Top Secret", "Confidential", "Secret" and "Ordinary", then "Top Secret" is determined as the sample sensitivity level. Then, the sample sensitivity data and the sample sensitivity level "Top Secret" determined in step 1504 are compared with the real sample training sensitivity data corresponding to the sample sensitivity data. If the matching result is that the two are completely matched, it means that there is no need to adjust the gradient for the time being. Otherwise, if the matching result is that the two are at least partially inconsistent, the comparison gradient is calculated based on the at least partially inconsistent part, and the adjustment direction for the next training is given.
[0060] Step 1506: Using the sample set, the first contrast gradient, and the second contrast gradient, the initial sensitive data identification model is trained until the preset training conditions are met, thus obtaining the sensitive data identification model.
[0061] The preset training conditions may include at least one of the following: the number of training sessions is equal to or greater than the preset number of training sessions, the first contrast gradient is less than or equal to the first preset contrast gradient, and the second contrast gradient is less than or equal to the second preset contrast gradient.
[0062] Additionally, in one or more possible embodiments, prior to step 1506, the data processing method may further include:
[0063] Step 1507: In the case that each type of sample resource includes abnormal sample resources in any two types, the sample attention distribution of the sample hidden features of each type of sample resource in any two types is reconstructed through the initial multilayer perceptron in the initial sensitive data identification model to obtain the sample probability value of the sample reconstruction feature corresponding to the abnormal sample resource.
[0064] Step 1508: Determine the sample sensitivity level of abnormal sample resources based on the sample probability value of the sample reconstruction features; wherein, abnormal sample resources include at least one of the following: deleted sample resources, modified sample resources, and replaced sample resources.
[0065] Step 1509: Compare the sample probability value with the preset sample probability value of the sample reconstruction feature corresponding to the abnormal sample resource to obtain the third comparison gradient.
[0066] Based on this, step 1506 above may specifically include:
[0067] The initial sensitive data identification model is obtained by using a sample set, a first contrast gradient, a second contrast gradient, and a third contrast echelon until the preset training conditions are met.
[0068] For example, the training resources for the initial sensitive data recognition model may include anomalous resources such as text or images, i.e., deleted, modified, or replaced resources. For instance, some text or images may be deliberately obscured or altered. Sometimes, to prevent leaks, resources containing sensitive data may be photographed, and sensitive data such as "confidential" may be obscured, replaced with misspelled words, or replaced with "Martian language," as well as obscured or altered labels. Thus, the initial multilayer perceptron in the initial sensitive data recognition model reconstructs these deliberately obscured or altered parts, obtaining sample probability values of the reconstructed features of the samples before they were obscured or altered. Then, pooling and fully connected layers are applied to train the model.
[0069] It should be noted that in step 1508, the sample sensitivity level of the sample abnormal resource can be determined by pooling and a fully connected layer.
[0070] Based on this, step 120, in one or more possible embodiments, may specifically include:
[0071] Step 1201: Using the sensitive data identification model, based on the association information between the preset type identifier and the preset mapping algorithm, obtain the mapping algorithm corresponding to each of the N type identifiers in the resource information;
[0072] Step 1202: Using the mapping algorithm corresponding to each type of identifier, map the resources and location vectors corresponding to each type of identifier in the N types of resources to obtain N mapping vectors. The vector format of the N mapping vectors corresponds to the input format of the encoder in the sensitive data identification model.
[0073] Step 1203: Extract features from each of the N mapping vectors using the encoder to obtain the hidden features of the N types of resources.
[0074] For example, taking text and images as an example, the resource to be identified, which contains text and images, is input into the trained sensitive data recognition model. First, the resource to be identified is encoded by the encoder in the same sensitive data recognition model, and the features of the text and images are extracted to obtain the hidden features (HT) of the text resource and the hidden features (HV) of the image resource, so that HT and HV can be used as the input of the same decoder.
[0075] Based on this, in another or more possible embodiments, the N type identifiers include a first type identifier and a second type identifier, the first type identifier corresponds to a first mapping algorithm, the second type identifier corresponds to a second mapping algorithm, and the N mapping vectors include multi-dimensional vectors and two-dimensional matrices. Based on this, step 1202 above may specifically include:
[0076] The first mapping algorithm maps the resource and location vector corresponding to the first type identifier among the N types of resources to a preset space to obtain a multi-dimensional vector.
[0077] Furthermore, the second mapping algorithm maps the resources corresponding to the second type of identifier to a one-dimensional vector, and converts the position vector of the resources corresponding to the second type of identifier into a two-dimensional matrix.
[0078] Specifically, the type identifier is used to mark the type of any resource among the resources to be identified. The first type identifier and the second type identifier respectively mark resources of different types among the resources to be identified. Since each type identifier in this embodiment can mark one type, for more accurate mapping, the mapping algorithm in this embodiment is determined based on the type identifier. That is, the mapping algorithm is used to map resources corresponding to any type of resource. Because the first mapping algorithm and the second mapping algorithm map resources of different types, the first mapping algorithm and the second mapping algorithm are also different.
[0079] For example, if the first type identifier is text, the first mapping algorithm can be word embedding, which can map words in the space to which the text resource belongs to a multidimensional vector in another space. And if the second type identifier is an image, the second mapping algorithm can be a linear projection of flattened patches algorithm, which can map each image patch in the image to a one-dimensional vector and transform the format [H, W, C] of the input image into a sequence of tokens (vectors) in the encoder's input format, i.e., a two-dimensional matrix [num_token, token_dim]. In this way, the data processed by Word Embedding and Linear Projection of Flattened Patches can be input to the Transformer encoder so that the HT corresponding to the text and the HV corresponding to the image can be calculated based on the Transformer encoder.
[0080] Therefore, the main function of feature extraction is to reduce the number of words or image fragments to be processed without damaging the core content of the resource, thereby reducing the dimension of the vector space, simplifying the calculation, and improving the speed and efficiency of data processing.
[0081] Furthermore, regarding step 130, in one or more possible embodiments, step 130 may specifically include:
[0082] Step 1301: Based on the hidden features of N types of resources, generate a first vector set of hidden features for each type of resource in the hidden features of N types of resources. The first vector set includes query vector, key vector and content vector.
[0083] Step 1302: Cross-interchange the query vectors in the vector sets of hidden features of any two types of resources to obtain a second vector set of hidden features of each type of resource in any two types. The second vector set includes the key vector and content vector of the first type of resource in any two types and the query vector corresponding to the resource of the other type.
[0084] Step 1303: Using a multi-head attention computation algorithm, process the second vector set of hidden features of each type of resource in any two types to obtain the first processing result;
[0085] Step 1304: Perform residual connection and normalization on the first processing result in sequence to obtain the second processing result;
[0086] Step 1305: Through the feedforward neural network in the sensitive data identification model, the second processing result is further processed by residual connection and normalization to obtain the attention-hidden features of each type of resource.
[0087] For example, in the same Transformer decoder, the inputs are HT and HV. Based on HT, a text query vector (QT), a text key vector (KT), and a text value vector (VT) are generated. Based on HV, an image query vector (QV), a text key vector (KV), and a text value vector (VV) are generated. Based on this, QT and QV are swapped to obtain HT(QT, KT, QV) and HV(QT, KV, VV). Multi-head attention calculations are then performed on (QT, KT, QV) and (QT, KV, VV) respectively. Then, given the results of the multi-head attention calculations, residual connections and normalization are sequentially applied to these results, followed by feedforward processing. Residual connections and normalization are then performed again. Finally, attention-hidden features are extracted, such as HT←V (image-to-text attention-hidden features) corresponding to HT, and HV←T (text-to-image attention-hidden features) corresponding to HV. It should be noted that the above attention formula (1) can be as follows:
[0088] (1)
[0089] Among them, Q (query) vector, K (key) key vector, V (value) value vector, It is the transpose of the K-vector matrix. It is the dimension of the K-vector matrix.
[0090] Then, relating to step 140, in one or more possible embodiments, step 140 may specifically include:
[0091] Step 1401: Merge the attention-hidden features of N types of resources to obtain the attention-hidden feature set;
[0092] Step 1402: The attention-hidden feature set is sequentially processed by pooling and full connection to obtain the sensitivity level of sensitive data in each type of resource.
[0093] Furthermore, step 1402 may specifically include:
[0094] In cases where no anomalous resources are included in each type of resource, the attention-hidden feature set is sequentially pooled and fully connected to obtain the sensitivity level of sensitive data in each type of resource.
[0095] Conversely, the data processing method in the embodiments of this application may further include:
[0096] In cases where each type of resource includes anomalous resources, the attention-hidden features of each type of resource are reconstructed using a multilayer perceptron in the sensitive data identification model to obtain the probability value of the reconstructed features corresponding to the anomalous resources.
[0097] The sensitivity level of abnormal resources is determined based on the probability value of the reconstruction features; wherein, abnormal resources include at least one of the following: deleted resources, modified resources, and replaced resources.
[0098] For example, HT←V and HV←T are added and merged, and different processing is applied to each type of resource depending on whether it includes abnormal resources. If no abnormal resources are included in each type of resource, the result of adding and merging HT←V and HV←T is used as the input to the pooling and fully connected layers, and processed by the pooling and fully connected layers in the sensitive data identification model. Figure 4 As shown, the outputs of the attention-hidden features of the image on the text and the attention-hidden features of the text on the image are merged, and then pooled and fully connected to output the probability of which sensitivity level, such as "top secret", "confidential", "secret" and "ordinary".
[0099] Conversely, such as Figure 5 As shown, if each type of resource includes anomalous resources, the result of adding and merging HT←V and HV←T is used as the input of a multilayer perceptron (MLP) to obtain the probability value of the reconstructed features corresponding to the anomalous resources. Then, based on the probability value of the reconstructed features, the sensitivity level of the anomalous resources is determined. In this way, the sensitivity level of each type of resource can be calculated through the sensitive data recognition model, such as outputting which sensitivity level it belongs to, such as "confidential" or "top secret". At the same time, it can also output the content before the masking or alteration corresponding to the alienated resources. Through the sequence-to-sequence multimodal approach of the sensitive data recognition model, multiple tokens (vocabularies, charts) are used for text and images, such as more than 6,000 commonly used Chinese tokens, so that sensitive data in text and images, whether images or speech, can be self-attentive and interactively attended to by text, images and speech. If there is masking or alteration, it should also be able to infer the masked or altered content as much as possible. That is, through the self-attention sensitive data recognition model, a multimodal multi-task recognition of sensitivity category judgment and prediction of masked or altered content is performed.
[0100] It should be noted that the multi-layer perceptron neural network (MLP) in the embodiments of this application refers to a multi-layer linear stack with hidden layers. The three layers of an MLP can also include an input layer, a hidden layer, and an output layer.
[0101] In addition, after step 140, the results obtained in step 140 can be displayed to the user in a preset manner, such as distinguishing and displaying sensitive data and marking sensitive data.
[0102] Thus, through a trained sensitive data identification model with a single structure, the classification level of classified text, text and images, images or audio can be accurately inferred. Furthermore, if classified content is obscured or altered by images, text, or audio, the model can accurately predict the content that has been obscured, altered, or replaced. Moreover, because the specific corresponding sensitivity level can be predicted through the multimodal and multi-task capabilities of a single sensitive data identification model, the method provided in this application can identify resources containing text, images, and audio to accurately obtain the sensitivity level of mixed sensitive data. It can also identify and classify sensitive data, formulate corresponding data security protection measures, and prevent the leakage of sensitive data.
[0103] Furthermore, by using multi-task training to predict masking and alteration, a sensitive data identification model with the function of identifying abnormal resources is obtained. In this way, the sensitive data identification model can measure the multimodal and multi-task nature of masking and alteration, identify the classified content that has been masked, altered, or replaced, and thus reconstruct the resource before it was masked or altered, and determine the sensitivity level of the resource before it was masked or altered. This not only restores the maliciously masked or altered content and provides evidence of malicious masking and alteration, but also improves the identification accuracy of sensitive data, saves a lot of manpower in the identification and judgment of sensitive data, effectively reduces the risk of sensitive data leakage, and strengthens the control capabilities of important data.
[0104] Based on the same inventive concept, this application also provides a data processing device. (Specifically combined with...) Figure 6 Please provide a detailed explanation.
[0105] Figure 6 This is a schematic diagram of the structure of a data processing apparatus provided in one embodiment of this application.
[0106] In some embodiments of this application, Figure 6 The data processing device shown can be set in, for example Figure 7 In the data processing device shown.
[0107] like Figure 6As shown, the data processing device 60 may specifically include:
[0108] The acquisition module 601 is used to acquire the resource to be identified and the resource information of the resource to be identified. The resource to be identified includes N types of resources. The resource information includes the type identifier and position vector of each type of resource in the N types of resources, where N is an integer greater than 1.
[0109] The processing module 602 is used to input the resource to be identified and resource information into the sensitive data identification model, and extract features of the resource to be identified through the sensitive data identification model to obtain hidden features of N types of resources;
[0110] The calculation module 603 is used to calculate the attention hidden features of each of the two types of resources in any two types of resources based on the hidden features of any two types of resources in N types of resources. The attention hidden features are used to characterize the attention distribution of the hidden features of one type of resource to the hidden features of the other type of resource in any two types of resources.
[0111] Output module 604 is used to output the recognition result of the resource to be identified from the sensitive data recognition model based on the attention-hidden features of N types of resources.
[0112] In this embodiment, the above method can identify both text-sensitive data and non-text-sensitive data such as images and audio in resources. It eliminates the need to use different identification methods for resources of various types. This improves the efficiency of sensitive data identification and avoids the problem of low identification accuracy caused by identifying multiple types of resources without any correlation. In addition, the above method can also determine the sensitivity level of sensitive data in the resource to be identified, making it convenient for users to process the resource to be identified based on the sensitivity level.
[0113] The data processing device 60 in the embodiments of this application will be described in detail below.
[0114] In one or more optional embodiments, the processing module 602 may be specifically used to obtain the mapping algorithm corresponding to each of the N types of identifiers in the resource information based on the association information between the preset type identifier and the preset mapping algorithm through the sensitive data identification model.
[0115] By using the mapping algorithm corresponding to each type of identifier, the resource and location vectors corresponding to each type of identifier in the N types of resources are mapped to obtain N mapping vectors. The vector format of the N mapping vectors corresponds to the input format of the encoder in the sensitive data identification model:
[0116] By extracting features from each of the N mapping vectors using an encoder, hidden features of N types of resources can be obtained.
[0117] In another or more alternative embodiments, the processing module 602 may be specifically used to, in the case where N type identifiers include a first type identifier and a second type identifier, the first type identifier corresponds to a first mapping algorithm, the second type identifier corresponds to a second mapping algorithm, and the N mapping vectors include multidimensional vectors and two-dimensional matrices, map the resources and location vectors corresponding to the first type identifier in the N types of resources to a preset space through the first mapping algorithm to obtain multidimensional vectors;
[0118] Furthermore, the second mapping algorithm maps the resources corresponding to the second type of identifier to a one-dimensional vector, and converts the position vector of the resources corresponding to the second type of identifier into a two-dimensional matrix.
[0119] In another or more alternative embodiments, the calculation module 603 may be specifically used to generate a first vector set of hidden features of each type of resource in the hidden features of N types of resources, based on the hidden features of N types of resources. The first vector set includes a query vector, a key vector, and a content vector.
[0120] By cross-interchanging the query vectors in the vector sets of hidden features of any two types of resources, a second vector set of hidden features of each type of resource in any two types is obtained. The second vector set includes the key vector and content vector of the first type of resource in any two types and the query vector corresponding to the resource of the second type.
[0121] The second vector set of hidden features of each type of resource in any two types is processed by the multi-head attention computation algorithm to obtain the first processing result;
[0122] The first processing result is sequentially subjected to residual connection and normalization to obtain the second processing result;
[0123] By using the feedforward neural network in the sensitive data identification model, the second processing result is further processed by residual connection and normalization to obtain the attention-hidden features of each type of resource.
[0124] In another or more alternative embodiments, the output module 604 may be specifically used to merge the attention-hidden features of N types of resources to obtain an attention-hidden feature set, provided that the identification result includes the sensitivity level of each type of resource.
[0125] The attention-hidden feature set is sequentially processed by pooling and fully connected layers to obtain the sensitivity level of sensitive data in each type of resource.
[0126] In another or more alternative embodiments, the data processing device 60 in this application embodiment may further include a first reconstruction module and a first determination module; wherein,
[0127] The first reconstruction module is used to reconstruct the attention-hidden features of each type of resource, including anomalous resources, through the multilayer perceptron in the sensitive data identification model, and obtain the probability value of the reconstructed features corresponding to the anomalous resources.
[0128] The first determining module is used to determine the sensitivity level of abnormal resources based on the probability value of the reconstructed features; wherein, abnormal resources include at least one of the following: deleted resources, modified resources, and replaced resources.
[0129] In another or more alternative embodiments, the data processing device 60 in this application embodiment may further include a training module for acquiring a sample set, the sample set including sample training resources and sample resource information of the sample training resources, the sample training resources including M types of sample resources, the sample resource information including a sample type identifier and a sample position vector for each type of sample resource in the M types, the sample position vector being used to represent the position of each type of sample resource in the sample training resources, where M is an integer greater than 1;
[0130] The sample set is input into the initial sensitive data recognition model. The encoder of the initial sensitive data recognition model extracts features from the sample training resources to obtain the hidden features of M types of sample resources.
[0131] Based on the sample hidden features of any two types of sample resources from the M types of sample hidden features, calculate the sample attention hidden features of each type of sample resource in the two types. The sample attention hidden features are used to characterize the sample attention distribution of the sample hidden features of one type of sample resource to the sample hidden features of the other type of sample resource in the two types of sample resources.
[0132] Based on the sample attention hiding features of M types of sample resources, the sample recognition results of the sample training resources are output from the initial sensitive data recognition model. The sample recognition results include the sample sensitive data of each type of sample resource in any two types and the sample sensitivity level corresponding to the sample sensitive data.
[0133] The first contrast gradient is obtained by comparing the sensitive data of the sample with the sensitive data of the sample training resources; and the second contrast gradient is obtained by comparing the sensitivity level of the sample with the sensitivity level of the sample training resources.
[0134] Using a sample set, a first contrastive gradient, and a second contrastive gradient, an initial sensitive data identification model is developed until the preset training conditions are met, thus obtaining the sensitive data identification model.
[0135] In another or more alternative embodiments, the data processing device 60 in this application embodiment may further include a second reconstruction module, a second determination module, and a comparison module; wherein,
[0136] The second reconstruction module is used to reconstruct the sample attention distribution of the sample hidden features of each type of sample resource in any two types, when each type of sample resource includes sample abnormal resources, through the initial multilayer perceptron in the initial sensitive data identification model, so as to obtain the sample probability value of the sample reconstruction feature corresponding to the sample abnormal resources.
[0137] The second determining module is used to determine the sample sensitivity level of the abnormal sample resource based on the sample probability value of the sample reconstruction feature; wherein the abnormal sample resource includes at least one of the following: deleted sample resource, modified sample resource, and replaced sample resource.
[0138] The comparison module is used to compare the sample probability value with the preset sample probability value of the sample reconstruction feature corresponding to the abnormal sample resource to obtain the third comparison gradient.
[0139] The training module is used to identify the sensitive data model using a sample set, a first contrast gradient, a second contrast gradient, and a third contrast echelon, until the preset training conditions are met, thus obtaining the sensitive data identification model.
[0140] Based on the same inventive concept, this application also provides a data processing device. (Specifically combined with...) Figure 7 Please provide a detailed explanation.
[0141] Figure 7 This is a schematic diagram of the structure of a data processing device provided in one embodiment of this application.
[0142] like Figure 7 As shown, the data processing device may include at least one of the following as described in the embodiments of this application: an electronic device, a server. The data processing device may include a processor 701 and a memory 702 storing computer program instructions.
[0143] Specifically, the processor 701 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0144] Memory 702 may include mass storage for data or instructions. For example, and not limitingly, memory 702 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 702 may include removable or non-removable (or fixed) media. Where appropriate, memory 702 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 702 is non-volatile solid-state memory. In a particular embodiment, memory 702 includes solid-state storage (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0145] The processor 701 implements any of the data processing methods described in the above embodiments by reading and executing computer program instructions stored in the memory 702.
[0146] In one example, the data processing device may further include a communication interface 703 and a bus 710. Wherein, as... Figure 7 As shown, the processor 701, memory 702, and communication interface 703 are connected through bus 710 and complete communication with each other.
[0147] The communication interface 703 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0148] Bus 710 includes hardware, software, or both, that couples components of a flow control device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 710 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0149] The data processing device can execute the data processing method described in the embodiments of this application, thereby achieving the combination Figures 1 to 6 The data processing methods and apparatus described.
[0150] Furthermore, in conjunction with the data processing methods in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the data processing methods in the above embodiments.
[0151] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0152] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0153] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0154] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A data processing method, characterized in that, include: Obtain the resource to be identified and its resource information. The resource to be identified includes N types of resources, and the N types of resources include at least two of the following: text, image, audio, and audio-visual. The resource information includes the type identifier and position vector of each type of resource among the N types of resources, where N is an integer greater than 1. The resource to be identified and the resource information are input into the sensitive data identification model. The sensitive data identification model is used to extract features from the resource to be identified to obtain the hidden features of the N types of resources. Based on the hidden features of the N types of resources, a first vector set of hidden features for each type of resource is generated, wherein the first vector set includes query vector, key vector and content vector. The query vectors in the vector sets of hidden features of any two types of resources are cross-interleaved to obtain a second vector set of hidden features of each type of resource in any two types. The second vector set includes the key vector and content vector of the first type and the query vector corresponding to the resource of the second type in any two types. The second vector set of hidden features of each type of resource in any two types is processed by the multi-head attention computation algorithm to obtain the first processing result; The first processing result is sequentially subjected to residual connection and normalization to obtain the second processing result; By using the feedforward neural network in the sensitive data identification model, the second processing result is further processed by residual connection and normalization to obtain the attention hidden features of each type of resource. The attention hidden features are used to characterize the attention distribution of the hidden features of one type of resource to the hidden features of the other type of resource in any two types of resources. Based on the attention-hiding features of the N types of resources, the identification result of the resource to be identified is output from the sensitive data identification model.
2. The method according to claim 1, characterized in that, The step involves inputting the resource to be identified and the resource information into a sensitive data identification model, and then using the sensitive data identification model to extract features from the resource to be identified to obtain the hidden features of the N types of resources, including: The sensitive data identification model obtains the mapping algorithm corresponding to each of the N types of identifiers in the resource information based on the association information between the preset type identifier and the preset mapping algorithm. By using the mapping algorithm corresponding to each type of identifier, the resources and location vectors corresponding to each type of identifier in the N types of resources are mapped to obtain N mapping vectors. The vector format of the N mapping vectors corresponds to the input format of the encoder in the sensitive data identification model. The encoder extracts features from each of the N mapping vectors to obtain the hidden features of the N types of resources.
3. The method according to claim 2, characterized in that, The N types of identifiers include a first type identifier and a second type identifier. The first type identifier corresponds to a first mapping algorithm, and the second type identifier corresponds to a second mapping algorithm. The N mapping vectors include multidimensional vectors and two-dimensional matrices. The method uses a mapping algorithm corresponding to each type identifier to map the resources and location vectors corresponding to each of the N types of resources to obtain N mapping vectors, including: Using the first mapping algorithm, the resource and location vector corresponding to the first type identifier in the N types of resources are mapped to a preset space to obtain the multidimensional vector; Furthermore, the resource corresponding to the second type identifier is mapped to a one-dimensional vector using the second mapping algorithm, and the location vector of the resource corresponding to the second type identifier is converted into a two-dimensional matrix.
4. The method according to claim 1, characterized in that, The identification result includes the sensitivity level of each type of resource; the step of outputting the identification result of the resource to be identified from the sensitive data identification model based on the attention hiding features of the N types of resources includes: Merge the attention-hidden features of the N types of resources to obtain an attention-hidden feature set; The attention-hidden feature set is sequentially processed by pooling and full connection to obtain the sensitivity level of sensitive data in each type of resource.
5. The method according to claim 1, characterized in that, After calculating the attention-hidden features of each type of resource in any two types, the method further includes: In the case where each type of resource includes anomalous resources, the attention-hidden features of each type of resource are reconstructed using the multilayer perceptron in the sensitive data identification model to obtain the probability value of the reconstructed features corresponding to the anomalous resources. The sensitivity level of the abnormal resource is determined based on the probability value of the reconstructed feature; wherein the abnormal resource includes at least one of the following: a deleted resource, a modified resource, or a replaced resource.
6. The method according to claim 1, characterized in that, Before inputting the resource to be identified and the resource information into the sensitive data identification model, the method further includes: Obtain a sample set, which includes sample training resources and sample resource information of the sample training resources. The sample training resources include M types of sample resources. The sample resource information includes a sample type identifier and a sample position vector for each type of sample resource in the M types. The sample position vector is used to represent the position of each type of sample resource in the sample training resources, where M is an integer greater than 1. The sample set is input into the initial sensitive data recognition model, and the encoder of the initial sensitive data recognition model is used to extract features from the sample training resources to obtain the hidden features of the M types of sample resources. Based on the sample hiding features of any two types of sample resources among the M types of sample resources, calculate the sample attention hiding features of each type of sample resource among the two types. The sample attention hiding features are used to characterize the sample attention distribution of the sample hiding features of one type of sample resource to the sample hiding features of the other type of sample resource among the two types of sample resources. Based on the sample attention hiding features of the M types of sample resources, the sample recognition result of the sample training resources is output from the initial sensitive data recognition model. The sample recognition result includes the sample sensitive data of each type of sample resource in any two types and the sample sensitivity level corresponding to the sample sensitive data. The sample sensitive data is compared with the sample training sensitive data corresponding to the sample training resources to obtain a first comparison gradient; and the sample sensitivity level is compared with the sample training sensitivity level corresponding to the sample training resources to obtain a second comparison gradient. The initial sensitive data identification model is trained using the sample set, the first contrast gradient, and the second contrast gradient until the preset training conditions are met, thus obtaining the sensitive data identification model.
7. The method according to claim 6, characterized in that, Before obtaining the sensitive data identification model by training the initial sensitive data identification model using the sample set, the first contrast gradient, and the second contrast gradient until the preset training conditions are met, the method further includes: In the case where each of the two types of sample resources includes abnormal sample resources, the sample attention distribution of the sample hidden features of each type of sample resource in the initial sensitive data identification model is reconstructed through the initial multilayer perceptron in the initial sensitive data identification model to obtain the sample probability value of the sample reconstruction feature corresponding to the abnormal sample resource. Based on the sample probability value of the reconstructed sample features, the sample sensitivity level of the abnormal sample resource is determined; wherein, the abnormal sample resource includes at least one of the following: deleted sample resource, modified sample resource, and replaced sample resource. The sample probability value is compared with the preset sample probability value of the sample reconstruction feature corresponding to the sample abnormal resource to obtain the third comparison gradient; The step of training the initial sensitive data identification model using the sample set, the first contrast gradient, and the second contrast gradient until preset training conditions are met to obtain the sensitive data identification model includes: The initial sensitive data identification model is trained using the sample set, the first contrast gradient, the second contrast gradient, and the third contrast echelon until the preset training conditions are met, thus obtaining the sensitive data identification model.
8. A data processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire the resource to be identified and the resource information of the resource to be identified. The resource to be identified includes N types of resources, and the N types of resources include at least two of text, image, audio, and audio-visual. The resource information includes the type identifier and position vector of each type of resource in the N types of resources, where N is an integer greater than 1. The processing module is used to input the resource to be identified and the resource information into the sensitive data identification model, and to extract features from the resource to be identified through the sensitive data identification model to obtain the hidden features of the N types of resources; The calculation module is used to generate a first vector set of hidden features for each type of resource among the hidden features of the N types of resources, the first vector set including query vectors, key vectors, and content vectors; cross-interchange the query vectors in the vector sets of hidden features of any two types of resources to obtain a second vector set of hidden features for each type of resource among the two types, the second vector set including the key vector and content vector of the first type and the query vector corresponding to the resource of the second type; process the second vector set of hidden features of each type of resource among the two types using a multi-head attention calculation algorithm to obtain a first processing result; perform residual connection and normalization processing on the first processing result in sequence to obtain a second processing result; and perform residual connection and normalization processing on the second processing result again through the feedforward neural network in the sensitive data identification model to obtain the attention hidden features of each type of resource, the attention hidden features being used to characterize the attention distribution of the hidden features of one type of resource to the hidden features of the other type of resource among the two types of resources; The output module is used to output the identification result of the resource to be identified from the sensitive data identification model based on the attention-hidden features of the N types of resources.
9. A computing device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the data processing method as described in any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium stores computer program instructions, which, when executed by a processor, implement the data processing method as described in any one of claims 1-7.
Citation Information
Patent Citations
Link risk detection method and device and storage medium
CN113221032A
Multi-mode-based sentiment classification method and device, equipment and storage medium
CN115240712A