Privacy computing platform based on data classification and grading
Through text data analysis and processing technology based on deep learning neural networks, the entered data is encoded word segmentation, word segmentation and semantic coding, which solves the inefficient data classification and grading problem in the existing technology, and realizes more efficient data classification and grading, while enhancing data privacy protection.
Patent Information
- Application Number
- CN202410940884.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2044-07-15
AI Technical Summary
The privacy calculations of existing data classification hierarchical depend on manual audits or simple rule matching, are inefficient and difficult to adapt to data diversity and complexity, resulting in low classification accuracy.
Text data analysis and processing technology based on deep learning neural networks is adopted to perform word segmentation, word segmentation and semantic encoding of the entered data. The data is automatically evaluated as personal data or industry data through the multi-grained semantic encoding characteristics fused by word granularity and word granularity.
It improves the automation level and efficiency of data processing, enhances data privacy protection, and achieves more accurate data classification and grading.
Smart Images

Figure CN118940765B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent data classification, and more specifically, to a privacy computing platform based on data classification and grading. Background Art
[0002] When processing large amounts of text data, enterprises and institutions often face the challenges of identifying, classifying, and grading sensitive data. This data may contain personal information, trade secrets, or industry-specific data, and needs to be properly graded, managed, and protected according to different sensitivity levels and privacy requirements.
[0003] However, existing privacy computing for data classification and grading usually relies on manual review or simple rule matching. These methods are not only inefficient but also difficult to adapt to data diversity and complexity, resulting in low data classification accuracy.
[0004] Therefore, a privacy computing platform based on data classification and grading is desired. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a privacy computing platform based on data classification and grading, which obtains input data and adopts text data analysis and processing technology based on deep learning neural networks to perform character segmentation, word segmentation and semantic encoding on the input data, so as to automatically evaluate whether the input data is personal data or industry data based on the multi-granularity semantic coding features after the fusion of the character granularity and word granularity of the input data. In this way, the semantic content of the text can be deeply understood, which helps to classify and grade the data more accurately, while improving the automation level and efficiency of data processing and strengthening data privacy protection.
[0006] According to one aspect of the present application, a privacy computing platform based on data classification and grading is provided, which includes:
[0007] Input data acquisition module, used to obtain input data;
[0008] An input data semantic coding module is used to perform character-based semantic coding and word-based semantic coding on the input data to obtain a sequence of input data character-based semantic coding feature vectors and a sequence of input data word-based semantic coding feature vectors;
[0009] A semantic coding feature fusion module, configured to input the sequence of the input data word-granularity semantic coding feature vectors and the sequence of the input data word-granularity semantic coding feature vectors into an energy measurement attention fusion module based on sequence annotation to obtain input data word-granularity fused semantic coding feature vectors and input data word-granularity fused semantic coding feature vectors;
[0010] A multi-granularity input data fusion module is used to input the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector into a balanced threshold feature vector adaptive fusion module to obtain a multi-granularity input data semantic coding feature vector as a multi-granularity input data semantic coding feature;
[0011] The classification result generation module is used to obtain a classification result based on the semantic coding features of the multi-granularity input data, and the classification result is used to indicate whether the input data is personal data or industry data.
[0012] In the above-mentioned privacy computing platform based on data classification and grading, the input data semantic encoding module includes: a character embedding encoding unit, which is used to perform character segmentation processing on the input data and then pass it through a character granularity semantic encoder containing a character embedding encoder to obtain a sequence of character granularity semantic encoding feature vectors of the input data; a word embedding encoding unit, which is used to perform word segmentation processing on the input data and then pass it through a word granularity semantic encoder containing a word embedding encoder to obtain a sequence of word granularity semantic encoding feature vectors of the input data.
[0013] In the above-mentioned privacy computing platform based on data classification and grading, the semantic coding feature fusion module includes: a vector serial number labeling unit, which is used to label the serial numbers of each input data word granularity semantic coding feature vector in the sequence of the input data word granularity semantic coding feature vector to obtain a sequence of input data word granularity semantic labeling serial numbers; a vector energy value calculation unit, which is used to calculate the energy function value of each input data word granularity semantic coding feature vector in the sequence of the input data word granularity semantic coding feature vector to obtain a sequence of input data word granularity semantic energy function values, wherein the energy function value of each input data word granularity semantic coding feature vector is correlated with the mean and variance of each input data word granularity semantic coding feature vector and the semantic labeling serial number of each input data word granularity. off; an energy function value activation unit, used to input the inverse of each input data word granularity semantic energy function value in the sequence of input data word granularity semantic energy function values into a Sigmoid function to obtain a sequence of input data word granularity semantic activation energy function values; a feature position point multiplication unit, used to use each input data word granularity semantic activation energy function value in the sequence of input data word granularity semantic activation energy function values as a weight, and perform position point multiplication on each input data word granularity semantic coding feature vector in the sequence of input data word granularity semantic coding feature vectors to obtain a sequence of input data word granularity enhanced semantic coding feature vectors; a fusion unit, used to fuse the sequence of input data word granularity enhanced semantic coding feature vectors to obtain the input data word granularity fused semantic coding feature vector.
[0014] In the above-mentioned privacy computing platform based on data classification and grading, the vector energy value calculation unit is used to: calculate the mean of the word granularity semantic coding feature vector of the input data to obtain the word granularity semantic mean of the input data; calculate the variance of the word granularity semantic coding feature vector of the input data to obtain the word granularity semantic variance of the input data; multiply the value obtained by adding the word granularity semantic variance of the input data and the regularization term parameter by four to obtain the first energy function value of the word granularity semantic of the input data; calculate the word granularity semantic encoding feature vector of the input data corresponding to the word granularity semantic encoding feature vector of the input data; The square of the difference between the granularity semantic annotation serial number and the input data word granularity semantic mean is calculated to obtain the input data word granularity semantic offset value; the modulated input data word granularity semantic variance obtained by multiplying the input data word granularity semantic variance by two is added to the value obtained by multiplying the regularization item parameter by two and the input data word granularity semantic offset value to obtain the input data word granularity semantic second energy function value; the input data word granularity semantic first energy function value is divided by the input data word granularity semantic second energy function value to obtain the input data word granularity semantic energy function value.
[0015] In the above-mentioned privacy computing platform based on data classification and grading, the multi-granularity input data fusion module includes: an input data granularity feature cascade unit, which is used to cascade the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector to obtain the input data word granularity word granularity cascade feature vector; a first threshold vector calculation unit, which is used to calculate the matrix multiplication of the input data word granularity word granularity cascade feature vector and the first transformation matrix, and then perform vector addition with the first bias vector to obtain a first bias adjustment feature vector, and input the first bias adjustment feature vector into s An igmoid activation function is used to obtain a first input data word granularity word granularity fusion threshold vector; an input data granularity feature addition unit is used to add the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector according to position to obtain an input data word granularity word granularity sum feature vector; a second threshold vector calculation unit is used to calculate the matrix multiplication of the input data word granularity word granularity sum feature vector and the second transformation matrix, and then perform vector addition with the second bias vector to obtain a second bias adjustment feature vector, and input the second bias adjustment feature vector into s an igmoid activation function is used to obtain a second input data word granularity word granularity fusion threshold vector; an input data granularity feature interaction unit is used to perform position point multiplication on the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector to obtain an input data word granularity word granularity interaction feature vector; a third threshold vector calculation unit is used to calculate the matrix multiplication of the input data word granularity word granularity interaction feature vector and the third transformation matrix, and then perform vector addition with the third bias vector to obtain a third bias adjustment feature vector, and input the third bias adjustment feature vector into si gmoi d activation function to obtain the third input data word granularity word granularity fusion threshold vector; a threshold mean vector calculation unit, used to calculate the positional mean of the first input data word granularity word granularity fusion threshold vector, the second input data word granularity word granularity fusion threshold vector and the third input data word granularity word granularity fusion threshold vector to obtain the input data word granularity word granularity mean threshold vector; a multi-granularity fusion unit, used to use the threshold value of each position in the input data word granularity word granularity mean threshold vector as a weight to calculate the weighted sum of the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector to obtain the multi-granularity input data semantic coding feature vector; wherein, the range of each threshold value in the first input data word granularity word granularity fusion threshold vector, the second input data word granularity word granularity fusion threshold vector and the third input data word granularity word granularity fusion threshold vector is greater than or equal to zero and less than or equal to one.
[0016] In the above-mentioned privacy computing platform based on data classification and grading, the multi-granularity fusion unit is used to: calculate the product of the feature value of each position in the input data word granularity fusion semantic coding feature vector and the corresponding threshold value in the input data word granularity word granularity mean threshold vector to obtain an input data word granularity fusion semantic modulation feature vector composed of multiple input data word granularity fusion semantic feature values; calculate the product of the feature value of each position in the input data word granularity fusion semantic coding feature vector and the corresponding threshold value in the input data word granularity word granularity mean threshold vector to obtain an input data word granularity fusion semantic modulation feature vector composed of multiple input data word granularity fusion semantic feature values; add the feature value of each position in the input data word granularity fusion semantic modulation feature vector and the feature value of each position in the input data word granularity fusion semantic modulation feature vector by position to obtain the multi-granularity input data semantic coding feature vector.
[0017] In the above-mentioned privacy computing platform based on data classification and grading, the classification result generation module is used to: input the semantic encoding feature vector of the multi-granularity input data into the classifier to obtain the classification result.
[0018] The above-mentioned privacy computing platform based on data classification and grading also includes a training module for training the character-granularity semantic encoder including the character embedding encoder, the word-granularity semantic encoder including the word embedding encoder, the energy measurement attention fusion module based on sequence labeling, the balanced threshold feature vector adaptive fusion module and the classifier.
[0019] In the above-mentioned privacy computing platform based on data classification and grading, the training module includes: a training data acquisition unit for acquiring training data, wherein the training data includes training input data and a true classification result, wherein the true classification result is the true value of the training input data as personal data or industry data; a training input data word-granularity semantic coding feature generation unit for performing word segmentation processing on the training input data and then passing it through the word-granularity semantic encoder including the word embedding encoder to obtain a sequence of training input data word-granularity semantic coding feature vectors; a training input data word-granularity semantic coding feature generation unit for performing word segmentation processing on the training input data and then passing it through the word-granularity semantic encoder including the word embedding encoder to obtain a sequence of training input data word-granularity semantic coding feature vectors; a training input data fusion semantic coding feature generation unit for respectively inputting the sequence of the training input data word-granularity semantic coding feature vectors and the sequence of the training input data word-granularity semantic coding feature vectors into the energy measurement attention fusion module based on sequence labeling to obtain training data. Training input data word granularity fusion semantic coding feature vectors and training input data word granularity fusion semantic coding feature vectors; training multi-granularity input data semantic coding unit, used to input the training input data word granularity fusion semantic coding feature vectors and the training input data word granularity fusion semantic coding feature vectors into the equalization threshold feature vector adaptive fusion module to obtain training multi-granularity input data semantic coding feature vectors; training classification result generation unit, used to input the training multi-granularity input data semantic coding feature vectors into the classifier to obtain training classification results; classification loss unit, used to calculate the cross entropy loss function value between the training classification results and the true classification results to obtain a classification loss function value; model training unit, used to train the word granularity semantic encoder containing the word embedding encoder, the word granularity semantic encoder containing the word embedding encoder, the energy measurement attention fusion module based on sequence labeling, the equalization threshold feature vector adaptive fusion module and the classifier based on the classification loss function value and through gradient descent direction propagation.
[0020] Compared with the existing technology, the privacy computing platform based on data classification and grading provided by this application obtains input data and uses text data analysis and processing technology based on deep learning neural networks to perform character segmentation, word segmentation and semantic encoding on the input data. In this way, the input data is automatically evaluated as personal data or industry data based on the multi-granularity semantic coding features after the fusion of the character granularity and word granularity of the input data. In this way, the semantic content of the text can be deeply understood, which helps to classify and grade the data more accurately, while improving the automation level and efficiency of data processing and strengthening data privacy protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0022] Figure 1 This is a block diagram of a privacy computing platform based on data classification and grading according to an embodiment of the present application.
[0023] Figure 2 This is a schematic diagram of the architecture of a privacy computing platform based on data classification and grading according to an embodiment of the present application.
[0024] Figure 3 This is a block diagram of a training module in a privacy computing platform based on data classification and grading according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0026] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in a different order and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0027] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0028] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0029] When processing large amounts of text data, enterprises and institutions often face the challenges of identifying, classifying, and grading sensitive data. This data may contain personal information, trade secrets, or industry-specific data, and needs to be properly graded, managed, and protected according to different sensitivity levels and privacy requirements.
[0030] However, existing privacy computing for data classification and grading usually relies on manual review or simple rule matching. These methods are not only inefficient but also difficult to adapt to data diversity and complexity, resulting in low data classification accuracy.
[0031] Therefore, in response to the above technical problems, the technical concept of this application is to obtain input data and use text data analysis and processing technology based on deep learning neural networks to perform character segmentation, word segmentation and semantic encoding on the input data, so as to automatically evaluate whether the input data is personal data or industry data based on the multi-granularity semantic encoding features after the fusion of the character granularity and word granularity of the input data. In this way, the semantic content of the text can be deeply understood, which helps to classify and grade the data more accurately, while improving the automation level and efficiency of data processing and strengthening data privacy protection.
[0032] Figure 1 This is a block diagram of a privacy computing platform based on data classification and grading according to an embodiment of the present application. Figure 2 Schematic diagram of the architecture of the privacy computing platform based on data classification and grading according to the embodiment of the present application. Figure 1 and Figure 2 As shown, in the privacy computing platform 100 based on data classification and grading, it includes: an input data acquisition module 110 for acquiring input data; an input data semantic encoding module 120 for performing word-granularity semantic encoding and term-granularity semantic encoding on the input data to obtain a sequence of input data word-granularity semantic encoding feature vectors and a sequence of input data term-granularity semantic encoding feature vectors; a semantic encoding feature fusion module 130 for inputting the sequence of input data word-granularity semantic encoding feature vectors and the sequence of input data term-granularity semantic encoding feature vectors into the energy measurement attention fusion based on sequence annotation. The module is used to obtain the word-granularity fusion semantic coding feature vector of the input data and the word-granularity fusion semantic coding feature vector of the input data; the multi-granularity input data fusion module 140 is used to input the word-granularity fusion semantic coding feature vector of the input data and the word-granularity fusion semantic coding feature vector of the input data into the equalization threshold feature vector adaptive fusion module to obtain the multi-granularity input data semantic coding feature vector as the multi-granularity input data semantic coding feature; the classification result generation module 150 is used to obtain the classification result based on the multi-granularity input data semantic coding feature, and the classification result is used to indicate that the input data is personal data or industry data.
[0033] In an embodiment of the present application, the input data acquisition module 110 is used to obtain input data. It should be understood that the input data may contain personal information identifiers or industry terms and special terms, and the semantic information and content in these information have an important influence and effect on the subsequent data classification. Based on this, in the technical solution of the present application, the input data is obtained, and the content of the input data is semantically analyzed and understood, so as to accurately determine the data type of the input data. In particular, in a specific embodiment of the present application, the input data can be obtained from a background database.
[0034] In an embodiment of the present application, the input data semantic coding module 120 is used to perform character-based semantic coding and word-based semantic coding on the input data to obtain a sequence of character-based semantic coding feature vectors and a sequence of word-based semantic coding feature vectors of the input data. Specifically, in an embodiment of the present application, the input data semantic coding module includes: a character embedding coding unit, which is used to perform character segmentation processing on the input data and then pass it through a character-based semantic encoder containing a character embedding encoder to obtain a sequence of character-based semantic coding feature vectors of the input data; a word embedding coding unit, which is used to perform word segmentation processing on the input data and then pass it through a word-based semantic encoder containing a word embedding encoder to obtain a sequence of word-based semantic coding feature vectors of the input data. It should be understood that, considering that the input data contains a large amount of text information, this text information is crucial for understanding and analyzing the semantics in the input data. The word embedding encoder can convert the input at the character level into a continuous vector representation to better capture the semantic relationship and contextual information between characters. Based on this, in the technical solution of the present application, after the input data is processed by word segmentation, a word granularity semantic encoder including a word embedding encoder is used to obtain a sequence of word granularity semantic encoding feature vectors of the input data, which helps to capture the fine semantic features and information of the more subtle word granularity of the input data in the input data, thereby improving the final word semantic understanding and classification accuracy of the input data. Similarly, the input data is also composed of multiple words or phrases, and each word has a semantic relationship and context information about the input data. Considering that the word embedding encoder can map words to a continuous vector space, it relies on better capturing the semantic similarity and correlation between each word. Therefore, in the technical solution of the present application, after the input data is processed by word segmentation, a word granularity semantic encoder including a word embedding encoder is used to divide the input text data into word units, and capture and understand the semantic meaning of each word in the input text data and the correlation between the context, thereby obtaining a sequence of word granularity semantic encoding feature vectors of the input data with semantic representation capabilities.
[0035] In an embodiment of the present application, the semantic coding feature fusion module 130 is used to input the sequence of the input data word granularity semantic coding feature vectors and the sequence of the input data word granularity semantic coding feature vectors into the energy measurement attention fusion module based on sequence annotation to obtain the input data word granularity fusion semantic coding feature vectors and the input data word granularity fusion semantic coding feature vectors. It should be understood that, considering that the semantic importance and criticality of each input data word granularity semantic coding feature vector and the input data word granularity semantic coding feature vector in the sequence of the input data word granularity semantic coding feature vectors and the sequence of the input data word granularity semantic coding feature vectors are different in their respective entire sequences. Based on this, in order to better capture the important information and significant semantic relevance in each sequence data, in the technical solution of the present application, the sequence of the input data word granularity semantic coding feature vectors and the sequence of the input data word granularity semantic coding feature vectors are respectively input into the energy measurement attention fusion module based on sequence annotation to obtain the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector. Specifically, first, each word-granularity semantic coding feature vector of the input data in the sequence of word-granularity semantic coding feature vectors is labeled with a serial number to learn the order and correlation between each word-granularity feature vector. Then, the variance and mean of each word-granularity semantic coding feature vector of the input data are calculated respectively, and the energy function value of each vector is calculated based on the variance mean and the serial number to calculate the score or weight of each feature vector to reflect its importance in the sequence. Finally, each vector is weighted and fused based on the energy function value to obtain the word-granularity fused semantic coding feature vector of the input data, which can improve the model's recognition and utilization of key semantic association features between different granularity sequences of the input data, thereby improving the model performance. In particular, the processing method of the sequence of the word-granularity semantic coding feature vector of the input data is the same as the processing method of the sequence of the word-granularity semantic coding feature vector of the input data.
[0036] Specifically, in an embodiment of the present application, the semantic coding feature fusion module includes: a vector serial number labeling unit, used to label the serial numbers of the individual input data word granularity semantic coding feature vectors in the sequence of the input data word granularity semantic coding feature vectors to obtain a sequence of input data word granularity semantic labeling serial numbers; a vector energy value calculation unit, used to calculate the energy function value of the individual input data word granularity semantic coding feature vectors in the sequence of the input data word granularity semantic coding feature vectors to obtain a sequence of input data word granularity semantic energy function values, wherein the energy function value of each input data word granularity semantic coding feature vector is related to the mean and variance of each input data word granularity semantic coding feature vector and the semantic labeling serial number of each input data word granularity; A quantity function value activation unit is used to input the inverse of each input data word granularity semantic energy function value in the sequence of input data word granularity semantic energy function values into a Sigmoid function to obtain a sequence of input data word granularity semantic activation energy function values; a feature position point multiplication unit is used to use each input data word granularity semantic activation energy function value in the sequence of input data word granularity semantic activation energy function values as a weight, and perform position point multiplication on each input data word granularity semantic coding feature vector in the sequence of input data word granularity semantic coding feature vectors to obtain a sequence of input data word granularity enhanced semantic coding feature vectors; a fusion unit is used to fuse the sequence of input data word granularity enhanced semantic coding feature vectors to obtain the input data word granularity fused semantic coding feature vector.
[0037] More specifically, in an embodiment of the present application, the vector energy value calculation unit is used to: calculate the mean of the word granularity semantic encoding feature vector of the input data to obtain the word granularity semantic mean of the input data; calculate the variance of the word granularity semantic encoding feature vector of the input data to obtain the word granularity semantic variance of the input data; multiply the value obtained by adding the word granularity semantic variance of the input data and the regularization item parameter by four to obtain the first energy function value of the word granularity semantic of the input data; calculate the word granularity semantic of the input data corresponding to the word granularity semantic encoding feature vector of the input data; The square of the difference between the semantic annotation serial number and the input data word granularity semantic mean is calculated to obtain the input data word granularity semantic offset value; the modulated input data word granularity semantic variance obtained by multiplying the input data word granularity semantic variance by two is added to the value obtained by multiplying the regularization item parameter by two and the input data word granularity semantic offset value to obtain the input data word granularity semantic second energy function value; the input data word granularity semantic first energy function value is divided by the input data word granularity semantic second energy function value to obtain the input data word granularity semantic energy function value.
[0038] In the embodiment of the present application, specifically, the vector energy value calculation unit is used to: process the input data word granularity semantic encoding feature vector using the following energy function calculation formula to obtain the input data word granularity semantic energy function value; wherein, the energy function calculation formula is:
[0039]
[0040] Among them, s i is the eigenvalue of the i-th position in the semantic encoding feature vector of the word granularity of the input data, M is the number of eigenvalues in the semantic encoding feature vector of the word granularity of the input data, μ is the semantic mean of the word granularity of the input data, σ 2 is the semantic variance of the word granularity of the input data, λ is the regularization term parameter, t is the semantic annotation sequence number of the word granularity of the input data, It is the semantic energy function value of the word granularity of the input data.
[0041] In an embodiment of the present application, the multi-granularity input data fusion module 140 is used to input the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector into the balanced threshold feature vector adaptive fusion module to obtain the multi-granularity input data semantic coding feature vector as the multi-granularity input data semantic coding feature. Accordingly, in order to comprehensively utilize the input data semantic coding feature information of different granularities, thereby improving the model's overall semantic understanding and representation ability of the input data, in the technical solution of the present application, the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector are input into the balanced threshold feature vector adaptive fusion module to obtain the multi-granularity input data semantic coding feature vector. In detail, first, the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector are cascaded, added by position and multiplied by position to obtain cascade features, sum features and interaction features to enrich the model's understanding of multi-granularity input information. Next, threshold vectors are calculated based on the concatenated features, the summed features, and the interactive features to control the contribution of each feature fusion method. The mean vectors obtained by averaging the threshold vectors are then used to perform weighted fusion on the word-granularity fused semantic coding feature vectors and the word-granularity fused semantic coding feature vectors of the input data, adaptively balancing the importance of feature vectors of different granularities and effectively integrating information of these different granularities to obtain a multi-granularity input data semantic coding feature vector.
[0042] Specifically, in the embodiment of the present application, the multi-granularity input data fusion module includes: an input data granularity feature cascade unit, which is used to cascade the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector to obtain an input data word granularity word granularity cascade feature vector; a first threshold vector calculation unit, which is used to calculate the matrix multiplication of the input data word granularity word granularity cascade feature vector and the first transformation matrix, and then perform vector addition with the first bias vector to obtain a first bias adjustment feature vector, and input the first bias adjustment feature vector into s An igmoid activation function is used to obtain a first input data word granularity word granularity fusion threshold vector; an input data granularity feature addition unit is used to add the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector according to position to obtain an input data word granularity word granularity sum feature vector; a second threshold vector calculation unit is used to calculate the matrix multiplication of the input data word granularity word granularity sum feature vector and the second transformation matrix, and then perform vector addition with the second bias vector to obtain a second bias adjustment feature vector, and input the second bias adjustment feature vector into s An igmoid activation function is used to obtain a second input data word granularity word granularity fusion threshold vector; an input data granularity feature interaction unit is used to perform position point multiplication on the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector to obtain an input data word granularity word granularity interaction feature vector; a third threshold vector calculation unit is used to calculate the matrix multiplication of the input data word granularity word granularity interaction feature vector and the third transformation matrix, and then perform vector addition with the third bias vector to obtain a third bias adjustment feature vector, and input the third bias adjustment feature vector into s An igmoid activation function is used to obtain a third input data word granularity word granularity fusion threshold vector; a threshold mean vector calculation unit is used to calculate the positional mean of the first input data word granularity word granularity fusion threshold vector, the second input data word granularity word granularity fusion threshold vector and the third input data word granularity word granularity fusion threshold vector to obtain the input data word granularity word granularity mean threshold vector; a multi-granularity fusion unit is used to use the threshold value of each position in the input data word granularity word granularity mean threshold vector as a weight to calculate the weighted sum of the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector to obtain the multi-granularity input data semantic coding feature vector; wherein, the range of each threshold value in the first input data word granularity word granularity fusion threshold vector, the second input data word granularity word granularity fusion threshold vector and the third input data word granularity word granularity fusion threshold vector is greater than or equal to zero and less than or equal to one.
[0043] In the embodiment of the present application, specifically, the multi-granularity input data fusion module is used to: input the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector into the balanced threshold feature vector adaptive fusion module, and process them according to the following fusion formula to obtain the multi-granularity input data semantic coding feature vector; wherein, the fusion formula is:
[0044]
[0045] Among them, v1 and v2 represent the feature values of each position in the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector respectively, t1, t2 and t3 are the feature values of each position in the first input data word granularity word granularity fusion threshold vector, the second input data word granularity word granularity fusion threshold vector and the third input data word granularity word granularity fusion threshold vector respectively, t1, t2 and t3 are all ∈ [0,1], v c is the characteristic value of each position in the multi-granularity input data semantic coding feature vector, T1, T2 and T3 are respectively the first input data word granularity word granularity fusion threshold vector, the second input data word granularity word granularity fusion threshold vector and the third input data word granularity word granularity fusion threshold vector, V1 and V2 represent the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector, concat(·,·) represents the cascade operation, W1, W2 and W3 are respectively the first transformation matrix, the second transformation matrix and the third transformation matrix, b1, b2 and b3 are respectively the first bias vector, the second bias vector and the third bias vector, It is the addition by position, ⊙ is the point product by position, and sigmoid(·) represents the activation function.
[0046] More specifically, in an embodiment of the present application, the multi-granularity fusion unit is used to: calculate the product of the feature value of each position in the input data word granularity fusion semantic coding feature vector and the corresponding threshold value in the input data word granularity word granularity mean threshold vector to obtain an input data word granularity fusion semantic modulation feature vector composed of multiple input data word granularity fusion semantic feature values; calculate the product of the feature value of each position in the input data word granularity fusion semantic coding feature vector and the corresponding threshold value in the input data word granularity word granularity mean threshold vector to obtain an input data word granularity fusion semantic modulation feature vector composed of multiple input data word granularity fusion semantic feature values; add the feature value of each position in the input data word granularity fusion semantic modulation feature vector and the feature value of each position in the input data word granularity fusion semantic modulation feature vector by position to obtain the multi-granularity input data semantic coding feature vector.
[0047] In an embodiment of the present application, the classification result generation module 150 is used to obtain a classification result based on the semantic coding features of the multi-granularity input data, and the classification result is used to indicate that the input data is personal data or industry data. Specifically, in an embodiment of the present application, the classification result generation module is used to: input the semantic coding feature vector of the multi-granularity input data into a classifier to obtain the classification result. That is, the multi-granularity input data semantic coding features obtained by adaptively fusion of the word-granularity fusion semantic coding feature vector of the input data and the word-granularity fusion semantic coding feature vector of the input data are used for classification processing, so as to automatically evaluate whether the input data is personal data or industry data. In this way, the semantic content of the text can be deeply understood, which helps to classify and grade the data more accurately, while improving the automation level and efficiency of data processing and strengthening data privacy protection.
[0048] It is worth mentioning that those skilled in the art should be aware that before applying a deep neural network model for inference, the deep neural network model must first be trained so that the deep neural network can implement specific functional capabilities.
[0049] Specifically, in the technical solution of the present application, the privacy computing platform based on data classification and grading also includes a training module for training the character-granularity semantic encoder including the character embedding encoder, the word-granularity semantic encoder including the word embedding encoder, the energy measurement attention fusion module based on sequence labeling, the balanced threshold feature vector adaptive fusion module and the classifier.
[0050] Figure 3 FIG is a block diagram of a training module in a privacy computing platform based on data classification and grading according to an embodiment of the present application. Figure 3As shown, the training module 200 includes: a training data acquisition unit 210, which is used to acquire training data, wherein the training data includes training input data and a true classification result, wherein the true classification result is the true value of the training input data as personal data or industry data; a training input data word granularity semantic coding feature generation unit 220, which is used to perform word segmentation processing on the training input data and then pass it through the word granularity semantic encoder including the word embedding encoder to obtain a sequence of training input data word granularity semantic coding feature vectors; a training input data word granularity semantic coding feature generation unit 230, which is used to perform word segmentation processing on the training input data and then pass it through the word granularity semantic encoder including the word embedding encoder to obtain a sequence of training input data word granularity semantic coding feature vectors; a training input data fusion semantic coding feature generation unit 240, which is used to input the sequence of the training input data word granularity semantic coding feature vectors and the sequence of the training input data word granularity semantic coding feature vectors into the energy measurement attention fusion module based on sequence labeling respectively to obtain the training input data word granularity semantic coding feature vectors. The word granularity fusion semantic coding feature vector and the training input data word granularity fusion semantic coding feature vector are input into the training input data word granularity fusion semantic coding feature vector; the training multi-granularity input data semantic coding unit 250 is used to input the training input data word granularity fusion semantic coding feature vector and the training input data word granularity fusion semantic coding feature vector into the equalization threshold feature vector adaptive fusion module to obtain the training multi-granularity input data semantic coding feature vector; the training classification result generation unit 260 is used to input the training multi-granularity input data semantic coding feature vector into the classifier to obtain the training classification result; the classification loss unit 270 is used to calculate the cross entropy loss function value between the training classification result and the true classification result to obtain the classification loss function value; the model training unit 280 is used to train the word granularity semantic encoder including the word embedding encoder, the word granularity semantic encoder including the word embedding encoder, the energy measurement attention fusion module based on sequence labeling, the equalization threshold feature vector adaptive fusion module and the classifier based on the classification loss function value and through gradient descent direction propagation.
[0051] It should be understood that in the embodiment of the present application, the sequence of the word-granularity semantic coding feature vectors of the training input data and the sequence of the word-granularity semantic coding feature vectors of the training input data encode semantic feature representations of the training input data based on different local semantic space segmentation scales of the source text, thereby causing differences in the text semantic feature distributions of the word-granularity fusion semantic coding feature vectors of the training input data and the word-granularity fusion semantic coding feature vectors based on the difference in attention weights of the local semantic space sequence energy measurement, so that the training multi-granularity input data semantic coding feature vectors obtained by adaptively fusion of the training input data word-granularity fusion semantic coding feature vectors through the balanced threshold feature vector have feature aggregation class representation imbalance, and make the probability density distribution convergence logic of the in-class features and out-of-class features after the clustering operation inconsistent.
[0052] Based on this, in this preferred embodiment, when the character granularity semantic encoder including the character embedding encoder, the word granularity semantic encoder including the word embedding encoder, the energy measurement attention fusion module based on sequence labeling, the balanced threshold feature vector adaptive fusion module and the classifier are trained based on the classification loss function value, wherein, in each iteration of the training, the semantic encoding feature vector of the training multi-granularity input data is iteratively optimized,
[0053] The iterative optimization process specifically includes the following steps: clustering the semantic coding feature vector of the training multi-granularity input data, for example, clustering based on the distance between eigenvalues, and determining the number of in-class features and the number of out-of-class features; dividing the total number of eigenvalues of the semantic coding feature vector of the training multi-granularity input data by the number of in-class features and the number of out-of-class features to obtain a class importance value and an out-of-class proportion value, and calculating the inverse of the class importance value to obtain a class constraint value; calculating each eigenvalue of the semantic coding feature vector of the training multi-granularity input data to obtain the class importance value; A power function with an exponent of the class constraint value is added to an exponential value with a natural function as the base and the class constraint value as the exponent, and then multiplied by the out-of-class proportion value to obtain a training multi-granularity entry data semantic coding modulation vector; the training multi-granularity entry data semantic coding feature vector is point multiplied with the class importance value to obtain a training multi-granularity entry data semantic coding ontology vector; the weighted sum of the training multi-granularity entry data semantic coding modulation vector and the training multi-granularity entry data semantic coding ontology vector is calculated with a weight hyperparameter to obtain an optimized training multi-granularity entry data semantic coding feature vector.
[0054] Specifically, in this preferred embodiment, the semantic encoding feature vector of the training multi-granularity input data is optimized to obtain an optimized semantic encoding feature vector of the training multi-granularity input data. The process formula is expressed as follows:
[0055]
[0056] Wherein, V is the semantic encoding feature vector of the training multi-granularity input data, n is the total number of eigenvalues of the semantic encoding feature vector of the training multi-granularity input data, k is the number of intra-class features of the semantic encoding feature vector of the training multi-granularity input data, ⊙ is the position point multiplication, β is the weight hyperparameter, It is added by position. represents an exponential operation on the feature values at each position in the feature vector, and V' is the optimized semantic encoding feature vector of the training multi-granularity input data.
[0057] Therefore, while performing clustering operations on the semantically encoded feature vectors of the training multi-granularity input data, the clustering importance measure of the eigenvalues of the semantically encoded feature vectors of the training multi-granularity input data is combined with the class constraint measure of the clustering operation to further perform responsive modulation with the out-of-class proportional factor, and based on the clustering scaling of the feature ontology representation of the feature vectors of the training multi-granularity input data, the simplicity and effectiveness of the feature purification of the feature set of the semantically encoded feature vectors of the training multi-granularity input data is established, and the logical inconsistency of the convergence of the probability density distribution caused by the clustering operation is suppressed, so as to improve the classification iteration effect of the feature set input classifier of the semantically encoded feature vectors of the training multi-granularity input data, that is, to improve the speed of classification training and the accuracy of classification results. In this way, the semantic content of the text can be deeply understood, which helps to classify and grade data more accurately, while improving the automation level and efficiency of data processing and strengthening data privacy protection.
[0058] In summary, the privacy computing platform 100 based on data classification and grading according to the embodiment of the present application is described. It obtains input data and uses text data analysis and processing technology based on deep learning neural networks to perform character segmentation, word segmentation, and semantic encoding on the input data. In this way, the input data is automatically evaluated as personal data or industry data based on the multi-granularity semantic coding features after the fusion of the character granularity and word granularity of the input data. In this way, the semantic content of the text can be deeply understood, which helps to classify and grade the data more accurately, while improving the automation level and efficiency of data processing and strengthening data privacy protection.
[0059] As described above, the privacy computing platform 100 based on data classification and grading according to the embodiment of the present application can be implemented in various terminal devices, such as a server for privacy computing based on data classification and grading. In one example, the privacy computing platform 100 based on data classification and grading according to the embodiment of the present application can be integrated into the terminal device as a software module and / or hardware module. For example, the privacy computing platform 100 based on data classification and grading can be a software module in the operating system of the terminal device, or it can be an application developed for the terminal device; of course, the privacy computing platform 100 based on data classification and grading can also be one of the many hardware modules of the terminal device.
[0060] Alternatively, in another example, the privacy computing platform 100 based on data classification and grading and the terminal device may also be separate devices, and the privacy computing platform 100 based on data classification and grading may be connected to the terminal device via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.
[0061] The foregoing is merely an example of the principles of the present disclosure, and various modifications may be made by those skilled in the art without departing from the scope of the present disclosure. The above embodiments are presented for the purpose of illustration and not limitation. The present disclosure may also take many forms other than those explicitly described herein.
Claims
1. A privacy computing platform based on data classification and grading, characterized by: include: Input data acquisition module, used to obtain input data; An input data semantic coding module is used to perform character-based semantic coding and word-based semantic coding on the input data to obtain a sequence of input data character-based semantic coding feature vectors and a sequence of input data word-based semantic coding feature vectors; A semantic coding feature fusion module, configured to input the sequence of the input data word-granularity semantic coding feature vectors and the sequence of the input data word-granularity semantic coding feature vectors into an energy measurement attention fusion module based on sequence annotation to obtain input data word-granularity fused semantic coding feature vectors and input data word-granularity fused semantic coding feature vectors; A multi-granularity input data fusion module is used to input the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector into a balanced threshold feature vector adaptive fusion module to obtain a multi-granularity input data semantic coding feature vector as a multi-granularity input data semantic coding feature; A classification result generating module, configured to obtain a classification result based on the semantic coding features of the multi-granularity input data, wherein the classification result is used to indicate whether the input data is personal data or industry data; The multi-granularity input data fusion module includes: first, respectively cascading the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector, adding them by position and multiplying them by position to obtain cascade features, sum features and interaction features to enrich the model's understanding of multi-granularity input information; then, calculating threshold vectors based on the cascade features, the sum features and the interaction features to control the contribution of each feature fusion method; then, using the mean vector after the mean calculation of each threshold vector, respectively weightedly fuse the input data word granularity fusion semantic coding feature vector and the input data word granularity fusion semantic coding feature vector to adaptively balance the importance of feature vectors of different granularities, and effectively integrate these different granularity information together to obtain a multi-granularity input data semantic coding feature vector.
2. The privacy computing platform based on data classification and grading according to claim 1 is characterized in that: The input data semantic encoding module includes: A word embedding encoding unit, configured to perform word segmentation processing on the input data and then pass the result through a word granularity semantic encoder including a word embedding encoder to obtain a sequence of word granularity semantic encoding feature vectors of the input data; A word embedding encoding unit is used to perform word segmentation processing on the input data and then pass it through a word granularity semantic encoder including a word embedding encoder to obtain a sequence of word granularity semantic encoding feature vectors of the input data.
3. The privacy computing platform based on data classification and grading according to claim 2 is characterized in that: The semantic coding feature fusion module includes: A vector serial number marking unit is used to mark each input data word granularity semantic coding feature vector in the sequence of input data word granularity semantic coding feature vectors with a serial number to obtain a sequence of input data word granularity semantic marking serial numbers; a vector energy value calculation unit, configured to calculate an energy function value of each recorded data word granularity semantic coding feature vector in the sequence of recorded data word granularity semantic coding feature vectors to obtain a sequence of recorded data word granularity semantic energy function values, wherein the energy function value of each recorded data word granularity semantic coding feature vector is related to a mean and a variance of each recorded data word granularity semantic coding feature vector and a sequence number of each recorded data word granularity semantic annotation; An energy function value activation unit, configured to input the reciprocal of each input data word granularity semantic energy function value in the sequence of input data word granularity semantic energy function values into a Sigmoid function to obtain a sequence of input data word granularity semantic activation energy function values; a feature positional point multiplication unit, configured to perform positional point multiplication on each input data word granularity semantic coding feature vector in the sequence of input data word granularity semantic activation energy function values using each input data word granularity semantic activation energy function value in the sequence of input data word granularity semantic activation energy function values as a weight, so as to obtain a sequence of input data word granularity enhanced semantic coding feature vectors; A fusion unit is used to fuse the sequence of the input data word granularity enhanced semantic coding feature vectors to obtain the input data word granularity fused semantic coding feature vector.
4. The privacy computing platform based on data classification and grading according to claim 3 is characterized in that: The vector energy value calculation unit is used to: Calculating the mean of the word granularity semantic encoding feature vectors of the input data to obtain the word granularity semantic mean of the input data; Calculating the variance of the word granularity semantic encoding feature vector of the input data to obtain the word granularity semantic variance of the input data; The value obtained by adding the word granularity semantic variance of the input data and the regularization term parameter is multiplied by four to obtain the first energy function value of the word granularity semantics of the input data; Calculating the square of the difference between the input data word granularity semantic annotation sequence number corresponding to the input data word granularity semantic encoding feature vector and the input data word granularity semantic mean to obtain an input data word granularity semantic offset value; The modulated recorded data word granularity semantic variance obtained by multiplying the recorded data word granularity semantic variance by two is added to a value obtained by multiplying the regularization term parameter by two and the recorded data word granularity semantic offset value to obtain a recorded data word granularity semantic second energy function value; The first energy function value of the word granularity semantics of the input data is divided by the second energy function value of the word granularity semantics of the input data to obtain the energy function value of the word granularity semantics of the input data.
5. The privacy computing platform based on data classification and grading according to claim 4 is characterized in that: The multi-granularity input data fusion module includes: An input data granularity feature cascade unit, configured to cascade the input data word granularity fusion semantic coding feature vector and the input data term granularity fusion semantic coding feature vector to obtain an input data word granularity word granularity cascade feature vector; A first threshold vector calculation unit is configured to calculate a first bias-adjusted feature vector by multiplying the concatenated feature vector of the input data word-granularity and word-granularity by a first transformation matrix, and then performing vector addition on the concatenated feature vector of the input data word-granularity and word-granularity and a first bias vector, and inputting the first bias-adjusted feature vector into a sigmoid activation function to obtain a first input data word-granularity and word-granularity fusion threshold vector; An input data granularity feature adding unit, configured to add the input data word granularity fusion semantic coding feature vector and the input data term granularity fusion semantic coding feature vector according to position to obtain an input data word granularity word granularity sum feature vector; A second threshold vector calculation unit is used to calculate the matrix multiplication of the word granularity sum feature vector of the input data and the second transformation matrix, and then perform vector addition with the second bias vector to obtain a second bias-adjusted feature vector, and input the second bias-adjusted feature vector into a sigmoid activation function to obtain a second word granularity fusion threshold vector of the input data; An input data granularity feature interaction unit, configured to perform positional point multiplication on the input data word granularity fusion semantic coding feature vector and the input data term granularity fusion semantic coding feature vector to obtain an input data word granularity term granularity interaction feature vector; A third threshold vector calculation unit is configured to calculate a third bias-adjusted feature vector by multiplying the word-granularity interaction feature vector of the input data by a third transformation matrix and then performing vector addition with a third bias vector, and input the third bias-adjusted feature vector into a sigmoid activation function to obtain a third word-granularity fusion threshold vector of the input data; a threshold mean vector calculation unit, configured to calculate the positional mean of the first input data word granularity word granularity fusion threshold vector, the second input data word granularity word granularity fusion threshold vector, and the third input data word granularity word granularity fusion threshold vector to obtain an input data word granularity word granularity mean threshold vector; a multi-granularity fusion unit, configured to calculate a weighted sum of the word-granularity fusion semantic coding feature vector of the input data and the word-granularity fusion semantic coding feature vector of the input data, using the threshold value of each position in the word-granularity mean threshold vector of the input data as a weight, to obtain the multi-granularity input data semantic coding feature vector; Among them, the range of each threshold value in the first input data word granularity word granularity fusion threshold vector, the second input data word granularity word granularity fusion threshold vector and the third input data word granularity word granularity fusion threshold vector is greater than or equal to zero and less than or equal to one.
6. The privacy computing platform based on data classification and grading according to claim 5 is characterized in that: The multi-granularity fusion unit is used to: Calculating the product of the feature value of each position in the input data word granularity fusion semantic coding feature vector and the corresponding threshold value in the threshold vector of one minus the input data word granularity word granularity mean to obtain an input data word granularity fusion semantic modulation feature vector composed of multiple input data word granularity fusion semantic feature values; Calculating the product of the feature value of each position in the input data word granularity fusion semantic coding feature vector and the corresponding threshold value in the input data word granularity word granularity mean threshold vector to obtain an input data word granularity fusion semantic modulation feature vector composed of multiple input data word granularity fusion semantic feature values; The feature values of each position in the word granularity fusion semantic modulation feature vector of the input data are added to the feature values of each position in the word granularity fusion semantic modulation feature vector of the input data to obtain the multi-granularity input data semantic coding feature vector.
7. The privacy computing platform based on data classification and grading according to claim 6, characterized in that: The classification result generating module is used to input the semantic encoding feature vector of the multi-granularity input data into a classifier to obtain the classification result.
8. The privacy computing platform based on data classification and grading according to claim 7 is characterized in that: It also includes a training module for training the character granularity semantic encoder including the character embedding encoder, the word granularity semantic encoder including the word embedding encoder, the sequence labeling-based energy measurement attention fusion module, the equalization threshold feature vector adaptive fusion module and the classifier.
9. The privacy computing platform based on data classification and grading according to claim 8, characterized in that: The training module includes: A training data acquisition unit is used to acquire training data, wherein the training data includes training input data and a true classification result, wherein the true classification result is a true value of the training input data being personal data or industry data; A training input data word granularity semantic coding feature generation unit is used to perform word segmentation processing on the training input data and then pass it through the word granularity semantic encoder including the word embedding encoder to obtain a sequence of word granularity semantic coding feature vectors of the training input data; A training input data word granularity semantic coding feature generation unit is used to perform word segmentation processing on the training input data and then pass it through the word granularity semantic encoder including the word embedding encoder to obtain a sequence of training input data word granularity semantic coding feature vectors; A training input data fusion semantic coding feature generation unit, configured to input the sequence of the training input data word granularity semantic coding feature vectors and the sequence of the training input data word granularity semantic coding feature vectors into the sequence-labeled energy measurement attention fusion module to obtain the training input data word granularity fusion semantic coding feature vectors and the training input data word granularity fusion semantic coding feature vectors; A training multi-granularity input data semantic coding unit is used to input the training input data word granularity fusion semantic coding feature vector and the training input data word granularity fusion semantic coding feature vector into the balanced threshold feature vector adaptive fusion module to obtain the training multi-granularity input data semantic coding feature vector; A training classification result generating unit, configured to input the semantically encoded feature vector of the training multi-granularity input data into the classifier to obtain a training classification result; A classification loss unit, configured to calculate a cross entropy loss function value between the training classification result and the true classification result to obtain a classification loss function value; A model training unit is used to train the character-granularity semantic encoder including the character embedding encoder, the word-granularity semantic encoder including the word embedding encoder, the sequence labeling-based energy measurement attention fusion module, the equalization threshold feature vector adaptive fusion module and the classifier based on the classification loss function value and through gradient descent direction propagation.
Citation Information
Patent Citations
Fused attention model-based Chinese text classification method
CN108595590A