Enterprise sensitive data desensitization method and system based on natural language
By detecting and classifying enterprise data formats, extracting text, image and numerical data features, combining sensitivity and utility for blurring, the problem of insufficient flexibility of desensitization methods in the prior art is solved, and flexible and privacy-protected data processing is achieved.
Patent Information
- Application Number
- CN202510375024.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
Existing enterprise data desensitization methods lack flexibility, are difficult to adapt to new types of sensitive data, and destroys data relevance, resulting in the desensitized data being unable to be used directly.
By detecting the format of enterprise data, the characteristic characters of text, images and numerical data are extracted, and the character sensitivity and utility of fuzzing is combined to mark the desensitized characteristic characters, and the data endpoint value is used to map the numerical data to realize data update.
It improves the desensitization flexibility of enterprise sensitive data, ensures the privacy protection and relevance of data, and adapts to data needs in different application scenarios.
Smart Images

Figure CN120296786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for desensitizing enterprise sensitive data based on natural language, and belongs to the technical field of data management. Background Art
[0002] Enterprise data refers to various digital information generated, collected, stored, and used by an enterprise during its operation. These data cover all aspects of the enterprise's business and are one of the important assets of the enterprise. For example, the data of an e-commerce enterprise includes user registration information (such as name, contact information, address, etc.), transaction data (such as order details, payment information, product reviews, etc.), product information (such as product name, price, inventory, etc.), and operation data (such as traffic statistics, marketing campaign effects, etc.), which contain a lot of sensitive information of users. In order to improve information security, it is necessary to desensitize enterprise data.
[0003] The existing enterprise data desensitization method is a rule-based data replacement method. According to the set rules, specific values are used to replace sensitive data. For example, some digits of the ID number are replaced with fixed symbols, or salary data is replaced with a fixed range of values. However, this method lacks flexibility. When new types of sensitive data appear, it is difficult to update the rules in a timely manner, and it will also destroy the original relevance of the data, resulting in the inability to directly use the desensitized data. Therefore, a method that can improve the flexibility of enterprise sensitive data desensitization is needed. Summary of the Invention
[0004] The present invention provides a method and system for desensitizing enterprise sensitive data based on natural language, and its main purpose is to solve the problem of low flexibility in desensitizing enterprise sensitive data.
[0005] To achieve the above object, a method for desensitizing enterprise sensitive data based on natural language provided by the present invention includes:
[0006] Obtain the original enterprise data to be processed, detect the data format corresponding to the original enterprise data, and based on the data format, classify the original enterprise data to obtain classified enterprise data, where the classified enterprise data includes: enterprise text data, enterprise image data, and enterprise numerical data;
[0007] Extract character features from the enterprise text data to obtain text feature characters, calculate the character sensitivity corresponding to the text feature characters, query the data application scenario corresponding to the enterprise text data, and combine the data application scenario and the text feature characters to evaluate the character utility corresponding to the text feature characters;
[0008] Combined with the character sensitivity and the character utility, mark the desensitized feature characters among the text feature characters, and perform semantic blurring processing on the desensitized feature characters to obtain blurred feature characters;
[0009] Mark the image desensitization areas corresponding to each image in the enterprise image data, analyze the regional attributes corresponding to the image desensitization areas, and based on the regional attributes, perform regional blurring processing on the image desensitization areas to obtain blurred image areas;
[0010] Identify the data endpoint values in the enterprise numerical data, calculate the desensitized mapping values corresponding to the enterprise numerical data based on the data endpoint values, and perform desensitization processing on the enterprise numerical data according to the desensitized mapping values to obtain target numerical data;
[0011] Use the target numerical data, the blurred image areas, and the blurred feature characters to perform data update processing on the classified enterprise data to obtain the enterprise desensitized data corresponding to the original enterprise data.
[0012] Optionally, the extracting character features from the enterprise text data to obtain text feature characters includes:
[0013] Perform text segmentation processing on the enterprise text data to obtain enterprise text segments;
[0014] Perform part-of-speech tagging processing on the enterprise text segments to obtain the part-of-speech of the segments, and count the segment word frequencies corresponding to the enterprise text segments;
[0015] Identify the segment entities in the enterprise text segments, and combine the segment entities, the part-of-speech of the segments, and the segment word frequencies to extract character features from the enterprise text data to obtain text feature characters.
[0016] Optionally, the calculating the character sensitivity corresponding to the text feature characters includes:
[0017] Perform standardization processing on the text feature characters to obtain standard feature characters;
[0018] Query the privacy information metrics corresponding to the original enterprise data, and calculate the privacy correlation degree between the standard feature characters and the privacy information metrics;
[0019] Evaluate the influence weight of the standard feature characters on the privacy information metrics;
[0020] Combine the influence weight and the privacy correlation degree to calculate the character sensitivity corresponding to the text feature characters.
[0021] Optionally, calculating the privacy correlation degree between the standard feature character and the privacy information metric includes:
[0022] Calculating the information entropy corresponding to the standard feature character and the privacy information metric respectively to obtain the character information entropy and the metric information entropy;
[0023] Calculating the joint information entropy between the standard feature character and the privacy information metric;
[0024] Combining the character information entropy, the metric information entropy, and the joint information entropy, the privacy correlation degree between the standard feature character and the privacy information metric can be calculated through the following formula:
[0025]
[0026] Wherein, A represents the privacy correlation degree between the standard feature character and the privacy information metric, C(a, b) represents the joint information entropy between the a-th character in the standard feature character and the b-th metric in the privacy information metric, B(a) represents the character information entropy corresponding to the a-th character in the standard feature character, D(b) represents the metric information entropy corresponding to the b-th metric in the privacy information metric, and a and b respectively represent the serial numbers corresponding to the standard feature character and the privacy information metric.
[0027] Optionally, combining the data application scenario and the text feature character to evaluate the character utility degree corresponding to the text feature character includes:
[0028] Performing vectorization processing on the data application scenario and the text feature character respectively to obtain an application scenario vector and a text character vector;
[0029] Extracting the feature vectors in the application scenario vector and the text character vector respectively to obtain a scenario feature vector and a character feature vector;
[0030] Calculating the vector similarity between the scenario feature vector and the character feature vector;
[0031] Evaluating the character utility degree corresponding to the text feature character according to the vector similarity.
[0032] Optionally, performing semantic fuzzification processing on the desensitized feature character to obtain a fuzzy feature character includes:
[0033] Performing semantic parsing on the desensitized feature character to obtain the feature character semantics;
[0034] Calculating the semantic similarity between the feature character semantics and each semantics in the preset semantic library;
[0035] When the semantic similarity is greater than a preset threshold, schedule the fuzzy semantics corresponding to the semantic of the feature character from the preset semantic library;
[0036] Use the fuzzy semantics to perform fuzzy processing on the semantic of the feature character to obtain a fuzzy feature character.
[0037] Optionally, marking the image desensitization area corresponding to each image in the enterprise image data includes:
[0038] Perform image denoising processing on each image in the enterprise image data to obtain a denoised enterprise image;
[0039] Perform image enhancement processing on the denoised enterprise image to obtain an enhanced enterprise image;
[0040] Identify the image connotation in the enhanced enterprise image and analyze the connotation metadata corresponding to the image connotation;
[0041] Parse the sensitive metadata in the connotation metadata, and based on the sensitive metadata, mark the image desensitization area corresponding to each image in the enterprise image data.
[0042] Optionally, the performing region blur processing on the image desensitization area based on the region attribute to obtain a blurred image area includes:
[0043] Extract the region color attribute, region texture attribute, and region shape attribute in the region attribute;
[0044] Based on the region color attribute, calculate the region color entropy corresponding to the image desensitization area;
[0045] Based on the region texture attribute, calculate the region texture entropy corresponding to the image desensitization area;
[0046] Based on the region shape attribute, calculate the shape complexity corresponding to the image desensitization area;
[0047] Combine the region color entropy, the region texture entropy, and the shape complexity to set the fuzzy priority corresponding to the image desensitization area;
[0048] Based on the fuzzy priority, perform region blur processing on the image desensitization area to obtain a blurred image area.
[0049] Optionally, the calculating the desensitization mapping value corresponding to the enterprise numerical data based on the data endpoint value includes:
[0050]
[0051] Where E represents the desensitization mapping value corresponding to the enterprise numerical data, and Fi represents the value corresponding to the i-th data in the enterprise numerical data, where i represents the serial number corresponding to the enterprise numerical data, minF represents the minimum value among the data endpoint values, maxF represents the maximum value among the data endpoint values, and H , max represents the upper limit value of the desensitized value range, H , min represents the lower limit value of the desensitized value range.
[0052] A natural language-based enterprise sensitive data desensitization system, characterized in that the system includes:
[0053] A data classification module for obtaining the original enterprise data to be processed, detecting the data format corresponding to the original enterprise data, and based on the data format, classifying the original enterprise data to obtain classified enterprise data, where the classified enterprise data includes: enterprise text data, enterprise image data, and enterprise numerical data;
[0054] A character utility evaluation module for extracting character features from the enterprise text data to obtain text feature characters, calculating the character sensitivity corresponding to the text feature characters, querying the data application scenarios corresponding to the enterprise text data, and combining the data application scenarios and the text feature characters to evaluate the character utility corresponding to the text feature characters;
[0055] A character fuzzing processing module for marking the desensitization feature characters in the text feature characters by combining the character sensitivity and the character utility, and performing semantic fuzzing processing on the desensitization feature characters to obtain fuzzy feature characters;
[0056] An image fuzzing processing module for marking the image desensitization area corresponding to each image in the enterprise image data, analyzing the area attributes corresponding to the image desensitization area, and based on the area attributes, performing area fuzzing processing on the image desensitization area to obtain a fuzzy image area;
[0057] A numerical desensitization processing module for identifying the data endpoint values in the enterprise numerical data, calculating the desensitization mapping value corresponding to the enterprise numerical data based on the data endpoint values, and performing desensitization processing on the enterprise numerical data according to the desensitization mapping value to obtain target numerical data;
[0058] A data update module for using the target numerical data, the fuzzy image area, and the fuzzy feature characters to perform data update processing on the classified enterprise data to obtain the enterprise desensitized data corresponding to the original enterprise data.
[0059] Compared with the problems described in the background art, the present invention can understand the data structure characteristics corresponding to the original enterprise data by detecting the data format corresponding to the original enterprise data, and classify the original enterprise data based on the data format, so as to group the data of the same type in the original enterprise data together, which is convenient for subsequent targeted analysis and processing of the data. The present invention extracts character features from the enterprise text data, can obtain representative characters in the enterprise text data, reduces the calculation amount during subsequent data processing, and provides a basis for the evaluation of the character utility corresponding to the subsequent text feature characters. The present invention combines the character sensitivity and the character utility to mark the desensitized feature characters in the text feature characters, thereby improving the marking accuracy of the desensitized feature characters, and performs semantic blurring processing on the desensitized feature characters, thereby accurately protecting the privacy of the enterprise text data. The present invention marks the image desensitized area corresponding to each image in the enterprise image data, can obtain the sensitive image positions of each image in the enterprise image data, and thus lays a foundation for subsequent area blurring processing of the image desensitized area. The present invention calculates the desensitized mapping value corresponding to the enterprise numerical data based on the data endpoint value, can obtain the value for desensitizing the enterprise numerical data, is convenient for subsequent desensitization processing of the enterprise numerical data, and thus is convenient for effectively hiding the enterprise numerical data. The present invention updates the classified enterprise data by using the target numerical data, the blurred image area and the blurred feature characters, thereby effectively desensitizing the classified enterprise data and improving the desensitization flexibility of the classified enterprise data. Therefore, the present invention proposes a method and system for desensitizing enterprise sensitive data based on natural language to improve the desensitization flexibility of enterprise sensitive data. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 FIG. is a schematic flowchart of a method for desensitizing enterprise sensitive data based on natural language provided by an embodiment of the present invention;
[0061] Figure 2 FIG. is a functional module diagram for implementing the system for desensitizing enterprise sensitive data based on natural language provided by an embodiment of the present invention.
[0062] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0064] An embodiment of the present application provides a method for desensitizing enterprise sensitive data based on natural language. The execution subject of the method for desensitizing enterprise sensitive data based on natural language includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for desensitizing enterprise sensitive data based on natural language can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.
[0065] Embodiment 1:
[0066] Referring to Figure 1 As shown, it is a schematic flowchart of a method for desensitizing enterprise sensitive data based on natural language provided by an embodiment of the present invention. In this embodiment, the method for desensitizing enterprise sensitive data based on natural language includes:
[0067] S1. Obtain the original enterprise data to be processed, detect the data format corresponding to the original enterprise data, and based on the data format, classify the original enterprise data to obtain classified enterprise data, where the classified enterprise data includes: enterprise text data, enterprise image data, and enterprise numerical data.
[0068] By detecting the data format corresponding to the original enterprise data in the present invention, the data structure characteristics corresponding to the original enterprise data can be understood, and based on the data format, classifying the original enterprise data can group the same type of data in the original enterprise data together, thereby facilitating subsequent targeted analysis and processing of the data. It should be explained that the original enterprise data is a general term for various digital information generated, collected, stored, and used by an enterprise during its operation. The data format is the organization and presentation method of the original enterprise data, which determines the storage, transmission, and processing methods of the data. The enterprise text data is the data content existing in the form of text in the classified enterprise data, such as documents, reports, emails, etc. The enterprise image data is the data part presented in the form of an image in the classified enterprise data, such as product pictures, screenshots of surveillance videos, etc. The enterprise numerical data is the data category expressed in the form of numbers in the classified enterprise data, such as specific numerical values in financial data, sales data, etc.; further, the detection of the data format corresponding to the original enterprise data can be implemented through a data format detection tool, such as the csvfingerprint tool, and the classification processing of the original enterprise data can be implemented through a classification function, such as the os.path.splitext() function.
[0069] S2. Extract character features from the enterprise text data to obtain text feature characters, calculate the character sensitivity corresponding to the text feature characters, query the data application scenarios corresponding to the enterprise text data, and evaluate the character utility corresponding to the text feature characters in combination with the data application scenarios and the text feature characters.
[0070] In the present invention, extracting character features from the enterprise text data can obtain representative characters in the enterprise text data, reduce the computational amount during subsequent data processing, and provide a basis for evaluating the character utility corresponding to the subsequent text feature characters. It should be noted that the text feature characters are representative text characters in the enterprise text data.
[0071] Specifically, the extracting character features from the enterprise text data to obtain text feature characters includes:
[0072] Perform text segmentation processing on the enterprise text data to obtain enterprise text segments;
[0073] Perform part-of-speech tagging on the enterprise text segments to obtain the part-of-speech of the segments, and count the word frequency of the enterprise text segments corresponding to the segments;
[0074] Identify the segment entities in the enterprise text segments, and combine the segment entities, the part-of-speech of the segments, and the word frequency of the segments to extract character features from the enterprise text data to obtain text feature characters.
[0075] It should be noted that the enterprise text segments are text words in the enterprise text data, the part-of-speech of the segments is the corresponding part-of-speech of the enterprise text segments, such as verbs, nouns, and adjectives, etc., the word frequency of the segments represents the corresponding occurrence times of the enterprise text segments, and the segment entities are the entity information in the enterprise text segments, such as enterprise names, product names, and personal names, etc.
[0076] Further, the text word segmentation processing of the enterprise text data can be implemented through a word segmentation tool, such as the Jieba word segmentation tool; the part-of-speech tagging processing of the enterprise text word segmentation can be implemented through the NLTK tool; the statistics of the word frequencies corresponding to the enterprise text word segmentation can be implemented through a statistics tool, and the statistics tool is compiled by C language; the recognition of the word segmentation entities in the enterprise text word segmentation can be implemented through the dictionary matching method; combining the word segmentation entities, the word segmentation part-of-speech, and the word segmentation word frequencies, character features are extracted from the enterprise text data to obtain text feature characters. For example, in the enterprise text data, through the recognized word segmentation entities such as specific product names, customer names, etc., combined with the word segmentation part-of-speech (such as key nouns, verbs, etc.) and the word segmentation word frequencies (words that appear frequently are often more representative), if a product name entity appears frequently and is of noun part-of-speech, it is a text feature character.
[0077] The present invention calculates the character sensitivity corresponding to the text feature characters, and can understand the privacy sensitivity degree corresponding to the text feature characters through the character sensitivity, thereby improving the marking accuracy of the desensitized feature characters in the subsequent text feature characters. It should be explained that the character sensitivity represents the privacy sensitivity degree corresponding to the text feature characters.
[0078] Specifically, the calculation of the character sensitivity corresponding to the text feature characters includes:
[0079] Perform normalization processing on the text feature characters to obtain standard feature characters;
[0080] Query the privacy information indicators corresponding to the original enterprise data, and calculate the privacy correlation degree between the standard feature characters and the privacy information indicators;
[0081] Evaluate the influence weight of the standard feature characters on the privacy information indicators;
[0082] Combine the influence weight and the privacy correlation degree to calculate the character sensitivity corresponding to the text feature characters.
[0083] It should be explained that the standard feature characters are the normalized form expressions corresponding to the text feature characters, such as the characters after removing special symbols; the privacy information indicators are the specific manifestations of the categories corresponding to the original enterprise data; the privacy correlation degree represents the correlation degree between the standard feature characters and the privacy information indicators; the influence weight represents the influence degree of the standard feature characters on the privacy information indicators.
[0084] Further, the standardization process of the text feature characters can be achieved through the lower() or upper() method of the string; the query of the privacy information metrics corresponding to the original enterprise data can be achieved through the enterprise data management platform; the evaluation of the influence weight of the standard feature characters on the privacy information metrics can be achieved through the analytic hierarchy process; multiplying the influence weight by the corresponding privacy correlation degree, the character sensitivity corresponding to the text feature characters is obtained.
[0085] Further, as an optional embodiment of the present invention, calculating the privacy correlation degree between the standard feature characters and the privacy information metrics includes:
[0086] Calculate the information entropy of the standard feature characters and the privacy information metrics respectively to obtain the character information entropy and the metric information entropy;
[0087] Calculate the joint information entropy between the standard feature characters and the privacy information metrics;
[0088] Combining the character information entropy, the metric information entropy and the joint information entropy, the privacy correlation degree between the standard feature characters and the privacy information metrics can be calculated through the following formula:
[0089]
[0090] Wherein, A represents the privacy correlation degree between the standard feature characters and the privacy information metrics, C(a, b) represents the joint information entropy between the a-th character in the standard feature characters and the b-th metric in the privacy information metrics, B(a) represents the character information entropy corresponding to the a-th character in the standard feature characters, D(b) represents the metric information entropy corresponding to the b-th metric in the privacy information metrics, and a and b respectively represent the serial numbers corresponding to the standard feature characters and the privacy information metrics.
[0091] It should be explained that the character information entropy and the metric information entropy are respectively the measures of the uncertainty of the self-information corresponding to the standard feature characters and the privacy information metrics, and the joint information entropy represents the measure of the overall information uncertainty when the standard feature characters and the privacy information metrics appear together. Further, calculate the occurrence probability corresponding to the standard feature characters and the privacy information metrics, and the corresponding information entropy can be calculated through the Shannon entropy algorithm according to the corresponding occurrence probability; the joint occurrence probability between the standard feature characters and the privacy information metrics can be calculated, and the joint information entropy can be calculated by using the above Shannon entropy algorithm.
[0092] By combining the data application scenario and the text feature characters, the character utility degree corresponding to the text feature characters is evaluated, and the usefulness of the text feature characters in achieving relevant business goals, completing analysis tasks, etc. in the data application scenario can be accurately understood. It should be noted that the data application scenario is the specific business scenario corresponding to the enterprise text data, and the character utility degree represents the measure of the value embodiment corresponding to the text feature characters. Further, the query of the data application scenario corresponding to the enterprise text data can be realized through a data query tool, such as a business intelligence (BI) tool. The BI tool can be connected to the data sources of the enterprise, including the database or data warehouse where the text data is located. Through the visualization interface and data analysis functions of these tools, the text data can be queried, analyzed, and visually displayed, helping users intuitively understand the distribution, trends, and relationships with other business data of the text data, so as to infer its data application scenario. For example, using BI tools such as Tableau and PowerBI, the enterprise's sales contract text data is associated with the sales performance data to determine the application scenario of the contract text data in the sales business.
[0093] Specifically, the combination of the data application scenario and the text feature characters to evaluate the character utility degree corresponding to the text feature characters includes:
[0094] Perform vectorization processing on the data application scenario and the text feature characters respectively to obtain an application scenario vector and a text character vector;
[0095] Extract the feature vectors in the application scenario vector and the text character vector respectively to obtain a scenario feature vector and a character feature vector;
[0096] Calculate the vector similarity between the scenario feature vector and the character feature vector;
[0097] Evaluate the character utility degree corresponding to the text feature characters according to the vector similarity.
[0098] It should be noted that the application scenario vector and the text character vector are the expression vectors corresponding to the data application scenario and the text feature characters respectively, the scenario feature vector and the character feature vector are the representative vectors in the application scenario vector and the text character vector respectively, and the vector similarity represents the similarity degree between the scenario feature vector and the character feature vector.
[0099] Furthermore, the vectorization processing of the data application scenario and the text feature characters can be achieved through the word2vec model; the extraction of the feature vectors in the application scenario vector and the text character vector can be achieved through a dimensionality reduction algorithm, such as the linear discriminant analysis method; the calculation of the vector similarity between the scenario feature vector and the character feature vector can be achieved through the cosine similarity algorithm; according to the vector similarity, evaluate the character utility of the text feature characters. For example, if the text feature character "high-quality product" has a high vector similarity with the vector representing product praise, it indicates that its utility in reflecting product reputation is relatively large and can provide effective information for market analysis; if "new function" has a high vector similarity with the innovation demand vector, then in the product R & D scenario, the character utility is high and can help determine the R & D direction.
[0100] S3. Combine the character sensitivity and the character utility, mark the desensitized feature characters in the text feature characters, and perform semantic blurring processing on the desensitized feature characters to obtain blurred feature characters.
[0101] In the present invention, by combining the character sensitivity and the character utility, the desensitized feature characters in the text feature characters are marked, thereby improving the marking accuracy of the desensitized feature characters, and semantic blurring processing is performed on the desensitized feature characters, thereby accurately protecting the privacy of the enterprise text data. It should be noted that the desensitized feature characters are the characters with high sensitivity in the text feature characters, and the blurred feature characters are the characters obtained after encrypting the sensitive characters in the desensitized feature characters. Furthermore, combined with the numerical levels of the character sensitivity and the character utility, the desensitized feature characters in the text feature characters are marked. For example, if a text feature character such as "ID number" has extremely high character sensitivity and low character utility in application scenarios such as market trend analysis of general enterprises, such characters should be marked as desensitized feature characters; another example is "employee's home address", which has high sensitivity and low utility except in specific employee care scenarios, and can also be determined as a desensitized feature character to ensure data security and compliance.
[0102] Specifically, the performing semantic blurring processing on the desensitized feature characters to obtain blurred feature characters includes:
[0103] Perform semantic parsing on the desensitized feature characters to obtain the semantic of the feature characters;
[0104] Calculate the semantic similarity between the semantic of the feature characters and each semantic in the preset semantic library;
[0105] When the semantic similarity is greater than the preset threshold, dispatch the blurred semantic corresponding to the semantic of the feature characters from the preset semantic library;
[0106] Perform fuzzification processing on the semantic meaning of the feature characters by using the fuzzy semantics to obtain fuzzy feature characters.
[0107] It should be explained that the semantic meaning of the feature characters is the character meaning explanation corresponding to the desensitized feature characters. The preset semantic library is a pre-constructed database used to replace similar semantics, which can be constructed by collecting a large amount of enterprise historical data. The semantic similarity represents the degree of similarity between the semantic meaning of the feature characters and each semantic in the preset semantic library. The preset threshold is the judgment standard value of the semantic similarity, which can be set to 0.8 or can be set according to the actual business scenario. The fuzzy semantics is the alternative semantic corresponding to the semantic meaning of the feature characters.
[0108] Furthermore, the semantic parsing of the desensitized feature characters can be realized by the semantic parsing method; the calculation of the semantic similarity between the semantic meaning of the feature characters and each semantic in the preset semantic library can be realized by the cosine similarity algorithm; the fuzzy semantics and the semantic meaning of the feature characters can be replaced by the direct replacement method to achieve the fuzzification processing of the semantic meaning of the feature characters and obtain fuzzy feature characters.
[0109] S4. Mark the image desensitization area corresponding to each image in the enterprise image data, and analyze the area attributes corresponding to the image desensitization area. Based on the area attributes, perform area fuzzification processing on the image desensitization area to obtain a fuzzy image area.
[0110] In the present invention, by marking the image desensitization area corresponding to each image in the enterprise image data, the sensitive image positions of each image in the enterprise image data can be obtained, which further lays a foundation for subsequent area fuzzification processing of the image desensitization area. It should be explained that the image desensitization area is the position where the sensitive information of each image in the enterprise image data is located.
[0111] Specifically, the marking of the image desensitization area corresponding to each image in the enterprise image data includes:
[0112] Perform image noise reduction processing on each image in the enterprise image data to obtain a noise-reduced enterprise image;
[0113] Perform image enhancement processing on the noise-reduced enterprise image to obtain an enhanced enterprise image;
[0114] Identify the image connotation in the enhanced enterprise image and analyze the connotation metadata corresponding to the image connotation;
[0115] Parse the sensitive metadata in the connotation metadata, and based on the sensitive metadata, mark the image desensitization area corresponding to each image in the enterprise image data.
[0116] It should be explained that the denoised corporate image refers to the image obtained after noise reduction processing is performed on each image in the corporate image data; the enhanced corporate image refers to the image that is further enhanced on the basis of the denoised corporate image; the image connotation is the embodiment of the specific content, meaning, etc. contained in the enhanced corporate image; the connotation metadata is the data corresponding to the image connotation and used to describe its relevant attributes; the sensitive metadata is the data involving sensitive information in the connotation metadata.
[0117] Furthermore, the image denoising processing of each image in the enterprise image data can be achieved through filtering algorithms such as Gaussian filtering, median filtering, etc.; the image enhancement processing of the denoised enterprise image can be achieved through contrast stretching; the recognition of the image connotation in the enhanced enterprise image can be achieved through a deep learning model, such as a convolutional neural network; the analysis of the connotation metadata corresponding to the image connotation can be achieved through a data mining method, such as an association rule mining method; the parsing of sensitive metadata in the connotation metadata can be achieved through a sensitive information recognition model, such as a rule-based text matching model; based on the sensitive metadata, the image desensitization area corresponding to each image in the enterprise image data can be manually marked.
[0118] The present invention can effectively protect the sensitive information in the enterprise image data and prevent the risk of information leakage by performing regional blurring processing on the image desensitized area based on the regional attributes. It should be explained that the regional attributes are the regional features corresponding to the image desensitized area, such as color, texture, shape and other characteristics, and the blurred image area is a specific part of the image desensitized area that is presented after blurring processing and has a visually blurred effect. Furthermore, the analysis of the regional attributes corresponding to the image desensitized area can be achieved through a support vector machine, such as taking the feature vector of the image desensitized area as input, and training an SVM model for classification and regression analysis. For example, the color histogram, texture features, etc. of the image desensitized area can be extracted as feature vectors, and the SVM model can be used to judge the sensitivity level and other attributes of the area.
[0119] In detail, based on the regional attributes, the image desensitization region is subjected to regional blurring processing to obtain a blurred image region, including:
[0120] Extracting regional color attributes, regional texture attributes and regional shape attributes from the regional attributes;
[0121] Based on the regional color attribute, calculating the regional color entropy corresponding to the image desensitization area;
[0122] Based on the regional texture attribute, calculating the regional texture entropy corresponding to the desensitized area of the image;
[0123] Based on the regional shape attribute, calculate the shape complexity corresponding to the image desensitization region;
[0124] Combined with the regional color entropy, the regional texture entropy, and the shape complexity, set the blur priority corresponding to the image desensitization region;
[0125] Based on the blur priority, perform regional blur processing on the image desensitization region to obtain a blurred image region.
[0126] It should be explained that the regional color attribute, the regional texture attribute, and the regional shape attribute are specific attributes in the regional attributes used to describe the characteristics of the image desensitization region in terms of color, texture, and shape respectively. The regional color entropy represents the color complexity corresponding to the image desensitization region, the regional texture entropy represents the texture complexity corresponding to the image desensitization region, the shape complexity represents the complexity corresponding to the image desensitization region, and the blur priority represents the priority degree when the image desensitization region is blurred.
[0127] Further, the regional color attribute, the regional texture attribute, and the regional shape attribute in the regional attributes can be extracted through an extraction function. The extraction function is compiled by a scripting language, such as the JS scripting language; according to the regional color attribute, a color histogram is constructed. Based on the color histogram, the probability of each color value appearing is determined. According to the probability and the corresponding calculation formula, the regional color entropy corresponding to the image desensitization region is calculated, such as the Shannon entropy formula; the calculation principle of the regional texture entropy is the same as that of the regional color entropy, and will not be elaborated here; based on the regional shape attribute, the regional perimeter and regional area corresponding to the image desensitization region are determined, and the ratio of the regional perimeter to the regional area is calculated to obtain the shape complexity corresponding to the image desensitization region; the weight coefficients corresponding to the regional color entropy, the regional texture entropy, and the shape complexity are determined respectively. The regional color entropy, the regional texture entropy, and the shape complexity are respectively normalized to obtain the normalization results. The normalization results are respectively multiplied by the corresponding weight coefficients to obtain the blur priority corresponding to the image desensitization region; based on the blur priority, the image desensitization region is subjected to regional blur processing using a blur algorithm to obtain a blurred image region. The blur algorithm includes the Gaussian blur algorithm.
[0128] S5. Identify the data endpoint values in the enterprise numerical data. Based on the data endpoint values, calculate the desensitization mapping value corresponding to the enterprise numerical data. According to the desensitization mapping value, perform desensitization processing on the enterprise numerical data to obtain the target numerical data.
[0129] According to the present invention, by calculating the desensitization mapping value corresponding to the enterprise numerical data based on the data endpoint values, the value for desensitizing the enterprise numerical data can be obtained, which facilitates subsequent desensitization processing of the enterprise numerical data, and further facilitates effective hiding of the enterprise numerical data. It should be noted that the data endpoint values are the maximum and minimum values in the enterprise numerical data, the desensitization mapping value is the desensitized replacement value in the enterprise numerical data, and the target numerical data is the data obtained after the enterprise numerical data is replaced by the desensitization mapping value. Further, the identification of the data endpoint values in the enterprise numerical data can be achieved by the box plot method; the desensitization processing of the enterprise numerical data can be achieved by the above-mentioned direct replacement method.
[0130] Specifically, calculating the desensitization mapping value corresponding to the enterprise numerical data based on the data endpoint values includes:
[0131]
[0132] where E represents the desensitization mapping value corresponding to the enterprise numerical data, F i represents the value corresponding to the i-th data in the enterprise numerical data, i represents the serial number corresponding to the enterprise numerical data, minF represents the minimum value among the data endpoint values, maxF represents the maximum value among the data endpoint values, H , max represents the upper limit value of the desensitization value range, and H , min represents the lower limit value of the desensitization value range.
[0133] Further, the upper limit value of the desensitization value range is the maximum value that the data is allowed to reach after desensitization processing, and the lower limit value of the desensitization value range is the minimum value that the data is allowed to reach after desensitization processing, which is usually determined according to the goals and requirements of data processing.
[0134] S6. Using the target numerical data, the blurred image area, and the blurred feature characters to perform data update processing on the classified enterprise data to obtain the enterprise desensitized data corresponding to the original enterprise data.
[0135] According to the present invention, by using the target numerical data, the blurred image area, and the blurred feature characters to perform data update processing on the classified enterprise data, the classified enterprise data is effectively desensitized, and the desensitization flexibility of the classified enterprise data is improved.
[0136] Compared with the problems described in the background art, the present invention can understand the data structure characteristics corresponding to the original enterprise data by detecting the data format corresponding to the original enterprise data, and classify the original enterprise data based on the data format, so as to group the data of the same type in the original enterprise data together, thereby facilitating subsequent targeted analysis and processing of the data. The present invention extracts character features from the enterprise text data, can obtain representative characters in the enterprise text data, reduces the calculation amount during subsequent data processing, and provides a basis for the evaluation of the character utility degree corresponding to the subsequent text feature characters. The present invention combines the character sensitivity and the character utility degree to mark the desensitization feature characters in the text feature characters, thereby improving the marking accuracy of the desensitization feature characters, and performs semantic blurring processing on the desensitization feature characters, thereby accurately protecting the privacy of the enterprise text data. The present invention can obtain the sensitive image positions of each image in the enterprise image data by marking the image desensitization area corresponding to each image in the enterprise image data, thereby laying a foundation for subsequent area blurring processing of the image desensitization area. The present invention can calculate the desensitization mapping value corresponding to the enterprise numerical data based on the data endpoint value, obtain the value for desensitizing the enterprise numerical data, facilitate subsequent desensitization processing of the enterprise numerical data, and thereby facilitate effectively hiding the enterprise numerical data. The present invention uses the target numerical data, the blurred image area, and the blurred feature characters to perform data update processing on the classified enterprise data, thereby effectively desensitizing the classified enterprise data and improving the desensitization flexibility of the classified enterprise data. Therefore, the present invention proposes a method for desensitizing enterprise sensitive data based on natural language to improve the desensitization flexibility of enterprise sensitive data.
[0137] Embodiment 2:
[0138] As Figure 2 shown, it is a functional module diagram of a system for desensitizing enterprise sensitive data based on natural language provided by an embodiment of the present invention.
[0139] The system 100 for desensitizing enterprise sensitive data based on natural language according to the present invention can be installed in an electronic device. According to the functions implemented, the system 100 for desensitizing enterprise sensitive data based on natural language can include a data classification module 101, a character utility degree evaluation module 102, a character blurring processing module 103, an image blurring processing module 104, a numerical desensitization processing module 105, and a data update module 106. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0140] In this embodiment, the functions of each module / unit are as follows:
[0141] The data classification module 101 is configured to obtain the original enterprise data to be processed, detect the data format corresponding to the original enterprise data, and classify the original enterprise data based on the data format to obtain classified enterprise data, where the classified enterprise data includes: enterprise text data, enterprise image data, and enterprise numerical data;
[0142] The character utility evaluation module 102 is configured to extract character features from the enterprise text data to obtain text feature characters, calculate the character sensitivity corresponding to the text feature characters, query the data application scenario corresponding to the enterprise text data, and evaluate the character utility corresponding to the text feature characters in combination with the data application scenario and the text feature characters;
[0143] The character fuzzy processing module 103 is configured to mark the desensitization feature characters in the text feature characters in combination with the character sensitivity and the character utility, and perform semantic fuzzy processing on the desensitization feature characters to obtain fuzzy feature characters;
[0144] The image fuzzy processing module 104 is configured to mark the image desensitization area corresponding to each image in the enterprise image data, analyze the area attributes corresponding to the image desensitization area, and perform area fuzzy processing on the image desensitization area based on the area attributes to obtain a fuzzy image area;
[0145] The numerical data desensitization processing module 105 is configured to identify the data endpoint values in the enterprise numerical data, calculate the desensitization mapping values corresponding to the enterprise numerical data based on the data endpoint values, and desensitize the enterprise numerical data according to the desensitization mapping values to obtain target numerical data;
[0146] The data update module 106 is configured to use the target numerical data, the fuzzy image area, and the fuzzy feature characters to perform data update processing on the classified enterprise data to obtain enterprise desensitized data corresponding to the original enterprise data.
[0147] Specifically, each module in the enterprise sensitive data desensitization system 100 based on natural language in this application embodiment uses the same technical means as those in the above Figure 1 described in an enterprise sensitive data desensitization method based on natural language, and can produce the same technical effects, which will not be elaborated here.
[0148] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for desensitizing enterprise sensitive data based on natural language, characterized in that, The method includes: Obtain the original enterprise data to be processed, detect the data format corresponding to the original enterprise data, and classify the original enterprise data based on the data format to obtain classified enterprise data, where the classified enterprise data includes: enterprise text data, enterprise image data, and enterprise numerical data; Extract character features from the enterprise text data to obtain text feature characters, calculate the character sensitivity corresponding to the text feature characters, query the data application scenarios corresponding to the enterprise text data, and evaluate the character utility corresponding to the text feature characters in combination with the data application scenarios and the text feature characters; Mark the desensitized feature characters in the text feature characters in combination with the character sensitivity and the character utility, and perform semantic fuzzification processing on the desensitized feature characters to obtain fuzzy feature characters; Mark the image desensitization areas corresponding to each image in the enterprise image data, analyze the area attributes corresponding to the image desensitization areas, and perform area fuzzification processing on the image desensitization areas based on the area attributes to obtain fuzzy image areas; Identify the data endpoint values in the enterprise numerical data, calculate the desensitized mapping values corresponding to the enterprise numerical data based on the data endpoint values, and desensitize the enterprise numerical data according to the desensitized mapping values to obtain target numerical data; Use the target numerical data, the fuzzy image areas, and the fuzzy feature characters to perform data update processing on the classified enterprise data to obtain the enterprise desensitized data corresponding to the original enterprise data.
2. The method for desensitizing enterprise sensitive data based on natural language according to claim 1, wherein, The extracting character features from the enterprise text data to obtain text feature characters includes: Perform text tokenization processing on the enterprise text data to obtain enterprise text tokens; Perform part-of-speech tagging processing on the enterprise text tokens to obtain token part-of-speech, and count the token frequencies corresponding to the enterprise text tokens; Identify the token entities in the enterprise text tokens, and extract character features from the enterprise text data in combination with the token entities, the token part-of-speech, and the token frequencies to obtain text feature characters.
3. The method for desensitizing enterprise sensitive data based on natural language according to claim 1, wherein The calculating the character sensitivity corresponding to the text feature characters includes: Perform normalization processing on the text feature characters to obtain standard feature characters; Query the privacy information metrics corresponding to the original enterprise data, and calculate the privacy correlation between the standard feature characters and the privacy information metrics; Evaluate the influence weight of the standard feature characters on the privacy information metrics; Calculate the character sensitivity corresponding to the text feature characters in combination with the influence weight and the privacy correlation.
4. The method for desensitizing enterprise sensitive data based on natural language according to claim 3, wherein The calculating the privacy correlation between the standard feature characters and the privacy information metrics includes: Calculate the information entropy corresponding to the standard feature characters and the privacy information metrics respectively to obtain character information entropy and metric information entropy; Calculate the joint information entropy between the standard feature characters and the privacy information metrics; Combining the character information entropy, the metric information entropy, and the joint information entropy, the privacy correlation between the standard feature characters and the privacy information metrics can be calculated through the following formula: Among them, A represents the privacy correlation degree between the standard feature character and the privacy information index, C(a, b) represents the joint information entropy between the a-th character in the standard feature character and the b-th index in the privacy information index, B(a) represents the character information entropy corresponding to the a-th character in the standard feature character, D(b) represents the index information entropy corresponding to the b-th index in the privacy information index, and a and b respectively represent the sequence numbers corresponding to the standard feature character and the privacy information index.
5. The method for desensitizing enterprise sensitive data based on natural language according to claim 1, wherein Evaluating the character utility degree corresponding to the text feature character by combining the data application scenario and the text feature character includes: Performing vectorization processing on the data application scenario and the text feature character respectively to obtain an application scenario vector and a text character vector; Extracting feature vectors from the application scenario vector and the text character vector respectively to obtain a scenario feature vector and a character feature vector; Calculating the vector similarity between the scenario feature vector and the character feature vector; Evaluating the character utility degree corresponding to the text feature character according to the vector similarity.
6. The method for desensitizing enterprise sensitive data based on natural language according to claim 1, wherein, Performing semantic blurring processing on the desensitized feature character to obtain a blurred feature character, including: Performing semantic parsing on the desensitized feature character to obtain the semantic of the feature character; Calculating the semantic similarity between the semantic of the feature character and each semantic in the preset semantic library; When the semantic similarity is greater than the preset threshold, dispatching the blurred semantic corresponding to the semantic of the feature character from the preset semantic library; Performing blurring processing on the semantic of the feature character by using the blurred semantic to obtain a blurred feature character.
7. The method for desensitizing enterprise sensitive data based on natural language according to claim 1, wherein Marking the image desensitization area corresponding to each image in the enterprise image data, including: Performing image noise reduction processing on each image in the enterprise image data to obtain a noise-reduced enterprise image; Performing image enhancement processing on the noise-reduced enterprise image to obtain an enhanced enterprise image; Identifying the image connotation in the enhanced enterprise image and analyzing the connotation metadata corresponding to the image connotation; Parsing the sensitive metadata in the connotation metadata, and based on the sensitive metadata, marking the image desensitization area corresponding to each image in the enterprise image data.
8. The method for desensitizing enterprise sensitive data based on natural language according to claim 1, characterized in that, Performing regional blurring processing on the image desensitization area based on the regional attribute to obtain a blurred image area, including: Extracting the regional color attribute, regional texture attribute, and regional shape attribute in the regional attribute; Calculating the regional color entropy corresponding to the image desensitization area based on the regional color attribute; Calculating the regional texture entropy corresponding to the image desensitization area based on the regional texture attribute; Calculating the shape complexity corresponding to the image desensitization area based on the regional shape attribute; Combining the regional color entropy, the regional texture entropy, and the shape complexity to set the fuzzy priority corresponding to the image desensitization area; Performing regional blurring processing on the image desensitization area based on the fuzzy priority to obtain a blurred image area.
9. The method for desensitizing enterprise sensitive data based on natural language according to claim 1, wherein Calculating the desensitization mapping value corresponding to the enterprise numerical data based on the data endpoint value, including: Among them, E represents the desensitized mapping value corresponding to the enterprise numerical data, and F i represents the value corresponding to the i-th data in the enterprise numerical data, i represents the serial number corresponding to the enterprise numerical data, minF represents the minimum value among the data endpoint values, maxF represents the maximum value among the data endpoint values, and H , max represents the upper limit value of the desensitized value range, and H , min represents the lower limit value of the desensitized value range.
10. An enterprise sensitive data desensitization system based on natural language, characterized in that, The system includes: A data classification module, which is used to obtain the original enterprise data to be processed, detect the data format corresponding to the original enterprise data, and classify the original enterprise data based on the data format to obtain classified enterprise data, where the classified enterprise data includes: enterprise text data, enterprise image data, and enterprise numerical data; A character utility evaluation module, which is used to extract character features from the enterprise text data to obtain text feature characters, calculate the character sensitivity corresponding to the text feature characters, query the data application scenarios corresponding to the enterprise text data, and evaluate the character utility corresponding to the text feature characters in combination with the data application scenarios and the text feature characters; A character blurring processing module, which is used to mark the desensitized feature characters in the text feature characters in combination with the character sensitivity and the character utility, and perform semantic blurring processing on the desensitized feature characters to obtain blurred feature characters; An image blurring processing module, which is used to mark the image desensitization area corresponding to each image in the enterprise image data, analyze the area attributes corresponding to the image desensitization area, and perform area blurring processing on the image desensitization area based on the area attributes to obtain a blurred image area; A numerical data desensitization processing module, which is used to identify the data endpoint values in the enterprise numerical data, calculate the desensitized mapping values corresponding to the enterprise numerical data based on the data endpoint values, and desensitize the enterprise numerical data according to the desensitized mapping values to obtain target numerical data; A data update module, which is used to perform data update processing on the classified enterprise data by using the target numerical data, the blurred image area, and the blurred feature characters to obtain enterprise desensitized data corresponding to the original enterprise data.
Citation Information
Patent Citations
Knowledge enhancement BERT-based word granularity Chinese semantic approximate adversarial sample generation method
CN115309898A
Self-adaptive industrial data desensitization and restoration method and system
CN115422597A
Sensitive data identification method and device, electronic equipment and storage medium
CN115618415A
Image processing method and device, electronic equipment and readable storage medium
CN116108464A
Desensitization processing method and device for traffic scene image and electronic equipment
CN118378297A