Data desensitization method and device and computing equipment
By combining the BERT model and random forest model to identify and classify structured and unstructured data, and using the third machine learning model to determine the desensitization strategy, the problem of low accuracy of data desensitization in the prior art is solved, and efficient identification and processing of sensitive information for multiple types of data is achieved.
Patent Information
- Application Number
- CN202510602772.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-05
AI Technical Summary
Existing data desensitization technologies are less accurate when processing both structured and unstructured data, and cannot effectively identify and process sensitive information under multiple data types.
Different machine learning models are used to process structured and unstructured data, and sensitive data are identified and fused separately. The BERT model and random forest model are used to identify sensitive features and levels and categories in unstructured and structured data, and the target desensitization strategy is determined through the third machine learning model.
Improve the accuracy and efficiency of data desensitization, and can accurately identify and process sensitive information in structured and unstructured data, avoid misjudgment, and ensure data security and practicality.
Smart Images

Figure CN120429894A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a data desensitization method, apparatus, and computing device. Background Art
[0002] With the rapid development of big data technology, data privacy and security issues are becoming increasingly important across various fields. Data masking has become a crucial means of protecting personal privacy and ensuring data security. Data masking involves desensitizing sensitive data so that it cannot be directly or indirectly identified with personal identities or key information, while still preserving the data's business value.
[0003] Currently, intelligent data desensitization technology is a common method for processing sensitive data. Intelligent data desensitization technology is mainly aimed at unstructured text data. After identifying sensitive data in text data, this intelligent data desensitization technology desensitizes the sensitive data based on predefined rules (such as character replacement, masking, or encryption) (such as uniformly replacing it with asterisks or fixed-length masking). As business scenarios become increasingly complex, the data faced often includes multiple data types, that is, unstructured data and structured data are intertwined. In this case, desensitizing only unstructured data is no longer sufficient, and structured data also needs to be desensitized. However, when desensitizing current data that includes both unstructured and structured data, the accuracy of existing desensitization technologies is low. Summary of the Invention
[0004] The embodiments of the present application provide a data desensitization method, apparatus, and computing device, which can accurately identify sensitive data in the data and perform desensitization, thereby improving the accuracy of data desensitization.
[0005] To achieve the above objectives, the present invention adopts the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides a data desensitization method, which includes: obtaining first original data; the first original data includes a structured data part and an unstructured data part; using the first original data as the input of a first machine learning model, and using the first machine learning model to obtain first sensitive data; wherein the first sensitive data is the sensitive data in the unstructured data part; using the first original data as the input of a second machine learning model, and using the second machine learning model to obtain second sensitive data; wherein the second sensitive data is the sensitive data in the structured data part; determining target sensitive data in the first original data; wherein the target sensitive data includes first sensitive data and second sensitive data; desensitizing the target sensitive data to obtain desensitized data.
[0007] Based on this solution, sensitive data is extracted from the structured and unstructured data in the first raw data using different machine learning models. The extracted sensitive data is then fused to obtain complete target sensitive data. The target sensitive data is then desensitized to obtain desensitized data. In this way, by processing data using different types of machine learning models, sensitive data can be accurately identified and desensitized even when the data contains both structured and unstructured data, thereby improving the accuracy of data desensitization.
[0008] In one possible implementation, the first original data is used as the input of the first machine learning model, and the first sensitive data is obtained by using the first machine learning model, including: using the first original data as the input of the first machine learning model, and the first machine learning model to obtain at least one word segmentation feature corresponding to the unstructured data part in the first original data, a first sensitivity level corresponding to each word segmentation feature, and a first category corresponding to each word segmentation feature; wherein the first category is used to characterize the result of dividing each word segmentation feature according to a preset word segmentation type; and determining the first sensitive data based on at least one word segmentation feature, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature.
[0009] Based on this solution, by parsing the word segmentation features, sensitivity levels and category information of the unstructured data in the first original data, these word segmentation features, sensitivity levels and category information are obtained, so that sensitive data can be accurately identified and misjudgment can be avoided.
[0010] In another possible implementation, the first original data is used as the input of the first machine learning model, and the first machine learning model is used to obtain at least one word segmentation feature corresponding to the unstructured data portion in the first original data, a first sensitivity level corresponding to each word segmentation feature, and a first category corresponding to each word segmentation feature, including: using the first original data as the input of the BERT model, and the BERT model is used to obtain at least one word segmentation feature corresponding to the unstructured data portion in the first original data, a first sensitivity level corresponding to each word segmentation feature, and a first category corresponding to each word segmentation feature; wherein the BERT model is used to identify sensitive data in the unstructured data portion.
[0011] Based on this solution, the BERT model can be used to automatically identify and extract word segmentation features and their sensitivity levels and categories in unstructured data. In this way, since the BERT model is more effective in identifying sensitive data in unstructured data, it can accurately and quickly identify sensitive data in the first original data, thereby improving the efficiency of identifying sensitive data.
[0012] In another possible implementation, the first original data is used as the input of the second machine learning model, and the second sensitive data is obtained by using the second machine learning model, including: using the first original data as the input of the second machine learning model, and using the second machine learning model to obtain at least one first field feature of the structured data part in the first original data, the second sensitivity level corresponding to each first field feature, and the second category corresponding to each first field feature; wherein the second category is used to characterize the result of dividing each first field feature according to a preset field type; and determining the second sensitive data based on at least one first field feature, the second sensitivity level corresponding to each first field feature, and the second category corresponding to each first field feature.
[0013] Based on this solution, by extracting the structured data from the first original data respectively, analyzing and determining the corresponding field characteristics, sensitivity levels and category information, based on these field characteristics, sensitivity levels and category information, sensitive data can be accurately identified to avoid misjudgment.
[0014] In another possible implementation, the first original data is used as the input of the second machine learning model, and the second machine learning model is used to obtain at least one first field feature of the structured data part in the first original data, the second sensitivity level corresponding to each first field feature, and the second category corresponding to each first field feature, including: using the first original data as the input of the random forest model, and the random forest model is used to obtain at least one first field feature, the second sensitivity level corresponding to each first field feature, and the second category corresponding to each first field feature; wherein the random forest model is used to identify sensitive data in the structured data part.
[0015] Based on this solution, the random forest model can be used to automatically identify and extract field features and their sensitivity levels and categories in structured data. In this way, since the random forest model is more effective in identifying sensitive data in structured data, it can accurately and quickly identify sensitive data in the first original data, thereby improving the efficiency of identifying sensitive data.
[0016] In another possible implementation, the target sensitive data is desensitized to obtain desensitized data; wherein the desensitized data includes the target sensitive data after desensitization, including: determining a target desensitization strategy corresponding to the target sensitive data; based on the target desensitization strategy, desensitizing the target sensitive data to obtain desensitized data.
[0017] Based on this solution, the target sensitive data is desensitized according to the target sensitive policy corresponding to the target sensitive data to obtain desensitized data. In this way, by selecting an appropriate desensitization policy according to the target sensitive data, that is, the target sensitive policy, the security of the desensitized data can be greatly improved.
[0018] In another possible implementation, determining the target desensitization strategy corresponding to the target sensitive data includes: taking the target sensitive data as the input of a third machine learning model, and using the third machine learning model to obtain the target desensitization strategy corresponding to the target sensitive data; wherein the third machine learning model is used to obtain the desensitization strategy corresponding to the sensitive data.
[0019] Based on this solution, a third machine learning model is used to intelligently recommend the optimal sensitivity strategy for target sensitive data. This can significantly improve the effect of data desensitization, better protect sensitive data, and retain the practicality of the data.
[0020] In another possible implementation, obtaining the first original data includes: obtaining the second original data; wherein the second original data is original data that has not been preprocessed; the second original data includes a structured data portion and an unstructured data portion; and obtaining the preprocessed first original data based on the second original data.
[0021] Based on this solution, the preprocessed first original data is obtained by preprocessing the directly acquired second original data, thereby ensuring the quality of the first original data and providing a high-quality data foundation for subsequent steps.
[0022] In the second aspect, an embodiment of the present application also provides a data desensitization device, which includes: an acquisition module, configured to: acquire first original data; the first original data includes a structured data part and an unstructured data part; the acquisition module is also configured to: use the first original data as the input of a first machine learning model, and use the first machine learning model to obtain first sensitive data; wherein the first sensitive data is the sensitive data in the unstructured data part; the acquisition module is also configured to: use the first original data as the input of a second machine learning model, and use the second machine learning model to obtain second sensitive data; wherein the second sensitive data is the sensitive data in the structured data part; the determination module is configured to: determine the target sensitive data in the first original data; wherein the target sensitive data includes the first sensitive data and the second sensitive data; the desensitization module is configured to: perform desensitization processing on the target sensitive data to obtain desensitized data.
[0023] Based on this solution, sensitive data is extracted from the structured and unstructured data in the first raw data using different machine learning models. The extracted sensitive data is then fused to obtain complete target sensitive data. The target sensitive data is then desensitized to obtain desensitized data. In this way, by processing data using different types of machine learning models, sensitive data can be accurately identified and desensitized even when the data contains both structured and unstructured data, thereby improving the accuracy of data desensitization.
[0024] In one possible implementation, the acquisition module is specifically configured to: take the first original data as the input of the first machine learning model, and use the first machine learning model to obtain at least one word segmentation feature corresponding to the unstructured data part in the first original data, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature; wherein the first category is used to characterize the result of dividing each word segmentation feature according to a preset word segmentation type; and determine the first sensitive data based on at least one word segmentation feature, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature.
[0025] In one possible implementation, the acquisition module is specifically configured to: use the first original data as the input of the BERT model, and use the BERT model to obtain at least one word segmentation feature corresponding to the unstructured data portion in the first original data, a first sensitivity level corresponding to each word segmentation feature, and a first category corresponding to each word segmentation feature; wherein the BERT model is used to identify sensitive data in the unstructured data portion.
[0026] In one possible implementation, the acquisition module is specifically configured to: use the first original data as input to the second machine learning model, and utilize the second machine learning model to obtain at least one first field feature of the structured data portion in the first original data, a second sensitivity level corresponding to each first field feature, and a second category corresponding to each first field feature; wherein the second category is used to represent a result of classifying each first field feature according to a preset field type;
[0027] The second sensitive data is determined based on at least one first field feature, a second sensitivity level corresponding to each first field feature, and a second category corresponding to each first field feature.
[0028] In one possible implementation, the acquisition module is specifically configured to: use the first original data as the input of the random forest model, and use the random forest model to obtain at least one first field feature, the second sensitivity level corresponding to each first field feature, and the second category corresponding to each first field feature; wherein the random forest model is used to identify sensitive data in the structured data part.
[0029] In one possible implementation, the determination module is specifically configured to: determine a target desensitization strategy corresponding to the target sensitive data; and desensitize the target sensitive data based on the target desensitization strategy to obtain desensitized data.
[0030] In one possible implementation, the determination module is specifically configured to: use the target sensitive data as the input of the third machine learning model, and use the third machine learning model to obtain the target desensitization strategy corresponding to the target sensitive data; wherein the third machine learning model is used to obtain the desensitization strategy corresponding to the sensitive data.
[0031] In one possible implementation, the acquisition module is specifically configured to: acquire second original data; wherein the second original data is original data that has not been preprocessed; the second original data includes a structured data portion and an unstructured data portion; and obtain the preprocessed first original data based on the second original data.
[0032] In a third aspect, an embodiment of the present application further provides a computing device comprising: a processor and a memory; the processor and the memory are coupled; the memory is used to store program instructions; and the processor is used to execute program instructions to perform any method as described in the first aspect above.
[0033] In a fourth aspect, an embodiment of the present application provides a chip, which is used to execute any method as described in the first aspect above.
[0034] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a computer, the method as described in any one of the first aspects is implemented.
[0035] In a sixth aspect, an embodiment of the present application provides a program product, comprising a computer program, which implements any method in the first aspect when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a flow chart of a data desensitization method provided in an embodiment of the present application;
[0037] Figure 2This is a schematic diagram of a data desensitization method provided in an embodiment of the present application;
[0038] Figure 3 This is a flow chart of another data desensitization method provided in an embodiment of the present application;
[0039] Figure 4 This is a schematic diagram of first original data and second original data provided in an embodiment of the present application;
[0040] Figure 5 This is a schematic diagram of the training process of a BERT model provided in an embodiment of the present application;
[0041] Figure 6 This is a schematic diagram of the training process of a random forest model provided in an embodiment of the present application;
[0042] Figure 7 This is a schematic diagram of the training process of a third machine learning model provided in an embodiment of the present application;
[0043] Figure 8 This is a schematic diagram of a target desensitization strategy provided in an embodiment of the present application;
[0044] Figure 9 This is a schematic diagram of a data desensitization device provided in an embodiment of the present application;
[0045] Figure 10 This is a schematic diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The following will describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. To facilitate the clear description of the technical solutions in the embodiments of the present application, the first, second, etc. descriptions in the embodiments of the present application are only used for illustration and to distinguish the described objects. There is no order, nor does it represent a special limitation on the number of devices in the embodiments of the present application, and it does not constitute any limitation on the embodiments of the present application.
[0047] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0048] It should be noted that many specific details are set forth in the following description to facilitate a full understanding of the present application. However, the present application can also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present application is not limited to the specific implementation methods disclosed below.
[0049] In the description of this application, it should be understood that the terms "upper", "lower", "horizontal", "bottom", "inner", "outer" (if any), etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended only to facilitate the description of this application and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore should not be understood as limiting this application. In this application, unless otherwise expressly specified or limited, a first feature being "upper" or "lower" than a second feature may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium.
[0050] In this application, unless otherwise expressly specified or limited, the terms "connected," "connected," "fixed," and the like should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integration; they can refer to direct connections or indirect connections through an intermediate medium; they can refer to internal communication between two elements or an interaction between two elements. However, the phrase "directly connected" indicates that the two connected entities are not connected through an intermediate structure, but are connected to form a whole through a connecting structure. Those skilled in the art can understand the specific meanings of the above terms in this application based on the specific circumstances.
[0051] In this application, references to "first," "second," and the like are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one of these features.
[0052] The following explains the professional terms mentioned in the embodiments of the present application to facilitate understanding by those skilled in the art.
[0053] Structured data, also known as row data, is data logically expressed and implemented by a two-dimensional table structure, strictly adheres to data format and length specifications, and is primarily stored and managed through a database. In contrast to structured data, unstructured data is not suitable for representation by a two-dimensional database table. This includes data in all formats, such as video data and text data. In the embodiments of this application, unstructured data refers to text data.
[0054] Data masking is a technology that protects sensitive information by modifying, replacing, or hiding sensitive information in original data to prevent it from being leaked during use. At the same time, the data masking process preserves its practicality, allowing the masked data to still be used in scenarios without affecting the original analysis results.
[0055] Incremental learning is an online learning method that gradually updates and enhances an existing model by adding new data without retraining the entire model.
[0056] Character replacement, masking, encryption, and generalization are common data desensitization strategies. Character replacement and masking are mainly used to process text information, encryption can process any type of data, and generalization can be used to process categorical or numeric fields.
[0057] Numeric fields are used to store numerical data during data storage and processing. These values can be integers (such as 1, 2, 3), floating-point numbers (such as 3.14, -0.5), or other data that conforms to a numeric format.
[0058] A categorical field is a field used to store limited, non-numeric, categorical data. This data is typically a text label or name that indicates the category or group to which an entity belongs. For example, in a student information table, the "Gender" field is a typical categorical field, with values typically being "Male" or "Female."
[0059] The embodiments of the present application are described below with reference to the accompanying drawings.
[0060] Figure 1 This is a flow chart of a data desensitization method provided in an embodiment of the present application.
[0061] Figure 2 This is a schematic diagram of a data desensitization method provided in an embodiment of the present application.
[0062] like Figure 1 and Figure 2 As shown, the data desensitization method may include the following steps:
[0063] S1: Acquire first original data.
[0064] The first original data may include a structured data portion and an unstructured data portion.
[0065] There are many ways to obtain the first original data, for example, directly collecting the unprocessed first original data from the source of data generation, or preprocessing the source data to obtain the first original data. Figure 3 One of the implementation modes is used as an example for description.
[0066] In one embodiment, step S1 includes steps S11 and S12.
[0067] Step S11: Acquire second original data.
[0068] The second original data is original data that has not been preprocessed; the second original data includes a structured data portion and an unstructured data portion.
[0069] The second raw data is directly derived from the data collection stage, retaining the original state of the data for data cleaning, conversion or structured processing. The second raw data includes a structured data portion and an unstructured data portion. In the embodiment of the present application, the unstructured data refers to text data.
[0070] The following combination Figure 4 , unstructured data will be described in the following examples using text data as an example.
[0071] Figure 4 This is a schematic diagram of first original data and second original data provided in an embodiment of the present application.
[0072] like Figure 4 As shown in (a) of FIG, the second original data includes a structured portion 21 and a text data portion 22. The structured portion 21 includes categorical fields, such as "Order ID," and numeric fields, such as "Order Time." The text data portion 22 includes, for example, "Details of User Purchases of XX Brand Cosmetics."
[0073] Step S12: obtaining pre-processed first original data based on the second original data.
[0074] Preprocessing can include cleaning, filtering, and format conversion. Data cleaning refers to removing erroneous, duplicate, or missing values from unprocessed data. Data filtering refers to selecting data that meets requirements based on specific criteria. Format conversion refers to converting data from one format to another.
[0075] For example, taking data cleaning in preprocessing as an example, Figure 4 The second original data shown in (b) is cleaned, such as removing duplicate order times, to obtain the preprocessed first original data.
[0076] S2: Using the first original data as input to a first machine learning model, and utilizing the first machine learning model to obtain first sensitive data in the first original data.
[0077] The first sensitive data is sensitive data in the unstructured data portion.
[0078] The following continues Figure 3 , taking the use of the first machine model to obtain the first sensitive data as an example for explanation.
[0079] In one embodiment, step S2 includes steps S21 - S22 .
[0080] Step S21: Using the first original data as the input of the first machine learning model, and using the first machine learning model, determine at least one word segmentation feature corresponding to the unstructured data portion in the first original data, a first sensitivity level corresponding to each word segmentation feature, and a first category corresponding to each word segmentation feature.
[0081] The first category is used to represent the result of dividing each word segmentation feature according to the preset word segmentation type. The first category is divided according to the preset word segmentation type, which can include entities (personal names, place names), time, common words, etc., and is not specifically limited here.
[0082] In one embodiment, the first neural network model simultaneously outputs third sensitive data while outputting the first sensitive data. The third sensitive data is sensitive data in the structured data portion. Specifically, while outputting at least one word segmentation feature corresponding to the unstructured data portion, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature, the first machine learning model also outputs at least one second field feature of the structured data portion in the first raw data, the third sensitivity level corresponding to each second field feature, and the third category corresponding to each second field feature.
[0083] The third category is divided according to preset field types. "Preset field types" can include basic data types such as integers, strings, and dates and times, as well as more complex data types with specific formats such as email addresses and phone numbers. These are not specifically limited here. The following example illustrates how the first neural network model outputs both the first sensitive data and the third sensitive data simultaneously.
[0084] Step S21 includes the following step S211.
[0085] Step S211: Using the first original data as the input of the BERT model, and using the BERT model, obtain at least one second field feature of the structured data part in the first original data, the third sensitivity level corresponding to each second field feature, and the third category corresponding to each second field feature, as well as at least one word segmentation feature corresponding to the unstructured data part in the first original data, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature.
[0086] Among them, the BERT model can be used not only to identify sensitive data in the unstructured data part, but also to identify sensitive data in the structured data part.
[0087] It should be noted that the field features output by the BERT model are described as the second field features, and the first field features are described in subsequent steps.
[0088] The bidirectional encoder representations from transformers (BERT) model is primarily used for natural language processing tasks, performing word embedding and classification through a bidirectional context-aware approach. The BERT model 10 is used to identify sensitive data in the first raw data.
[0089] The BERT model 10 is a machine learning model that has been trained on a large-scale dataset before use.
[0090] Further, combined with Figure 5 The training method of the above-mentioned BERT model 10 is explained.
[0091] Figure 5 This is a schematic diagram of the training process of a BERT model provided in an embodiment of the present application.
[0092] like Figure 5 As shown, the BERT model 10 training method includes the following steps:
[0093] Step S501: Obtain a first training set.
[0094] The first training set includes at least one sample data, features, sensitivity levels, and categories corresponding to the sample data. The sample data includes a structured data portion and an unstructured data portion.
[0095] Step S502: Using at least one sample data as input to the BERT model, and using the features, sensitivity level, and category corresponding to the sample data as outputs of the BERT model, the BERT model is trained.
[0096] Step S502 includes steps S5021 to S5024.
[0097] Step S5021: Input the structured data part and the unstructured data part of at least one sample data into the BERT model, and output the prediction features, prediction sensitivity levels and prediction categories corresponding to the structured data part and the unstructured data part respectively.
[0098] Step S5022: Based on a preset loss function, determine the loss value between the predicted features, predicted sensitivity levels, and predicted categories and the features, sensitivity levels, and categories corresponding to the sample data.
[0099] Exemplarily, based on a preset loss function, a first loss value between the predicted feature and the feature corresponding to the present data is determined, a second loss value between the predicted sensitivity level and the sensitivity level corresponding to the sample data is determined, and a third loss value between the predicted category and the category corresponding to the sample data is determined. Based on the first loss value, the second loss value and the third loss value, the loss value between the predicted feature, the predicted sensitivity level and the predicted category and the features, sensitivity levels and categories corresponding to the sample data is determined.
[0100] Step S5023: When the loss value is greater than the preset loss value, adjust the model parameters in the BERT model so that the loss value is less than the preset loss value.
[0101] Step S5024: When the loss value is less than the preset loss value, it is determined that the current BERT model training is completed to obtain a trained BERT model.
[0102] Exemplarily, when the loss value is less than the preset loss value, it means that the output prediction data (predicted features, predicted sensitivity levels, and predicted categories) are infinitely close to the features, sensitivity levels, and categories corresponding to the sample data, indicating that the current first BERT model training is completed to obtain the trained BERT model 10.
[0103] After the training of the BERT model 10 is completed, the BERT model 10 is used to obtain at least one word segmentation feature corresponding to the unstructured data part in the first original data, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature, as well as at least one second field feature of the structured data part in the first original data, the third sensitivity level corresponding to each second field feature, and the third category corresponding to each second field feature.
[0104] Following the above example, Figure 4Specifically, using BERT model 10, we obtain corresponding word segmentation features such as "user" feature, "purchase" feature, "XX brand" feature, "cosmetics" feature, and "details table" feature. The first sensitivity level corresponding to the "user" feature is 0.36, and the corresponding first category is "name"; the first sensitivity level corresponding to the "purchase" feature is 0.36, and the corresponding first category is "common vocabulary"; the first sensitivity level corresponding to the "XX brand" feature is 0.96, and the corresponding first category is "name"; the first sensitivity level corresponding to the "cosmetics" feature is 0.46, and the corresponding first category is "category"; the first sensitivity level corresponding to the "details table" feature is 0.26, and the corresponding first category is "common vocabulary". In addition, using BERT model 10, the second field feature "xxx" is obtained, the corresponding third sensitivity level is 0.6, and the corresponding third category is "string"; the second field feature "1234mmm!", the corresponding third sensitivity level is 0.6, and the corresponding third category is "password"; the second field feature "2025.4.5", the corresponding third sensitivity level is 0.7, and the corresponding third category is "date"; the third sensitivity level corresponding to the "customer service" feature is 0.33, and the corresponding third category is "string"; the third sensitivity level corresponding to the "Xiao Li" feature is 0.8, and the corresponding third category is "string"; the third sensitivity level corresponding to the "attitude" feature is 0.45, and the corresponding third category is "string", and the third sensitivity level corresponding to the "especially good" feature is 0.45, and the corresponding third category is "string". Optionally, the first original data is used as the input of the BERT model 10, and at least one second field feature corresponding to the first original data, the third sensitivity level corresponding to each second field feature, and the third category corresponding to each second field feature, as well as at least one word segmentation feature corresponding to the unstructured data part in the first original data, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature are used as the output of the BERT model 10, and incremental learning is performed on the BERT model 10 to obtain an updated BERT model 10.
[0105] Step S22: Determine the third sensitive data in the first original data based on at least one second field feature, the third sensitivity level corresponding to each second field feature, and the third category corresponding to each second field feature, and determine the first sensitive data in the first original data based on at least one word segmentation feature, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature.
[0106] In one example, the first sensitive data and the third sensitive data are determined based on the sensitivity level and the preset sensitivity level.
[0107] Continuing with the above example, if the preset sensitivity level is 0.5, based on the preset sensitivity level, the second field feature "1234mmm!", the second field feature "xxx", the second field feature "2025.4.5", and the second field feature "Xiao Li" in the first original data are determined as the third sensitive data, and the "XX brand" feature in the first original data is determined as the first sensitive data.
[0108] S3: Use the first original data as input of the second machine learning model, and use the second machine learning model to obtain second sensitive data in the first original data.
[0109] The second sensitive data is the sensitive data in the structured data portion.
[0110] The following continues Figure 3 The following is explained by taking the acquisition of the second sensitive data as an example.
[0111] In one embodiment, step S3 includes steps S31 and S32.
[0112] Step S31: Using the first original data as the input of the second machine learning model, and utilizing the first machine learning model, obtain at least one first field feature of the structured data portion in the first original data, a first sensitivity level corresponding to each first field feature, and a first category corresponding to each first field feature.
[0113] The first category is used to represent the result of dividing each first field feature according to the preset field type.
[0114] Step S31 includes step S311.
[0115] Step S311: using the first original data as input to the random forest model, and utilizing the random forest model to obtain at least one first field feature, a first sensitivity level corresponding to each first field feature, and a first category corresponding to each first field feature.
[0116] Random forest is an ensemble learning model that builds multiple decision trees and averages their results for prediction. It excels at processing structured data, such as numerical and categorical fields. The Random Forest model11 is used to identify sensitive data within structured data.
[0117] The random forest model 11 is a machine learning model that has been trained on a large-scale dataset before use.
[0118] Further, combined with Figure 6The training method of the random forest model 11 is described.
[0119] Figure 6 This is a schematic diagram of the training process of a random forest model provided in an embodiment of the present application.
[0120] like Figure 6 As shown, the random forest model 11 training method includes the following steps:
[0121] Step S601: Obtain a second training set.
[0122] The second training set includes at least one sample data, and the features, sensitivity levels, and categories corresponding to the sample data.
[0123] The sample data includes structured data and unstructured data.
[0124] Step S602: Using the structured data portion of at least one sample data as input to the random forest model, and using the features, sensitivity levels, and categories corresponding to the sample data as outputs of the random forest model, the random forest model is trained.
[0125] Step S602 includes steps S6021 to S6024.
[0126] Step S6021: Input the structured data portion of at least one sample data into the random forest model, and output the corresponding prediction features, prediction sensitivity level, and prediction category.
[0127] Step S6022: Based on a preset loss function, determine the loss value between the predicted features, predicted sensitivity levels, and predicted categories and the features, sensitivity levels, and categories corresponding to the sample data.
[0128] Step S6023: When the loss value is greater than the preset loss value, adjust the model parameters in the random forest model so that the loss value is less than the preset loss value.
[0129] Step S6024: When the loss value is less than the preset loss value, it is determined that the training of the current random forest model is completed to obtain a trained random forest model.
[0130] The above steps S6021 to S6024 can refer to the specific contents of the above steps S5021 to S5024, which will not be repeated here.
[0131] After the random forest model 11 is trained, the random forest model 11 is used to obtain at least one first field feature of the structured data part in the first original data, a first sensitivity level corresponding to each first field feature, and a first category corresponding to each first field feature.
[0132] Continuing with the above example, Figure 4 For example, the first data in the random forest model is used to illustrate the above. Specifically, using the random forest model, the first field feature "xxx" is obtained, with a corresponding first sensitivity level of 0.9 and a corresponding first category of "string"; the first field feature "1234mmm!" is obtained, with a corresponding first sensitivity level of 0.8 and a corresponding first category of "password"; the first field feature "2025.4.5" is obtained, with a corresponding first sensitivity level of 0.3 and a corresponding first category of "date"; the first field feature "customer service" is obtained, with a corresponding first sensitivity level of 0.33 and a corresponding first category of "string"; the first field feature "Xiao Li" is obtained, with a corresponding first sensitivity level of 0.9 and a corresponding first category of "string"; the first field feature "attitude" is obtained, with a corresponding first sensitivity level of 0.45 and a corresponding first category of "string"; the first field feature "especially good" is obtained, with a corresponding first sensitivity level of 0.49 and a corresponding first category of "string".
[0133] Optionally, the first original data is used as the input of the random forest model 11, and at least one first field feature corresponding to the first original data, the first sensitivity level corresponding to each first field feature, and the first category corresponding to each first field feature are used as the output of the random forest model. The random forest model is incrementally learned to obtain an updated random forest model 11.
[0134] Step S32: Determine the second sensitive data in the first original data based on at least one first field feature, the first sensitivity level corresponding to each first field feature, and the first category corresponding to each first field feature.
[0135] Step S32 may refer to the similar contents of the above-mentioned step S22 and will not be described again here.
[0136] Continuing with the above example, if the preset sensitivity level is 0.5, based on the preset sensitivity level, the first field feature "1234mmm!", the first field feature "xxx" and the first field feature "Xiao Li" in the first original data are determined as the second sensitive data.
[0137] S4: Determine target sensitive data in the first original data.
[0138] The target sensitive data includes first sensitive data and second sensitive data.
[0139] There are various ways to determine the target sensitive data in the first original data. For example, the target sensitive data in the first original data can be determined by replacing or concatenating xxx. One example is described below.
[0140] In one example, based on the second sensitive data, the third sensitive data is replaced to obtain the target sensitive data.
[0141] Continuing with the above example, the second sensitive data is replaced with the third sensitive data. That is, the first field feature "1234mmm!", the first field feature "xxx", the first field feature "2025.4.5", and the first field feature "Xiao Li" are replaced with the second field feature "1234mmm!", the second field feature "xxx", and the second field feature "Xiao Li" to obtain the target sensitive data, that is, the target sensitive data includes the features "1234mmm!", "xxx", "Xiao Li", and "XX brand".
[0142] S5: Desensitize the target sensitive data to obtain desensitized data.
[0143] Among them, the desensitized data includes the target sensitive data after desensitization processing.
[0144] In one embodiment, the continued combination Figure 3 As shown, step S5 may include steps S51-S52.
[0145] S51: Determine the target desensitization strategy corresponding to the target sensitive data.
[0146] A targeted desensitization strategy is a specific desensitization method for each field in the target sensitive data. The most appropriate desensitization method is adopted based on the characteristics and sensitivity of each field or word segmentation to achieve personalized desensitization.
[0147] There are many ways to determine the target desensitization strategy. For example, a personalized desensitization strategy can be obtained based on a third machine learning model, or a personalized desensitization strategy can be determined based on predefined rules, or a data desensitization strategy can be determined based on a specific algorithm, such as a hash algorithm. Figure 3 As shown, one of the implementation modes is specifically described.
[0148] In one embodiment, step S51 includes step S511.
[0149] Step S511: Use the target sensitive data as the input of the third machine learning model, and use the third machine learning model to obtain the target desensitization strategy corresponding to the target sensitive data.
[0150] Among them, the third machine learning model 12 is used to obtain the desensitization strategy corresponding to the sensitive data.
[0151] The third machine learning model 12 is a machine learning model 12 that has been trained on a large-scale dataset before use. Optionally, the third machine learning model 12 can be a KNN recommendation model. The K-nearest neighbors (KNN) algorithm is an instance-based learning algorithm that can be used for classification and regression tasks. In an embodiment of the present application, the KNN model is used to recommend the optimal desensitization strategy based on historical execution records and contextual information.
[0152] Further, combined Figure 7 The training method of the third machine learning model 12 is explained.
[0153] Figure 7 This is a schematic diagram of the training process of a third machine learning model provided in an embodiment of the present application.
[0154] like Figure 7 As shown, the third machine learning model 12 training method includes the following steps:
[0155] Step S701: Obtain a third training set. The third training set includes at least one historical masking record and at least one historical masking strategy corresponding to the historical masking record.
[0156] Step S702: extracting desensitized data and target historical desensitization strategy from historical desensitization records.
[0157] Step S703: Use historical sensitive data as input of the third machine learning model and use the target historical desensitization strategy as output of the third machine learning model 12 to train the third machine learning model.
[0158] In one example, at least one historical sensitive data is input into the initial machine learning model 12, and the corresponding predicted desensitization strategy is output. Then, the loss values of the target historical desensitization strategy and the predicted desensitization strategy are determined using a preset loss function. Based on the loss values, the parameters in the initial machine learning model are continuously adjusted. When the above loss values are respectively less than the preset loss values, it means that the training of the initial machine learning model at this time is completed, and the third machine learning model 12 is obtained.
[0159] After the training of the third machine learning model 12 is completed, the third machine learning model 12 is used to obtain the target desensitization strategy corresponding to the target sensitive data.
[0160] Figure 8 This is a schematic diagram of a target desensitization strategy provided in an embodiment of the present application.
[0161] Continuing with the above example, Figure 8 As shown, using the third machine learning model 12, the desensitization strategy for the second field feature "1234mmm!" is "masking", the desensitization strategy for the second field feature "xxx" is "replacement", the desensitization strategy for the second field feature "Xiao Li" is deletion, and the desensitization strategy for the word segmentation feature "XX brand" is generalization.
[0162] Optionally, the target sensitive data is used as the input of the third machine learning model 12, and the target desensitization strategy corresponding to the target sensitive data is used as the output of the third machine learning model 12. The third machine learning model 12 is incrementally learned to obtain an updated third machine learning model 12.
[0163] S52: Based on the target desensitization strategy, desensitize the target sensitive data to obtain desensitized data corresponding to the first original data.
[0164] Among them, the desensitized data includes the target sensitive data after desensitization processing.
[0165] Following the above example, continue to combine Figure 8 As shown, "1234mmm!" is masked as "*******", "xxx" is replaced with "abc", "Xiao Li" is deleted, and "XX brand" is generalized to "a certain sports brand" to obtain desensitized data.
[0166] It should be noted that the above desensitization strategies are just illustrative of several strategies and are not specifically limited.
[0167] In summary, regardless of whether the data includes structured data and unstructured data, sensitive data can be accurately identified and desensitized, thereby improving the accuracy of data desensitization.
[0168] Corresponding to the aforementioned embodiments of the data desensitization method, the present application also provides embodiments of a data desensitization device.
[0169] Figure 9 This is a schematic diagram of a data desensitization device provided in an embodiment of the present application.
[0170] like Figure 9 As shown, the data desensitizing device 900 may include an acquisition module 910 , a determination module 920 , a desensitizing module 930 , and an update module 940 .
[0171] The acquisition module 910 is configured to: acquire first original data; the first original data includes a structured data part and an unstructured data part; the acquisition module 910 is also configured to: use the first original data as the input of the first machine learning model, and use the first machine learning model to obtain first sensitive data in the first original data; wherein the first sensitive data includes an unstructured sensitive data part; the acquisition module 910 is also configured to: use the first original data as the input of the second machine learning model, and use the second machine learning model to obtain second sensitive data in the first original data; wherein the second sensitive data includes a second structured sensitive data part; the determination module 920 is configured to: determine the target sensitive data in the first original data; wherein the target sensitive data includes the first sensitive data and the second sensitive data; the desensitization module 930 is configured to: perform desensitization processing on the target sensitive data to obtain desensitized data; wherein the desensitized data includes the target sensitive data after desensitization processing.
[0172] In one possible implementation, the acquisition module 910 is specifically configured to: use the first original data as the input of the first machine learning model, and use the first machine learning model to obtain at least one word segmentation feature corresponding to the unstructured data part in the first original data, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature; wherein the first category is used to characterize the result of dividing each word segmentation feature according to a preset word segmentation type; based on at least one word segmentation feature, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature, determine the first sensitive data in the first original data.
[0173] In one possible implementation, the acquisition module 910 is specifically configured to: use the first original data as the input of the BERT model, and use the BERT model to obtain at least one word segmentation feature corresponding to the unstructured data part in the first original data, the first sensitivity level corresponding to each word segmentation feature, and the first category corresponding to each word segmentation feature; wherein the BERT model is used to identify sensitive data in the first original data.
[0174] In one possible implementation, the acquisition module 910 is specifically configured to: use the first original data as input to the second machine learning model, and utilize the second machine learning model to obtain at least one first field feature of the structured data portion in the first original data, a first sensitivity level corresponding to each first field feature, and a second category corresponding to each first field feature; wherein the second category is used to represent the result of classifying each first field feature according to a preset field type;
[0175] Based on at least one first field feature, a second sensitivity level corresponding to each first field feature, and a second category corresponding to each first field feature, second sensitive data in the first original data is determined.
[0176] In one possible implementation, the acquisition module 910 is specifically configured to: use the first original data as the input of the random forest model, and use the random forest model to obtain at least one first field feature, the second sensitivity level corresponding to each first field feature, and the second category corresponding to each first field feature; wherein the random forest model is used to identify sensitive data in the structured data part.
[0177] In a possible implementation, the determination module 920 is specifically configured to: determine a target desensitization strategy corresponding to the target sensitive data; and perform desensitization processing on the target sensitive data based on the target desensitization strategy to obtain desensitized data.
[0178] In one possible implementation, the determination module 920 is specifically configured to: use the target sensitive data as the input of the third machine learning model, and use the third machine learning model to obtain the target desensitization strategy corresponding to the target sensitive data; wherein the third machine learning model is used to obtain the desensitization strategy corresponding to the sensitive data.
[0179] In one possible implementation, the acquisition module 910 is specifically configured to: acquire second original data; wherein the second original data is original data that has not been preprocessed; the second original data includes a structured data portion and an unstructured data portion; and obtain the preprocessed first original data based on the second original data.
[0180] In one possible implementation, the update module 940 is configured to: use the sample sensitive data as the input of the third machine learning model and the sample desensitization strategy as the output of the third machine learning model, and update the third machine learning model; wherein the sample sensitive data corresponds to the sample desensitization strategy.
[0181] Figure 10 This is a schematic diagram of a computing device provided in an embodiment of the present application.
[0182] like Figure 10 As shown, the computing device 1000 includes a processor 1001 and a memory 1002. By way of example, the computing device 1000 may further include a communication interface 1003 and a communication bus 1004.
[0183] The processor 1001, the memory 1002, and the communication interface 1003 communicate with each other via a communication bus 1004. The communication interface 1003 may include a transmitter and a receiver for communicating with other devices or a communication network, and may be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet interface (GE).
[0184] In some embodiments, the processor 1001 is configured to execute a program 1005, specifically, the relevant steps in the above-mentioned inference task execution method embodiment. Specifically, the program 1005 may include program code, which includes computer-executable instructions.
[0185] For example, processor 1001 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of the present application. The computing device 1000 may include one or more processors of the same type, such as one or more CPUs, or different types of processors, such as one or more CPUs and one or more ASICs. The CPU may be a single-core CPU (single-CPU) or a multi-core CPU (multi-CPU).
[0186] In some embodiments, the memory 1002 is used to store the program 1005. The memory 1002 may include a high-speed random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory.
[0187] Program 1005 may be specifically called by processor 1001 to enable computing device 1000 to perform the inference task execution method operation.
[0188] Some embodiments of the present application provide a computer-readable storage medium, which stores at least one executable instruction. When the executable instruction runs on a computing device 1000, the computing device 1000 executes the reasoning task execution method in the above embodiment.
[0189] The executable instructions may be specifically used to enable the computing device 1000 to perform the inference task execution method operations.
[0190] For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0191] Some embodiments of the present application provide a chip system for use in a server. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via circuits. The interface circuits are configured to receive signals from the server's memory and send signals to the processors. The signals include computer instructions stored in the memory. When the server processor executes the computer instructions, the server performs the steps of the inference task execution method described in the above method embodiment.
[0192] The beneficial effects that can be achieved by the readable storage medium provided in some embodiments of the present application can be referred to the beneficial effects in the corresponding reasoning task execution method provided above, and will not be repeated here.
[0193] It should be noted that, in the application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0194] Each embodiment in this specification is described in a related manner. Similar portions between the embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so their description is relatively simple. For related portions, refer to the description of the method embodiments.
[0195] The logic and / or steps represented in the flowchart or otherwise described herein may be considered, for example, as an ordered list of executable instructions for implementing logical functions, and may be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device).
[0196] For the purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with an instruction execution system, apparatus, or device.
[0197] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic device, and a portable compact disc read-only memory (CDROM).
[0198] In addition, the computer readable medium can even be paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in the computer memory. It should be understood that various parts of the present application can be implemented in hardware, software, firmware or a combination thereof.
[0199] In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following technologies known in the art can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc. The above-described embodiments are merely specific embodiments of the present application and are not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made based on the technical solutions of the present application shall be included within the scope of protection of the present application.
Claims
1. A data desensitization method, characterized in that: The data desensitization method includes: Acquire first original data; the first original data includes a structured data portion and an unstructured data portion; Using the first original data as input to a first machine learning model, and utilizing the first machine learning model to obtain first sensitive data; wherein the first sensitive data is sensitive data in the unstructured data portion; Using the first original data as input to a second machine learning model, and utilizing the second machine learning model to obtain second sensitive data; wherein the second sensitive data is sensitive data in the structured data portion; Determining target sensitive data in the first original data; wherein the target sensitive data includes the first sensitive data and the second sensitive data; Desensitizing the target sensitive data to obtain desensitized data.
2. The method according to claim 1, characterized in that The step of taking the first original data as input of a first machine learning model and using the first machine learning model to obtain first sensitive data includes: The first machine learning model uses the first raw data as input to obtain, using the first machine learning model, at least one word segmentation feature corresponding to the unstructured data portion in the first raw data, a first sensitivity level corresponding to each word segmentation feature, and a first category corresponding to each word segmentation feature; wherein the first category is used to represent a result of classifying each word segmentation feature according to a preset word segmentation type; The first sensitive data is determined based on the at least one word segmentation feature, the first sensitivity level corresponding to each of the word segmentation features, and the first category corresponding to each of the word segmentation features.
3. The method according to claim 2, characterized in that The step of using the first original data as input to the first machine learning model and utilizing the first machine learning model to obtain at least one word segmentation feature corresponding to the unstructured data portion in the first original data, a first sensitivity level corresponding to each word segmentation feature, and a first category corresponding to each word segmentation feature includes: The first original data is used as the input of the BERT model, and the BERT model is used to obtain at least one word segmentation feature corresponding to the unstructured data part in the first original data, a first sensitivity level corresponding to each of the word segmentation features, and a first category corresponding to each of the word segmentation features; wherein the BERT model is used to identify sensitive data in the unstructured data part.
4. The method according to claim 1, wherein The step of using the first original data as input to a second machine learning model and utilizing the second machine learning model to obtain second sensitive data includes: The second machine learning model uses the first original data as input to obtain, using the second machine learning model, at least one first field feature of the structured data portion in the first original data, a second sensitivity level corresponding to each first field feature, and a second category corresponding to each first field feature; wherein the second category is used to represent the result of classifying each first field feature according to a preset field type; The second sensitive data is determined based on the at least one first field feature, the second sensitivity level corresponding to each of the first field features, and the second category corresponding to each of the first field features.
5. The method according to claim 4, characterized in that The step of using the first original data as input to the second machine learning model and utilizing the second machine learning model to obtain at least one second field feature of the structured data portion in the first original data, a second sensitivity level corresponding to each second field feature, and a second category corresponding to each second field feature includes: The first original data is used as the input of a random forest model, and the random forest model is used to obtain the at least one first field feature, the second sensitivity level corresponding to each of the first field features, and the second category corresponding to each of the first field features; wherein the random forest model is used to identify sensitive data in the structured data portion.
6. The method according to claim 1, characterized in that The desensitizing the target sensitive data to obtain desensitized data includes: Determine a target desensitization strategy corresponding to the target sensitive data; Based on the target desensitization strategy, the target sensitive data is desensitized to obtain the desensitized data.
7. The method according to claim 6, characterized in that Determining the target desensitization strategy corresponding to the target sensitive data includes: The target sensitive data is used as the input of the third machine learning model, and the third machine learning model is used to obtain the target desensitization strategy corresponding to the target sensitive data; wherein the third machine learning model is used to obtain the desensitization strategy corresponding to the sensitive data.
8. The method according to any one of claims 1 to 7, characterized in that: The obtaining of the first original data includes: Acquire second original data; wherein the second original data is original data that has not been preprocessed; the second original data includes a structured data portion and an unstructured data portion; Based on the second original data, the preprocessed first original data is obtained.
9. A data desensitization device, characterized in that: The data desensitization device includes: The acquisition module is configured to: acquire first original data; the first original data includes a structured data portion and an unstructured data portion; The acquisition module is further configured to: use the first original data as input to a first machine learning model, and use the first machine learning model to obtain first sensitive data; wherein the first sensitive data is sensitive data in the unstructured data portion; the acquisition module is further configured to: use the first original data as input to a second machine learning model, and use the second machine learning model to obtain second sensitive data; wherein the second sensitive data is sensitive data in the structured data portion; the determination module is configured to: determine target sensitive data in the first original data; wherein the target sensitive data includes the first sensitive data and the second sensitive data; The desensitizing module is configured to: perform desensitizing processing on the target sensitive data to obtain desensitized data.
10. A computing device, characterized in that The computing device includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor is used to execute the computer instructions, the computing device executes the data desensitization method as described in any one of claims 1 to 8.