Method and device for processing plaintext element information and electronic equipment

By performing multi-level column name filtering and confidence calculation on structured data in the storage engine, sensitive information is identified and processed, the risk of high false alarm rates and sensitive data leakage in the prior art is solved, and efficient and accurate identification and processing of sensitive data is achieved.

CN120217440APending Publication Date: 2025-06-27DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510354691.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art detects sensitive data in user plaintext element information, the false alarm rate is high, and there is a risk of sensitive data leakage.

Method used

By obtaining the data in the storage engine, multi-level column name filtering of column data in the structured data, data that does not meet multiple filtering conditions are selected, and the confidence of the data is calculated. If the confidence level is higher than the preset threshold, it is identified as the target sensitive data.

Benefits of technology

It realizes efficient and accurate identification and processing of sensitive information present in the storage engine, reducing the risk of sensitive data breaches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217440A_ABST
    Figure CN120217440A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a plaintext element information processing method and device and electronic equipment, and the method comprises the steps: obtaining data in a storage engine, and carrying out the multi-level column name filtering of column data in structured data; data which do not meet two or more filtering conditions such as a regular filtering condition, a format rule filtering condition, a verification algorithm filtering condition, a column name filtering condition and an abstract filtering condition are screened out, the confidence coefficient of the data meeting the filtering conditions is calculated, and if the confidence coefficient is higher than a preset confidence coefficient threshold value, the data are filtered out; if yes, indicating that the data is target sensitive data containing sensitive information, and then generating and outputting a plaintext element information processing result after integrating the target sensitive data, therefore, the data user and the data maintenance party can maintain the plaintext element information based on the position description information of the target sensitive data contained in the plaintext element information processing result, so that the sensitive information can be efficiently and accurately identified and processed, and the leakage of the sensitive information in the plaintext element information is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method, apparatus, and electronic device for processing plaintext element information. Background Art

[0002] With the development of big data technology, all data generated by users when using the Internet will be recorded. Among the data generated by users, there are some plaintext element information. The plaintext element is information expressed in natural language or conventional symbols that has not been encrypted and can be directly understood and read by people. Among them, in different application scenarios, the types of plaintext element information are different. For example, taking a news event as an example, the plaintext element information may include: time, place, person, event process, reason, result, etc. Taking the financial scenario as an example, the plaintext element information may include: transaction time, transaction platform, both parties to the transaction, transaction amount, transaction remarks information, and so on.

[0003] Since users will generate some plaintext element information, and the plaintext element information will contain sensitive information such as the user's mobile phone number, ID card number, bank card number, etc., according to the requirements of the Data Security Law of the People's Republic of China, in the process of operation, enterprises have the obligation and responsibility to take corresponding technical measures to protect the sensitive information of users. Therefore, how to detect and encrypt and protect the sensitive data in the plaintext element information of users has become the key to ensuring user privacy and security.

[0004] In the prior art, the solution for detecting the plaintext element information of user sensitive data is: using the method of regular expressions for detection. In the actual application process, the method of using regular expressions for detection will have a high false positive rate, thus there is a risk of sensitive data leakage. Summary of the Invention

[0005] In view of this, the embodiments of this application provide a method, apparatus, and electronic device for processing plaintext element information, which reduces the risk of sensitive data leakage.

[0006] In a first aspect, the embodiments of this application provide a method for processing plaintext element information, where the method includes:

[0007] Obtain the data in the storage engine, where the data in the storage engine includes: structured data;

[0008] Perform first-level column name filtering and second-level column name filtering on the column data in the structured data, and screen out the target data columns that meet the first-level column name filtering conditions and the second-level column name filtering conditions. Among them, the first-level column name filtering is: identify sensitive information for the column names in the data table structure of the structured data. If there is sensitive information in the column names of the data table structure, it is determined that the data columns with the sensitive information meet the first-level column name filtering conditions; the second-level column name filtering is to identify sensitive information for the column names that do not belong to the preset column name whitelist according to the preset column name whitelist. If there is sensitive information in the column names that do not belong to the preset column name whitelist, it is determined that the data columns with the sensitive information meet the second-level column name filtering conditions;

[0009] Extract the values of the data stored in the target data columns, and sequentially detect the first sensitive information, the second sensitive information, and the third sensitive information. According to whether the values of the data meet two or more of the following filtering conditions: regular filtering conditions, format rule filtering conditions, verification algorithm filtering conditions, third-level column name filtering conditions, and summary filtering conditions, calculate the confidence levels of the values of the data. Among them, the first sensitive information, the second sensitive information, and the third sensitive information belong to three categories of sensitive information with different sensitive information priorities. The third-level column name filtering condition is: the column name of the data column to which the first sensitive information belongs contains a first preset character, the column name of the data column to which the second sensitive information belongs contains a second preset character, or the data column to which the third sensitive information belongs contains a third preset character. The summary filtering condition is: for the case where the values of the data stored in the target data columns are a piece of text, determine whether the values of the data in the data column to which the first sensitive information belongs contain the first preset character, determine whether the values of the data in the data column to which the second sensitive information belongs contain the second preset character, and determine whether the values of the data in the data column to which the third sensitive information belongs contain the third preset character;

[0010] Determine the data with a confidence level higher than the preset confidence level threshold as the target sensitive data, and integrate the target sensitive data to generate and output the plaintext element information processing result. Among them, the plaintext element information processing result contains the position description information of each target sensitive data.

[0011] In some possible embodiments, the data in the storage engine further includes: unstructured data, and the method further includes:

[0012] For the values of each item of data in the unstructured data, detect the first sensitive information, the second sensitive information, and the third sensitive information in sequence, and calculate the confidence level of the value of each item of data according to whether the value of each item of data meets two or more filtering conditions among the regular filtering condition, the format rule filtering condition, the verification algorithm filtering condition, and the digest filtering condition;

[0013] Determine the unstructured data with a confidence level higher than the preset confidence level threshold as the target sensitive data.

[0014] In some possible embodiments, the method further includes:

[0015] Calculate the column name confidence level of each item of data according to whether each item of data in the structured data meets the third-level column name filtering condition, and calculate the digest confidence level of each item of data according to whether each item of data in the structured data or each item of data in the unstructured data meets the digest filtering condition;

[0016] Input the preprocessed sensitive data with the column name confidence level meeting the preset column name confidence level threshold and / or the digest confidence level meeting the preset digest confidence level threshold into the sensitive information detection model, and the sensitive information detection model outputs the model confidence level of the preprocessed sensitive data based on the input data;

[0017] Calculate the comprehensive confidence level of each preprocessed sensitive data according to the column name confidence level, the digest confidence level, and the model confidence level;

[0018] If the comprehensive confidence level is greater than the preset comprehensive confidence level threshold, determine the preprocessed sensitive data greater than the preset comprehensive confidence level threshold as the target sensitive data.

[0019] In some possible embodiments, the sensitive information detection model is trained based on the following method:

[0020] Perform perturbation processing on the target sensitive data to generate training sample data, input the training sample data into the initial neural network model, and train the initial neural network model until the difference between the recognition result output by the initial neural network model based on the training sample data and the target sensitive data is maximized.

[0021] In some possible embodiments, the performing perturbation processing on the target sensitive data to generate training sample data includes:

[0022] Add noise and data transformation to the target sensitive data to generate and store the training sample data.

[0023] In a second aspect, an embodiment of the present application provides an apparatus for processing plaintext element information. The apparatus includes:

[0024] An acquisition module, configured to acquire data in a storage engine, where the data in the storage engine includes: structured data;

[0025] A first-level filtering module, configured to perform first-level column name filtering and second-level column name filtering on column data in the structured data, and screen out target data columns that meet the first-level column name filtering condition and the second-level column name filtering condition. The first-level column name filtering is: identifying sensitive information for column names in the data table structure of the structured data. If there is sensitive information in the column names of the data table structure, it is determined that the data column with the sensitive information meets the first-level column name filtering condition; the second-level column name filtering is to identify sensitive information for column names that do not belong to the preset column name whitelist according to the preset column name whitelist. If there is sensitive information in the column names that do not belong to the preset column name whitelist, it is determined that the data column with the sensitive information meets the second-level column name filtering condition;

[0026] A second-level filtering module, configured to extract the values of the data stored in the target data columns, and sequentially detect the first sensitive information, the second sensitive information, and the third sensitive information. According to whether the values of the data meet two or more filtering conditions among the regular filtering condition, the format rule filtering condition, the verification algorithm filtering condition, the third-level column name filtering condition, and the digest filtering condition, calculate the confidence level of the values of the data; where the first sensitive information, the second sensitive information, and the third sensitive information belong to three categories of sensitive information with different sensitive information priorities. The third-level column name filtering condition is: the column name of the data column where the first sensitive information belongs contains a first preset character, the column name of the data column where the second sensitive information belongs contains a second preset character, or the data column where the third sensitive information belongs contains a third preset character. The digest filtering condition is: for the case where the values of the data stored in the target data columns are a piece of text, determine whether the values of the data in the data column where the first sensitive information belongs contain the first preset character, determine whether the values of the data in the data column where the second sensitive information belongs contain the second preset character, and determine whether the values of the data in the data column where the third sensitive information belongs contain the third preset character;

[0027] An output module, configured to determine the data with a confidence level higher than a preset confidence level threshold as target sensitive data, integrate the target sensitive data, and generate and output a plaintext element information processing result, where the plaintext element information processing result includes position description information of each target sensitive data.

[0028] In some possible embodiments, the data in the storage engine further includes: unstructured data, and the second-level filtering module is further configured to:

[0029] For each value of the data in the unstructured data, sequentially detect the first sensitive information, the second sensitive information, and the third sensitive information, and calculate the confidence level of each value of the data according to whether each value of the data meets two or more filtering conditions among the regular filtering condition, the format rule filtering condition, the verification algorithm filtering condition, and the digest filtering condition;

[0030] Determine the unstructured data with a confidence level higher than the preset confidence level threshold as the target sensitive data.

[0031] In some possible embodiments, the second-level filtering module is further configured to:

[0032] Calculate the column name confidence level of each item of the data according to whether each item of the data in the structured data meets the third-level column name filtering condition, and calculate the digest confidence level of each item of the data according to whether each item of the data in the structured data or each item of the unstructured data meets the digest filtering condition;

[0033] Input the preprocessed sensitive data with the column name confidence level meeting the preset column name confidence level threshold and / or the digest confidence level meeting the preset digest confidence level threshold into the sensitive information detection model, and the sensitive information detection model outputs the model confidence level of the preprocessed sensitive data based on the input data.

[0034] In some possible embodiments, the device further includes:

[0035] A third-level filtering module, configured to calculate the comprehensive confidence level of each preprocessed sensitive data according to the column name confidence level, the digest confidence level, and the model confidence level;

[0036] If the comprehensive confidence level is greater than the preset comprehensive confidence level threshold, determine the preprocessed sensitive data greater than the preset comprehensive confidence level threshold as the target sensitive data.

[0037] In some possible embodiments, the device further includes:

[0038] A model training module, configured to perform perturbation processing on the target sensitive data to generate training sample data, input the training sample data into the initial neural network model, and train the initial neural network model until the difference between the recognition result output by the initial neural network model based on the training sample data and the target sensitive data is maximized;

[0039] In some possible embodiments, the perturbation processing of the target sensitive data to generate training sample data includes:

[0040] Adding noise and data transformation to the target sensitive data, and generating and storing the training sample data.

[0041] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes: a processor; and a memory storing a program; where the program includes instructions that, when executed by the processor, cause the processor to execute the method for processing plaintext element information described in the first aspect.

[0042] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method for processing plaintext element information described in the first aspect.

[0043] Advantages of the present application:

[0044] An embodiment of the present application provides a method, apparatus, and electronic device for processing plaintext element information. Among them, the method obtains data in a storage engine, performs multi-level column name filtering on column data in structured data, and then screens out data that does not meet two or more filtering conditions such as regular filtering conditions, format rule filtering conditions, verification algorithm filtering conditions, column name filtering conditions, and digest filtering conditions, and calculates the confidence of the data that meets the filtering conditions. If the confidence is higher than a preset confidence threshold, it indicates that the data is target sensitive data containing sensitive information. Then, after integrating these target sensitive data, a processing result of plaintext element information is generated and output, so that data users and data maintainers can maintain the plaintext element information based on the location description information of the target sensitive data included in the processing result of plaintext element information, so as to reduce the leakage of sensitive information in the plaintext element information.

[0045] By selecting the embodiment of the present application, in addition to detecting sensitive data using a simple regular expression detection method, multiple detection algorithms such as format rule detection, verification algorithm detection, column name detection, and digest detection are introduced for detection, and then by calculating the confidence of each item of data, a comprehensive sensitive data evaluation can be performed on each item of data, which can efficiently and accurately identify and process sensitive information existing in the storage engine and reduce the risk of sensitive data leakage. Description of the Drawings

[0046] In the following description of exemplary embodiments in conjunction with the drawings, more details, features, and advantages of the present application are disclosed. In the drawings:

[0047] Figure 1Shows a schematic flowchart of a method for processing plaintext element information provided by an embodiment of the present application;

[0048] Figure 2 Shows another schematic flowchart of a method for processing plaintext element information provided by an embodiment of the present application;

[0049] Figure 3 Shows a schematic structural diagram of a device for processing plaintext element information provided by an embodiment of the present application;

[0050] Figure 4 Shows a structural block diagram of an exemplary electronic device that can be used to implement an embodiment of the present application. Detailed implementation manners

[0051] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.

[0052] It should be understood that the various steps recorded in the method embodiments of the present application can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.

[0053] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions executed by these devices, modules or units.

[0054] It should be noted that the modifications of "one" and "multiple" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0055] For the convenience of clearly understanding the technical problems that the present application actually needs to solve, the common existing sensitive data monitoring methods are described in detail here:

[0056] 1) The sensitive information detection method based on regular expressions extracts all data that conforms to the regular expression rules, so the data recall rate is high. The recognition accuracy of sensitive information is related to the rules of the regular expression, and the rules are not strict enough. For information types with relatively loose formats, such as mobile phone numbers, it will result in a high false positive rate.

[0057] 2) The sensitive information detection method based on deep learning models. Since the cost of deep learning models is high, a large amount of positive and negative sample data is required for model training, and it is necessary to let the deep learning model learn what kind of data belongs to sensitive information. The recognition recall rate and accuracy of this method depend on a large amount of real sensitive training data. In the case of insufficient data volume, the recall rate and accuracy will be poor. Using a large amount of real data to train the model is likely to have the risk of data leakage. At the same time, during the detection process, the data to be detected often has a very large volume. Using the model for all detections will consume a large amount of GPU resources and time, making it difficult to implement.

[0058] Through comprehensive analysis, it can be seen that the existing sensitive information detection methods are prone to the risk of sensitive data leakage. In view of this, the present application provides a method, device, and electronic device for processing plaintext element information. First, the present application provides a method for processing plaintext element information. This method is applied to any electronic device with the function of processing plaintext element information, including but not limited to personal mobile terminals, computers, or servers, etc. As Figure 1 shown, the method includes the following steps:

[0059] S11. Obtain the data in the storage engine, where the data in the storage engine includes: structured data;

[0060] S12. Perform first-level column name filtering and second-level column name filtering on the column data in the structured data to screen out the target data columns that meet the first-level column name filtering conditions and the second-level column name filtering conditions;

[0061] Among them, the first-level column name filtering is: perform sensitive information recognition on the column names in the data table structure of the structured data. If there is sensitive information in the column names in the data table structure, it is determined that the data column with the sensitive information meets the first-level column name filtering conditions; the second-level column name filtering is to perform sensitive information recognition on the column names that do not belong to the preset column name whitelist according to the preset column name whitelist. If there is sensitive information in the column names that do not belong to the preset column name whitelist, it is determined that the data column with the sensitive information meets the second-level column name filtering conditions;

[0062] S13. Extract the values of the data stored in the target data column, and sequentially detect the first sensitive information, the second sensitive information, and the third sensitive information. Calculate the confidence level of the values of the data according to whether the values of the data meet two or more of the following filtering conditions: regular filtering condition, format rule filtering condition, verification algorithm filtering condition, third-level column name filtering condition, and summary filtering condition;

[0063] Among them, the first sensitive information, the second sensitive information, and the third sensitive information belong to three categories of sensitive information with different sensitive information priorities. The third-level column name filtering condition is that the column name of the data column to which the first sensitive information belongs contains a first preset character, the column name of the data column to which the second sensitive information belongs contains a second preset character, or the data column to which the third sensitive information belongs contains a third preset character. The summary filtering condition is that for the case where the values of the data stored in the target data column are a piece of text, determine whether the values of the data in the data column to which the first sensitive information belongs contain the first preset character, determine whether the values of the data in the data column to which the second sensitive information belongs contain the second preset character, and determine whether the values of the data in the data column to which the third sensitive information belongs contain the third preset character;

[0064] S14. Determine the data with a confidence level higher than the preset confidence level threshold as target sensitive data, and integrate the target sensitive data to generate and output the processing result of the plaintext element information;

[0065] Among them, the processing result of the plaintext element information contains the position description information of each piece of the target sensitive data.

[0066] In the embodiment of the present application, by obtaining the data in the storage engine, performing multi-level column name filtering on the column data in the structured data, then screening out the data that does not meet two or more of the following filtering conditions: regular filtering condition, format rule filtering condition, verification algorithm filtering condition, column name filtering condition, and summary filtering condition, and calculating the confidence level of the data that meets the filtering conditions. If the confidence level is higher than the preset confidence level threshold, it indicates that the data is target sensitive data containing sensitive information. Then, after integrating these target sensitive data, generate and output the processing result of the plaintext element information, so that the data user and the data maintainer can maintain the plaintext element information based on the position description information of the target sensitive data contained in the processing result of the plaintext element information, so as to reduce the leakage of sensitive information in the plaintext element information.

[0067] By selecting the embodiments of the present application, in addition to detecting sensitive data by using a simple regular expression detection method, multiple detection algorithms such as format rule detection, verification algorithm detection, column name detection, and abstract detection are introduced for detection. Then, by calculating the confidence levels of various data items, a comprehensive sensitive data assessment can be performed on various data items, enabling efficient and accurate identification and processing of sensitive information existing in the storage engine and reducing the risk of sensitive data leakage.

[0068] The following will elaborate on the above steps S11 to S14 in combination with specific examples:

[0069] In the embodiments of the present application, the storage engine is a component in a database management system (DBMS) responsible for managing data storage and retrieval. It is like the "engine" of the database, determining how data is stored, organized, and how it is accessed and operated. Among them, various data types can be supported in the storage engine, and the specific data types depend on the type of the database managed by the storage engine. Exemplarily, if the database managed by the storage engine is a relational database, the storage engine will store the data in data tables. The data tables are composed of rows and columns. Each row represents a record, and each column represents an attribute of the record. This structure enables the data to have a clear logical relationship, facilitating operations such as querying, inserting, updating, and deleting by the storage engine. If the database managed by the storage engine is a non-relational database, the storage engine will store the data in an unstructured database in the form of text files, images, audio, videos, etc., such as in the Hadoop distributed file system.

[0070] Based on this, when performing step S11, various types of data under the storage engine are read or traversed. However, since the storage forms of structured data and unstructured data are different, the identification methods for sensitive information in structured data and unstructured data are different. Specifically, for structured data, step S12 needs to be executed before step S13. For unstructured data, step S12 does not need to be executed, and step S13 is directly executed. When performing step S13, the step of filtering the third-level column names can also be omitted. It can be understood that for unstructured data, since unstructured data does not store data rows and data columns, column name filtering is not required, and each item of data in the unstructured data can be separately targeted for sensitive information identification. In the embodiments of the present application, when performing step S11, the specific data format of the data in the storage engine can be obtained to distinguish whether the obtained data is structured data or unstructured data.

[0071] Further, it can be understood that step S12 is only applicable to structured data. Specifically, each data column in the structured data is screened one by one to determine whether the target data column that meets the first-level column name filtering condition and the second-level column name filtering condition is included. Among them, the first-level column name filtering condition can be understood as identifying sensitive information for all column names in the data table structure in the structure. If there is sensitive information in the column names in the data table structure, it can be determined that the data column with sensitive information meets the first-level column name filtering condition. Among them, to determine whether there is sensitive information in the column names in the data table structure, the attributes of the set sensitive information can be matched with the column names. If the matching degree is greater than the preset matching degree threshold, it can be determined that there is sensitive information in the column name.

[0072] Exemplarily, taking the user information table as an example, if it includes 11 data columns, and the column names of each data column are: user name, user ID, user gender, user age, user education level, user ID number, user mobile phone number, user bank card number, user's first registration time, user's most recent usage time, user's most recent usage project name. If information such as "name, gender, age, education, ID number, mobile phone number, bank card number" is preset as sensitive information, then the column names of the above 11 data columns can be matched with the 7 items of sensitive information "user name, user gender, user age, user education level, user ID number, user mobile phone number, user bank card number". If one or more of them match successfully, it indicates that there is sensitive information in the column names in the data table structure, meeting the first-level column name filtering condition. Then, after marking the data columns that meet the first-level column name filtering condition, they are separately recorded or stored, that is, the data columns with sensitive information are extracted from the data table for further sensitive information identification.

[0073] Further, in the process of executing step S12, the operation of the second-level column name filtering for the column data in the structured data is similar to that of the first-level column name filtering. The difference is that the first-level column name filtering is to identify sensitive information for all column names as a whole to determine whether there is a data column containing sensitive information, while the second-level column name filtering is to identify sensitive information for the data columns that do not belong to the preset column name whitelist. There are two links in this process: 1) First, determine whether the data column does not belong to the data column in the preset column name whitelist. 2) Judge the sensitive information of the column names of the data columns that do not belong to the preset column name whitelist.

[0074] Specifically, if a data column belongs to the data columns in the preset column name whitelist, it indicates that the data column belongs to the recognized publicly available information, and the information in this data column is by default not sensitive information. If a data column does not belong to the data columns in the preset column name whitelist, it indicates that the data in this data column does not belong to the information that is publicly available by default, and the data in this data column cannot be publicly disclosed at will. Exemplarily, the preset column name whitelist can be set according to the actually publicly available information. Exemplarily, the information in the preset column name whitelist can be: legal representative's name, company name, unified social credit code, establishment time, legal representative's contact information, company address. If the column name of a data column is the legal representative's name, it indicates that the column name of this data column belongs to the column names in the preset column name whitelist.

[0075] Further, for the data columns that do not belong to the preset column name whitelist, the process such as the first-level column name filtering is performed again, and the column names of the data columns that do not belong to the preset column name whitelist are matched with the pre-set sensitive information. If the match is successful, and the column name belongs to one or more of the pre-set sensitive information, it indicates that there is sensitive information in this data column, and it can be determined that this data column meets the second-level column name filtering condition, and this data column is marked as the target data column.

[0076] Up to this point, the data columns with sensitive information can be screened out from the huge amount of data stored in the storage engine, without the need to identify sensitive information for the values of the data belonging to X rows and Y columns (similar to the data in a cell in a data table), improving the recall rate of sensitive information identification and being beneficial to improving the efficiency of sensitive information identification.

[0077] Based on the target data columns screened out in the above step S12, it indicates that there is a large amount of sensitive information in the target data columns. If the data in the target data columns is not desensitized, there will be a problem of sensitive information leakage. Based on this, step S13 is executed to check each item of data in each target data column one by one to determine whether it contains specific sensitive information. As a preferred implementation manner, the first sensitive information, the second sensitive information, and the third sensitive information depend on the priority of the sensitive information. The higher the priority of the sensitive information, the more likely it can be determined as the first sensitive information, the second sensitive information, and the third sensitive information for priority screening.

[0078] Among them, the categories of the first sensitive information, the second sensitive information, and the third sensitive information are different, and the specific categories can be flexibly set according to the actual application process. As a preferred implementation manner, the first sensitive information, the second sensitive information, and the third sensitive information are the three categories of sensitive information with the highest priority of sensitive information. The priority of the sensitive information can be flexibly set according to actual experience, and this application does not make strict restrictions.

[0079] Exemplarily, taking the above 7 items of sensitive information such as "user name, user gender, user age, user education level, user ID number, user mobile phone number, user bank card number" as an example, according to the priority of data security, it can be determined that the ID number, mobile phone number, and bank card number are sensitive information with the highest priority, and the user name, user gender, and user education level are sensitive information with a relatively high priority. Then, the mobile phone number is taken as the first sensitive information, the ID number is taken as the second sensitive information, and the bank card number is taken as the third sensitive information.

[0080] Then, each piece of data in the target data column is checked one by one to determine whether the value of each piece of data specifically belongs to a mobile phone, ID number, or bank card number. Specifically, taking the first sensitive information as the mobile phone number as an example, the data in the target data column is checked to see if it belongs to a mobile phone number. Exemplarily, if the column name of a certain data column is: Contact Information, and this data column contains 100 contact information, and the specific contact information can be: contact address, contact landline number, contact mobile phone number. At this time, two or more filtering methods such as regular filtering, format rule filtering, verification algorithm filtering, third-level column name filtering, and summary filtering can be adopted for screening to screen out the mobile phone numbers existing in these 100 contact information.

[0081] Specifically, the regular filtering condition is to filter based on a regular expression. Exemplarily, a mobile phone number regular expression is written according to the structure of 11 consecutive digits. If the value of a certain piece of data conforms to this mobile phone number regular expression, it can be determined that this piece of data may be a mobile phone number.

[0082] The format rule filtering condition is to filter based on the data format of sensitive information. For example, the mobile phone number is 11 digits, or a 14-digit string containing the area code "+XX". At this time, if the data format of a certain piece of data is 11 digits, or a 14-digit string containing the area code "+XX", it can be determined that this piece of data may be a mobile phone number. Another example is that the ID number is 18 digits. If the data format of a certain piece of data is 18 digits, it indicates that this piece of data may be an ID number. Or, according to the characteristics of specific sensitive information, such as the first 3-digit number segment of the mobile phone number, the first 6-digit card BIN of the bank card, the first 6 digits of the ID number are the provincial and municipal numbers, and the middle 8 digits are the date of birth, etc.

[0083] The verification algorithm filtering condition is to filter based on the verification algorithm constructed from sensitive information. Among them, different types of sensitive information correspond to different verification algorithms. Exemplarily, taking the second sensitive information as the ID number as an example, the ID number verification algorithm is an algorithm constructed based on the ID verification code ISO 7064:1983.MOD 11-2. Specifically, the weighted sum of the first 17 digits of the data can be calculated, and the corresponding weighting factors are 7, 9, 10, 5, 8, 4, 2, 1, 6, 3, 7, 9, 10, 5, 8, 4, 2 in sequence. Multiply each digit by its corresponding weighting factor, and then add up all the products to obtain the total sum. For example, for the ID number 320106aaaabbbbb001X, assuming the calculation process is: 3*7 + 2*9 + 0*10 + 1*5 + 0*8 + 6*4 + a*2 + a*1 + a*6 + a*3 + b*7 + 1*9, the calculation result is 189. Then divide the calculation result by 11 to obtain the remainder, that is, 189*11 = 17 ······ 2, the remainder is 2, and then find the corresponding verification code according to the remainder. Among them, the corresponding relationship between the remainder and the verification code is: 0 corresponds to 1, 1 corresponds to 0, 2 corresponds to X, 3 corresponds to 9, 4 corresponds to 8, 5 corresponds to 7, 6 corresponds to 6, 7 corresponds to 5, 8 corresponds to 4, 9 corresponds to 3, (10)X corresponds to 2. Therefore, when the remainder is 2, if the last digit of the ID number is X, it indicates that this item of data is a valid ID number and is sensitive information. Similarly, for bank card numbers, there is also a corresponding verification algorithm. Based on the corresponding verification algorithm, it can be determined whether each item of data is a valid bank card number. If it is a valid bank card number, then the bank card number is sensitive information.

[0084] The third-level column name filtering condition is a secondary verification of the column name based on the first-level column name filtering condition and the second-level column name filtering condition to determine whether the first-level column name filtering and the second-level column name filtering are true and valid, so as to ensure the accuracy of data column screening. Specifically, the third-level column name filtering is to verify each item of data in the target data column to check whether each item of data belongs to the data column based on which the first-level column name filtering is performed.

[0085] Specifically, the third-level column name filtering condition is as follows: the column name of the data column to which the first sensitive information belongs contains the first preset character. For example, if the first sensitive information is a mobile phone number, the first preset characters are: mobile phone, phone, hand, machine, etc. If the word "mobile phone" exists in the column name of the data column to which the first sensitive information belongs, it indicates that there is a misjudgment during the first-level column name filtering and the second-level column name filtering. The column name of the data column to which the second sensitive information belongs contains the second preset character. For example, if the second sensitive information is an ID number, the second preset characters are: body, or part. If the column name of the data column to which the second sensitive information belongs does not contain "body" or "part", but only contains "number", it indicates that there is a misjudgment during the first-level column name filtering and the second-level column name filtering. Similarly, if the data column to which the third sensitive information belongs contains the third preset character, for example, the bank card number does not contain the three special characters of "bank", "card", it indicates that there is a misjudgment during the first-level column name filtering and the second-level column name filtering.

[0086] The abstract filtering condition is applicable to unstructured data or text data. Specifically, when the values of the data stored in the target data column are a piece of text, the same processing idea as the third-level column name filtering is adopted to check whether the text contains the first preset character, the second preset character, and the third preset character. Specifically, for the values of the data in the data column to which the first sensitive information belongs, it is determined whether each piece of data in this data column contains the first preset character. For the values of the data in the data column to which the second sensitive information belongs, it is determined whether each piece of data in this data column contains the second preset character. Similarly, for the values of the data in the data column to which the third sensitive information belongs, it is determined whether each piece of data in this data column contains the third preset character. Among them, the specific characters of the first preset character, the second preset character, and the third preset character can be flexibly set according to the actual application. Exemplarily, the first preset characters include: +86, hand, machine, mobile phone number, cellphone, mobilephone, etc. The second preset characters include: body, part, Idfication and other information associated with the ID card. The third preset characters include: bank, card, bank, account and other information associated with the bank card. The range of the third preset characters is relatively wide and can specifically be Chinese and English characters associated with each character in the first sensitive information, the second sensitive information, and the third sensitive information.

[0087] In the embodiment of the present application, during the execution of step S13, the regular confidence level can be set for each piece of data according to the matching degree of the regular expression, and the format rule confidence level can be set for each piece of data according to the matching degree of the format rule filtering, and so on. The confidence level under each filtering condition can be determined according to the matching degree of each filtering condition. Exemplarily, if the matching degree of the regular expression of a certain piece of data is 80%, the regular confidence level of 0.8 is set for this piece of data.

[0088] Further, since the higher the confidence level, the greater the probability that the data item is sensitive data. Based on this, step S14 is executed to determine the data items with confidence levels greater than a set confidence threshold as target sensitive data. The preset confidence threshold may vary according to different filtering conditions or may be a unified value. It can be flexibly set according to actual experience, and this application does not make strict limitations. As a possible implementation manner, the confidence levels calculated for each item can be weighted and summed to obtain a weighted confidence level, and then it is determined whether the weighted confidence level is greater than the preset confidence threshold. If it is greater, it indicates that the data item is sensitive data. In this way, the confidence results obtained from multiple sensitive information detection dimensions of a data item can be comprehensively used to accurately determine whether the data item is sensitive data, so as to improve the accuracy of sensitive data detection.

[0089] In some possible embodiments, the data in the storage engine further includes: unstructured data. As described above, this unstructured data cannot perform data column name filtering. Therefore, the related operations of the first-level column name filtering, the second-level column name filtering, and the third-level column name filtering can be directly skipped, and it is directly determined whether the value of the unstructured data satisfies two or more of the following filtering conditions: regular filtering condition, the format rule filtering condition, the verification algorithm filtering condition, and the digest filtering condition. The confidence level of each data value is calculated, so as to perform sensitive data detection on the unstructured data. The specific execution process can refer to the detailed description of step S13 above, and will not be elaborated here.

[0090] On this basis, in some possible embodiments, the method provided in this application, in the process of executing step S13, may specifically include the following steps:

[0091] S13-1. Calculate the column name confidence level of each data according to whether each data of the structured data satisfies the third-level column name filtering condition, and calculate the digest confidence level of each data according to whether each data of the structured data or each data of the unstructured data satisfies the digest filtering condition;

[0092] S13-2. Input the preprocessed sensitive data with the column name confidence level higher than the preset column name confidence threshold and / or the digest confidence level higher than the preset digest confidence threshold into the sensitive information detection model, and the sensitive information detection model outputs the model confidence level of the preprocessed sensitive data based on the input data;

[0093] S13-3. Calculate the comprehensive confidence level of each preprocessed sensitive data according to the column name confidence level, the digest confidence level, and the model confidence level;

[0094] S13-4. If the comprehensive confidence level is greater than a preset comprehensive confidence level threshold, determine the preprocessed sensitive data greater than the preset comprehensive confidence level threshold as the target sensitive data.

[0095] Among them, when performing step S13-1, the content of calculating the column name confidence level and the abstract confidence level of each item of data in step S13 above can be referred to, which will not be elaborated here. In this way, the preprocessed sensitive data with a high column name confidence level and a high abstract confidence level can be initially screened out. This preprocessed sensitive data belongs to the data with a high probability of being sensitive data. At this time, this preprocessed sensitive data can be input into the sensitive information detection model, and the sensitive information detection model is used to further check this preprocessed sensitive data again, and output the model confidence level of the preprocessed sensitive data. Among them, the sensitive information detection model is a neural network model trained based on a deep learning model, which can automatically learn the characteristics of non-sensitive data, and then identify the input data based on the characteristics of non-sensitive data to determine whether the input data has the characteristics of non-sensitive data. If there are more characteristics of non-sensitive data, it indicates that the probability of the input household being sensitive data is lower, and then the model confidence level of the input data can be output. This model confidence level represents the probability that the preprocessed sensitive data is sensitive data.

[0096] Further, when performing step S13-3, the weighted confidence level can be calculated in a weighted summation manner based on the column name confidence level, the abstract confidence level, and the model confidence level as the comprehensive confidence level of the preprocessed sensitive data. When performing step S13-4, if this comprehensive confidence level is greater than the preset comprehensive confidence level threshold, the preprocessed sensitive data greater than the preset comprehensive confidence level threshold can be determined as the target sensitive data. Among them, the preset comprehensive confidence level threshold is also flexibly set based on actual experience or requirements, and this application does not make strict restrictions.

[0097] In some possible embodiments, the above sensitive information detection model can be pre-trained in the following manner:

[0098] Perform perturbation processing on the target sensitive data to generate training sample data, input the training sample data into the initial neural network model, and train the initial neural network model until the difference between the recognition result output by the initial neural network model based on the training sample data and the target sensitive data is maximized.

[0099] Specifically, as an implementation manner, the perturbation processing of the target sensitive data can be:

[0100] Add noise and data transformation to the target sensitive data to generate and store the training sample data.

[0101] In the embodiments of the present application, the training sample data can be desensitized in advance based on sensitive data, such as adding noise, changing the positions between numbers, or replacing the symbols of characters at specific positions, so that the sensitive data becomes non-sensitive data. Additionally, it can also be to perturb the target sensitive data obtained by performing step S14, such as adding noise, changing the information in the sensitive data, etc., to generate a large amount of non-sensitive data. Specifically, on the premise of ensuring that it can pass the existing detection rules, the numbers at certain positions in the target sensitive data can be randomly modified to lose the corresponding relationship with the original user, and then encrypted and stored as training sample data.

[0102] Then, input the non-sensitive data into the initial neural network model, instruct the initial neural network model to learn the data features of the non-sensitive data, and then output the corresponding recognition result. If the difference between the output recognition result and the target sensitive data is larger, it indicates that the neural network model can more accurately determine whether it is non-sensitive data based on the data features. On the contrary, if the analysis result shows that the features of the non-sensitive data are less, it indicates that the difference between the input data and the target sensitive data is smaller, and the probability that the corresponding input data is sensitive data is greater.

[0103] Adopting the embodiments of the present application can effectively overcome the problems existing in the second point in the above description of the prior art, thereby reducing the problem of sensitive information leakage in the model training stage.

[0104] When performing step S14, integrating the target sensitive data can be understood as, when there are multiple target sensitive data, summarizing the multiple target sensitive data, and obtaining the specific storage locations of each target sensitive data, or specifically the locations in the storage engine, determining the location information as the location description information of the target sensitive data, and after summarizing it with the target sensitive data, obtaining a sensitive data summary table. This sensitive data summary table can be regarded as a processing result of plaintext element information, and finally output to the corresponding data operation and maintenance personnel to remind the data operation and maintenance personnel that there is plaintext element information that has not been encrypted and that encryption measures need to be taken to encrypt this plaintext element information to avoid sensitive information leakage.

[0105] As another implementation manner, all data sources can be traversed, and a small amount of detected sensitive data is desensitized and displayed to the data operation and maintenance personnel, along with the specific locations of the existing sensitive data and the probabilities corresponding to the sensitive data, to prompt the operation and maintenance personnel of the risk of data leakage.

[0106] To facilitate understanding of the method provided in the present application, it can be understood in combination with the Figure 2 method flow shown as follows:

[0107] By performing the first-level column name filtering and the second-level column name filtering in step S11, the first-level column name filtering can be understood as column name filtering at the general level, and the second-level column name filtering can be understood as column name filtering at the business whitelist level. Then, the obtained filtered results are used to further identify sensitive data in each item of data under the three types of data columns of mobile phone number filtering, ID card filtering, and bank card filtering. Specifically, mobile phone number filtering is relatively simple and involves regular filtering, format rule filtering, third-level column name filtering, and summary filtering. While ID cards and bank cards are relatively more complex. In addition to the three types of filtering included in mobile phone numbers, a verification algorithm filtering is added to further accurately determine whether the plaintext element information currently obtained is really a real ID card number or bank card number.

[0108] Furthermore, the filtered results of each item of data are input into the sensitive information detection model, and steps S13-2 and S13-3 are executed for model filtering to calculate the model confidence. Finally, the comprehensive confidence is calculated, and the detection results of each item of preprocessed sensitive data with the comprehensive confidence greater than the preset comprehensive confidence threshold are counted as each target sensitive data. A small amount of target sensitive data is extracted, desensitized, output, and displayed. During the process, the sensitive data needs to be perturbed and encrypted for storage to train the sensitive information detection model and reduce the leakage of sensitive information.

[0109] By adopting the embodiment of the present application, the processing flow of plaintext element information is optimized. Tools such as regular expressions, data format verification, and deep learning models are used to comprehensively evaluate multiple dimensions such as the column names, regular expressions, and summaries of data columns, improving the sensitive information detection ability and finally obtaining a detection scheme with high accuracy and high recall rate. Moreover, by adopting the method of first performing policy detection and then model detection, the amount of data that the sensitive information detection model needs to process can be reduced, thereby reducing resource consumption and helping to expand the application scope of this method.

[0110] Based on the method provided in the first aspect, in the second aspect, the embodiment of the present application provides a device for processing plaintext element information, where, as Figure 3 shown, the device 30 includes:

[0111] An acquisition module 301, configured to acquire data in a storage engine, where the data in the storage engine includes: structured data;

[0112] The first-level filtering module 302 is used to perform first-level column name filtering and second-level column name filtering on the column data in the structured data, and screen out target data columns that meet the first-level column name filtering condition and the second-level column name filtering condition. Among them, the first-level column name filtering is: identifying sensitive information for the column names in the data table structure of the structured data. If there is sensitive information in the column names in the data table structure, it is determined that the data column with the sensitive information meets the first-level column name filtering condition; the second-level column name filtering is to identify sensitive information for the column names that do not belong to the preset column name whitelist according to the preset column name whitelist. If there is sensitive information in the column names that do not belong to the preset column name whitelist, it is determined that the data column with the sensitive information meets the second-level column name filtering condition;

[0113] The second-level filtering module 303 is used to extract the values of the data stored in the target data columns, and sequentially detect the first sensitive information, the second sensitive information, and the third sensitive information. According to whether the values of the data meet two or more filtering conditions among the regular filtering condition, the format rule filtering condition, the verification algorithm filtering condition, the third-level column name filtering condition, and the summary filtering condition, calculate the confidence level of the values of the data. Among them, the first sensitive information, the second sensitive information, and the third sensitive information belong to three categories of sensitive information with different sensitive information priorities. The third-level column name filtering condition is: the column name of the data column to which the first sensitive information belongs contains a first preset character, the column name of the data column to which the second sensitive information belongs contains a second preset character, or the data column to which the third sensitive information belongs contains a third preset character. The summary filtering condition is: for the case where the values of the data stored in the target data columns are a piece of text, determine whether the values of the data in the data column to which the first sensitive information belongs contain the first preset character, determine whether the values of the data in the data column to which the second sensitive information belongs contain the second preset character, and determine whether the values of the data in the data column to which the third sensitive information belongs contain the third preset character;

[0114] The output module 304 is used to determine the data with a confidence level higher than the preset confidence level threshold as the target sensitive data, and integrate the target sensitive data to generate and output the processing result of the plaintext element information, where the processing result of the plaintext element information includes the position description information of the target sensitive data.

[0115] In some possible embodiments, the data in the storage engine further includes: unstructured data, and the second-level filtering module is further used for:

[0116] For the values of each piece of data in the unstructured data, detect the first sensitive information, the second sensitive information, and the third sensitive information in sequence, and calculate the confidence level of the value of each piece of data according to whether the value of each piece of data meets two or more of the regular filtering condition, the format rule filtering condition, the verification algorithm filtering condition, and the digest filtering condition;

[0117] Determine the unstructured data with the confidence level higher than the preset confidence level threshold as the target sensitive data.

[0118] In some possible embodiments, the second-level filtering module is further configured to:

[0119] Calculate the column name confidence level of each piece of data according to whether each piece of data in the structured data meets the third-level column name filtering condition, and calculate the digest confidence level of each piece of data according to whether each piece of data in the structured data or each piece of data in the unstructured data meets the digest filtering condition;

[0120] Input the preprocessed sensitive data with the column name confidence level meeting the preset column name confidence level threshold and / or the digest confidence level meeting the preset digest confidence level threshold into the sensitive information detection model, and the sensitive information detection model outputs the model confidence level of the preprocessed sensitive data based on the input data.

[0121] In some possible embodiments, the device further includes:

[0122] The third-level filtering module is configured to calculate the comprehensive confidence level of each piece of the preprocessed sensitive data according to the column name confidence level, the digest confidence level, and the model confidence level;

[0123] If the comprehensive confidence level is greater than the preset comprehensive confidence level threshold, determine the preprocessed sensitive data greater than the preset comprehensive confidence level threshold as the target sensitive data.

[0124] In some possible embodiments, the device further includes:

[0125] The model training module is configured to perform perturbation processing on the target sensitive data to generate training sample data, input the training sample data into the initial neural network model, and train the initial neural network model until the difference between the recognition result output by the initial neural network model based on the training sample data and the target sensitive data is maximized;

[0126] In some possible embodiments, the performing perturbation processing on the target sensitive data to generate training sample data includes:

[0127] Add noise and data transformation to the target sensitive data, and generate and store the training sample data.

[0128] Among them, in this application, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0129] The names of the messages or information exchanged between multiple devices in the embodiments of this application are only for illustrative purposes and do not limit the scope of these messages or information.

[0130] In a third aspect, an exemplary embodiment of this application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiments of this application.

[0131] An exemplary embodiment of this application further provides a non-transitory computer-readable storage medium storing a computer program, where when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of this application.

[0132] An exemplary embodiment of this application further provides a computer program product, including a computer program, where when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of this application.

[0133] Reference Figure 4 , the structural block diagram of the electronic device 400 that can be used as the server or client of this application will now be described. It is an example of a hardware device that can be applied to various aspects of this application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of this application described and / or claimed herein.

[0134] As Figure 4As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM 402) or a computer program loaded from a storage unit 408 into a random access memory (RAM 403). In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output interface (I / O interface 405) is also connected to the bus 404.

[0135] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device capable of inputting information into the electronic device 400. The input unit 406 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 407 can be any type of device capable of presenting information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include but is not limited to a magnetic disk, an optical disk. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0136] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above. For example, in some embodiments, the method for processing the foregoing plaintext element information can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. In some embodiments, the computing unit 401 can be configured to execute the method for processing the foregoing plaintext element information by any other appropriate means (for example, by means of firmware).

[0137] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine, or entirely on a remote machine or server.

[0138] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0139] As used in the present application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, an optical disc, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0140] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0141] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0142] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The client - server relationship is generated by computer programs that run on respective computers and have a client - server relationship with each other.

Claims

1. A method for processing plaintext element information, characterized in that: The method comprises: Acquire data in a storage engine, wherein the data in the storage engine includes: structured data; Performing first-level column name filtering and second-level column name filtering on the column data in the structured data, and screening out target data columns that meet the first-level column name filtering conditions and the second-level column name filtering conditions, wherein the first-level column name filtering is: performing sensitive information identification on the column names in the data table structure in the structured data, and if sensitive information exists in the column names in the data table structure, then determining that the data column containing the sensitive information meets the first-level column name filtering conditions; the second-level column name filtering is performing sensitive information identification on the column names that do not belong to the preset column name whitelist according to a preset column name whitelist, and if sensitive information exists in the column names that do not belong to the preset column name whitelist, then determining that the data column containing the sensitive information meets the second-level column name filtering conditions; Extract the values ​​of each item of data stored in the target data column, detect the first sensitive information, the second sensitive information and the third sensitive information in turn, and calculate the confidence of the value of each item of data according to whether the value of each item of data satisfies two or more of the following filtering conditions: regular filtering condition, format rule filtering condition, verification algorithm filtering condition, third-level column name filtering condition and summary filtering condition; wherein the first sensitive information, the second sensitive information and the third sensitive information belong to three categories of sensitive information with different priorities of sensitive information, and the third-level column name filtering condition is: the column name of the data column to which the first sensitive information belongs contains the first preset character, the column name of the data column to which the second sensitive information belongs contains the second preset character, or the data column to which the third sensitive information belongs contains the third preset character, and the summary filtering condition is: for the case where the value of each item of data stored in the target data column is a piece of text, determine whether the value of each item of data in the data column to which the first sensitive information belongs contains the first preset character, determine whether the value of each item of data in the data column to which the second sensitive information belongs contains the second preset character, and determine whether the value of each item of data in the data column to which the third sensitive information belongs contains the third preset character; Determine that the data with a confidence level higher than a preset confidence threshold is target sensitive data, and integrate the target sensitive data to generate and output a plaintext element information processing result, wherein the plaintext element information processing result includes location description information of the target sensitive data.

2. The method according to claim 1, characterized in that: The data in the storage engine also includes: unstructured data, and the method further includes: For the values ​​of each item of data in the unstructured data, the first sensitive information, the second sensitive information, and the third sensitive information are detected in turn, and the confidence of the value of each item of data is calculated according to whether the value of each item of data satisfies two or more of the regular filtering condition, the format rule filtering condition, the verification algorithm filtering condition, and the summary filtering condition; The unstructured data whose confidence level is higher than the preset confidence level threshold is determined as the target sensitive data.

3. The method according to claim 1 or 2, characterized in that: The method further comprises: Calculating the column name confidence of each item of the data according to whether each item of the structured data satisfies the third-level column name filtering condition, and calculating the summary confidence of each item of the data according to whether each item of the structured data or each item of the unstructured data satisfies the summary filtering condition; Inputting the pre-processed sensitive data whose column name confidence satisfies a preset column name confidence threshold and / or whose summary confidence satisfies a preset summary confidence threshold into a sensitive information detection model, and the sensitive information detection model outputs a model confidence of the pre-processed sensitive data based on the input data; Calculate the comprehensive confidence of each of the pre-processed sensitive data according to the column name confidence, the summary confidence, and the model confidence; If the comprehensive confidence is greater than a preset comprehensive confidence threshold, the pre-processed sensitive data greater than the preset comprehensive confidence threshold is determined to be the target sensitive data.

4. The method according to claim 3, characterized in that The sensitive information detection model is trained based on the following method: The target sensitive data is disturbed to generate training sample data, the training sample data is input into an initial neural network model, and the initial neural network model is trained until the difference between the recognition result output by the initial neural network model based on the training sample data and the target sensitive data is maximized.

5. The method according to claim 4, characterized in that The perturbation processing is performed on the target sensitive data to generate training sample data, including: Noise and data transformation are added to the target sensitive data to generate and store the training sample data.

6. A device for processing plaintext element information, characterized in that: The device comprises: An acquisition module, used to acquire data in a storage engine, wherein the data in the storage engine includes: structured data; The first-level filtering module is used to perform first-level column name filtering and second-level column name filtering on the column data in the structured data, and screen out target data columns that meet the first-level column name filtering conditions and the second-level column name filtering conditions, wherein the first-level column name filtering is: performing sensitive information identification on the column names in the data table structure in the structured data, if there is sensitive information in the column names in the data table structure, then determining that the data column containing the sensitive information meets the first-level column name filtering conditions; the second-level column name filtering is to perform sensitive information identification on the column names that do not belong to the preset column name whitelist according to the preset column name whitelist, if there is sensitive information in the column names that do not belong to the preset column name whitelist, then determining that the data column containing the sensitive information meets the second-level column name filtering conditions; a second-level filtering module, for extracting the values ​​of each data item stored in the target data column, detecting the first sensitive information, the second sensitive information and the third sensitive information in turn, and calculating the confidence of the value of each data item according to whether the value of each data item satisfies two or more of the following filtering conditions: a regular filtering condition, a format rule filtering condition, a verification algorithm filtering condition, a third-level column name filtering condition and a summary filtering condition; wherein the first sensitive information, the second sensitive information and the third sensitive information belong to three categories of sensitive information with different priorities of sensitive information, and the third-level column name filtering condition is: the column name of the data column to which the first sensitive information belongs contains a first preset character, the column name of the data column to which the second sensitive information belongs contains a second preset character, or the data column to which the third sensitive information belongs contains a third preset character, and the summary filtering condition is: for the case where the value of each data item stored in the target data column is a piece of text, determining whether the value of each data item in the data column to which the first sensitive information belongs contains the first preset character, determining whether the value of each data item in the data column to which the second sensitive information belongs contains the second preset character, and determining whether the value of each data item in the data column to which the third sensitive information belongs contains the third preset character; An output module is used to determine that the data with a confidence level higher than a preset confidence threshold is target sensitive data, and integrate the target sensitive data to generate and output a plaintext element information processing result, wherein the plaintext element information processing result includes location description information of each target sensitive data.

7. The device according to claim 6, characterized in that The data in the storage engine also includes: unstructured data, and the second-level filtering module is also used to: For the values ​​of each item of data in the unstructured data, the first sensitive information, the second sensitive information, and the third sensitive information are detected in turn, and the confidence of the value of each item of data is calculated according to whether the value of each item of data satisfies two or more of the regular filtering condition, the format rule filtering condition, the verification algorithm filtering condition, and the summary filtering condition; The unstructured data whose confidence level is higher than the preset confidence level threshold is determined as the target sensitive data.

8. The device according to claim 6, characterized in that The second-level filtering module is also used for: Calculating the column name confidence of each item of the data according to whether each item of the structured data satisfies the third-level column name filtering condition, and calculating the summary confidence of each item of the data according to whether each item of the structured data or each item of the unstructured data satisfies the summary filtering condition; Inputting the pre-processed sensitive data whose column name confidence satisfies a preset column name confidence threshold and / or whose summary confidence satisfies a preset summary confidence threshold into a sensitive information detection model, and the sensitive information detection model outputs a model confidence of the pre-processed sensitive data based on the input data; The device also includes: A third-level filtering module, used to calculate the comprehensive confidence of each of the pre-processed sensitive data according to the column name confidence, the summary confidence, and the model confidence; If the comprehensive confidence is greater than a preset comprehensive confidence threshold, determining the pre-processed sensitive data greater than the preset comprehensive confidence threshold as the target sensitive data; The device also includes: A model training module, used to perform perturbation processing on the target sensitive data to generate training sample data, input the training sample data into an initial neural network model, and train the initial neural network model until the difference between the recognition result output by the initial neural network model based on the training sample data and the target sensitive data is maximized; The perturbation processing is performed on the target sensitive data to generate training sample data, including: Noise and data transformation are added to the target sensitive data to generate and store the training sample data.

9. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing a program; wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 5.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1-5.