Data desensitization method and device, equipment and storage medium

By using general regular expressions and preset desensitization rules in the data desensitization method, the problems of low data desensitization efficiency and limited scenarios in the prior art are solved, and efficient desensitization processing for various data representation formats is achieved.

CN120030594APending Publication Date: 2025-05-23ZHONGAN ONLINE P&C INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510123807.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When desensitizing data, the prior art has limitations in usage scenarios and low desensitization efficiency, making it difficult to effectively process multiple data representation formats.

Method used

By obtaining the original data, a common regular expression is determined to match multiple data representation formats, sensitive keywords are identified and data desensitization is performed according to preset desensitization rules.

Benefits of technology

It improves the diversity and efficiency of data desensitization scenarios, and can be batched through a common regular expression regardless of the data representation format of the original data, achieving more efficient data desensitization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030594A_ABST
    Figure CN120030594A_ABST
Patent Text Reader

Abstract

The invention provides a data desensitization method and device, equipment and a storage medium, and the method comprises the steps: obtaining at least one piece of original data, the original data being key-value pair data comprising keywords and values; determining a general regular expression, wherein the general preset regular expression is used for matching data in any target data representation format in the at least one data representation format; matching the at least one piece of original data according to the general regular expression, and determining a value corresponding to the to-be-desensitized data including the sensitive keyword in the at least one piece of original data; and performing data desensitization processing on the corresponding value of the to-be-desensitized data according to a preset desensitization rule. The diversity of data desensitization scenes and the data desensitization efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data processing technology, and in particular to a data desensitization method, device, equipment and storage medium. Background Art

[0002] In many industries such as finance and healthcare, sensitive data such as personal privacy and business secrets are required to be desensitized to ensure information security.

[0003] Currently, when desensitizing data, it is usually supported to desensitize one piece of data in the string format of JS object (JavaScript Object Notation, JSON) at a time. Therefore, there are limitations in usage scenarios and low desensitization efficiency. Summary of the invention

[0004] The present application provides a data desensitization method, device, equipment and storage medium, which can improve the diversity of data desensitization scenarios and the efficiency of data desensitization.

[0005] In a first aspect, the present application provides a data desensitization method, which includes: obtaining at least one piece of original data, the original data being key-value pair data including a keyword (key) and a value (value); determining a universal regular expression, the universal preset regular expression being used to match data in any target data representation format in at least one data representation format; matching at least one piece of original data according to the universal regular expression, and determining a value corresponding to the data to be desensitized that includes a sensitive keyword in at least one piece of original data; and performing data desensitization processing on the corresponding value of the data to be desensitized according to a preset desensitization rule.

[0006] In the second aspect, the present application provides a data desensitization device, including: a first acquisition module, used to acquire at least one piece of original data, the original data is key-value pair data including keywords and values; a first determination module, used to determine a general regular expression, the general preset regular expression is used to match data in any target data representation format in at least one data representation format; a matching determination module, used to match at least one piece of original data according to the general regular expression, and determine the value corresponding to the data to be desensitized that includes sensitive keywords in at least one piece of original data; a data desensitization module, used to perform data desensitization processing on the corresponding value of the data to be desensitized according to a preset desensitization rule.

[0007] In a third aspect, the present application provides an electronic device, comprising: a processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory, and executing the method in the first aspect or its various implementations.

[0008] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program enables a computer to execute the method in the first aspect or its various implementations.

[0009] In a fifth aspect, the present application provides a computer program product, comprising computer program instructions, which enable a computer to execute the method in the first aspect or its various implementations.

[0010] In a sixth aspect, the present application provides a computer program, which enables a computer to execute the method in the first aspect or its various implementations.

[0011] In the technical solution of the present application, since the universal preset regular expression can match data in any data representation format of at least one data representation format, no matter at least one piece of original data is based on data in one data representation format of at least one data representation format, or data in multiple data representation formats of at least one data representation format, the electronic device can batch process them through a regular expression, namely, the universal regular expression, thereby not only achieving more efficient data desensitization, but also improving the diversity of data desensitization scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The following is an introduction to the drawings required for describing the embodiments.

[0013] Figure 1 A flowchart of a data desensitization method provided in an embodiment of the present application;

[0014] Figure 2 A schematic diagram of a data desensitization method provided in an embodiment of the present application;

[0015] Figure 3 A schematic diagram of another data desensitization method provided in an embodiment of the present application;

[0016] Figure 4 A schematic diagram of another data desensitization method provided in an embodiment of the present application;

[0017] Figure 5 A schematic diagram of another data desensitization method provided in an embodiment of the present application;

[0018] Figure 6 A schematic diagram of another data desensitization method provided in an embodiment of the present application;

[0019] Figure 7 A schematic diagram of a data desensitization device 700 provided in an embodiment of the present application;

[0020] Figure 8A schematic block diagram of an electronic device 800 provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0022] In addition, the original data, test data, sensitive keyword information, desensitization rules and other data involved in this application have been fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0023] It should be understood that the technical solution of the present application can be applied to the following scenarios, but is not limited to:

[0024] In one embodiment, the technical solution of the present application can be applied to data desensitization scenarios, for example, it can be applied to scenarios where log data, page display data and other data are desensitized, and the present application does not impose any restrictions on this.

[0025] In one embodiment, the solution provided by the present application can be executed by any electronic device with data processing capabilities. For example, the electronic device can be a server, specifically an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. For another example, the electronic device can be a terminal device, specifically a tablet computer, a laptop computer, or a desktop computer. For another example, the electronic device can be a combination of a server and a terminal device, wherein the server and the terminal device in the combination can communicate wirelessly or wired, and the present application does not impose specific restrictions on the electronic device.

[0026] After introducing the application scenarios of the embodiments of the present application, the technical solution of the present application will be described in detail below:

[0027] It should be noted that all the technical solutions of the present application can be combined in any way to form optional embodiments of the present application. For example, one embodiment can refer to the content of another embodiment, and the execution order of each step in an embodiment can be adjusted, which will not be elaborated here.

[0028] Figure 1 This is a flow chart of a data desensitization method provided in an embodiment of the present application. The method can be executed by the above electronic device, but is not limited thereto. Figure 1 As shown, the method may include the following steps:

[0029] S110: Acquire at least one piece of original data, where the original data is a key-value pair data including a keyword and a value;

[0030] S120: Determine a universal regular expression, where the universal preset regular expression is used to match data in any target data representation format in at least one data representation format;

[0031] S130: matching at least one piece of original data according to a general regular expression, and determining a value corresponding to the data to be desensitized that includes a sensitive keyword in the at least one piece of original data;

[0032] S140: Perform data desensitization processing on the corresponding values ​​of the data to be desensitized according to the preset desensitization rules.

[0033] It should be noted that the electronic device may execute S120 first and then S110, or may execute S110 first and then S120, and this application does not impose any limitation on this.

[0034] In one embodiment, the regular expressions in the present application, such as general regular expressions, sub-regular expressions, and verification regular expressions, can be Perl compatible regular expressions (Perl Compatible Regular Expression), which can be implemented in any programming language such as Java, Python, C#, JavaScript, etc.

[0035] Taking the usage in Java scenario as an example, the regular expressions in this application may involve the following modes, but are not limited to: (exp), which means grouping and capturing the text matched by the group; (?:exp), which means grouping but not capturing the text matched by the group; (?=exp), which means zero-width assertion, matching the text until ex p; .*exp, which means greedy mode, which can match as much text as possible; .? exp, which means reluctant mode, which belongs to the minimum matching mode. In addition, in the entire regular match, all captured brackets will be grouped (group) and numbered, starting from 1, and numbered 0 means matching the entire original data. This will be introduced in the subsequent embodiments.

[0036] In one embodiment, the data representation format can be understood as a data display form or a data format. Exemplarily, at least one data representation format can include at least one of the following, but is not limited to: a specific data representation format, a specific data representation format with escape processing, and a data representation format in which the boundary characters (i.e., boundaries) of the corresponding key-value pairs do not conform to a specific format.

[0037] The specific format refers to the format of the boundary characters of the key-value pair, for example, it can be a double quote format, that is, double quotes are used as boundary characters. The boundary characters of the key-value pair corresponding to the data in the specific data representation format conform to the specific format (corresponding to obvious boundaries), for example, the boundary characters of its keywords and values ​​can be double quotes. Similarly, if the boundary characters of the corresponding key-value pair do not conform to the specific format (corresponding to no obvious boundaries), the boundary characters of the corresponding key-value pair can be spaces, can be blanks (that is, no boundaries), but not double quotes.

[0038] In addition, the boundary characters of the keywords and values ​​in the key-value pairs corresponding to the data in the specific data representation format that undergoes escape processing can be double quotes and slashes "\", or can be double quotes and other symbols, and this application does not impose any restrictions on this.

[0039] Specifically, the specific data representation format may be a standard JSON format, the specific data representation format that undergoes escape processing may be an escaped JSON format, and the data representation format whose corresponding key-value pair boundary does not conform to the specific format may be a non-standardized key-value pair format.

[0040] For example, data in standard JSON format can be:

[0041] {"name":"pj","mobileNo":"12222222222","address":"<>?,. / / !@#$%^\"&*()_+","email":"12345678@gm.com"}.

[0042] The data in the escaped JSON format can be:

[0043] {\"name\":\"pj\",\"mobnsileNo\":\"12222222222\",\"addrtivess\":\"<>?,. / / !@#$%^\\\"&*(DTO)_+\",\"email\":\"12345678@q.com\"}".

[0044] The non-standard key-value format can be:

[0045] SensitiveDTO(name=pj,mobileNo=12222222222,address=<>?. / / !@#$%^\"&*()_+,email=12345678@q.com).

[0046] In addition, the data representation format can be an array representation format or a non-array representation format. The above examples of the standard JSON format, the escaped JSON format, and the non-standardized key-value pair format are all in the non-array representation format. The data in the array representation format can be as shown in the following three pieces of data:

[0047] "emailList":["12345678@gm.com","123456789@gm.com"];

[0048] emailList=[12345678@q.com,123456789@gm.com];

[0049] "\"emailList\":[\"12345678@gm.com\",\"123456789@gm.com\"]".

[0050] In one embodiment, at least one piece of original data may be data in at least one of the above data representation formats. For example, at least one piece of original data may be data in JSON format, or a mixture of data in JSON format and data in escaped JSON format, or a mixture of data in JSON format and data in array representation format.

[0051] For example, taking the JSON format as an example, assuming the original data is {"name":"pj"}, the keyword in the original data is: name, the boundary of the keyword is: "", the left boundary of the keyword is: ", the right boundary of the keyword is: ", the value is: pj, the boundary of the value is: "", the left boundary of the value is: ", the right boundary of the value is: ", and the separator used to separate keywords and values ​​is: :.

[0052] It can be understood that since the universal regular expression in the present application supports matching data in any data representation format in at least one data representation format, and the data representation format of the original data can be at least one data representation format, it is not only possible to achieve data desensitization for a variety of data desensitization scenarios, but also to perform batch processing on at least one piece of original data through a regular expression, i.e., a universal regular expression, that is, at least one piece of original data can be matched through the same regular expression, i.e., the universal regular expression, thereby achieving more efficient data desensitization.

[0053] In one embodiment, after obtaining the original data, the electronic device can first determine whether the original data is key-value pair data, that is, data in key-value format. If the original data is not key-value pair data, it can be converted into key-value pair data first.

[0054] It can be understood that in data desensitization scenarios, if the specific semantics of the data or the defined types contained in the data cannot be determined, and desensitization is performed only through the data representation format, there will be many scenarios where desensitization is difficult. For example, taking the desensitization of telephone numbers as an example, telephone numbers in region 1 usually start with the number 1, followed by the operator identifier, a total of 11 digits; when a string is identified that meets this feature, it can be regarded as a telephone number; however, in other regions other than region 1, the characteristics of the telephone number may be different from those in region 1, for example, the number of digits included is not 11, so it is more difficult to identify. For another example, in a piece of text, it is more difficult to determine whether a word is a name, especially in complex scenarios where the name coexists with other words in its context.

[0055] In the technical solution of the present application, data desensitization can be performed based on the key-value pair format, that is, the original data in the present application is key-value pair data, and the key-value pair data has clear keywords and values, and does not need to rely on the data content. Therefore, the corresponding value can be desensitized by identifying the keyword. For example, the corresponding value "pj" can be desensitized by identifying the keyword "name", so that the value that needs to be desensitized in the original data can be accurately and quickly determined, which can reduce the difficulty of desensitization and the adaptability to various scenarios, and improve the flexibility and versatility of the desensitization method.

[0056] In one embodiment, before determining the universal regular expression, the electronic device may first obtain a configuration file, which may include sensitive keyword information and desensitization rules, wherein the sensitive keyword information includes sensitive keywords, and the sensitive keywords are used to locate values ​​that need to be desensitized in the original data, and the values ​​that need to be desensitized are values ​​corresponding to the sensitive keywords in the original data (i.e., key-value pair data).

[0057] Exemplarily, sensitive keywords may include at least one of the following: sensitive keywords preset by the system, sensitive keywords selected or input by the user based on the client; desensitization rules may include at least one of the following: desensitization rules preset by the system, desensitization rules selected or input by the user based on the client.

[0058] For example, the sensitive keyword information in the configuration file may be as follows:

[0059] {"fieldName":"emailList","format":"EMAIL","iterable":"true"}.

[0060] Among them, "fieldName":"emailList" indicates that the sensitive keyword is "emailList"; "format":"EMAIL" indicates the identifier of the corresponding desensitization rule, and the specific content of the desensitization rule can be included in the desensitization rule part of the configuration file; "iterable":"true" indicates that the data representation format corresponding to the sensitive keyword is an array representation format, and similarly, "iterable":"false" indicates that the data representation format corresponding to the sensitive keyword is a non-array representation format.

[0061] The desensitization rules in the configuration file can be as follows:

[0062] RULE_PHONE_NO(6,'*',3,2).

[0063] The parameters represented by "6,'*',3,2" are: paddingStarLength (supplementary number length), paddingStar (supplementary number), beforeIndex (retain the first n digits), and afterIndex (retain the last n digits).

[0064] For example, taking the telephone number 12222222222 as an example, assuming that its corresponding key-value pair data is: {"mobileNo":"12222222222"}, if the sensitive keyword is: mobileNo, then the final desensitized result can be: {"mobileNo":"122******22"}.

[0065] In the above content, users can customize sensitive keywords (i.e. desensitized fields) and desensitization rules (i.e. the specific processing method of desensitization), improve the flexibility and visualization of data desensitization and meet the personalized customization needs of users. Moreover, since each sensitive keyword and its desensitization rules can be configured, fine-grained desensitization of data can also be achieved.

[0066] In one embodiment, the electronic device may first perform an initialization process. Specifically, it may first read a configuration file to obtain sensitive keywords and desensitization rules; then, a general regular expression may be obtained based on the sensitive keywords, and the general expression and desensitization rules may be stored in a cache (for example, they may be stored in a memory to improve data acquisition efficiency); thereafter, the original data may be desensitized based on the regular expression and the desensitization rules.

[0067] Alternatively, the electronic device may also obtain sensitive keywords or desensitization rules through other channels, for example, it may directly obtain sensitive keywords or desensitization rules from the client. This application does not limit the method of obtaining sensitive keywords and desensitization rules.

[0068] Exemplarily, before determining the universal regular expression, the electronic device may first determine the delimiters and boundary characters involved in the original data; then, compare the delimiters and boundary characters with the original data and sensitive keywords; if there is a first designated character consistent with the delimiter or boundary character in the original data or sensitive keywords, the first designated character in the original data or sensitive keywords is adjusted to a second designated character inconsistent with the delimiter or boundary character, so as to avoid the universal regular expression from mistakenly identifying the characters in the original data as boundary characters or delimiters.

[0069] Specifically, the second specified character includes at least one of the following: a number, a character whose frequency of appearance in the specified regular expression is lower than a frequency threshold. The specified regular expression may be any regular expression used in any business scenario.

[0070] It is understandable that the characters involved in matching data in regular expressions are generally not numbers. Therefore, replacing the original data or sensitive keywords with numbers or characters with low frequency of occurrence can further prevent the general regular expression from mistakenly identifying the characters in the original data as boundary characters, separators or even other characters.

[0071] For example, assuming that the keyword in the original data is: name (first name), and the boundary character is: (); then the electronic device can adjust the keyword in the original data to name1first name2; or, assuming that the value in the original data is: di / zhi, and the boundary character is: / ; then the electronic device can adjust the value in the original data to di&zhi.

[0072] The following describes the process of determining a general regular expression based on sensitive keywords:

[0073] In one embodiment, the electronic device may first obtain sensitive keyword information including sensitive keywords; then, determine a general regular expression template corresponding to the general regular expression; finally, according to the sensitive keyword information, fill the sensitive keywords into the general regular expression template to obtain the general regular expression.

[0074] Exemplarily, the types and number of data representation formats that the general regular expression supports matching can be determined based on the types and number of data representation formats involved in the original data, that is, the types and number of data representation formats that the general regular expression supports matching can be consistent with the types and number of data representation formats involved in the original data. In other words, in addition to being compatible with all data representation formats at the same time, the general regular expression can also adjust the compatible data representation format according to the data that needs to be desensitized, that is, the data representation format involved in the original data, thereby ensuring that the length and complexity of the general regular expression are reduced and the matching efficiency is improved on the basis of satisfying the matching function of the original data.

[0075] For example, assuming that at least one piece of original data includes data in standard JSON format and data in escaped JSON format, the data representation formats that the universal regular expression supports matching may be JSON format and escaped JSON format; if at least one piece of original data includes data in standard JSON format, data in escaped JSON format, and data in non-standardized key-value pair format, the data representation formats that the universal regular expression supports matching may be JSON format, escaped JSON format, and non-standardized key-value pair format.

[0076] Exemplarily, the above-mentioned determination of the general regular expression template corresponding to the general regular expression may include: determining, for the target data representation format, a sub-regular expression for matching data in the target data representation format; and combining at least one sub-regular expression to obtain the general regular expression template.

[0077] For example, Figure 2 As shown, the value with obvious boundaries may refer to the value corresponding to the data in JSON format, and the value without obvious boundaries may refer to the value corresponding to the data in non-standardized key-value pair format. The electronic device can combine regular expressions for matching different data representation formats, specifically combining different fields involved in the different regular expressions, such as the left boundary of the keyword, the keyword, the right boundary of the keyword, the connection between the keyword and the value, i.e., the separator, the value, the left boundary of the value, and the right boundary of the value, to obtain a general regular expression that can simultaneously support matching data in different data representation formats. For example, the general regular expression template can be:

[0078] "((?:\\?")|(?:[",'{\[\s(]))("+"key1|key2"+")(\\?"?\s*)([:|=])(((\s*\\?["'])(.*?(?=(?:\\"\s*,|"\s*,| \\"\s*}|"\s*}|'\s*,|'\s*})))|((\s*\\?)(.*?(?=(?:\s*,|\s*}|\s*\)))))((?:(?:\\")|")?'?\s*,?\)?} )".

[0079] Exemplarily, the general regular expression template may be a regular expression carrying historical sensitive keywords, or may be a regular expression with blank spaces at positions corresponding to sensitive keywords, and this application does not impose any restrictions on this.

[0080] For example, a general regular expression template carries historical sensitive keywords, which may be: "((?:\\?"")|(?:[",'{\[\s(]))("+"key1|key2"+")(\\?"?\s*)([:|=])(((\s*\\?["'])(.*?(?=(?:\\"\s*,|"\s*,|\\"\s*}|"\s*}|'\s*,|'\s*}))))|((\s*\\?")(.*?(?=(?:\s*,|\s*}|\s*\))))))((?:(?:\\")|")?'?\s*,?\)?}?)", where key1 and key2 may be historical sensitive keywords. If the sensitive keywords are key4 and key5, the electronic device may use key4 and key5 to replace the above k ey1 and key2, the obtained general regular expression can be: "((?:\\?"")|(?:[",'{\[\s(]))("+"key4|key5"+")(\\?"?\s*)([:|=])(((\s*\\?["'])(.*?(?=(?:\\"\s*,|"\s*,|\\"\s*}|"\s*}|'\s*,|'\s*}))))|((\s*\\?)(.*?(?=(?:\s*,|\s*}|\s*\))))))((?:(?:\\")|")?'?\s*,?\)?}?)"; that is, the above-mentioned filling of sensitive keywords into the general regular expression template can refer to: using sensitive keywords to replace historical sensitive keywords or corresponding fields in the general regular expression template.

[0081] For example, the position corresponding to the sensitive keyword in the general regular expression template is empty, which can be: "((?:\\?"")|(?:[",'{\[\s(]))("+"|"+")(\\?"?\s*)([:|=])(((\s*\\?["'])(.*?(?=(?:\\"\s*,|"\s*,|\\"\s*}|"\s*}|'\s*,|'\s*}))))|((\s*\\?)(.*?(?=(?:\s*,|\s*}|\s*\))))))((?:(?:\\")|")?'?\s*,?\)?}?)", if the sensitive keywords are key4 and key5, the electronic device can directly fill key4 and key5 into the corresponding positions, and obtain The general regular expression can be: "((?:\\?"")|(?:[",'{\[\s(]))("+"key4|key5"+")(\\?"?\s*)([:|=])(((\s*\\?["'])(.*?(?=(?:\\"\s*,|"\s*,|\\"\s*}|"\s*}|'\s*,|'\s*}))))|((\s*\\?)(.*?(?=(?:\s*,|\s*}|\s*\))))))((?:(?:\\")|")?'?\s*,?\)?}?)"; that is, the above-mentioned filling of sensitive keywords into the general regular expression template can refer to: filling the sensitive keywords into the corresponding positions in the general regular expression template.

[0082] Exemplarily, the sensitive keyword information may include a data representation format corresponding to the sensitive keyword; the electronic device may fill the sensitive keyword into a universal regular expression template corresponding to the data representation format corresponding to the sensitive keyword to obtain a universal regular expression.

[0083] It can be understood that there can be multiple general regular expressions (for example, the first general regular expression and the second general regular expression in the subsequent embodiments), and each general regular expression supports different data representation formats (that is, the data representation formats that can be matched); accordingly, there can also be multiple general regular expression templates, and one general regular expression template can correspond to one general regular expression; in addition, different sensitive keywords can be used to match data in different data representation formats (for example, some sensitive keywords are only used to desensitize data in array representation format, and some sensitive keywords are only used to desensitize data in non-array representation format). Therefore, the sensitive keyword information can be used to indicate the data representation format that the sensitive keyword can be used to match, and then the general regular expression template corresponding to the sensitive keyword can be selected according to the sensitive keyword information, and then the sensitive keyword can be filled into the appropriate general regular expression template to obtain the general regular expression corresponding to the data representation format that can match the sensitive keyword, that is, the general regular expression suitable for the original data can be determined, thereby ensuring the desensitization effect.

[0084] In one embodiment, before matching at least one piece of original data, the electronic device may first determine whether the at least one piece of original data includes a sensitive keyword; then, in response to the at least one piece of original data including the sensitive keyword, the electronic device may match the at least one piece of original data according to a general regular expression to determine the value corresponding to the data to be desensitized that includes the sensitive keyword in the at least one piece of original data; or, in response to the at least one piece of original data not including the sensitive keyword, it may be determined that the subsequent matching process cannot match the sensitive keyword and the corresponding value, and therefore, the process of determining the general regular expression and matching the original data may not be executed, and the original data may be directly returned, thereby implementing optimized processing for the desensitization process and avoiding invalid execution of the desensitization step.

[0085] Exemplarily, the electronic device may determine whether at least one piece of original data includes a sensitive keyword according to a regular expression, which may be implemented in any of the following ways:

[0086] In a first approach, the electronic device may first assemble a third verification regular expression for verifying whether at least one piece of original data includes the sensitive keyword based on the sensitive keyword; and then determine whether at least one piece of original data includes the sensitive keyword based on the third verification regular expression.

[0087] For example, assuming that the sensitive keywords are key1, key2, and key3, the third verification regular expression can be:

[0088] "[\",'{\[\s(](key1|key2|key3)(\\?\"?'?\s*)([:|=])".

[0089] Among them, [\",'{\[\s(] is the start match of the keyword. If the characters that appear are: ",'{[(\s, then any subsequent characters can be considered as proof of the start of the keyword. \" and "' and 0 to multiple spaces or blanks can be used as markers for the end of the keyword. According to the matching rules of the regular expression, the connection identifier of the keyword and value, that is, the separator, can be: or =. If the separator is matched, it means that the original data needs to be desensitized.

[0090] It can be understood that the data representation format can be divided into an array representation format and a non-array representation format. For these two data representation formats, the boundary features in the data are different. Therefore, the electronic device can match them separately through different general regular expressions to achieve desensitization, that is, the general regular expression is divided into two types: a first general regular expression for matching data in the array representation format, and a second general regular expression for matching data in the non-array representation format. That is to say, the general regular expression may include at least one of the above two types, namely, the first general regular expression and / or the second general regular expression; therefore, the electronic device can divide the sensitive keywords into two types: the first sensitive keywords corresponding to the data in the array representation format, The second sensitive keyword corresponding to the data in the non-array representation format, the sensitive keyword may include at least one of the above two types, that is, the sensitive keyword includes: the first sensitive keyword and / or the second sensitive keyword; then, different verification regular expressions can be constructed according to the first sensitive keyword and the second sensitive keyword, and the original data can be verified respectively; if the original data involves or includes the first sensitive keyword, the first general regular expression can be used to match it for desensitization, and if the original data involves or includes the second sensitive keyword, the second general regular expression can be used to match it for desensitization, so that on the basis of comprehensive verification of the original data, it can be ensured that the appropriate general regular expression is quickly selected to improve the desensitization efficiency. The specific method is shown in the following method 2.

[0091] Method 2: assemble a first verification regular expression based on the first sensitive keyword, and assemble a second verification regular expression based on the second sensitive keyword; determine whether at least one piece of original data includes the first sensitive keyword based on the first verification regular expression, and determine whether at least one piece of original data includes the second sensitive keyword based on the second verification regular expression.

[0092] Correspondingly, the above-mentioned matching of at least one original data according to a general regular expression to determine the value corresponding to the data to be desensitized that includes sensitive keywords in at least one original data may include: in response to at least one original data including a first sensitive keyword, matching at least one original data according to a first general regular expression to determine the value corresponding to the data to be desensitized that includes the first sensitive keyword in at least one original data; and / or, in response to at least one original data including a second sensitive keyword, matching at least one original data according to a second general regular expression to determine the value corresponding to the data to be desensitized that includes the second sensitive keyword in at least one original data.

[0093] For example, the first sensitive keyword is key3, and the second sensitive keywords are key1 and key2. Then the first verification regular expression can be:

[0094] "[\",'{\[\s(](key3)(\\?\"?'?\s*)([:|=])".

[0095] The second validation regular expression can be:

[0096] "[\",'{\[\s(](key1|key2)(\\?\"?'?\s*)([:|=]).

[0097] In one embodiment, in combination with the above, Figure 3 As shown, the electronic device can first read the configuration file to obtain sensitive keywords (i.e., sensitive keywords); then, combine the sensitive keywords to obtain a sensitive keyword array; assemble the sensitive keyword array to obtain a regular expression for verifying whether the original data contains sensitive keywords. Afterwards, if the original data contains sensitive keywords, the set of non-array type sensitive keywords can be assembled to obtain a general regular expression for locating the data to be desensitized that contains non-array type sensitive keywords in the original data; and / or, the set of array type sensitive keywords can be assembled to obtain a general regular expression for locating the data to be desensitized that contains array type sensitive keywords in the original data.

[0098] The following is an introduction to the process of matching and desensitizing the original data according to the general regular expression, namely S130 and S140:

[0099] In one embodiment, in combination with the above, Figure 4As shown, after determining that there are sensitive keywords in the original data, it is possible to first verify whether there are sensitive keywords of the array type (i.e., array representation format) in the original data; afterwards, different general regular expressions can be used to extract the values of the data to be desensitized of the array type and the values of the data to be desensitized of the non-array type, i.e., sensitive values, in the original data; then, according to the corresponding desensitization rules, the sensitive values can be replaced with the desensitized content (i.e., replacement value or desensitized value) to obtain the desensitized original data. Specifically, in combination with the above embodiments, as Figure 5 shown, taking the data to be desensitized as a telephone number as an example, for the non-array type, the corresponding regular expression (i.e., the regular expression for replacing the sensitive value) can be: ((?:mobileNo)\?\"?\s*)(.*[:|=])(((\s*\?[\"'].{3})(.*?)(.{2})(?=(?:\"\s*,|\"\s*,|\"\s*}|\"\s*}|'\s*,|'\s*})))|((\s*\?.{3})(.*?)(.{2})(?=(?:\s*,|\s*}|\s*\)))))(.*), and the corresponding replacement value, i.e., desensitized value, can be: $1$2$5$9******$7$11$12; for the array type, the corresponding regular expression can be: (.{3})([\s\S]*)(.{2}), and the corresponding replacement value, i.e., desensitized value, can be: $1******$3.

[0100] Among them, the electronic device can replace the value of the data to be desensitized according to the general regular expression to achieve the desensitization of the data to be desensitized; it can also generate a corresponding regular expression according to the desensitization rule to replace the value of the data to be desensitized and other desensitized data to achieve the desensitization of the data to be desensitized.

[0101] Exemplarily, taking the data representation format of the original data as an array representation format as an example, assuming that the first sensitive keyword is key3, the first general regular expression is: ((?:\\?")|(?:[",'{\[\s(]))(key3)(\\?"?\s*)([:|=])("?\[(.*?)(?=],|]}|]\))])([,)}]); wherein, ((?:\\?")|(?:[",'{\[\s(])), indicates matching the part on the left of the keyword; (\\?"?\s*), indicates matching the part on the right of the keyword; "?\[, indicates matching the part on the left of the value; (?=],|]}|]\)), is a zero-width assertion, which is used to truncate the content of the value to the right boundary of the value. Finally, a more detailed desensitization can be achieved by saving the contents of group0 (i.e., the group numbered 0) and group5 (i.e., the group numbered 5). Among them, the content of group6 (i.e., the grouping numbered 6) is the value of the data to be desensitized. By judging whether group6 is empty, it can be determined whether the value needs to be desensitized. The content of group5 includes the left and right boundaries of the value. Using the value of group5 for desensitization can avoid losing the left boundary of the first content and the right boundary of the last content in the original data; the content of group0 is the original key-value pair data of the original data. In addition, group3 (i.e., the grouping numbered 3) can be parsed to obtain the data to be used (i.e., the specific desensitization object: value). Specifically, it can be judged whether the content of group5 has obvious boundaries. The e value starting with [" or [\" has obvious boundaries, otherwise it has no obvious boundaries. For values ​​with obvious boundaries, (,\s|,|\[|\[\s)(\\"|")([\s\S]*?)(\\"|")(?=,|\]) can be used for parsing, otherwise (,\s|,|\[|\[\s)([\s\S]*?)(?=,|\]) can be used for parsing. After parsing the value of the data to be desensitized, the value can be desensitized according to the corresponding desensitization rule. For example, the desensitization rule can be: retain the first and last fields of the value, and replace the unreserved fields with defined characters. Finally, the replaced value can be used to replace the original value based on a non-regular expression method to obtain the desensitized original data.

[0102] Exemplarily, taking the data representation format of the original data as a non-array representation format as an example, assuming that the second sensitive keywords are key1 and key2, the second general regular expression is: "((?:\\?"")|(?:[",'{\[\s(]))("+"key1|key2"+")(\\?"?\s*)([:|=])(((\s*\\?["'])(.*?(?=(?:\\"\s*,|"\s*,|\\"\s*}|"\s*}|'\s*,|'\s*}))))|((\s*\\?)(.*?(?=(?:\s*,|\s*}|\s*\))))))((?:(?:\\")|")?'?\s*,? \)?}? )"; "((?:\\?\")|(?:[\",'{\[\s(]))(" means taking one of \" or ",'{[( and a space to match the start of the keyword; \\?\"?\s* is used to match the end of the keyword; : or = means connecting the keyword and the value; the value can be analyzed according to the boundary or no boundary. For the boundary scenario, \s*\?[\"'] means that the value may start with an existing space, with a single quote or double quote as the boundary, and the boundary may have an escape symbol \ as the start mark of the value; [\s\S]* means matching all fields until the right boundary starting point of the value is determined, and the content corresponding to the line break can be matched; ? means that the right boundary of the value may not actually exist. For example, the value is an empty string. For "(?=(?:\\\"\s*,|\"\s*,|\\\"\s*}|\"\s*}|'\s*,|'\s*}))". (?= represents a zero-width assertion, which can limit the matching range of the previous [\s\S]*. The right boundary of the value can be matched as \"\s*, or "\s* or "\s*} or \"\s*} or '\s*, or '\s*}. For unbounded string parsing, there may not be a string on the left side of the value, but there may be spaces. Like bounded values, the night can match the content of the newline character, and then limit the right boundary of the value through a zero-width assertion. For "(?=(?:\s*,|\s*}|\s*\)))", with \s*, or \s *} or \s*\) indicates the start of the right boundary. REGULAR_RIGHT_START = "(? = (?:\\\"\s*,|\"\s*,|\\\"\s*}|\"\s*}|'\s*,|'\s*}))" and IRREGULAR_RIGHT_START = "(? = (?:\s*,|\s*}|\s*\)))" are only used to indicate the end of the tag value, not to really match the right boundary of the value. Therefore, there is also a right boundary format match: "((?:(?:\\\")|\")?'?\s*,?\)?}?)", this format is compatible with both bounded and unbounded types, so it can match the entire keyword and value content.Among them, the 8th group in the regular match, group8, matches the value in the bounded scenario, and the 11th group matches the value in the unbounded scenario. The match is successful only when one of the two groups has a value. At this time, the keywords and values ​​that need to be desensitized can be located. If there are no special requirements for the desensitization rules, group8 or group11 can be quickly replaced with fixed desensitization values ​​to achieve desensitization for the value. If more detailed desensitization rules are required for desensitization, you can continue with the following steps. You can first get a list of group0, which is the original value of all key-value pairs that need to be desensitized. The corresponding form can be: "key":"value". Assume that the desensitization rule is the one in the embodiment corresponding to the above telephone number: ((?:key1)\\?"?\s*)(.*[:|=])(((\s*\\?["'].{3})(.*?)(.{2})(?=(?:\\"\s*,|"\s*,|\\"\s*}|"\s*}|'\s*,|'\s*})))|((\s*\\?.{3})(.*?)(.{2})(?=(?:\s*,|\s*}|\s*\)))))(.*), desensitizing the value using this rule can obtain: matcher.replaceAll("$1$2$5$9******$7$11$12"), the middle desensitized value * can also be freely replaced with other symbols. Among them, 6, 3, and 26 come from the first, second, and third digits in the desensitization rule respectively.

[0103] Exemplarily, as shown in Table 1, the original data, i.e., the data to be desensitized, and the regular expressions and replacement values ​​required for desensitizing the drink are listed respectively.

[0104] Table 1

[0105]

[0106] In one embodiment, before using the universal regular expression to match the original data, test data can be obtained first, and the test data is key-value pair data including keywords and values; then, the data features of the test data are extracted through a preset recognition model, and the test data is character labeled and segmented according to the data features to obtain a first keyword and a first value corresponding to the test data; then, the test data can be matched according to the universal regular expression to obtain a second keyword and a second value corresponding to the test data; thereafter, the first keyword and the second keyword can be compared, and the first value and the second value can be compared; in response to inconsistency between the first keyword and the second keyword and / or inconsistency between the first value and the second value, the universal regular expression is updated.

[0107] The data features represent separator information, character distribution information, boundary character information of separators used to separate keywords and values ​​in the test data, and format features of a data representation format corresponding to the test data.

[0108] For example, at least one of the following items in the general regular expression can be updated according to data features and / or character annotation and word segmentation results: separators, boundary characters, and matching patterns included in the general regular expression to achieve automatic correction. Alternatively, it can also be updated manually.

[0109] For example, the first delimiter involved in the test data can be determined based on the delimiter information in the data features or the delimiters in the character annotation and word segmentation results; then, check whether the second delimiter contained in the general regular expression is consistent with the first delimiter; if inconsistent (it can be explained that the delimiter description in the general regular expression is incorrect, for example, some delimiters are omitted), the first delimiter is used to adjust the second delimiter in the general regular expression.

[0110] Similarly, the first boundary character involved in the test data can be determined based on the boundary character information in the data features or the boundary characters in the character annotation and word segmentation results; then, check whether the second boundary character contained in the general regular expression is consistent with the first boundary character; if inconsistent (which can be explained that the boundary character description in the general regular expression is incorrect, for example, some boundary characters are omitted), the first boundary character is used to adjust the second boundary character in the general regular expression.

[0111] In addition, the matching mode (or matching strategy) in the general regular expression may be adjusted according to the character distribution information and format characteristics in the data characteristics, for example, the non-greedy matching strategy may be adjusted to a greedy matching strategy.

[0112] It is understandable that in some original data, the boundary between keywords and values ​​is relatively obvious. For example, for the JSON string: {"name":"pj"}, the boundary between keywords and values ​​is: "", so the value of name can be accurately located as pj through a general regular expression; while in some original data, its value may be the same as the character or separator corresponding to the boundary, so there will be problems with positioning or replacement errors. Therefore, before using the general regular expression, you can verify the general regular expression according to the correct positioning results of keywords and values ​​to improve the accuracy of desensitization.

[0113] Exemplarily, after the general regular expression is updated, the above-mentioned process of extracting data features of the test data through the preset recognition model can be repeatedly executed until the similarity between the first keyword and the second keyword is greater than the first similarity threshold and the similarity between the first value and the second value is greater than the second similarity threshold.

[0114] Exemplarily, training data involving various types of separators and boundary characters (referring to characters corresponding to the boundaries between keywords and values) can be obtained, and the training data is key-value pair data including keywords and values; then, the training data is input into a recognition model, and the training data features of the training data (used to characterize separator information, character distribution information, boundary character information of separators in the training data, and format features of the data representation format corresponding to the training data) are extracted through the recognition model, and character annotation and word segmentation are performed on the training data according to the training data features to obtain training keywords and training values ​​corresponding to the training data; thereafter, the recognition model can be trained by comparing the training keywords and training values ​​with the actual keywords and actual values ​​corresponding to the training data.

[0115] In the process of training or using the recognition model, when character annotation and word segmentation are performed on the training data, the priority of the separator can be combined. For example, assuming that the separators include: comma, space, quotation mark, equal sign, the priority corresponding to the comma can be set to the highest, so that the comma can be segmented as a separator first.

[0116] The above-mentioned character tagging and word segmentation can be specifically implemented based on the word segmentation and sequence tagging algorithm (for example, CRF, BiLSTM-CRF) in natural language processing (Natural Language Processing, NLP), but is not limited thereto.

[0117] Exemplarily, the recognition model may include: a backbone network, a neck network and a head network; wherein the backbone network is used to extract data features of the test data; the neck network is used to perform character annotation on the test data according to the data features; and the head network is used to perform word segmentation on the test data according to the character annotation results to obtain the word segmentation results, namely the first keyword and the first value.

[0118] In one embodiment, in combination with the above, Figure 6As shown, the electronic device can first perform boundary recognition on the test data according to the recognition model, that is, the above-mentioned feature extraction, annotation and segmentation, to obtain the annotation and segmentation results of the test data, that is, the first keyword and the first value; according to the matching result of the test data of the general regular expression, the second keyword and the second value are obtained; then, through the corresponding comparison of the above-mentioned first keyword and the first value and the second keyword and the second value, the verification of the general regular expression is realized; thereafter, the desensitization service can be started, and the sensitive keywords and desensitization rules can be obtained according to the configuration file (including system configuration data and / or user configuration data); the verification regular expression is assembled according to the desensitization keyword to determine whether the original data contains the sensitive keyword; when it is determined that it exists, the general regular expression can be assembled according to the sensitive keyword (similar to the above-mentioned assembly of the general regular expression according to the general regular expression template and the sensitive keyword to obtain the general regular expression), so as to determine that the desensitization service is successfully started, and then, the subsequent desensitization process can be carried out.

[0119] Figure 7 A schematic diagram of a data desensitization device 700 provided in an embodiment of the present application.

[0120] like Figure 7 As shown, the data desensitizing device 700 includes: a first acquisition module 701, a first determination module 702, a matching determination module 703, a data desensitizing module 704, a first judgment module 705, a second acquisition module 706, a model processing module 707, a test matching module 708, a result comparison module 709, a data update module 710, a second determination module 711, a character comparison module 712, and a character adjustment module 713.

[0121] Exemplarily, the first acquisition module 701 is used to acquire at least one piece of original data, where the original data is key-value pair data including keywords and values; the first determination module 702 is used to determine a universal regular expression, where the universal preset regular expression is used to match data in any target data representation format in at least one data representation format; the matching determination module 703 is used to match at least one piece of original data according to the universal regular expression, and determine the value corresponding to the data to be desensitized that includes sensitive keywords in at least one piece of original data; the data desensitization module 704 is used to perform data desensitization processing on the corresponding value of the data to be desensitized according to preset desensitization rules.

[0122] Exemplarily, the first determination module 702 is specifically used to obtain sensitive keyword information including sensitive keywords; determine a general regular expression template corresponding to the general regular expression; and fill the sensitive keywords into the general regular expression template according to the sensitive keyword information to obtain the general regular expression.

[0123] Exemplarily, the first determination module 702 is specifically configured to determine, with respect to the target data representation format, a sub-regular expression for matching data in the target data representation format; and combine at least one sub-regular expression to obtain a general regular expression template.

[0124] Exemplarily, the sensitive keywords include at least one of the following: system preset sensitive keywords, sensitive keywords selected or input by the user based on the client; the desensitization rules include at least one of the following: system preset desensitization rules, desensitization rules selected or input by the user based on the client.

[0125] Exemplarily, the first judgment module 705 is used to determine whether at least one piece of original data includes sensitive keywords; the matching determination module 703 is specifically used to match at least one piece of original data according to a general regular expression in response to at least one piece of original data including sensitive keywords, and determine the value corresponding to the data to be desensitized that includes the sensitive keywords in the at least one piece of original data.

[0126] Exemplarily, the sensitive keyword includes: a first sensitive keyword corresponding to the data in the array representation format and / or a second sensitive keyword corresponding to the data in the non-array representation format; the general regular expression includes: a first general regular expression for matching the data in the array representation format and / or a second general regular expression for matching the data in the non-array representation format; a first judgment module 705 is specifically used to assemble a first verification regular expression according to the first sensitive keyword, and to assemble a second verification regular expression according to the second sensitive keyword; to judge whether at least one piece of original data includes the first sensitive keyword according to the first verification regular expression, and to judge whether at least one piece of original data includes the second sensitive keyword according to the second verification regular expression; a matching determination module 703 is specifically used to match the at least one piece of original data according to the first general regular expression in response to at least one piece of original data including the first sensitive keyword, and determine the value corresponding to the data to be desensitized that includes the first sensitive keyword in at least one piece of original data; and / or, in response to at least one piece of original data including the second sensitive keyword, to match the at least one piece of original data according to the second general regular expression, and determine the value corresponding to the data to be desensitized that includes the second sensitive keyword in at least one piece of original data.

[0127] Exemplarily, the second acquisition module 706 is used to acquire test data, which is key-value pair data including keywords and values; the model processing module 707 is used to extract data features of the test data through a preset recognition model, and perform character annotation and word segmentation on the test data according to the data features to obtain a first keyword and a first value corresponding to the test data; the test matching module 708 is used to match the test data according to a universal regular expression to obtain a second keyword and a second value corresponding to the test data; the result comparison module 709 is used to compare the first keyword and the second keyword, and to compare the first value and the second value; the data update module 710 is used to update the universal regular expression in response to the inconsistency between the first keyword and the second keyword and / or the inconsistency between the first value and the second value; the data features of the test data are continuously extracted through the preset recognition model until the similarity between the first keyword and the second keyword is greater than the first similarity threshold and the similarity between the first value and the second value is greater than the second similarity threshold; wherein the data features characterize the delimiter information, character distribution information, boundary character information of the delimiter used to separate keywords and values ​​in the test data, and the format features of the data representation format corresponding to the test data.

[0128] Exemplarily, the data updating module 710 is specifically used to update at least one of the following items in the general regular expression according to data features and / or character annotation and word segmentation results: a separator, a boundary character, and a matching pattern included in the general regular expression.

[0129] Exemplarily, the second determination module 711 is used to determine the delimiter and boundary characters contained in the at least one original data; the character comparison module 712 is used to compare the delimiter and the boundary characters with the original data and the sensitive keyword; the character adjustment module 713 is used to adjust the first specified character in the original data or the sensitive keyword to a second specified character inconsistent with the delimiter or the boundary character if there is a first specified character consistent with the delimiter or the boundary character in the original data or the sensitive keyword.

[0130] Exemplarily, the second specified character includes at least one of the following: a number, and a character whose appearance frequency in the specified regular expression is lower than a frequency threshold.

[0131] It should be understood that the embodiments of the device are similar to the embodiments of the above method, and the contents and effects thereof can refer to the contents and effects of the above method, which will not be described in detail in this application. Figure 7 The device 700 shown can execute the above method embodiments, and the above and other operations and / or functions of each module in the device 700 are respectively for implementing the corresponding processes in the above methods, which will not be repeated here for the sake of brevity.

[0132] The above describes the device 700 of the embodiment of the present application from the perspective of the functional module in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or a combination of hardware and software modules in the decoding processor to perform. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in conjunction with its hardware.

[0133] Figure 8 A schematic block diagram of an electronic device 800 provided in an embodiment of the present application.

[0134] like Figure 8 As shown, the electronic device 800 may include:

[0135] The memory 810 and the processor 820, the memory 810 is used to store the computer program and transmit the program code to the processor 820. In other words, the processor 820 can call and run the computer program from the memory 810 to implement the method in the embodiment of the present application.

[0136] For example, the processor 820 may be configured to execute the above method embodiments according to instructions in the computer program.

[0137] In some embodiments of the present application, the processor 820 may include but is not limited to:

[0138] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0139] In some embodiments of the present application, the memory 810 includes but is not limited to:

[0140] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SL DRAM) and direct memory bus random access memory (DR RAM).

[0141] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 810 and executed by the processor 820 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0142] like Figure 8 As shown, the electronic device may also include:

[0143] The transceiver 830 may be connected to the processor 820 or the memory 810 .

[0144] The processor 820 may control the transceiver 830 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices. The transceiver 830 may include a transmitter and a receiver. The transceiver 830 may further include an antenna, and the number of antennas may be one or more.

[0145] It should be understood that the various components in the electronic device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0146] The present application also provides a computer storage medium on which a computer program is stored, and when the computer program is executed by a computer, the computer can perform the method of the above method embodiment. In other words, the present application embodiment also provides a computer program product containing instructions, and when the instructions are executed by a computer, the computer can perform the method of the above method embodiment.

[0147] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instruction is loaded and executed on a computer, the computer can be made to perform the corresponding flow in each method in the embodiment of the present application in whole or in part, and generate the functions that can be realized by each method in the embodiment of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instruction can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instruction can be transmitted from a website site, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, a data center, etc. that contains one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0148] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0149] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the system, device or module can be electrical, mechanical or other forms.

[0150] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. For example, each functional module in each embodiment of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

Claims

1. A data desensitization method, characterized in that: include: Acquire at least one piece of original data, where the original data is key-value pair data including a keyword and a value; Determine a universal regular expression, where the universal preset regular expression is used to match data in any target data representation format in at least one data representation format; Matching the at least one piece of original data according to the general regular expression, and determining a value corresponding to the data to be desensitized that includes the sensitive keyword in the at least one piece of original data; The data to be desensitized is desensitized according to the preset desensitization rules.

2. The method according to claim 1, characterized in that: Determining the general regular expression includes: Acquire sensitive keyword information including the sensitive keyword; Determine a general regular expression template corresponding to the general regular expression; According to the sensitive keyword information, the sensitive keyword is filled into the general regular expression template to obtain the general regular expression.

3. The method according to claim 2, characterized in that The determining of the general regular expression template corresponding to the general regular expression comprises: For the target data representation format, determining a sub-regular expression for matching data in the target data representation format; At least one of the sub-regular expressions is combined to obtain the general regular expression template.

4. The method according to claim 2, characterized in that: The sensitive keywords include at least one of the following: sensitive keywords preset by the system, sensitive keywords selected or input by the user based on the client; The desensitization rule includes at least one of the following: a desensitization rule preset by the system, a desensitization rule selected or input by the user based on the client.

5. The method according to any one of claims 1 to 4, characterized in that: Before matching the at least one piece of original data according to the universal regular expression to determine the value corresponding to the data to be desensitized that includes the sensitive keyword in the at least one piece of original data, the method further includes: Determining whether the at least one piece of original data includes the sensitive keyword; The matching the at least one piece of original data according to the general regular expression to determine a value corresponding to the data to be desensitized that includes the sensitive keyword in the at least one piece of original data includes: In response to the at least one piece of original data including the sensitive keyword, the at least one piece of original data is matched according to the universal regular expression to determine a value corresponding to the data to be desensitized that includes the sensitive keyword in the at least one piece of original data.

6. The method according to claim 5, characterized in that The sensitive keywords include: a first sensitive keyword corresponding to the data in the array representation format and / or a second sensitive keyword corresponding to the data in the non-array representation format; The general regular expression includes: a first general regular expression for matching data in an array representation format and / or a second general regular expression for matching data in a non-array representation format; The determining whether the at least one piece of original data includes the sensitive keyword comprises: Assembling a first verification regular expression according to the first sensitive keyword, and assembling a second verification regular expression according to the second sensitive keyword; Determining whether the at least one piece of original data includes the first sensitive keyword according to the first verification regular expression, and determining whether the at least one piece of original data includes the second sensitive keyword according to the second verification regular expression; The matching the at least one piece of original data according to the general regular expression to determine a value corresponding to the data to be desensitized that includes the sensitive keyword in the at least one piece of original data includes: In response to the at least one piece of original data including the first sensitive keyword, matching the at least one piece of original data according to the first general regular expression to determine a value corresponding to the data to be desensitized that includes the first sensitive keyword in the at least one piece of original data; and / or, In response to the at least one piece of original data including the second sensitive keyword, the at least one piece of original data is matched according to the second general regular expression to determine a value corresponding to the data to be desensitized that includes the second sensitive keyword in the at least one piece of original data.

7. The method according to any one of claims 1 to 4, characterized in that: Before matching the at least one piece of original data according to the universal regular expression to determine the value corresponding to the data to be desensitized that includes the sensitive keyword in the at least one piece of original data, the method further includes: Acquire test data, wherein the test data is key-value pair data including a keyword and a value; Extracting data features of the test data through a preset recognition model, and performing character annotation and word segmentation on the test data according to the data features to obtain a first keyword and a first value corresponding to the test data; Match the test data according to the general regular expression to obtain a second keyword and a second value corresponding to the test data; comparing the first keyword with the second keyword, and comparing the first value with the second value; In response to the first keyword and the second keyword being inconsistent and / or the first value and the second value being inconsistent, updating the universal regular expression; Continue to extract the data features of the test data by using the preset recognition model until the similarity between the first keyword and the second keyword is greater than a first similarity threshold and the similarity between the first value and the second value is greater than a second similarity threshold; The data feature is used to characterize at least one of the following: separator information, character distribution information, boundary character information of a separator used to separate keywords and values ​​in the test data, and format features of a data representation format corresponding to the test data.

8. The method according to claim 7, characterized in that The updating of the universal regular expression comprises: At least one of the following items in the general regular expression is updated according to the data feature and / or the character annotation and word segmentation result: a separator, a boundary character, and a matching pattern included in the general regular expression.

9. The method according to any one of claims 1 to 4, characterized in that: Before determining the universal regular expression, the method further includes: Determine the separator and boundary characters included in the at least one piece of original data; Compare the separator and the boundary character with the original data and the sensitive keyword; If there is a first designated character in the original data or the sensitive keyword that is consistent with the separator or the boundary character, the first designated character in the original data or the sensitive keyword is adjusted to a second designated character that is inconsistent with the separator or the boundary character.

10. The method according to claim 9, characterized in that The second designated character includes at least one of the following: a number, and a character whose appearance frequency in the designated regular expression is lower than a frequency threshold.

11. A data desensitization device, characterized in that: include: A first acquisition module, used to acquire at least one piece of original data, wherein the original data is a key-value pair data including a keyword and a value; A first determination module is used to determine a universal regular expression, where the universal preset regular expression is used to match data in any target data representation format in at least one data representation format; A matching determination module, used to match the at least one piece of original data according to the general regular expression, and determine a value corresponding to the data to be desensitized that includes sensitive keywords in the at least one piece of original data; The data desensitization module is used to perform data desensitization processing on the corresponding values ​​of the data to be desensitized according to preset desensitization rules.

12. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1-10 by executing the executable instructions.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

14. A computer program product comprising instructions, characterized in that When the computer program product is executed on an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 10.