Data Desensitization and Restoration Method and Device

By matching the selected character set during the data desensitization process and performing format retention encrypted FPE processing, the problem of inconsistent format after data desensitization in the prior art is solved, and the ability of maintenance personnel to efficiently diagnose data problems is realized.

CN114491624BActive Publication Date: 2025-05-27DINGTALK (CHINA) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111667179.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-05-27
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The prior art cannot effectively ensure that maintenance personnel can efficiently diagnose data problems during data desensitization. Traditional encryption methods cause desensitization data to be inconsistent with the original data format and cannot ensure length consistency.

Method used

By matching the alternative character set with the characters to be desensitized in the data to be desensitized, the target character set is determined, and the format retained encrypted FPE processing is performed according to the target character set, desensitized data that is consistent with the content length, format, and character type are consistent.

Benefits of technology

Without direct contact with the data to be desensitized, confusing content consistent with the content length, format and character type of the data to be desensitized is provided, effectively improving the efficiency of maintenance personnel in processing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114491624B_ABST
    Figure CN114491624B_ABST
Patent Text Reader

Abstract

This specification provides a data desensitization and restoration method and apparatus. The data desensitization method includes: matching an alternative character set with the characters to be desensitized in the data to be desensitized, and determining the character set with successful matching as the target character set, where each alternative character set corresponds to one or more types of characters; performing format-preserving encryption (FPE) processing on the characters to be desensitized according to the target character set to obtain the desensitized data corresponding to the data to be desensitized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of information security, and in particular to a data desensitization and restoration method and device. Background Art

[0002] With the rapid development of the internet and the advent of the big data era, the information security of every citizen is becoming increasingly important. How to carefully design information privacy protection methods to achieve a balance between information freedom and privacy protection has gradually become a common challenge faced by major companies.

[0003] In related technologies, maintenance personnel typically encrypt documents containing sensitive user information to desensitize them before accessing them to troubleshoot data issues. However, the desensitized data obtained through traditional encryption methods is completely different from the original data, making it impossible for maintenance personnel to diagnose data issues. Even format-preserving encryption only ensures that the length of the data before and after encryption is consistent, which still cannot guarantee that maintenance personnel can effectively identify data issues. Summary of the Invention

[0004] In view of this, this specification provides a data desensitization and restoration method and device to address the deficiencies in the related art.

[0005] Specifically, this specification is implemented through the following technical solutions:

[0006] According to a first aspect of an embodiment of this specification, a data desensitization method is provided, the method comprising:

[0007] Matching the candidate character sets with the characters to be desensitized in the data to be desensitized, and determining the character sets that successfully match as the target character sets, wherein each candidate character set corresponds to one or more types of characters;

[0008] The characters to be desensitized are subjected to format-preserving encryption (FPE) processing according to the target character set to obtain desensitized data corresponding to the data to be desensitized.

[0009] According to a second aspect of the embodiments of this specification, a data restoration method is provided, the method comprising:

[0010] Matching candidate character sets with characters to be restored in the data to be restored, and determining a character set with successful matching as a target character set, wherein each candidate character set corresponds to one or more types of characters;

[0011] Performing format-preserving encryption (FPE) inverse processing on the characters to be restored according to the target character set to obtain original data corresponding to the data to be restored.

[0012] According to a third aspect of the embodiments of this specification, a data desensitization device is provided, the method comprising:

[0013] a to-be-demensitized character matching unit, configured to match candidate character sets with the to-be-demensitized characters in the to-be-demensitized data, and determine the character set that successfully matches as a target character set, wherein each candidate character set corresponds to one or more types of characters;

[0014] The character encryption unit performs format-preserving encryption (FPE) processing on the characters to be desensitized according to the target character set to obtain desensitized data corresponding to the data to be desensitized.

[0015] According to a fourth aspect of the embodiments of this specification, a data restoration device is provided, characterized in that the device includes:

[0016] a to-be-restored character matching unit, configured to match candidate character sets with the to-be-restored characters in the to-be-restored data, and determine the character set that successfully matches as a target character set, wherein each candidate character set corresponds to one or more types of characters;

[0017] The character restoration unit is used to perform format-preserving encryption (FPE) inverse processing on the characters to be restored according to the target character set to obtain original data corresponding to the data to be restored.

[0018] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in the first and second aspects are implemented.

[0019] According to the sixth aspect of the embodiments of this specification, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method described in the first and second aspects are implemented.

[0020] In the technical solution provided in this specification, target character sets that match the characters to be desensitized in the data to be desensitized are obtained from the alternative character sets, and then format-preserving encryption FPE processing is performed on the characters to be desensitized according to the target character sets to obtain desensitized data corresponding to the data to be desensitized. This allows maintenance personnel to provide obfuscated content with the same content length, format, and character type as the data to be desensitized as desensitized data without directly contacting the data to be desensitized, thereby effectively improving the efficiency of maintenance personnel in processing data.

[0021] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a flowchart of a data desensitization method shown in an exemplary embodiment of this specification;

[0023] Figure 2 is a schematic diagram of a group FPE solution shown in an exemplary embodiment of this specification;

[0024] Figure 3 This is a flowchart of a data restoration method shown in an exemplary embodiment of this specification;

[0025] Figure 4 is a schematic structural diagram of an electronic device shown in an exemplary embodiment of this specification;

[0026] Figure 5 This is a schematic structural diagram of a data desensitization device shown in an exemplary embodiment of this specification;

[0027] Figure 6 It is a structural diagram of a data restoration device shown in an exemplary embodiment of this specification. DETAILED DESCRIPTION

[0028] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this specification. Rather, they are merely examples of apparatus and methods consistent with certain aspects of this specification, as detailed in the appended claims.

[0029] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. As used in this specification and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0030] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information without departing from the scope of this specification. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0031] Figure 1This is a flow chart of a data desensitization method shown in an exemplary embodiment of this specification. Figure 1 As shown, the method may include the following steps:

[0032] S101, matching candidate character sets with characters to be desensitized in the data to be desensitized, and determining a successfully matched character set as a target character set, wherein each candidate character set corresponds to one or more types of characters.

[0033] Desensitization refers to the process of modifying sensitive information using specific rules to reliably protect private data. Desensitization can be used to modify real data for testing purposes, provided it involves customer security data or commercially sensitive data without violating system rules.

[0034] The data to be desensitized can be represented as a string consisting of one or more types of characters that needs to be desensitized, and each character in the string serves as the character to be desensitized. The candidate character set is one or more predefined character sets, wherein each candidate character set contains one or more types of characters, and each type of character is unique in the candidate character set.

[0035] When the data to be desensitized contains a large number of character types that are not predefined in the candidate character set, the candidate character set can be set to perform targeted processing on the undefined character types.

[0036] In one embodiment, the alternative character set may include a known character set and an extended character set. The known character set includes one or more predefined types of characters, and the extended character set corresponds to unpredefined characters. Regardless of whether the target character set is a known character set or an extended character set, the characters to be desensitized are desensitized according to the target character set. Since the characters to be desensitized that participate in the desensitization process in the extended character set are all in the same confusing character space, the character types of the desensitized data obtained below are confused, which is not conducive to maintenance personnel to carry out related maintenance work.

[0037] For example, consider a candidate character set consisting of the English character set A1, the extended character set B, and the data X to be desensitized, containing the content "tom@abc.com." The English character set A1, as a known character set, includes the characters "a"-"z" and "A"-"Z," while the extended character set includes characters not found in known character sets, such as "@" and ".". After desensitization, the data to be desensitized, X, may produce desensitized data Y similar to "asd#zxc?wFk" or "qwe$qns+Oir." This means that characters like "@" and ".", which do not contain sensitive information but indicate the data to be desensitized as an "email address," are desensitized along with other characters. This prevents maintenance personnel from knowing that the data to be desensitized was originally an email address, making it difficult to perform targeted maintenance work.

[0038] In another embodiment, the above-mentioned extended character set is replaced by a reserved character set, that is, the alternative character set may include a known character set and a reserved character set. The known character set includes one or more predefined types of characters, and the reserved character set corresponds to unpredefined characters. When the target character set is a known character set, the characters to be desensitized are desensitized according to the target character set; when the target character set is a reserved character set, the characters to be desensitized are not desensitized. This allows maintenance personnel to actively select the character content that needs to be desensitized while retaining the complete format (content) of the content that is not of concern.

[0039] For example, consider a candidate character set consisting of the English character set A1, the reserved character set C, and the unencrypted data X containing "tom@abc.com." The English character set A1, as a known character set, includes characters "a"-"z" and "A"-"Z," while the reserved character set includes characters not found in known character sets, such as "@" and ".". After desensitization, unencrypted data X might produce desensitized data Y similar to "asd@zxc.wFk" or "qwe@qns.Oir." Based on features like "@" and "." in desensitized data Y, maintenance personnel can easily infer that unencrypted data X was originally an email address, allowing them to conduct further maintenance work.

[0040] The characters to be desensitized in the desensitized data are traversed, and a character set matching the character type of the currently traversed character to be desensitized is searched in the candidate character set to determine the target character set for the character to be desensitized. This allows different types of characters to be desensitized to correspond to different target character sets.

[0041] For example: there is an alternative character set containing English character set A1 and numeric character set A2, and the content of the data X to be desensitized is "12abc3". The English character set A1 contains characters "a"-"z", "A"-"Z", and the numeric character set A2 contains characters "0"-"9". If the alternative character set needs to be matched with the data to be desensitized X, the characters "1", "2", and "3" match the numeric character set A2, and the characters "a", "b", and "c" match the English character set A1. In other words, the English character set A1 is the target character set for the characters "a", "b", and "c" to be desensitized in the data to be desensitized X, and the numeric character set A2 is the target character set for the characters "1", "2", and "3" to be desensitized in the data to be desensitized X.

[0042] Those skilled in the art can imagine that the above-mentioned data to be desensitized can be obtained not only directly through traditional text format, but also through text content extracted from storage files such as Word documents, JSON (JavaScript Object Notation), XML files, database data, etc., and this manual does not limit this.

[0043] S102, performing format-preserving encryption (FPE) processing on the characters to be desensitized according to the target character set to obtain desensitized data corresponding to the data to be desensitized.

[0044] Format Preserving Encryption (FPE) is an encryption method that ensures that ciphertext and plaintext have the same format and length. Unlike traditional encryption algorithms, which often expand data, changing its length and / or type, FPE ensures that the encrypted data maintains a certain degree of length and type consistency with the original data before encryption. This technology allows maintenance personnel to effectively perform operations such as reviewing complete document styles and conducting document format compatibility research during routine maintenance without accessing sensitive user information.

[0045] Perform FPE on the desensitized characters corresponding to different target character sets to obtain encrypted desensitized characters. For example, if the target character set corresponding to the desensitized characters "1", "2", and "3" is a numeric character set, then performing FPE encryption on the desensitized characters based on the numeric character set will produce desensitized characters similar to "5", "2", and "7". Similarly, performing FPE encryption on the desensitized characters "a", "b", and "c" based on the English character set will produce desensitized characters similar to "G", "x", and "Q".

[0046] When faced with data to be desensitized that contains multiple character types, we can first reorganize the desensitized characters corresponding to the same target character set based on the idea of ​​"grouping", encrypt the desensitized characters according to the target character set, and then restore the reorganized desensitized characters to the desensitized data in the original order corresponding to the data to be desensitized.

[0047] In one embodiment, position information of the characters to be desensitized in the data to be desensitized is recorded, the characters to be desensitized corresponding to the same target character set in the data to be desensitized are reorganized, and FPE processing is performed on each group of reorganized characters to be desensitized to obtain corresponding groups of desensitized characters, and the positions of the desensitized characters are restored according to the position information corresponding to the corresponding characters to be desensitized to obtain the desensitized data.

[0048] For example: there is an alternative character set containing English character set A1 and digital character set A2, and the content of a data to be desensitized X is "12abc3". The English character set A1 contains characters "a"-"z", "A"-"Z", and the digital character set A2 contains characters "0"-"9". Record the position information of each character to be desensitized in the data to be desensitized X, that is, the position information of character "1" is 1, the position information of character "2" is 2, the position information of character "a" is 3... and so on. Reorganize the characters corresponding to the English character set A1 and the digital character set A2 to obtain "123" and "abc". Among them, the position information of the above "123" can be expressed as "126", and the position information of the above "abc" can be expressed as "345". Assume that the reorganized characters to be desensitized are processed by FPE to obtain desensitized characters similar to "591" and "Rpg". Since the position information of the two is fixed and does not change with the FPE processing, the desensitized characters "591" and "Rpg" can be restored according to the position information to obtain "59Rpg1" as the desensitized data Y, and the desensitized data Y is consistent with the data to be desensitized X with the content of "12abc3" in both form and length.

[0049] In addition to being represented by increasing numbers, the above-mentioned position information can also be represented based on any ordered symbols or identifiers. In addition, it can also be stored in the form of an array, a linked list, etc., which is not limited in this specification.

[0050] Traditional FPE solutions usually follow the conversion idea, that is, converting the FPE problem on the complex domain to the integer domain, and then using some proposed integer FPE algorithm to solve it.

[0051] In one embodiment, for the data to be desensitized that simultaneously involves character sets such as numbers, English letters, Chinese characters, and emoticons, an integer domain FPE embodiment is uniformly constructed. And the FPE processing is performed on the data to be desensitized through this FPE instance. This will result in the inability to independently scramble the elements of each character set during FPE processing, and thus the original format of the data to be desensitized may be lost. For example: A data to be desensitized X with the content of "123 and abc" can obtain a desensitized data Y similar to "Qiaowan D Pei 1A2" after FPE processing based on the same FPE embodiment, that is, it cannot be guaranteed that the character types of the characters to be desensitized and the corresponding desensitized characters are the same.

[0052] In another embodiment, for the data to be desensitized that simultaneously involves character sets such as numbers, English letters, Chinese characters, and emoticons, integer domain FPE embodiments corresponding to different character sets are respectively constructed. That is, the FPE instance corresponding to the target character set is determined, and the FPE processing is performed on the data to be desensitized through the determined FPE instance. Among them, each alternative character set corresponds to a different FPE instance. This enables the elements of each character set to be independently scrambled and retains the original format of the data to be desensitized to the greatest extent. For example: A data to be desensitized X with the content of "123 and abc" can obtain a desensitized data Y similar to "369 of bkb" after FPE processing based on three FPE embodiments respectively constructed according to English letters, numbers, and Chinese characters. Among them, "123" in the data to be desensitized X corresponds to "369" after FPE processing, "and" corresponds to "of", and "abc" corresponds to "bkb", that is, it is guaranteed that the character types of the characters to be desensitized and the corresponding desensitized characters are the same.

[0053] In the specific process of FPE processing, each character to be desensitized actually undergoes multiple transformations, that is, the character to be desensitized is first mapped to a value in a preset transformation domain, and then the corresponding desensitized character is obtained through a series of transformations in the preset transformation domain.

[0054] In one embodiment, according to the values corresponding to each character in the target character set in the preset transformation domain respectively, the character to be desensitized is converted to the first value in the preset transformation domain, and FPE processing is performed on the first value to transform the first value into a second value, and the character to be desensitized is mapped to the character in the target character set corresponding to the second value.

[0055] It should be noted that there are two mainstream implementations of FPE technology, FF1 and FF3, based on the specific encryption algorithm implementation given in the draft standard SP800-38G[7] released by the National Institute of Standards and Technology (NIST) for FPE. The main processes of these algorithms are similar, and their core is the Feistel network. The Feistel network is a symmetric structure used in block ciphers. Each encryption step is called a round, and the encryption process is a cycle of several rounds. Moreover, FF1 and FF3 are both linear changes implemented based on the 128-bit Advanced Encryption Standard (AES) algorithm. However, FF2 was abandoned when it was designed because it could not meet the expected 128-bit security strength. The difference between FF1 and FF3 is that FF1 undergoes 10 rounds of iteration, while FF3 undergoes 8 rounds of iteration. FF3's performance is slightly better than FF1, but FF1 has higher security strength.

[0056] The following example uses the group FPE solution to implement data desensitization. Figure 2 The technical solutions involved in the above embodiments in this application are elaborated in detail. Figure 2 FIG. 1 is a schematic diagram of a group FPE solution shown in an exemplary embodiment. Figure 2 As shown, the scheme may include the following steps:

[0057] 1. Construct an alternative character set and the corresponding FPE instance.

[0058] Depend on Figure 2 As shown, candidate character set 221 is constructed from a known character set consisting of numbers (number), English letters (enLetter), 3500 commonly used Chinese characters (usual3500CH), Japanese characters (Japanese), and Korean characters (Korean), as well as a reserved character set (null) containing characters outside of these character sets. The positions of these candidate character sets are sequentially numbered as "0," "1," "2," "3," "4," and "-1." A unique integer domain FPE instance is constructed for each candidate character set and stored in a mapping table between character sets and FPE instances.

[0059] 2. Traverse the data to be desensitized, determine the target character set corresponding to each character to be desensitized in the data to be desensitized, and record the target character set in the form of a position array.

[0060] Suppose there is a data to be desensitized: "1.hello world!2.你好啊,世界!3.こんにちは、世界!4. " as Figure 2 the plaintext 210 in Figure 2 is desensitized. By traversing the plaintext 210, it can be seen that the character set of this data to be desensitized, that is, the plaintext character set 211 in

[0061] simultaneously involves numbers, English letters, 3500 common Chinese characters, Japanese characters and Korean characters. Through the partitioning operation 220, each group of string groups are obtained: ["1234", "こんにちは", "helloworld",

[0062] "你好啊世界世界"], etc. In addition, characters such as "!" and "、" that are not matched by the pre-defined character set will be uniformly classified into the reserved character set "STOP WORD".

[0063] 3. Recombine the characters to be desensitized in the data to be desensitized that correspond to the same target character set respectively.

[0064] Obtain the original integer mapping of each group of strings on the corresponding target character set: [1, 2, 3, 4], [18, 82, 42, 32, 46], [7, 4, 11, 11, 14, 22, 14, 17, 11, 3], [6472, 1365, 5432, 5313], [156, 779, 592, 13, 2207, 13, 2207].

[0065] 4. Each group of characters to be desensitized is subjected to FPE processing in its respective transform domain.

[0066] See Figure 2 For the FPE encryption operation 230 in the integer domain in, each group of characters to be desensitized performs an integer domain FPE transformation on the above-mentioned original integer mapping in its respective transform domain, obtaining the transformed target integer mapping: [4, 9, 5, 6], [101, 104, 53, 177, 131], [16, 0, 20, 10, 16, 21, 6, 17, 21, 10], [5311, 10965, 10328, 6555], [1765, 1261, 996, 922, 2042, 3149, 1601].

[0067] 5. Map the result obtained after FPE processing back to the characters to be desensitized in each target character set, obtaining a group of desensitized strings.

[0068] See Figure 2 For the mapping operation 240 in, map the above-mentioned target integer mapping obtained after the FPE encryption operation 230 in the integer domain back to the characters to be desensitized in each target character set again, obtaining the following group of desensitized strings: ["4956", "オキぶヶト", "qaukqvgrvk", "残扯布尧烁豹杏"]. It can be seen that the above group of desensitized strings is completely different from the content of each group of strings before FPE processing.

[0069] 6. According to the position information corresponding to the corresponding characters to be desensitized, perform position restoration on the group of desensitized strings to obtain desensitized data.

[0070] Based on the group of desensitized strings obtained from the above mapping operation 240 and the target position information array 242, traverse the data to be desensitized, iterate the group of desensitized strings corresponding to each target character set into the data to be desensitized. If the corresponding character to be desensitized belongs to the reserved character set, directly add it to the data to be desensitized, and splice to obtain the desensitized data: "4.qaukqvgrvk! 9.残扯布,尧烁! 5.オキぶヶト、豹杏! 6. " as Figure 2 the ciphertext 250 in. Among them, the ciphertext character set 251 is consistent with the plaintext character set 211, that is, each desensitized character in the ciphertext 250 still corresponds to the character type of each character to be desensitized in the plaintext 210.

[0071] In the process of transforming the first value into the second value, the correspondence between the first value and the second value can be changed according to different actual usage scenarios and user needs. For example, in an application scenario where the specific content of the document is not important but only the format is important, the correspondence between the first value and the second value can be random and uncertain, thereby ensuring the security of the desensitized data. For another example, in scenarios such as data diagnosis where attention needs to be paid to the specific content of the document, which involves the process of appending partial content to the overall content and content comparison, the correspondence between the first value and the second value can be fixed and unique, so as to more effectively locate character-level data problems.

[0072] In another embodiment, a seed corresponding to the data to be desensitized is generated based on the master key maintained by a preset key maintenance object and the subkey corresponding to the document where the data to be desensitized is located, and a fixed mapping relationship for the data to be desensitized is determined based on the seed, and FPE processing is performed on the first value, and the first value is transformed into the second value according to the fixed mapping relationship.

[0073] Taking the group FPE solution to achieve data desensitization as an example, combined with Figure 2 The technical solutions involved in the above embodiments in this application are elaborated in detail. Figure 2 FIG. 1 is a schematic diagram of a group FPE solution shown in an exemplary embodiment. Figure 2 As shown, the scheme may include the following steps:

[0074] 1. Construct an alternative character set and the corresponding FPE instance.

[0075] 2. Traverse the data to be desensitized, determine the target character set corresponding to each character to be desensitized in the data to be desensitized, and record the target character set in the form of a position array.

[0076] 3. Reorganize the characters to be desensitized that correspond to the same target character set in the data to be desensitized.

[0077] The above three steps are basically the same as above, and will not be described in detail in this manual.

[0078] 4. Each group of characters to be desensitized is encrypted with a specific mapping in its respective transformation domain.

[0079] Unlike the previous embodiment, which randomly determines the mapping relationship between the first value and the second value based on the Feistel network, this embodiment uses the master key and subkey to replace the original key 231 to construct a certain integer domain for mapping. That is, it traverses the characters to be desensitized, confirms the target character set corresponding to the characters to be desensitized and the position of the information appearing in the target character set, and when the target character set appears for the first time, uses the master key and subkey to construct a random mapping of a certain integer domain. For example:

[0080] Digital integer domain: [1,6,3,4,9,2,7,5,0,8];

[0081] English alphabet integer domain: [1,6,23,11,9,17,12,5,0,15,7,24,20,19,18,14,25,2,13,8,16,4,3,22,21,10];

[0082] Chinese integer domain: [3552,3543,3476,1890,3469,728,2348,392,2930,1887,+3,741more];

[0083] Japanese integer domain: [26,173,23,144,94,49,175,129,153,190,+193more];

[0084] Korean integer domain: [10227,10959,4965,5733,3469,10994,2348,10292,2930,5946,+11,162more].

[0085] Those skilled in the art will appreciate that due to space constraints, some of the integer fields above describe the total number of integers by using the word "more".

[0086] Combine Figure 2 In the integer domain FPE encryption operation 230, each plaintext character is assigned a random mapping coordinate based on its coordinate in the character set space. For example, the number 1, whose coordinate in the numeric character set is 1, corresponds to the first bit in the numeric integer domain, i.e., the number 6. Similarly, the number 2 is mapped to the number 3, and the number 0 is mapped to the number 1.

[0087] 5. Map the results obtained after FPE processing back to the characters to be desensitized in each target character set to obtain a desensitized character string group.

[0088] 6. According to the position information corresponding to the corresponding characters to be desensitized, the position of the desensitized character string group is restored to obtain desensitized data.

[0089] The above two steps are basically the same as above, and will not be described in detail in this manual.

[0090] The aforementioned preset key maintenance object may be a key management center in a specific online server or a specific local device, which is not limited in this specification.

[0091] The above-mentioned master key can be a string set manually or generated by a specific random generation algorithm, and the number and length of the master key can be changed accordingly according to the actual application scenario and the performance of the device of the above-mentioned key maintenance object. This manual does not limit this.

[0092] The above-mentioned subkey can be in the form of a number or a random string as a unique identifier of the document containing the data to be desensitized, and stored in the above-mentioned key maintenance object or other specific storage medium, or it can be specific information generated based on the inherent attributes of the document containing the data to be desensitized. This specification does not limit this.

[0093] In some usage scenarios, there may be a need to perform FPE processing on multiple character sets such as numeric characters and English characters in the same FPE embodiment.

[0094] In one embodiment, an obfuscation granularity configuration instruction can be received, and the number of character types corresponding to the candidate character set can be set according to the obfuscation granularity configuration instruction. This allows maintenance personnel to achieve autonomous and controllable obfuscation granularity according to their own needs, achieving high scalability of the technical solution.

[0095] For example: there is a data X to be desensitized with the content of "123 and abc". A maintenance personnel initiates two obfuscation granularity configuration instructions Z1 and Z2 respectively, wherein Z1 sets the alternative character set to contain only one character type, and Z2 sets the alternative character set to contain two character types. Specifically, Z1 sets three alternative character sets to contain Chinese, numeric and English alphabetic character sets respectively, while Z2 sets two alternative character sets to contain an independent character set containing Chinese character set and a mixed character set containing both numeric character set and English character set respectively. Then, after FPE processing by the obfuscation granularity configuration instruction Z2, desensitized data Y1 similar to "584 plus uTq", "753 Hao jkL" can be obtained, while after FPE processing by the obfuscation granularity configuration instruction Z2, desensitized data Y2 similar to "5qW plus 16q", "qwe Hao 91L" can be obtained. Obviously, compared with the data to be desensitized X, the desensitized data Y2 obtained by the obfuscation granularity configuration instruction Z2 may be transformed into desensitized characters with character types of English letters or numbers after desensitization processing.

[0096] The above-mentioned operation of receiving the obfuscation granularity configuration instruction can be implemented by customizing the configuration interface in a specific SDK (Software Development Kit), and this specification does not limit this.

[0097] Through the above embodiments, it can be seen that the method of data desensitization in this specification, by introducing the technical means of the reserved character set, makes it possible to maintain the uniformity of the character type of the desensitized characters when desensitizing data to be desensitized that simultaneously involves multiple character sets, thereby avoiding the influence of the character type not predefined in the reserved character set on the overall format of the desensitized data. At the same time, unlike the traditional FPE process that simply judges the type of content to be obfuscated and selects a specified character set for desensitization, this specification utilizes the grouped FPE scheme to desensitize the characters to be desensitized corresponding to the target character set, so that the format and length of the desensitized data can be kept consistent at the character level, and maintenance personnel can obtain a better operating experience when viewing the complete document style, document format compatibility research, etc. On the other hand, through the combination of the master key and the subkey, the data to be desensitized and each character in the desensitized data have a definite correspondence, which solves the two major problems in the traditional FPE technology: the inability to process the data to be desensitized containing only a single character and the context in the data to be desensitized affecting the content of the desensitized data, and provides a more suitable operating environment for maintenance personnel to process the specific content of the document and data diagnosis. In addition, the introduction of obfuscation granularity diversifies the content of alternative character sets, allowing technicians to freely adjust the display effect of desensitized data.

[0098] Corresponding to the aforementioned embodiments of the data desensitization method, this specification also provides an embodiment of data restoration.

[0099] Figure 3 This is a flow chart of a data restoration method shown in an exemplary embodiment of this specification. Figure 3 As shown, the method may include the following steps:

[0100] S301, matching candidate character sets with characters to be restored in the data to be restored, and determining a successfully matched character set as a target character set, wherein each candidate character set corresponds to one or more types of characters;

[0101] The so-called "restoration" refers to the reverse operation of desensitized data. The data to be restored can be represented as a string consisting of one or more types of characters that need to be restored, and each character in the string is used as the character to be restored. The alternative character set is one or more predefined character sets, where each alternative character set contains one or more types of characters, and each type of character is unique in the alternative character set.

[0102] When the data to be restored contains a large number of character types that are not predefined in the candidate character set, the candidate character set can be set to perform targeted processing on the unpredefined character types.

[0103] In one embodiment, the alternative character set may include a known character set and an extended character set. The known character set includes one or more predefined types of characters, and the extended character set corresponds to unpredefined characters. Regardless of whether the target character set is a known character set or an extended character set, the characters to be restored are restored according to the target character set. Since the characters to be restored in the extended character set that participate in the restoration process are all in the same obfuscated character space, the character types of the restored data obtained below are confused, which is not conducive to maintenance personnel to carry out related maintenance work. The specific process is similar to the above-mentioned data desensitization method, and this manual will not go into details about it.

[0104] In another embodiment, the extended character set is replaced by a reserved character set, that is, the alternative character set may include a known character set and a reserved character set. The known character set includes one or more predefined types of characters, and the reserved character set corresponds to unpredefined characters. When the target character set is a known character set, the characters to be restored are restored according to the target character set; when the target character set is a reserved character set, the characters to be restored are not restored. The specific process is similar to the above-mentioned data desensitization method, and this manual will not elaborate on it one by one.

[0105] The characters to be restored in the data to be restored are traversed, and a character set matching the character type of the currently traversed character to be restored is searched in the candidate character set to determine the target character set for the character to be restored. This allows different types of characters to be restored to correspond to different target character sets. The specific process is similar to the data desensitization method described above and will not be described in detail in this manual.

[0106] Those skilled in the art can imagine that, in addition to being directly obtained through traditional text formats, the above-mentioned data to be restored can also be text content extracted from storage files such as Word documents, JSON (JavaScript Object Notation), XML files, and database data, and this manual does not limit this.

[0107] S302 , performing format preserving encryption (FPE) inverse processing on the characters to be restored according to the target character set to obtain restored data corresponding to the data to be restored.

[0108] The characters to be restored corresponding to different target character sets are respectively subjected to FPE processing to obtain the encrypted restored characters. The specific process is similar to the above-mentioned data desensitization method, and this manual will not elaborate on it one by one.

[0109] When faced with data to be restored that contains multiple character types, we can first reorganize the characters to be restored that correspond to the same target character set based on the idea of ​​"grouping", encrypt the characters to be restored according to the target character set, and then restore the reorganized characters to the restored data in the original order corresponding to the data to be restored.

[0110] In one embodiment, the position information of the characters to be restored in the data to be restored is recorded, the characters to be restored in the data to be restored that correspond to the same target character set are reorganized, and each group of reorganized characters to be restored is subjected to FPE processing to obtain corresponding groups of restored characters. The restored characters are then positionally restored based on the position information corresponding to the corresponding characters to be restored to obtain the restored data. The specific process is similar to the above-mentioned data desensitization method and will not be described in detail in this specification.

[0111] In addition to being represented by increasing numbers, the above-mentioned position information can also be represented based on any ordered symbols or identifiers. In addition, it can also be stored in the form of an array, a linked list, etc., which is not limited in this specification.

[0112] In one embodiment, for data to be restored that includes character sets such as numbers, English letters, Chinese characters, and emoticons, a unified integer domain FPE instance is constructed. This FPE instance is then used to perform FPE processing on the data to be restored. This results in the inability to independently obfuscate elements of each character set during FPE processing, potentially losing the original format of the data to be restored. The specific process is similar to the aforementioned data desensitization method and will not be detailed in this specification.

[0113] In another embodiment, integer domain FPE implementations corresponding to different character sets are constructed for data to be restored that simultaneously includes numbers, English letters, Chinese characters, emoticons, and other character sets. Specifically, an FPE instance corresponding to the target character set is determined, and the data to be restored is subjected to FPE processing using the determined FPE instance. Each candidate character set corresponds to a different FPE instance. This allows the elements of each character set to be independently obfuscated, preserving the original format of the data to be restored to the greatest extent possible. The specific process is similar to the aforementioned data desensitization method, and will not be further elaborated in this specification.

[0114] In the specific process of FPE processing, each character to be restored actually undergoes multiple transformations, that is, the character to be restored is first mapped to a value in a preset transformation domain, and then the corresponding restored character is obtained through a series of transformations in the preset transformation domain.

[0115] In one embodiment, the character to be restored is converted to a first value in the preset transformation domain according to the corresponding values ​​of each character in the target character set in the preset transformation domain, the first value is subjected to FPE processing to transform the first value into a second value, and the character to be restored is mapped to the character corresponding to the second value in the target character set.

[0116] The following example uses the group FPE solution to achieve data restoration. Figure 2 The technical solutions involved in the above embodiments in this application are elaborated in detail. Figure 2 FIG. 1 is a schematic diagram of a group FPE solution shown in an exemplary embodiment. Figure 2 As shown, the scheme may include the following steps:

[0117] 1. Construct an alternative character set and the corresponding FPE instance.

[0118] 2. Traverse the data to be restored, determine the target character set corresponding to each character to be restored in the data to be restored, and record the target character set in the form of a position array.

[0119] 3. Reorganize the characters to be restored that correspond to the same target character set in the data to be restored.

[0120] 4. Each group of characters to be restored is processed by FPE in its own transform domain.

[0121] 5. Map the results obtained after FPE processing back to the characters to be restored in each target character set to obtain a restored string group.

[0122] 6. Perform position recovery on the restored character string group according to the position information corresponding to the corresponding characters to be restored to obtain restored data.

[0123] During the specific execution of the above steps, except that this embodiment restores the data to be restored as ciphertext 210 to obtain the restored data 250, the rest of the contents are similar to the above data desensitization method, and this manual will not elaborate on them one by one.

[0124] During the process of converting the first value to the second value, the correspondence between the first value and the second value can be changed based on different actual usage scenarios and user needs. For example, in an application scenario where the specific content of the document is not important but only the format is important, the correspondence between the first value and the second value can be random and uncertain, thereby ensuring the security of the restored data. For another example, in scenarios such as data diagnosis where attention needs to be paid to the specific content of the document, involving the process of appending partial content to the overall content and content comparison, the correspondence between the first value and the second value can be fixed and unique, so as to more effectively locate character-level data problems.

[0125] In another embodiment, a seed corresponding to the data to be restored is generated based on the master key maintained by a preset key maintenance object and the subkey corresponding to the document where the data to be restored is located, and a fixed mapping relationship for the data to be restored is determined based on the seed, and the first value is FPE processed, and the first value is transformed into the second value according to the fixed mapping relationship.

[0126] Taking the data restoration using the group FPE solution as an example, combined with Figure 2 The technical solutions involved in the above embodiments in this application are elaborated in detail. Figure 2 FIG. 1 is a schematic diagram of a group FPE solution shown in an exemplary embodiment. Figure 2 As shown, the scheme may include the following steps:

[0127] 1. Construct an alternative character set and the corresponding FPE instance.

[0128] 2. Traverse the data to be restored, determine the target character set corresponding to each character to be restored in the data to be restored, and record the target character set in the form of a position array.

[0129] 3. Reorganize the characters to be restored that correspond to the same target character set in the data to be restored.

[0130] 4. Each group of characters to be restored is encrypted using a specific mapping in its respective transformation domain.

[0131] 5. Map the results obtained after FPE processing back to the characters to be restored in each target character set to obtain a restored string group.

[0132] 6. Perform position recovery on the restored character string group according to the position information corresponding to the corresponding characters to be restored to obtain restored data.

[0133] During the specific execution of the above steps, except that this embodiment restores the data to be restored as ciphertext 210 to obtain the restored data 250, the rest of the contents are similar to the above data desensitization method, and this manual will not elaborate on them one by one.

[0134] The aforementioned preset key maintenance object may be a key management center in a specific online server or a specific local device, which is not limited in this specification.

[0135] The above-mentioned master key can be a string set manually or generated by a specific random generation algorithm, and the number and length of the master key can be changed accordingly according to the actual application scenario and the performance of the device of the above-mentioned key maintenance object. This manual does not limit this.

[0136] The above-mentioned subkey can be in the form of a number or a random string as a unique identifier of the document containing the data to be restored, and stored in the above-mentioned key maintenance object or other specific storage medium, or it can be specific information generated based on the inherent properties of the document containing the data to be restored. This manual does not limit this.

[0137] As can be seen from the above embodiments, the data restoration method disclosed herein, through the introduction of a reserved character set technique, maintains the uniformity of the restored characters' types when restoring data to be restored that simultaneously involves multiple character sets, thus preventing the impact of undefined character types within the reserved character set on the overall format of the restored data. Furthermore, unlike traditional FPE processing, which simply determines the type of obfuscated content and selects a specific character set for restoration, this disclosure utilizes a grouped FPE scheme to restore the characters to be restored that correspond to the target character set, ensuring that the format and length of the data to be restored remain consistent at the character level.

[0138] Figure 4 This is a schematic structural diagram of an electronic device in an exemplary embodiment. Figure 4 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course may also include hardware required for other functions. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a data desensitization device and a data restoration device at the logical level. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0139] Corresponding to the aforementioned embodiments of the data desensitization method, this specification also provides embodiments of a data desensitization device.

[0140] Please refer to Figure 5 , Figure 5 FIG. 1 is a schematic diagram of the structure of a data desensitization device shown in an exemplary embodiment. Figure 5 As shown, in a software implementation, the data desensitization device may include:

[0141] The to-be-demensitized character matching unit 501 is configured to match candidate character sets with the to-be-demensitized characters in the to-be-demensitized data, and determine the character set that successfully matches as the target character set, wherein each candidate character set corresponds to one or more types of characters;

[0142] The character encryption unit 502 is used to perform format-preserving encryption (FPE) processing on the characters to be desensitized according to the target character set to obtain desensitized data corresponding to the data to be desensitized.

[0143] Optionally, the candidate character set includes a known character set and a reserved character set, the known character set includes one or more predefined types of characters, and the character encryption unit 502 further includes:

[0144] A reserved character set processing unit 503 is configured to perform FPE processing on the to-be-desensitized characters according to the target character set when the target character set is a known character set;

[0145] In the case that the target character set is a reserved character set, no desensitization processing is performed on the characters to be desensitized.

[0146] A position information processing unit 504 is used to record the position information of the character to be desensitized in the data to be desensitized;

[0147] Recombining the characters to be desensitized corresponding to the same target character set in the data to be desensitized, and performing FPE processing on each group of characters to be desensitized formed by the recombination to obtain corresponding groups of desensitized characters;

[0148] According to the position information corresponding to the corresponding character to be desensitized, the position of the desensitized character is restored to obtain the desensitized data.

[0149] Optionally, the character encryption unit 502 is specifically used to: determine the FPE instance corresponding to the target character set, and perform FPE processing on the data to be desensitized through the determined FPE instance; wherein each alternative character set corresponds to a different FPE instance.

[0150] Optionally, the character encryption unit 502 is specifically configured to: convert the character to be desensitized into a first value in the preset transformation domain according to the corresponding values ​​of each character in the target character set in the preset transformation domain;

[0151] performing FPE processing on the first value to transform the first value into a second value;

[0152] Map the character to be desensitized to a character in the target character set corresponding to the second value.

[0153] Optionally, the character encryption unit 502 further includes:

[0154] A key management unit 505 is configured to generate a seed corresponding to the data to be desensitized based on a master key maintained by a preset key maintenance object and a subkey corresponding to the document containing the data to be desensitized, and determine a fixed mapping relationship for the data to be desensitized based on the seed;

[0155] Perform FPE processing on the first value, and transform the first value into the second value according to the fixed mapping relationship.

[0156] The obfuscation granularity configuration unit 506 is configured to receive an obfuscation granularity configuration instruction;

[0157] The number of character types corresponding to the candidate character set is set according to the obfuscation granularity configuration instruction.

[0158] Corresponding to the aforementioned embodiment of the data restoration method, this specification also provides an embodiment of a data restoration device.

[0159] Please refer to Figure 6 , Figure 6 FIG. 1 is a schematic diagram of the structure of a data restoration device shown in an exemplary embodiment. Figure 6 As shown, in a software implementation, the data restoration device may include:

[0160] The character matching unit 601 is configured to match candidate character sets with characters to be restored in the data to be restored, and determine the character set that successfully matches as a target character set, wherein each candidate character set corresponds to one or more types of characters;

[0161] The character decryption unit 602 is configured to perform format preserving encryption (FPE) inverse processing on the characters to be restored according to the target character set, so as to obtain restored data corresponding to the data to be restored.

[0162] Optionally, the candidate character set includes a known character set and a reserved character set, the known character set includes one or more predefined types of characters, and the character encryption unit 602 further includes:

[0163] A reserved character set processing unit 603 is configured to perform FPE processing on the to-be-restored character according to the target character set when the target character set is a known character set;

[0164] In the case that the target character set is a reserved character set, no restoration process is performed on the characters to be restored.

[0165] A position information processing unit 604 is configured to record position information of the character to be restored in the data to be restored;

[0166] Reorganizing the characters to be restored corresponding to the same target character set in the data to be restored, and performing FPE processing on each group of characters to be restored formed by the reorganization to obtain corresponding groups of restored characters;

[0167] The restored characters are positionally restored according to the position information corresponding to the corresponding characters to be restored, so as to obtain the restored data.

[0168] Optionally, the character encryption unit 602 is specifically configured to: determine an FPE instance corresponding to the target character set, and perform FPE processing on the data to be restored through the determined FPE instance; wherein each candidate character set corresponds to a different FPE instance.

[0169] Optionally, the character encryption unit 602 is specifically configured to: convert the character to be restored into a first value in the preset transformation domain according to the corresponding values ​​of each character in the target character set in the preset transformation domain;

[0170] performing FPE processing on the first value to transform the first value into a second value;

[0171] Map the character to be restored to a character in the target character set corresponding to the second value.

[0172] Optionally, the character encryption unit 602 further includes:

[0173] A key management unit 605 is configured to generate a seed corresponding to the data to be restored based on a master key maintained by a preset key maintenance object and a subkey corresponding to the document containing the data to be restored, and to determine a fixed mapping relationship for the data to be restored based on the seed;

[0174] Perform FPE processing on the first value, and transform the first value into the second value according to the fixed mapping relationship.

[0175] The number of character types corresponding to the candidate character set is set according to the obfuscation granularity configuration instruction.

[0176] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0177] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0178] Embodiments of the subject matter and functional operations described in this specification may be implemented in the following: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier to be executed by a data processing device or to control the operation of the data processing device. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information and transmit it to a suitable receiver device for execution by the data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0179] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform the corresponding functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0180] Computers suitable for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or the computer will be operably coupled to such mass storage devices to receive data from them or to transmit data to them, or both. However, a computer does not necessarily have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.

[0181] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0182] Although this specification includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of specific embodiments of specific inventions. Certain features described in multiple embodiments within this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may work in certain combinations as described above and even initially claimed as such, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of the sub-combination.

[0183] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.

[0184] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order shown or sequential sequence to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.

[0185] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.

Claims

1. A data desensitization method, characterized in that, the method includes: matching an alternative character set with the characters to be desensitized in the data to be desensitized, and determining the character set with successful matching as the target character set, where each alternative character set corresponds to one or more types of characters, and each alternative character set corresponds to a different format-preserving encryption FPE instance respectively; performing format-preserving encryption FPE processing on the characters to be desensitized according to the target character set to obtain the desensitized data corresponding to the data to be desensitized; the method further includes: recording the position information of the characters to be desensitized in the data to be desensitized; recombining the characters to be desensitized corresponding to the same target character set in the data to be desensitized respectively, and performing FPE processing on each group of recombined characters to be desensitized respectively to obtain the corresponding groups of desensitized characters; restoring the positions of the desensitized characters according to the position information corresponding to the characters to be desensitized to obtain the desensitized data; the performing format-preserving encryption FPE processing on the characters to be desensitized according to the target character set includes: determining the FPE instance corresponding to the target character set, and performing FPE processing on the characters to be desensitized through the determined FPE instance.

2. The method according to claim 1, characterized in that, the alternative character set includes a known character set and a reserved character set, the known character set includes one or more types of predefined characters, and the reserved character set corresponds to undefined characters; the performing format-preserving encryption FPE processing on the characters to be desensitized according to the target character set includes: when the target character set is a known character set, performing FPE processing on the characters to be desensitized according to the target character set; when the target character set is a reserved character set, not performing desensitization processing on the characters to be desensitized.

3. The method according to claim 1, characterized in that, the performing format-preserving encryption FPE processing on the characters to be desensitized according to the target character set includes: converting the characters to be desensitized to the first value in the preset transform domain according to the values corresponding to the respective characters in the target character set in the preset transform domain; performing FPE processing on the first value to transform the first value into a second value; mapping the characters to be desensitized to the characters in the target character set corresponding to the second value.

4. The method according to claim 3, characterized in that, further includes: generating a seed corresponding to the data to be desensitized according to the main key maintained by the preset key maintenance object and the sub key corresponding to the document where the data to be desensitized is located, and determining a fixed mapping relationship for the data to be desensitized based on the seed; performing FPE processing on the first value, and transforming the first value into the second value according to the fixed mapping relationship.

5. The method according to claim 1, characterized in that, further includes: receiving a confusion granularity configuration instruction; setting the number of character types corresponding to the alternative character set according to the confusion granularity configuration instruction.

6. A data restoration method, characterized in that, The method includes: Matching the alternative character sets with the characters to be restored in the data to be restored, and determining the character set with successful matching as the target character set, where each alternative character set corresponds to one or more types of characters, and each alternative character set corresponds to a different format-preserving encryption FPE instance; Performing format-preserving encryption FPE inverse processing on the characters to be restored according to the target character set to obtain the original data corresponding to the data to be restored; The method further includes: Recording the position information of the characters to be restored in the data to be restored; Recombining the characters to be restored corresponding to the same target character set in the data to be restored respectively, and performing FPE processing on each group of recombined characters to be restored respectively to obtain the corresponding groups of restored characters; Performing position restoration on the restored characters according to the position information corresponding to the corresponding characters to be restored to obtain the restored data; The performing format-preserving encryption FPE inverse processing on the characters to be restored according to the target character set includes: Determining the FPE instance corresponding to the target character set, and performing FPE processing on the characters to be restored through the determined FPE instance.

7. A data desensitization device, Characterized in that, The device includes: A to-be-desensitized character matching unit, configured to match the alternative character sets with the to-be-desensitized characters in the to-be-desensitized data, and determine the character set with successful matching as the target character set, where each alternative character set corresponds to one or more types of characters, and each alternative character set corresponds to a different format-preserving encryption FPE instance; A character encryption unit, configured to perform format-preserving encryption FPE processing on the to-be-desensitized characters according to the target character set to obtain the desensitized data corresponding to the to-be-desensitized data; The device further includes: A position information processing unit, configured to record the position information of the to-be-desensitized characters in the to-be-desensitized data; Recombining the to-be-desensitized characters corresponding to the same target character set in the to-be-desensitized data respectively, and performing FPE processing on each group of recombined to-be-desensitized characters respectively to obtain the corresponding groups of desensitized characters; Performing position restoration on the desensitized characters according to the position information corresponding to the corresponding to-be-desensitized characters to obtain the desensitized data; The performing format-preserving encryption FPE processing on the to-be-desensitized characters according to the target character set includes: Determining the FPE instance corresponding to the target character set, and performing FPE processing on the to-be-desensitized characters through the determined FPE instance.

8. A data restoration device, Characterized in that, The device includes: A to-be-restored character matching unit, configured to match the alternative character sets with the to-be-restored characters in the to-be-restored data, and determine the character set with successful matching as the target character set, where each alternative character set corresponds to one or more types of characters, and each alternative character set corresponds to a different format-preserving encryption FPE instance; A character restoration unit, configured to perform format-preserving encryption FPE inverse processing on the to-be-restored characters according to the target character set to obtain the original data corresponding to the to-be-restored data; The device further includes: A location information processing unit for recording the location information of the character to be restored in the data to be restored; Recombining the characters to be restored corresponding to the same target character set in the data to be restored respectively, and performing FPE processing on each group of recombined characters to be restored respectively to obtain corresponding groups of restored characters; Restoring the positions of the restored characters according to the location information corresponding to the corresponding characters to be restored to obtain the restored data; The format-preserving encryption FPE inverse processing of the characters to be restored according to the target character set includes: Determining an FPE instance corresponding to the target character set, and performing FPE processing on the characters to be restored through the determined FPE instance.

9. A computer-readable storage medium, on which a computer program is stored, Characterized in that, When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, Characterized in that, When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Streaming encryption method for text data reservation format

    CN110704854A

  • Implementation mode of extensible format preserving encryption method

    CN112597480A