Data desensitization method, device, medium and system
Through the combination of FF3 and the Guomi algorithm SM4, the data is classified and encrypted, which solves the data leakage problem caused by a single encryption algorithm in the existing technology, and achieves data format consistency and high security.
Patent Information
- Application Number
- CN202510328714.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-01
AI Technical Summary
Existing data desensitization technologies often use a single encryption algorithm, which leads to poor security in desensitization results and is prone to data leakage.
The FF3 encryption algorithm and the Guoxin algorithm SM4 are used to classify and process the desensitized data, generate random factors and keys, ensure the consistency of data format and encryption strength, and generate keys through the Guoxin algorithm SM3 to process pseudo-random values, and combine mapping relationships and regular expressions for data classification and encryption.
While retaining the original format and structure of data, it maximizes data privacy, improves data security, prevents data leakage, and adapts to the data exchange and analysis needs of multiple business scenarios.
Smart Images

Figure CN120234828A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data desensitization, and more particularly, to a data desensitization method, device, medium, and system. Background Art
[0002] To achieve the availability and invisibility of data, the application of desensitization technology is particularly important. Data desensitization is widely used in scenarios such as development testing, data provision, data analysis, production applications, data exchange, and compliance requirements. However, traditional desensitization technologies (such as information truncation, mask masking, local obfuscation, display hiding, offset rounding, and data encryption) can damage the internal format of the data, making it impossible to adapt to the database or pass the parameter verification logic.
[0003] In addition, the desensitization of existing solutions often uses a single encryption algorithm, resulting in poor security of the desensitization results and still being prone to data leakage. Summary of the Invention
[0004] The main objective of this application is to provide a data desensitization method, device, medium, and system to at least solve the problem that the desensitization of existing solutions often uses a single encryption algorithm, resulting in poor security of the desensitization results and still being prone to data leakage.
[0005] To achieve the above objective, according to one aspect of this application, a data desensitization method is provided. The method includes:
[0006] Obtain the data to be desensitized, and classify the array of the data to be desensitized in the form of a mapping relationship to obtain array categories, where the array categories include addresses, ID numbers, names, and bank card numbers;
[0007] Obtain a pseudo-random value, process the pseudo-random value using the national cryptography algorithm SM3 to obtain a key, and generate a random factor based on the array categories in the form of a mapping relationship;
[0008] Perform desensitization processing on the random factor, the key, and the array of the data to be desensitized using the FF3 encryption algorithm and the national cryptography algorithm SM4 to obtain the desensitization result.
[0009] Optionally, obtaining a pseudo-random value includes:
[0010] Generate a seed value based on the system random parameters in the form of a mapping relationship, where the system random parameters include the running time of the system, and the seed value is a random number;
[0011] Process the seed value using a pseudo-random number generator PRN to obtain the pseudo-random value.
[0012] Optionally, the FF3 encryption algorithm and the national encryption algorithm SM4 are used to desensitize the array of the random factor, the key, and the data to be desensitized, and the desensitization result is obtained, including:
[0013] The FF3 encryption algorithm is used to normalize the random factor to obtain a normalization result;
[0014] The national encryption algorithm SM4 is used to process the key and the normalization result to obtain an intermediate value;
[0015] The desensitization result is determined according to the array of the data to be desensitized and the intermediate value.
[0016] Optionally, using the national encryption algorithm SM4 to process the key and the normalization result to obtain an intermediate value, including:
[0017] A unique identifier is generated based on the array category in the form of a mapping relationship;
[0018] The national encryption algorithm SM4 is used to perform bitwise exclusive OR processing on the unit reverse string of the normalization result and the unique identifier based on the unit reverse string of the key to obtain an exclusive OR processing result;
[0019] It is determined that the intermediate value is the unit reverse string of the exclusive OR processing result.
[0020] Optionally, according to the array of the data to be desensitized and the intermediate value, determining the desensitization result, including:
[0021] The radix string of the size of the target unit reverse string is converted into a preset radix value, and the target unit reverse string is at least part of the unit reverse string in the array of the data to be desensitized;
[0022] The sum value of the preset radix value and the intermediate value is normalized to obtain the desensitization result.
[0023] Optionally, classifying the array of the data to be desensitized in the form of a mapping relationship to obtain an array category, including:
[0024] Determining a regular expression corresponding to the array of the data to be desensitized in the form of a mapping relationship;
[0025] Determining the array category according to the corresponding regular expression in the form of a mapping relationship.
[0026] Optionally, after using the FF3 encryption algorithm and the national encryption algorithm SM4 to desensitize the array of the random factor, the key, and the data to be desensitized to obtain a desensitization result, the method further includes
[0027] In the case where it is determined that data restoration is required, perform the reverse operation of the desensitization step to obtain the data to be desensitized. The desensitization step represents using the FF3 encryption algorithm, the national cryptography algorithm SM3, and the national cryptography algorithm SM4 to desensitize the array of the data to be desensitized based on the array category to obtain the desensitization result.
[0028] According to another aspect of the present application, there is provided a data desensitization device, including:
[0029] A first acquisition unit, configured to acquire data to be desensitized, and classify the array of the data to be desensitized in a mapping relationship manner to obtain an array category, where the array category includes address, ID number, name, and bank card number;
[0030] A second acquisition unit, configured to acquire a pseudo-random value, process the pseudo-random value using the national cryptography algorithm SM3 to obtain a key, and generate a random factor based on the array category in a mapping relationship manner;
[0031] A first processing unit, configured to use the FF3 encryption algorithm and the national cryptography algorithm SM4 to desensitize the random factor, the key, and the array of the data to be desensitized to obtain the desensitization result.
[0032] According to another aspect of the present application, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the methods.
[0033] According to another aspect of the present application, there is provided a data desensitization system, which includes: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the methods.
[0034] Applying the technical solution of the present application, by classifying the data to be desensitized according to its attributes (such as address, ID number, name, and bank card number), it is ensured that while the desensitized data retains its original format and structure, the privacy of the data is maximally protected. By generating a key by processing a pseudo-random value based on the national cryptography algorithm SM3, this process provides additional security. Using the FF3 encryption algorithm and the national cryptography algorithm SM4 for desensitization processing can not only maintain the format consistency of the data but also ensure the encryption strength, meeting the high requirements for data security, and improving the security compared with the existing solutions. Thus, it solves the problem that the existing solutions often use only one encryption algorithm for desensitization, resulting in poor security of the desensitization result and still being prone to data leakage. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0036] Figure 1 FIG. shows a schematic flow chart of a data desensitization method provided according to an embodiment of this application;
[0037] Figure 2 FIG. shows a schematic diagram of a data desensitization framework provided according to an embodiment of this application;
[0038] Figure 3 FIG. shows a schematic diagram of the implementation principle of the supervisability of a data desensitization framework provided according to an embodiment of this application;
[0039] Figure 4 FIG. shows a schematic diagram of the implementation principle of the controllability of a data desensitization framework provided according to an embodiment of this application;
[0040] Figure 5 FIG. shows a block diagram of the structure of a data desensitization device provided according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will describe this application in detail with reference to the drawings and in combination with the embodiments.
[0042] In order to enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0043] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so as to implement the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0044] For the convenience of description, some nouns or terms related to the embodiments of this application are described below:
[0045] Data desensitization: Generalize, replace, encrypt, truncate, obfuscate, etc. sensitive data under given rules to eliminate or hide sensitive information and prevent the data from being leaked or misused without authorization. However, it is necessary to ensure that the processed data retains sufficient value and usability to meet business requirements or data analysis and statistics, etc.
[0046] Format-Preserving Encryption (FPE): A special symmetric cryptography technique. FPE can not only provide data confidentiality but also ensure the consistency of data format and structure before and after encryption. FPE has the advantages of not requiring changes to the database paradigm and being transparent to upper-layer applications, making FPE suitable for data desensitization.
[0047] FF3 encryption algorithm (Format-Preserving Encryption Feistel-based Format 3): It is a Format-Preserving Encryption (FPE) algorithm. The FF3 algorithm is based on the Feistel network structure and is particularly suitable for encrypting equal-length strings with specific formats, such as phone numbers, license plate numbers, ID card numbers, etc., while keeping the encrypted data format consistent with the original data format. In this way, sensitive information can be securely stored and processed without modifying the database structure and application programs.
[0048] Guomi algorithm SM3: A digest algorithm, that is, a hash algorithm. The output of the SM3 algorithm is a 256-bit fixed-length digest, which is used to ensure the integrity and consistency of data. The design of the SM3 algorithm has strong collision resistance and unidirectionality, which means it is difficult to reverse the original data from the digest information. At the same time, it is very difficult to find two different inputs that produce the same digest, making SM3 an important part of security applications such as data integrity and digital signatures.
[0049] Guomi Algorithm SM4: SM4 is a block cipher algorithm that uses a 128-bit key and a 128-bit data block size for symmetric encryption and decryption. The SM4 algorithm supports multiple operation modes (such as ECB, CBC, CFB, OFB, and CTR, etc.), and can be applied to various scenarios such as data encryption, file security, and network communication confidentiality. Since SM4 is completely independently developed, its key scheduling and round function design are different from international standard algorithms such as AES, thus ensuring the autonomy and security of the algorithm.
[0050] As introduced in the background art, in order to achieve the availability and invisibility of data, the application of data desensitization technology is particularly important. Data desensitization is widely used in scenarios such as development and testing, data provision, data analysis, production applications, data exchange, and compliance requirements. However, traditional desensitization technologies (such as information truncation, mask masking, local confusion, display hiding, offset rounding, and data encryption) will damage the internal format of the data, making it impossible to adapt to the database or pass the parameter verification logic. To solve the problem that the desensitization of existing solutions often uses an encryption algorithm, resulting in poor security of the desensitization results and still being prone to data leakage, the embodiments of the present application provide a data desensitization method, device, medium, and system.
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0052] In this embodiment, a data desensitization method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0053] Figure 1 It is a flowchart of a data desensitization method provided according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0054] Step S101, obtain the data to be desensitized, and classify the array of the above-mentioned data to be desensitized in the form of a mapping relationship to obtain array categories, and the above-mentioned array categories include address, ID number, name, and bank card number;
[0055] In an embodiment of the present application, classifying the array of the above-mentioned data to be desensitized in the form of a mapping relationship to obtain array categories includes:
[0056] Determine the regular expression corresponding to the array of the above-mentioned data to be desensitized in the form of a mapping relationship;
[0057] Determine the above array category according to the corresponding regular expression in the form of a mapping relationship.
[0058] By automatically identifying different types of sensitive data through regular expressions, such as ID card numbers, bank card numbers, phone numbers, etc., automatic classification of data can be achieved without manual intervention, improving the efficiency and automation of data desensitization. As a powerful tool for text pattern matching, regular expressions can accurately match data in a specific format, so data can be accurately classified into the corresponding array categories. Regular expressions can clarify the format of data, such as the 18 digits of an ID card number, the fixed number of digits and prefix of a bank card number, etc., which ensures that the result after data desensitization still conforms to the original data format and structure. This is particularly important for systems that need to comply with strict data format regulations, such as the verification rules for bank card numbers in the financial system. The desensitization result that retains the format helps to meet data protection and compliance requirements without modifying the existing system structure. The establishment of the mapping relationship enables the desensitization strategy to be designed specifically according to the type and format of the data.
[0059] Step S102, obtain a pseudo-random value, process the pseudo-random value using the national cryptography algorithm SM3 to obtain a key, and generate a random factor based on the above array category in the form of a mapping relationship;
[0060] Among them, obtaining the pseudo-random value includes:
[0061] Generate a seed value based on the system random parameters in the form of a mapping relationship. The above system random parameters include the running time of the system, and the seed value is a random numerical value;
[0062] Process the above seed value using a pseudo-random number generator PRN to obtain the above pseudo-random value.
[0063] Taking the system's current running time and other system random parameters as the basis of the seed value can ensure that the seed value for each encryption or desensitization operation is unique and unpredictable. Since the running time is a constantly changing parameter, its inclusion as part of the seed value adds additional non-repeatability to the generation of pseudo-random values, thereby enhancing the security of the encryption and desensitization processes. The pseudo-random values obtained by processing the seed value through PRN can generate a numerical sequence with statistical randomness characteristics, which is crucial for encryption algorithms. The design of the pseudo-random number generator ensures that the output sequence is random in the cryptographic sense. Even if an attacker knows part of the sequence or the working principle of the algorithm, it is difficult to predict subsequent values, which increases the difficulty for the attacker to recover the original data through pattern analysis. Dynamicity and flexibility: Since the seed value depends on random parameters such as the system running time, the pseudo-random values generated each time are dynamically changing, which provides flexibility for encryption and desensitization strategies. In the dynamic desensitization scenario, the randomly changing pseudo-random values can prevent the same data from being desensitized to the same result at different time points, thereby protecting data privacy and preventing data correlation exposure due to data duplication.
[0064] Step S103: Use the FF3 encryption algorithm and the national cryptography algorithm SM4 to desensitize the array of the above-mentioned random factor, the above-mentioned key, and the above-mentioned data to be desensitized, and obtain the above-mentioned desensitization result.
[0065] In the above steps, by classifying the data to be desensitized according to its attributes (such as address, ID number, name, and bank card number), while ensuring that the desensitized data retains its original format and structure, it maximally protects data privacy. The key is generated by processing the pseudo-random value based on the national cryptography algorithm SM3. This process provides additional security. Using the FF3 encryption algorithm and the national cryptography algorithm SM4 for desensitization can not only maintain data format consistency but also ensure encryption intensity, meeting the high requirements of data security, and improving security compared to existing solutions. Thus, it solves the problem that existing solutions often use only one encryption algorithm for desensitization, resulting in poor security of the desensitization result and still being prone to data leakage.
[0066] Among them, using the FF3 encryption algorithm and the national cryptography algorithm SM4 to desensitize the array of the above-mentioned random factor, the above-mentioned key, and the above-mentioned data to be desensitized to obtain the desensitization result includes: normalizing the above-mentioned random factor using the above-mentioned FF3 encryption algorithm to obtain a normalization result; processing the above-mentioned key and the above-mentioned normalization result using the above-mentioned national cryptography algorithm SM4 to obtain an intermediate value; determining the above-mentioned desensitization result based on the array of the above-mentioned data to be desensitized and the above-mentioned intermediate value.
[0067] A specific application scenario of desensitizing the above-mentioned random factor, the above-mentioned key, and the array of the above-mentioned data to be desensitized by using the FF3 encryption algorithm and the national cryptographic algorithm SM4: In the financial industry, banks need to desensitize customer data to ensure that sensitive information such as bank card numbers and ID card numbers is not leaked during the processes of development, testing, analysis, and external data exchange. Suppose a bank has a database table containing customer bank card numbers, and these bank card numbers need to be desensitized when sent to an external audit or data analysis team. The system first reads the bank card number data from the database and identifies and classifies it through regular expressions, determining that the array category of the data to be desensitized is "bank card number". The system generates a random seed value based on system parameters such as the current time, processes it through a pseudo-random number generator (PRN) to obtain a pseudo-random value, and then uses the FF3 encryption algorithm to normalize this pseudo-random value to ensure that its format and length match the bank card number field. The advantage of this step is that the normalized random factor can adapt to the input requirements of the FF3 algorithm while maintaining randomness, providing a secure and format-matching random input for the subsequent encryption process. The system uses the SM3 algorithm to generate a key based on the system running time. Then, the national cryptographic algorithm SM4 is used to process the generated key and the normalized random factor to obtain an intermediate value. The use of the SM4 algorithm ensures the security and strength of the encryption process because it is an encryption algorithm certified by the National Cryptography Administration, providing strong data protection capabilities. Finally, the system uses the obtained intermediate value to encrypt the array of bank card number data to obtain the desensitized new bank card numbers. Due to the use of format-preserving encryption (FF3) and the national cryptographic algorithm SM4, the desensitized bank card numbers not only maintain the same format and length as the original data but are also encrypted in the cryptographic sense, protecting the customer's bank card information.
[0068] Corresponding beneficial effects: Using the national cryptographic algorithm SM4 in combination with the FF3 algorithm for encryption improves the security of the desensitized data because even if an attacker knows the desensitization algorithm, it is difficult to recover the original data from the ciphertext. The desensitized bank card numbers still maintain their original format and length, which means that the desensitized data can be directly used for system testing, analysis, or external data exchange without additional data verification or format adjustment. The format-preserving feature ensures that the desensitized data can continue to meet the business requirements of internal system applications. For example, the desensitized bank card numbers can still be verified through the Luhn algorithm to support the normal operation of business logic. Since the encryption algorithm is based on national cryptographic standards, the entire desensitization process is technically independently controllable, reducing the dependence on external algorithms and enhancing the national information security level. The normalized random factor is new each time, which increases the dynamic randomness of the encryption process, causing the same bank card number to produce different ciphertexts in different encryption sessions and further protecting data privacy.
[0069] Among them, the above-mentioned national secret algorithm SM4 is used to process the above-mentioned key and the above-mentioned normalization result to obtain an intermediate value, including: generating a unique identifier in the form of a mapping relationship based on the above-mentioned array category; using the above-mentioned national secret algorithm SM4, based on the unit reverse string of the above-mentioned key, performing bit-by-bit exclusive OR processing on the unit reverse string of the above-mentioned normalization result and the above-mentioned unique identifier to obtain an exclusive OR processing result; determining the above-mentioned intermediate value as the unit reverse string of the above-mentioned exclusive OR processing result.
[0070] Specifically, by performing reverse string processing based on the unique identifier of the array category and then performing an exclusive OR operation with the normalization result, the security of data encryption can be significantly enhanced. The exclusive OR operation is a basic cryptographic operation that can achieve information confusion, making the relationship between the ciphertext and the plaintext unpredictable. And the reverse string processing increases the complexity of the processing, making it difficult for an attacker to restore the plaintext information even if the ciphertext is intercepted, because both the reverse operation and the exclusive OR result are completed based on the dynamically generated unique identifier. The combined use of reverse string and exclusive OR operation ensures the irreversibility of the desensitization result, that is, without the key and the unique identifier, the original data cannot be restored from the desensitization result. This is particularly important for data desensitization because one of the purposes of data desensitization is to ensure that once the data is desensitized, it cannot be easily restored, thus protecting the privacy of sensitive information. In some business scenarios, such as when the correlation between data needs to be retained, a deterministic encryption method is necessary. For example, by performing encryption processing based on the unique identifier generated from the same array category, it can be ensured that the same type of sensitive data produces consistent desensitization results under the same key and random factor.
[0071] Among them, according to the above-mentioned array of data to be desensitized and the above-mentioned intermediate value, the above-mentioned desensitization result is determined, including: converting a radix-based string of the size of the target unit reverse string into a preset radix value, where the target unit reverse string is at least part of the unit reverse string in the above-mentioned array of data to be desensitized; performing normalization processing on the sum value of the above-mentioned preset radix value and the above-mentioned intermediate value to obtain the above-mentioned desensitization result.
[0072] Specifically, through the conversion of radix numbers, the data formats before and after desensitization are ensured to be consistent. For example, when desensitizing a bank card number, the bank card number can be converted to a decimal number, and after encryption, it is converted back to a string of the same format. In this way, even if the data is encrypted, its display format (such as fixed number of digits, digital type) is retained, which is crucial for maintaining the availability of data and compatibility with existing systems. Using a preset radix value for processing can reduce the computational complexity in the encryption and decryption processes. Generally, directly operating at the string level (such as XOR) may involve additional character encoding conversions, while operating through numerical values can simplify the calculation process and improve the processing speed. Especially when dealing with a large amount of data, the advantages of this method are more obvious. Converting the reversed string of the target unit into a preset radix value can effectively expand the numerical domain and increase the randomness and security of encryption. Through this step, even seemingly simple data (such as consecutive digits in an ID number) will exhibit complexity and unpredictability during the conversion and encryption / decryption processes, thus increasing the difficulty of cracking.
[0073] In addition, after using the FF3 encryption algorithm and the national cryptographic algorithm SM4 to desensitize the array of the above-mentioned random factor, the above-mentioned key, and the above-mentioned data to be desensitized to obtain a desensitization result, the above-mentioned method further includes, in the case of determining that data needs to be restored, performing the reverse operation of the desensitization step to obtain the above-mentioned data to be desensitized. The above-mentioned desensitization step represents using the above-mentioned FF3 encryption algorithm, the above-mentioned national cryptographic algorithm SM3, and the above-mentioned national cryptographic algorithm SM4 to desensitize the array of the above-mentioned data to be desensitized based on the above-mentioned array category to obtain the above-mentioned desensitization result.
[0074] Specifically, this reverse operation mechanism allows for the adjustment and update of desensitization policies without accessing the original data. For example, if it is found that the desensitization intensity of a certain type of data is insufficient, the desensitization step can be re-executed and the encryption parameters can be adjusted without having to process all the data from the beginning, improving efficiency. The reverse operation ensures the integrity of the desensitized data when restored, that is, the desensitization process does not permanently damage the original information. This is crucial for the flexible use of data in different business scenarios, especially when the data needs to be shared among multiple systems or business processes.
[0075] This application constructs the retention format encryption algorithm SM-FF3 based on the national cryptographic algorithm, achieving the autonomy and control of the underlying algorithm. At the same time, it gives the adaptation strategies for desensitizing complex fields such as the desensitization time, bank card number, and ID card number when using SM-FF3 (combining SM3, SM4, and FF3). The desensitized data in this application resembles real data, supports various complex association queries, and maintains business consistency, with good applicability in scenarios such as static desensitization and dynamic desensitization. Compared with traditional implementation schemes such as transformation, modification, and masking, this application not only provides cryptographic security and good adaptability but also solves the problem of desensitization traceability, so that when risks and accidents occur, it can trace back from the desensitized data to the real data source and risk points, improving the overall controllability and scalability of the scheme.
[0076] Construct the FPE algorithm SM-FF3 based on the FF3 algorithm. SM-FF3 is based on the interactive Feistel structure, and the two core operations of the algorithm are the encrypted compression mapping of the string value and the generation and superposition of the pseudo-random offset. This application makes the following extensions to FF3: truncation and padding are implemented based on the national cryptographic algorithm SM3 to adapt to the encryption network, and pseudo-random offset and superposition are implemented based on the national cryptographic algorithm SM4 to ensure the security strength. The encryption core components of SM-FF3 are all implemented based on the national cryptographic algorithm, ensuring the security and autonomy of the algorithm.
[0077] For traceability considerations, this scheme makes full use of the encryption and decryption characteristics of SM-FF3 to give a security traceability scheme based on the key. Traceability does not depend on the original copy and global matching, and the storage and time overhead are very low. In addition, the scheme can ensure the high precision of the traceability result. Based on the technical idea of traceability, this application also realizes the dynamic adjustment of the desensitization strategy, that is, when changing the desensitization method due to business requirements, the desensitization method of relevant fields can be changed in fine-grained manner and does not depend on the original copy, improving the flexibility of desensitization.
[0078] For applicability considerations, this application designs a desensitization scheme applicable to various scenarios based on SM-FF3. For example, deterministic encryption mapping is used for the primary and foreign keys to support join queries and order-preserving desensitization; for another example, combined with the data classification standard, different desensitization methods with different security strengths are proposed to balance the requirements of availability and security; and for yet another example, the method of custom encoding and special mapping is used to complete the format-preserving desensitization of complex fields such as bank card number, time, and address in addition to desensitizing a single type of field.
[0079] NUM(X): Convert the binary string X to decimal, similar to Java Integer.parseInt(X, 2).
[0080] REV(X): Reverse the string X.
[0081] REVB(X): Reverse the string X in bytes.
[0082] [X] a : Convert the string X into a bit string of length a.
[0083] NUM radix(X) : Convert the radix string X to decimal, similar to Java Integer.parseInt(X, radix).
[0084] STR m radix(x) : Convert the numerical value x into a radix string of length m, which is the inverse operation of NUMradix(X).
[0085] As Figure 2 shown, the principle of the data desensitization framework is demonstrated. At the encryption algorithm layer, module S110 is based on the cryptographic components of the national cryptographic extension FF3 to achieve enhanced security and algorithm autonomy and control; module S120 constructs an encryption mapping based on the cryptographic components to implement the SM-FF3 algorithm. Since SM-FF3 can only process numerical strings, this application sets the encoding strategy for format strings to handle actual business fields; at the desensitization adaptation layer, module S210 divides the desensitized fields into three categories and gives the application method of SM-FF3 taking the financial system database as an example, including the construction of different types of desensitization strategies, strategies for maintaining relevance, etc. In addition, considering the security defect that the ciphertext space of SM-FF3 is too small, module S220 implements a security extension strategy to balance security and efficiency; at the desensitization traceability layer, modules S310 and S320 handle the desensitization traceability problem.
[0086] Algorithm 1: KDF(rnd) → K;
[0087] Input: System random parameter rnd; Output: 128-bit key K;
[0088]
[0089] Algorithm 2: PRF K (X, s) → Y m ;
[0090] Input: String X, key K, shared parameter s (i.e., unique identifier); Output: 128-bit ciphertext block;
[0091]
[0092] The pseudo-random generation function PRN of the S110 module is used to derive the key and the random factor T. PRN must use a cryptographically secure generation method, such as the SecureRandom class in Java; the domestic cryptographic algorithms SM3 and SM4 are implemented by calling the official library; the encryption network is constructed based on the interactive Feistel; and the KDF implementation is as shown in Algorithm 1: The seed value rnd is random parameters such as the current system time and running state. The rnd is input into PRN to obtain the pseudo-random value seed, and then a 128-bit key K is derived through SM3; the implementation of PRF is as shown in Algorithm 2: That is, use SM3 to perform truncated expansion on X, and then use SM4 to encrypt the message. The original AES-CBC is not used in the PRF function mainly to adapt to different lengths of X and reduce the number of block cipher calls to improve efficiency. Secondly, due to the one-way property of SM3, the security of the PRF function can also be improved.
[0093] Algorithm 3: SM-FF3.Enc(X, K, T) → Y;
[0094] Input: array X, key K, random factor T; Output: array Y;
[0095]
[0096] The S120 module mainly completes encryption mapping and encoding mapping. The core of the encryption mapping is to map an array to another array of the same length and ensure the irreversibility of the mapping without the key K. The core of the encoding mapping is to use a suitable encoding table to map a string to an array that can be processed by SM-FF3.
[0097] The encryption mapping consists of the encryption and decryption algorithms of SM-FF3. The encryption of SM-FF3 is as shown in Algorithm 3. Compared with the standard SP800-38G, in this application, an additional extended truncation is performed based on SM3 in the second step to prevent the reuse of the T value that occurs in direct truncation and at the same time adapt to various types of input T; in the eighth step, instead of directly calling the block encryption, the PRF function extended by the national cryptography is used to handle the length adaptation and efficiency problems that occur when the array X is too long; in the ninth step, this application repeats the modulo operation during the NUM calculation to improve the calculation efficiency. At the same time, for security reasons, the input X needs to satisfy both radixn≥106 and n≥2. The decryption of SM-FF3 is the same as the encryption, and each step corresponds to the inverse operation of the encryption operation. The specific implementation is as shown in Algorithm 4, so Algorithm 4 will not be elaborated here.
[0098] Algorithm 4: SM-FF3.Dec(Y, K, T) → X;
[0099] Input: array Y, key K, random factor T; Output: array X;
[0100]
[0101] The encoding mapping customizes the encoding table for various types of fields. For example, the alphabetic password "root" can be encoded as [17, 14, 14, 19] using the alphabet table; if mixed with numbers, using the encoding table {0, 1,.., 9, a, b,.., z}, "r00t" can be encoded as the array [27, 0, 0, 29]; then SM-FF3 maps the array to another legal array and decodes it inversely to obtain the ciphertext string. For example, the encoding array of "r00t" is mapped by SM-FF3 to [4, 34, 15, 8], and after inverse encoding, the ciphertext "4yf8" is obtained. To handle more complex encodings, this application adopts three encoding methods: custom encoding set, ASCII encoding, and UTF-8 encoding (removing redundant characters), which can meet the encoding and decoding mapping of almost all strings.
[0102] The desensitization adaptation layer combines business logic and database characteristics to give the application solution of SM-FF3. Among them, the S210 module processes the application problems of SM-FF3 and divides the desensitized fields into three types: numerical type, single format type, and special format type; the S220 module adapts to the security constraints of SM-FF3 and provides security enhancement functions.
[0103] The S210 module converts the corresponding numerical value into an array. For example, 50000 is encoded as [5, 0, 0, 0, 0] and input into SM-FF3, and then the mapped array [0, 8, 9, 3, 4] is decoded into 8934. However, the leading 0 in this process is lost, which may cause the encryption and decryption to be irreversible. Therefore, the S220 module will count and calculate an upper limit of the number of digits. During the encryption and decryption process, the array is padded with 0 at the beginning to a fixed length. On the one hand, it can ensure the reversibility of encryption and decryption, and on the other hand, it can expand the encryption space to increase security.
[0104] For single format fields such as mobile phone numbers, passwords, and addresses, after encoding with the alphanumeric table, the S210 module can call for encryption. However, to balance usability and security, the S220 module will determine the intensity of single format desensitization according to the data value, such as whether to desensitize all provinces-cities-counties of the address number or only desensitize cities-counties, and whether to desensitize the last four digits or the last eight digits of the mobile phone number. And the weakly desensitized fields can still provide statistical and retrieval value. For example, after the address is desensitized, statistics can still be made on customers who live permanently in other provinces.
[0105] For special fields such as bank card numbers, ID card numbers, and account opening times, customized desensitization strategies are required. For example, a bank card number consists of an issuer identification (BIN), a personal identification number, and a check digit. The S210 module only uses SM-FF3 encryption for the identification code part, and then generates a check digit according to the Luhn algorithm to ensure that the bank card number can pass the business verification. The same principle applies to the desensitization of ID card numbers. The S210 module only processes the 5th to 17th digits and generates a check digit through an algorithm. For the desensitization of time types, the S210 module maps the date to a numerical value representing the number of days from a certain year, month, and day to that day, then encrypts it in the value type, and appends the ciphertext numerical value as the number of days after a certain year, month, and the first day of that month to obtain the desensitized date. This method can avoid considering complex situations such as leap February and months with different numbers of days.
[0106] For maintaining relevance, this application utilizes the characteristics of SM-FF3 deterministic encryption. After reading structured and semi-structured data, the S220 module records the association relationship between a certain field and other fields in the same data table, and ensures that the same T is used in the desensitization of fields with corresponding relationships. Although the Id field is desensitized, because the same plaintext corresponds to the same ciphertext, the table join query remains valid.
[0107] For maintaining more complex business relevance, the S220 module is implemented based on weak desensitization and low-bit encryption of numerical fields. Some information is retained for retrieval or statistics. Low-bit encryption encrypts the low bits of the numerical value in a reserved format and performs an order-preserving offset on the high bits to achieve weak order-preserving desensitization of the numerical value. For example, the numerical value 50000 can be split into "5" and "0000". The "5" is input into a certain monotonically increasing function, and the "0000" is input into SM-FF3, and the outputs are concatenated to obtain the desensitized value. The S220 module can also implement retrieval according to ciphertext or dynamic desensitization. For example, to retrieve a certain customer's information according to the ID card number, the ID card number can be encrypted and desensitized with the key K, then retrieve the desensitized data item according to the ciphertext, and decrypt it within the trusted area to obtain the original data, thereby avoiding information leakage during the query and communication processes.
[0108] The padding strategy provided by the S220 module realizes the expansion of the message space. Since the input of SM-FF3 needs to satisfy radixn≥106 and the length is greater than 2, for the inputs that do not meet the conditions, S220 adopts two expansion methods: forward expansion and random insertion. Forward expansion is used for value type desensitization. At the beginning bit of the encoding array, 0 is expanded to the specified length, expanding the encryption space and hiding the order of magnitude. At the same time, the 0 at the high position is automatically discarded after decryption to ensure correctness; random insertion is used for desensitization of variable-length single-format fields. Random characters are inserted at the specified position of the encoding array and then input into SM-FF3. For example, the password string "admin" can be padded as "awd3mJiUn" and then encrypted. This hides the original length of the string and expands the message space. At the same time, the random characters are removed after decryption to ensure correctness. Due to the avalanche effect of the SM-FF3 cipher, both of these expansion methods are safe and effective.
[0109] The Twist(parameter T) selection strategy implemented by the S220 module is used to control the randomness of encryption. The S220 module first reads the database fields and the associated relationships. For strongly private fields such as passwords, the S220 module uses the Id of each data item as T and inputs it into SM-FF3. In this way, the same password will not be desensitized to the same result, preventing the leakage of equivalent relationships. For example, the same balance of 50,000 can be desensitized to different values to prevent reverse inference. In addition, the association between T and Id does not require additional storage of T, and it also ensures that T will not be reused (the uniqueness of the primary key);
[0110] For the related fields involved in primary keys, foreign keys or full-table queries, the S220 module appropriately tolerates the leakage of equivalent relationships and generates T = SM3(database_name||table_name||field_name||seed) for field desensitization. On the basis of hiding sensitive information, the desensitized data table still maintains usability.
[0111] The desensitization traceability layer mainly realizes the supervision and controllability of desensitized data. The supervision can trace the suspicious original data through the desensitized data and locate the risk points or clarify the relevant responsible parties. The controllability dynamically adjusts the desensitization method through desensitization traceability to meet the needs of compliance and business changes.
[0112] The realization of supervision is as Figure 3 shown, corresponding to Figure 2The S310 module in it. In step 1, according to the data classification standard, the data owner identifies and matches the sensitive data in the database and file system. For the identified sensitive data, in step 2, it is copied and passed into the desensitization adaptation layer. Since the desensitized data may be used by multiple units, in step 3, different desensitization keys K1, K2, and K3 are used to desensitize the data. This can not only improve security but also facilitate the customization of different desensitization strategies. In step 4, when production supervision discovers abnormal behavior data, but because the key data has been desensitized, it is impossible to directly locate the real abnormal information, and it is also difficult to clarify the risk outbreak point. At this time, in step 5, suspicious data is captured and desensitized using K1, K2, and K3. As Figure 3 shown, at this time, the desensitization using K2 is successful. The problem can be verified and analyzed based on the desensitized data, and at the same time, it can be clarified that the risk outbreak point is in Unit B.
[0113] The realization of controllability is as Figure 4 shown and corresponds to Figure 2 the S320 module in it. In step 1, according to the data classification standard, the data owner identifies and matches the sensitive data in the database and file system. For the identified sensitive data, in step 2, it is copied and passed into the desensitization module. In step 3, the selected strategy is used to desensitize the data and send it to the using unit. However, with the requirements of business changes or compliance control, etc., some desensitized fields are difficult to meet the usability requirements. At this time, the using unit submits a desensitization conversion application in step 4 and sends the desensitized fields to the desensitization adaptation layer. Subsequently, the desensitization module retrieves and uses the key to desensitize the data, and in step 5, a desensitization scheme is reselected in combination with the application content in step 4, and finally the converted desensitized fields are sent to the using unit. S320 realizes fine-grained conversion and does not rely on the backup of the original data set.
[0114] Two detailed optimizations are also involved in the implementation of S310 and S320. First, each time the desensitization adaptation layer is called, S310 and S320 use a HashMap to record key-value pairs (using unit, desensitization key). When S310 and S320 need to desensitize the data, the desensitization key corresponding to the unit is quickly retrieved in the HashMap to improve efficiency. Second, when it is necessary to determine whether the decryption is valid. Only some flag fields can be decrypted. If multiple legal results are obtained when decrypting the flag fields at the same time, it is determined that the decryption is valid, otherwise it must be invalid. This improves the determination efficiency and effectively avoids possible decryption ambiguities.
[0115] In order to enable those skilled in the art to more clearly understand the technical solution of this application, the implementation process of the data desensitization method of this application will be described in detail below in combination with specific embodiments.
[0116] This embodiment relates to a specific data desensitization method, including:
[0117] Obtain the data to be desensitized, and classify the array of data to be desensitized in the form of a mapping relationship to obtain array categories, where the array categories include addresses, ID numbers, names, and bank card numbers;
[0118] Specifically, determine the regular expression corresponding to the array of data to be desensitized in the form of a mapping relationship; and determine the array category according to the corresponding regular expression in the form of a mapping relationship.
[0119] Obtain a pseudo-random value, process the pseudo-random value using the national cryptography algorithm SM3 to obtain a key, and generate a random factor based on the array category in the form of a mapping relationship;
[0120] Specifically, generate a seed value based on the system random parameters in the form of a mapping relationship. The system random parameters include the running time of the system, and the seed value is a random numerical value; process the seed value using a pseudo-random number generator PRN to obtain a pseudo-random value.
[0121] Perform normalization processing on the random factor using the FF3 encryption algorithm to obtain a normalization result;
[0122] Process the key and the normalization result using the national cryptography algorithm SM4 to obtain an intermediate value;
[0123] Specifically, generate a unique identifier based on the array category in the form of a mapping relationship; use the national cryptography algorithm SM4 to perform bitwise exclusive OR processing on the unit reverse string of the normalization result and the unique identifier based on the unit reverse string of the key to obtain an exclusive OR processing result; determine the intermediate value as the unit reverse string of the exclusive OR processing result.
[0124] Determine the desensitization result according to the array of data to be desensitized and the intermediate value.
[0125] Specifically, convert the radix string of the size of the target unit reverse string into a preset radix value. The target unit reverse string is at least part of the unit reverse string in the array of data to be desensitized; perform normalization processing on the sum value of the preset radix value and the intermediate value to obtain the desensitization result.
[0126] In the case of determining that data needs to be restored, perform the reverse operation of the desensitization step to obtain the data to be desensitized. The desensitization step means using the FF3 encryption algorithm, the national cryptography algorithm SM3, and the national cryptography algorithm SM4 to perform desensitization processing on the array of data to be desensitized based on the array category to obtain the desensitization result.
[0127] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0128] The embodiments of the present application also provide a data desensitization device. It should be noted that the data desensitization device in the embodiments of the present application can be used to execute the data desensitization method provided by the embodiments of the present application. The device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0129] The following introduces the data desensitization device provided by the embodiments of the present application.
[0130] Figure 5 is a structural block diagram of a data desensitization device provided according to an embodiment of the present application. As Figure 5 shown, the device includes:
[0131] A first acquisition unit 51, configured to acquire data to be desensitized, and classify the array of the data to be desensitized in the form of a mapping relationship to obtain an array category, where the array category includes address, ID number, name, and bank card number;
[0132] A second acquisition unit 52, configured to acquire a pseudo-random value, process the pseudo-random value using the national cryptography algorithm SM3 to obtain a key, and generate a random factor based on the array category in the form of a mapping relationship;
[0133] A first processing unit 53, configured to perform desensitization processing on the random factor, the key, and the array of the data to be desensitized using the FF3 encryption algorithm and the national cryptography algorithm SM4 to obtain the desensitization result.
[0134] In the above device, by classifying the data to be desensitized according to its attributes (such as address, ID number, name, and bank card number), it is ensured that while the desensitized data retains its original format and structure, the privacy of the data is maximally protected. By processing the pseudo-random value based on the national cryptography algorithm SM3 to generate a key, this process provides additional security. Using the FF3 encryption algorithm and the national cryptography algorithm SM4 for desensitization processing can not only maintain the format consistency of the data but also ensure the encryption strength, meeting the high requirements of data security. Compared with the existing solutions, the security is improved, thus solving the problem that the existing solutions often use only one encryption algorithm for desensitization, resulting in poor security of the desensitization result and still being prone to data leakage.
[0135] In an embodiment of the present application, the second acquisition unit includes a first processing module and a second processing module. The first processing module is configured to generate a seed value based on system random parameters in a mapping relationship. The system random parameters include the running time of the system, and the seed value is a random number. The second processing module is configured to process the seed value using a pseudo-random number generator PRN to obtain the pseudo-random value.
[0136] In an embodiment of the present application, the first processing unit includes a third processing module, a fourth processing module, and a fifth processing module. The third processing module is configured to normalize the random factor using the FF3 encryption algorithm to obtain a normalization result. The fourth processing module is configured to process the key and the normalization result using the national cryptographic algorithm SM4 to obtain an intermediate value. The fifth processing module is configured to determine the desensitization result according to the array of the data to be desensitized and the intermediate value.
[0137] In an embodiment of the present application, the fourth processing module includes a first processing sub-module, a second processing sub-module, and a third processing sub-module. The first processing sub-module is configured to generate a unique identifier based on the array category in a mapping relationship. The second processing sub-module is configured to perform a bitwise exclusive OR operation on the inverted unit string of the normalization result and the unique identifier based on the inverted unit string of the key using the national cryptographic algorithm SM4 to obtain an exclusive OR processing result. The third processing sub-module is configured to determine that the intermediate value is the inverted unit string of the exclusive OR processing result.
[0138] In an embodiment of the present application, the fifth processing module includes a fourth processing sub-module and a fifth processing sub-module. The fourth processing sub-module is configured to convert a radix string of a target inverted unit string into a preset radix value. The target inverted unit string is the inverted unit string of at least a part of the array of the data to be desensitized. The fifth processing sub-module is configured to normalize the sum value of the preset radix value and the intermediate value to obtain the desensitization result.
[0139] In an embodiment of the present application, the first acquisition unit includes a first determination module and a second determination module. The first determination module is configured to determine a regular expression corresponding to the array of the data to be desensitized in a mapping relationship. The second determination module is configured to determine the array category according to the corresponding regular expression in a mapping relationship.
[0140] In an embodiment of the present application, the above device further includes a second processing unit, which is used to perform the inverse operation of the desensitization step to obtain the above data to be desensitized after performing desensitization processing on the above array of random factors, the above key, and the above data to be desensitized by using the FF3 encryption algorithm and the national encryption algorithm SM4 to obtain a desensitization result, where the desensitization step is to perform desensitization processing on the above array of data to be desensitized based on the above array category by using the above FF3 encryption algorithm, the above national encryption algorithm SM3, and the above national encryption algorithm SM4 to obtain the above desensitization result.
[0141] The above data desensitization device includes a processor and a memory. The above first acquisition unit, second acquisition unit, first processing unit, etc. are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above program units stored in the memory. The above modules are all located in the same processor; or, the above each module is located in different processors in any combination form.
[0142] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem that the desensitization of the existing solution often uses one encryption algorithm, resulting in poor security of the desensitization result and still being prone to data leakage, can be solved.
[0143] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0144] An embodiment of the present invention provides a computer-readable storage medium, where the above computer-readable storage medium includes a stored program, and when the above program runs, it controls the device where the above computer-readable storage medium is located to execute the above data desensitization method.
[0145] An embodiment of the present invention provides a processor, where the above processor is used to run a program, and when the above program runs, it executes the above data desensitization method.
[0146] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the following steps: obtaining the data to be desensitized, classifying the array of the data to be desensitized in the form of a mapping relationship to obtain array categories, where the array categories include address, ID number, name, and bank card number; obtaining a pseudo-random value, processing the pseudo-random value using the national cryptography algorithm SM3 to obtain a key, and generating random factors based on the array categories in the form of a mapping relationship; using the FF3 encryption algorithm and the national cryptography algorithm SM4 to perform desensitization processing on the random factors, the key, and the array of the data to be desensitized to obtain the desensitization result. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0147] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps: obtaining the data to be desensitized, classifying the array of the data to be desensitized in the form of a mapping relationship to obtain array categories, where the array categories include address, ID number, name, and bank card number; obtaining a pseudo-random value, processing the pseudo-random value using the national cryptography algorithm SM3 to obtain a key, and generating random factors based on the array categories in the form of a mapping relationship; using the FF3 encryption algorithm and the national cryptography algorithm SM4 to perform desensitization processing on the random factors, the key, and the array of the data to be desensitized to obtain the desensitization result.
[0148] The present application also provides a data desensitization system, which includes: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the above methods.
[0149] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0150] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0151] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0152] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0153] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0154] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0155] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0156] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0157] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0158] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A data desensitization method, characterized in that: include: Obtain the data to be desensitized, and classify the array of the data to be desensitized in a mapping relationship to obtain array categories, where the array categories include address, ID card number, name, and bank card number; Obtain a pseudo-random value, and process the pseudo-random value using the national secret algorithm SM3 to obtain a key, and generate a random factor based on the array category in a mapping relationship; The FF3 encryption algorithm and the national secret algorithm SM4 are used to desensitize the random factor, the key and the array of the data to be desensitized to obtain the desensitization result.
2. The method according to claim 1, characterized in that: Get pseudo-random values, including: Generate a seed value based on a system random parameter in a mapping relationship, wherein the system random parameter includes the running time of the system, and the seed value is a random value; The seed value is processed using a pseudo-random number generator PRN to obtain the pseudo-random value.
3. The method according to claim 1, characterized in that The random factor, the key and the array of the data to be desensitized are desensitized by using the FF3 encryption algorithm and the national secret algorithm SM4 to obtain a desensitization result, including: The random factor is normalized using the FF3 encryption algorithm to obtain a normalized result; The key and the normalization result are processed using the national secret algorithm SM4 to obtain an intermediate value; The desensitization result is determined according to the array of the data to be desensitized and the intermediate value.
4. The method according to claim 3, characterized in that: The key and the normalization result are processed using the national secret algorithm SM4 to obtain an intermediate value, including: Generate a unique identifier based on the array category in a mapping relationship manner; Using the national secret algorithm SM4, based on the unit inversion string of the key, the unit inversion string of the normalization result and the unique identifier are subjected to bit-by-bit XOR processing to obtain an XOR processing result; The intermediate value is determined to be a unit inverse character string of the XOR processing result.
5. The method according to claim 3, characterized in that: Determining the desensitization result according to the array of the data to be desensitized and the intermediate value includes: Convert a radix string of a target unit inverted string size into a preset base value, wherein the target unit inverted string is at least a portion of the unit inverted string in the array of the data to be desensitized; The sum of the preset base value and the intermediate value is normalized to obtain the desensitization result.
6. The method according to claim 1, characterized in that The array of the data to be desensitized is classified in a mapping relationship to obtain array categories, including: Determine the regular expression corresponding to the array of the data to be desensitized in a mapping manner; The array category is determined according to the corresponding regular expression in a mapping relationship.
7. The method according to any one of claims 1 to 6, characterized in that After using the FF3 encryption algorithm and the national secret algorithm SM4 to desensitize the random factor, the key and the array of the data to be desensitized to obtain the desensitization result, the method further includes: When it is determined that the data needs to be restored, the inverse operation of the desensitization step is performed to obtain the data to be desensitized. The desensitization step is characterized by using the FF3 encryption algorithm, the national secret algorithm SM3 and the national secret algorithm SM4 to desensitize the array of the data to be desensitized based on the array category to obtain the desensitization result.
8. A data desensitization device, characterized in that: include: A first acquisition unit is used to acquire the data to be desensitized, and classify the array of the data to be desensitized in a mapping relationship to obtain array categories, where the array categories include address, ID card number, name and bank card number; A second acquisition unit is used to acquire a pseudo-random value, and process the pseudo-random value using the national secret algorithm SM3 to obtain a key, and generate a random factor based on the array category in a mapping relationship; The first processing unit is used to use the FF3 encryption algorithm and the national secret algorithm SM4 to desensitize the random factor, the key and the array of the data to be desensitized to obtain the desensitization result.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
10. A data desensitization system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 7.
Citation Information
Cited By
Clinical data management system, clinical data privacy protection method, equipment and medium
CN121393705A