Personal information data desensitization processing method
Through multi-threaded parallel processing and dynamic task allocation methods, the problem that the existing technology cannot meet the needs of diversified user identification and data desensitization is solved, and efficient and secure data desensitization processing is achieved, reducing operational costs and development and maintenance difficulties.
Patent Information
- Application Number
- CN202510695902.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
In the scenario where banks cooperate with third-party institutions, the existing technology cannot meet the diverse user identification and data desensitization needs, resulting in increased operating costs and difficulty in development and maintenance.
Using multi-threaded parallel processing and dynamic task allocation methods, efficient data desensitization processing is achieved by confirming the file structure and configuring the data analysis model. Specific steps include multi-threaded file reading, data analysis and desensitization processing, thread safety and exception processing, and support for multiple file formats and dynamic resource allocation.
It significantly improves the processing efficiency of large-scale data files, ensures the security and privacy protection of sensitive information, and reduces operational costs and development and maintenance difficulties.
Smart Images

Figure CN120217448A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method for desensitizing personal information data. Background Art
[0002] With the development of the bank's retail business, the data center has accumulated a large amount of customer personal information collected during account opening and business handling, including name, ID number, ID type, mobile phone number, bank card number, etc., all of which belong to sensitive data. To develop business, banks usually need to cooperate with external third-party institutions. In this process, in order to facilitate data innovation applications and meet compliance requirements, it is necessary to perform sufficient data desensitization processing to reduce the data sensitivity level.
[0003] In terms of customer information processing, a mainstream method currently used by banks is to assign a unique identification code to a single customer in the ECIF system. When interacting with external third-party institutions, this code is used for customer identity recognition. In this way, only the information recognition at the "person" dimension of the customer can be satisfied, and it cannot be refined to the card level, and it cannot be satisfied in many scenarios that require bank card interaction. The essential reason is that the ECIF system is based on the customer life cycle management of enterprise customer information and cannot penetrate into specific businesses to elaborate on the relationship between "person" and "card". In addition, in terms of data desensitization, common algorithms such as replacement, encryption, hashing, and obfuscation can all complete data desensitization and interaction in specific scenarios, but there are deficiencies in the integrity of the data desensitization link.
[0004] Currently, there are various differences in the scenarios of cooperation between banks and third-party institutions. Relying on single-user identification and traditional desensitization methods cannot meet the requirements or requires customized development. However, customized services not only increase the operating costs of enterprises but also cause great difficulties in the development and operation and maintenance of service providers. Summary of the Invention
[0005] In order to achieve efficient information desensitization processing, this application provides a method for desensitizing personal information data.
[0006] The method for desensitizing personal information data provided by this application adopts the following technical solutions: A method for desensitizing personal information data, confirming the file structure and configuring the data parsing model, and setting the source file template and the output file adaptation template; Reading files in multiple threads, including obtaining the basic information of the file, obtaining the total number of bytes of the file, the total number of lines of the file, and the maximum number of threads; and calculating the number of bytes processed by a single thread to determine the start position and end position of the bytes for each thread; Data parsing and desensitization processing, with multiple threads performing data parsing and desensitization processing.
[0007] By adopting the above technical solutions, the processing range of each thread is dynamically allocated according to the total number of bytes, total number of lines, and maximum number of threads of the file, optimizing resource utilization. Through multi-threaded parallel processing, the processing efficiency of large-scale data files is significantly improved; multiple file formats (such as fixed-length format or delimiter format) are supported to meet the requirements of different data sources; through data parsing and desensitization processing, the security and privacy protection of sensitive information are ensured.
[0008] Preferably, determining the start position and end position of each thread includes the following steps: Roughly calculate the end position of each thread; the rough end position is equal to the total number of bytes of the file divided by the maximum number of threads; Precisely calculate the start position and end position of each group of threads. Let the thread coordinates be t(x,y), where x is the start coordinate and y is the end position. Then the corresponding x and y coordinate values of f(t) are ((t)b, (t + 1)b). t starts from 0 by default, max(t) = C, and t + 1 is less than C; Where C is the number of threads, and C is set to twice the number of hardware cores; b is the rough processing length of each thread, and b = Z / C, where Z is the total number of bytes of the file.
[0009] By adopting the above technical solutions, through the combination of rough calculation and precise calculation, it is ensured that the data block range processed by each thread is accurate, achieving the effects of optimizing thread task allocation, reducing calculation overhead, and improving overall processing efficiency.
[0010] Preferably, the number of bytes processed by a single thread is equal to the average number of bytes processed by each thread plus the remaining bytes of the data in the line where the file avg bit is located. The average number of bytes processed by each thread is equal to the total number of bytes of the file divided by the maximum number of threads.
[0011] By adopting the above technical solutions, by calculating the average number of bytes processed by each thread and considering the integrity of the file line data, it is ensured that the task volume of each thread is balanced, avoiding data breakage or loss caused by improper data block segmentation.
[0012] Preferably, the multi-threaded file reading further includes verifying the integrity of each thread, including: ; In F(Cxs), F is the function name and Cxs is the independent variable. Here, C is the total number of threads, xs is the thread number of the current position to be calculated, that is, Cxs is the corrected value after integrity verification of the thread end position xs under the constraint of the total number of threads C; xs / (NIO( / n)) is the calculated file processing position allocated to each thread; NIO( / n) is the delimiter of the data line or data block.
[0013] By adopting the above technical solution, by verifying whether the end position of each thread is a complete line or data block, data breakage is prevented, and the accuracy and consistency of data processing in a multi-threaded environment are ensured.
[0014] Preferably, it further includes the length t(x) of the fuzzy computing thread processing unit group, ; wherein, t(buf) is the size of the buffer, x is the number of threads, and (Floor) is the floor function.
[0015] By adopting the above technical solution, the size of the data block processed by each thread is dynamically adjusted according to the buffer size and the number of threads, optimizing resource utilization.
[0016] Preferably, it further includes an accurate positioning end flag, and the accurate positioning end flag process includes reading byte blocks in a fixed size after completing file stream positioning and verifying each end identifier in the byte block one by one until a valid end identifier is found.
[0017] By adopting the above technical solution, by reading byte blocks in a fixed size and verifying the end identifier, it is ensured that the data block processed by each thread is complete, reducing redundant operations of data reading and improving processing efficiency.
[0018] Preferably, the start position of the Nth subsequent thread is set to one more than the end position of its previous thread, i.e., the (N - 1)th thread.
[0019] By adopting the above technical solution, by setting the start position of the subsequent thread to one more than the end position of the previous thread, the continuity of data processing is ensured, avoiding data omission or duplicate processing.
[0020] Preferably, the data parsing and desensitization processing includes the following steps: Loading the template configuration in the data parsing model when the application starts, and parsing the fields to be desensitized in each line of data according to the template configuration; Generating ciphertext data through a hybrid obfuscation algorithm; Storing the plaintext and the ciphertext data using key-value pairs; Distributed caching of the plaintext and ciphertext data.
[0021] By adopting the above technical solution, through template-based configuration, the automation of data parsing is realized; a hybrid obfuscation algorithm is used to desensitize sensitive fields to ensure data security; the plaintext and ciphertext data are stored using key-value pairs to support fast query and verification; through distributed caching, high-concurrency processing in a multi-threaded environment is supported.
[0022] Preferably, it also includes thread safety and exception handling: The AtomicLong in an efficient lock-free manner is used to implement thread-safe integer operations. It is judged whether the AtomicLong readCount for reading the number of file lines is equal to the AtomicLong writeCount for writing the number of lines. If they are equal, the file processing is completed normally; if not, the file processing is abnormal.
[0023] By adopting the above technical solution, a safe operation in a multi-threaded environment is realized through a lock-free mechanism, avoiding data competition; by comparing the number of read lines and the number of written lines, it is quickly detected whether the data processing is complete, so that when an abnormality is found, the exception handling process can be entered in a timely manner to ensure the accuracy and integrity of data processing.
[0024] In summary, the present application includes at least one of the following beneficial technical effects: 1. Significantly improve data processing efficiency through multi-threaded parallel processing and dynamic task allocation; 2. Adopt a hybrid desensitization algorithm and a distributed cache mechanism to ensure the security and consistency of sensitive information; 3. Ensure the accuracy and integrity of data processing through a thread safety mechanism and exception handling; 4. Support multiple file formats and dynamic resource allocation to adapt to different scales of data processing requirements. Brief Description of the Drawings
[0025] Figure 1 is the flowchart of the implementation of the present application; Figure 2 is the flowchart of step S1; Figure 3 is the flowchart of step S2; Figure 4 is the flowchart of step S3. Detailed Description of the Embodiment
[0026] The following further describes the present application in detail Figure 1 with reference to the attached drawings.
[0027] The embodiment of the present application discloses a method for desensitizing personal information data.
[0028] Referring to Figure 1 , a method for desensitizing personal information data includes the following steps: S1: Confirm the file structure and configure the data parsing model.
[0029] S2: Read the file in multiple threads.
[0030] S3: Data parsing and desensitization processing.
[0031] S4: Thread Safety and Exception Handling.
[0032] Refer to Figure 2 , S1: Confirm the file structure and configuration data parsing model, including the following steps: S11: File structure confirmation. Confirm the structural characteristics of each line of data in the file. The file can be in fixed-length format or delimiter format (such as separated by "|"). To parse each field in each line of data, a set of data parsing models needs to be configured.
[0033] S12: Source file template configuration: Configure the template of the source file, specifying the delimiter for each line of data and the attributes of each field. For example: FILE::SPLIT[|]::[a:str:4][b:str:4][c:str:6][papers:str:19][card:str:19][phone:str:12][sum:str:5]; Among them, "SPLIT[|]" indicates that the delimiter for each line of data in the source file is "|", and the content in [] is the corresponding field attributes for each line of data. Taking [a:str:4] as an example, a is the field name, str means padding with spaces on the right (long means padding with 0 on the left, decimal means supplementing decimal places), and 4 is the field length.
[0034] S13: Output file adaptation template configuration: FILE::SPLIT[|]::[a:str:4][b:str:4][c:str:6][papers:str:19][card:str:19][phone:str:12][sum:str:5]; Among them, "SPLIT[|+|]" indicates that the delimiter for each line of data in the output file is "|+|", and the content in [] is the corresponding field attributes and desensitization configuration for each line of data. Taking [papers:str:19:Mapping(PAPERS)] as an example, papers represents the keyword for the ID number, and Mapping(PAPERS) means desensitizing the sensitive data ID number in this line of data.
[0035] Refer to Figure 3 , S2: Read the file in multiple threads, including the following steps: S21: Obtain the basic information of the file, including the total number of bytes sum of the file, the total number of lines count of the file, and the maximum number of threads maxThreadNum. Assume that the total number of bytes sum of the file = 100000, the total number of lines count of the file = 1000, and the maximum number of threads maxThreadNum = 10.
[0036] S22: Calculate the number of bytes processed by a single thread. According to the formula: avg = sum / maxThreadNum, obtain the average number of bytes processed by each thread. The number of bytes processed by a single thread is equal to the average number of bytes processed by each thread plus the remaining bytes of the data in the line where the file avg bit is located, that is, obtain the start and end positions of the bytes processed by each thread.
[0037] S23: Determine the start and end positions of the bytes for each thread.
[0038] Roughly calculate the end position. The rough end position is equal to the total number of bytes of the file divided by the maximum number of threads, that is, int i = sum / maxThreadNum; Fuzzily calculate the length of the processing unit group of the first thread group. The default value of the t(x) function is t(core), that is, determine the size of the data block processed by each thread according to the number of CPU cores. For example, x0(0, t(x)), x0 is a thread group, indicating starting from position 0 and processing a data block with a length of t(x).
[0039] Precisely calculate the start and end positions of each group of threads. Let the thread coordinates be t(x, y), where x is the start coordinate and y is the end position. Then the corresponding x and y coordinate values of f(t) are ((t)b, (t + 1)b). t starts from 0 by default, max(t) = C, and t + 1 is less than C; where C is the number of threads, and C is set to twice the number of hardware cores; b is the rough processing length of each thread, and b = Z / C, where Z is the total number of bytes of the file.
[0040] In each thread, the data block range corresponding to x(0) to x is x(0)(x0a, xb); the data block range corresponding to x to xs is x(xb, xsb); the data block range corresponding to xs to F(Cxs) is xs(x:xsb, xsb); the data block range corresponding to F(Cxs) to max(CoreF(NIO)) is Cxs(xs:xsb, Cxsb).
[0041] Round up to calculate the size of the byte block processed by each group of threads in t(x). The relevant formula is: ; where t(buf) is the size of the buffer, x is the number of threads, and (Floor) is the floor function. By dividing the buffer size by x * 2 and taking the integer, the size of the data block processed by each thread in t(x) is obtained.
[0042] S24: Verify the integrity of each thread. Each thread calculates whether the current end position is a complete line or complete data to prevent data fragmentation. The relevant formula is: ; Among them, F(Cxs) is used to ensure that the end position processed by each thread is a complete data line or data block, preventing data breakage. C is the total number of threads, xs is the end position of the current thread, and Cxs is the corrected value after integrity verification of the thread end position xs under the constraint of the total number of threads C. xs / (NIO( / n)) is the file processing position calculated by each allocated thread, and NIO( / n) is the delimiter of the data line or data block (such as the newline character \n). If xs is an integer multiple of NIO( / n), then Cxs is taken; otherwise, xs is adjusted to the end position of the next complete data line or data block.
[0043] Then, read in 1024-bit chunks and verify 1024 bytes. When the end marker is verified, make a mark (such as marking it as e); Next, each group of threads obtains the corresponding start marker based on the calculation results of the previous group of threads. That is, the start position of the n+1 thread is the end position of the n thread + 1.
[0044] Finally, the byte ranges of each thread as shown in the following list can be obtained: thread1->0-9000 (byte) thread2->9001-18012 (byte) thread3->18013-27453 (byte) ...... Refer to Figure 4 , S3, data parsing and desensitization processing, including the following steps: S31: When the application starts, load the template configuration in the data parsing model and parse the fields specified for desensitization in each line of data. Through template-based configuration, automate data parsing, reduce manual intervention, and improve processing efficiency. Support multiple file formats (such as fixed-length format or delimiter format) to adapt to different data source requirements. Ensure that the parsing of each field is strictly executed according to the template definition to avoid data parsing errors.
[0045] S32: Generate ciphertext data through a desensitization algorithm.
[0046] Adopt a hybrid obfuscation algorithm, such as triple DES combined with the date format $Y$m$d, to desensitize sensitive fields. For example, after desensitizing the ID number field "papers", generate ciphertext data. Desensitize sensitive information through complex encryption algorithms (such as triple DES) to prevent data leakage. Support generating ciphertext in combination with dynamic factors such as date format to enhance the randomness and irreversibility of desensitization. While ensuring security, the algorithm design takes performance into account and is suitable for large-scale data processing.
[0047] S33: Plaintext and ciphertext data storage.
[0048] Use a key-value pair map to store plaintext and ciphertext data in the format <bcp:p:plain_txt cipher_txt> and <bcp:c:cipher_txt plain_txt>.
[0049] Among them, plain_txt is the plaintext data and cipher_txt is the ciphertext data.
[0050] By storing plaintext and ciphertext data through key-value pairs, it is possible to quickly query and verify the mapping relationship between plaintext and ciphertext, ensure the one-to-one correspondence between plaintext and ciphertext, avoid data loss or misalignment, and support subsequent further processing and analysis of plaintext and ciphertext data.
[0051] S34: Distributed caching of plaintext and ciphertext data.
[0052] Use Redis to implement a distributed lock to determine whether the plaintext-ciphertext mapping relationship exists. If it exists, do not insert; if it does not exist, perform a batch hset insertion into the cache. Through the distributed caching and lock mechanism, it supports high-concurrency processing in a multi-threaded environment and avoids data conflicts. Ensure the uniqueness of the plaintext-ciphertext mapping relationship to prevent duplicate insertions or data redundancy. Through batch insertion operations, reduce the number of cache operations and improve system performance.
[0053] S4: Thread safety and exception handling.
[0054] Use AtomicLong in an efficient lock-free manner to implement thread-safe integer operations. Through the lock-free mechanism, achieve safe operations in a multi-threaded environment and avoid data competition and lock overhead. After the file processing is completed, determine whether the number of lines read from the file AtomicLong readCount is equal to the number of lines written AtomicLong writeCount to quickly detect whether the data processing is complete and ensure the accuracy of data processing. If they are equal, the file is processed normally; if they are not equal, the file processing is abnormal and enters exception handling to avoid data loss or error diffusion.
[0055] The implementation principle of a personal information data desensitization processing method in an embodiment of this application is as follows: Read the file in parallel through multiple threads, divide the large file into multiple data blocks, and each thread is responsible for processing one data block. The total number of bytes and lines of the file are pre-calculated, and combined with the maximum number of threads, the byte range processed by each thread is dynamically allocated. By accurately calculating the start and end positions of each thread, ensure that the data block processed by each thread is a complete line or data block and avoid data fragmentation. Specifically, adjust the end position of the thread by verifying whether the end position is a complete line to ensure data integrity.
[0056] The above are all preferred embodiments of this application, and do not limit the protection scope of this application. Therefore, all equivalent changes made according to the structure, shape, and principle of this application shall be covered within the protection scope of this application.
Claims
1. A method for desensitizing personal information data, characterized in that: Including: Confirm the file structure and configuration data parsing model, and set the source file template and output file adaptation template; Read files in multiple threads, including obtaining basic file information, obtaining the total number of bytes of the file, the total number of lines of the file, and the maximum number of threads; and calculating the number of bytes processed by a single thread, and determining the start position and end position of the bytes for each said thread; Data parsing and desensitization processing, and multiple said threads perform data parsing and desensitization processing.
2. The personal information data desensitization processing method according to claim 1, characterized in that: Determining the start position and end position of the bytes for each thread includes the following steps: Roughly calculate the end position of each thread; the rough end position is equal to the total number of bytes of the file divided by the maximum number of threads; Precisely calculate the start position and end position of each group of threads. Let the thread coordinates be t(x,y), where x is the start coordinate and y is the end position. Then the corresponding x and y coordinate values of f(t) are ((t)b, (t + 1)b), t defaults to start from 0, max(t)=C, and t + 1 is less than C; Where C is the number of threads, and C is set to twice the number of hardware cores; b is the rough processing length of each thread, and b = Z / C, where Z is the total number of bytes of the file.
3. A method for desensitizing personal information data according to claim 1, characterized in that: The number of bytes processed by a single thread is equal to the average number of bytes processed by each thread plus the number of remaining bytes in the data line where the file avg bit is located. The average number of bytes processed by each thread is equal to the total number of bytes of the file divided by the maximum number of threads.
4. A method for desensitizing personal information data according to claim 1, characterized in that: The multi-threaded file reading also includes verifying the integrity of each thread, including: ; In F(Cxs), F is the function name and Cxs is the independent variable. Here, C is the total number of threads, and xs is the thread number of the position to be calculated currently. That is, Cxs is the corrected value after the integrity check of the end position of the thread xs under the constraint of the total number of threads C; xs / (NIO( / n)) is the file processing position allocated to each thread; NIO( / n) is the delimiter of the data line or data block.
5. A method for desensitizing personal information data according to claim 1, characterized in that: It also includes the length t(x) of the fuzzy computing thread processing unit group, ; Where t(buf) is the size of the buffer, x is the number of threads, and (Floor) is the floor function.
6. The personal information data desensitization processing method according to claim 1, wherein: It also includes precisely positioning the end flag. The process of precisely positioning the end flag includes reading byte blocks in a fixed size after completing the file stream positioning and verifying each end identifier in the byte block one by one until a valid end identifier is found.
7. A personal information data desensitization processing method according to claim 6, characterized in that: The start position of the Nth subsequent thread is set to one more than the end position of its previous thread, that is, the (N - 1)th thread.
8. A method for desensitizing personal information data according to claim 1, characterized in that: The data parsing and desensitization processing includes the following steps: Load the template configuration in the data parsing model when the application starts, and parse the fields specified for desensitization in each line of data according to the template configuration; Generate ciphertext data through a hybrid obfuscation algorithm; Use key-value pairs to store the plaintext and the ciphertext data; Distributed cache the plaintext and ciphertext data.
9. A personal information data desensitization processing method according to claim 1, characterized in that: It also includes thread safety and exception handling: Use AtomicLong in an efficient lock-free manner to implement thread-safe integer operations. Judge whether the number of lines read by AtomicLong readCount is equal to the number of lines written by AtomicLong writeCount. If they are equal, the file is normally processed and completed; if not, the file processing is abnormal.
Citation Information
Patent Citations
Post-processing method and device for data desensitization interruption of a file
CN113961968A
Multi-thread file analysis method and device
CN114238213A
Data parallel desensitization processing method
CN116561795A
Data backup method, system and device and medium
CN119739561A
Systems and methods for ingesting data files using multi-threaded processing
US20220283986A1