Incremental Desensitization Method and System for Heterogeneous Data Sources

Structured feature vectors are generated through stream sampling and dynamic feature analysis, and combined with self-balancing binary search tree index and non-cryptographic hashing, the applicability and efficiency of incremental desensitization of heterogeneous data sources is solved, and efficient and accurate data processing is achieved.

CN120124106BActive Publication Date: 2025-07-11HANGZHOU ANQUAN DIGITAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510600666.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-11
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Existing data desensitization technologies face problems such as limited scope of application, poor adaptability, performance bottlenecks and low computing efficiency when dealing with heterogeneous data sources, especially in terms of incremental desensitization.

Method used

Stream sampling and dynamic feature analysis are used to generate structured feature vectors, combined with the self-balanced binary search tree index strategy adjusted by dynamic window and the fingerprint channel of non-cryptographic hash and semantic weights, and data checksum incremental desensitization is performed through the probability bit mapping matrix.

Benefits of technology

It realizes efficient incremental desensitization of heterogeneous data sources, and can automatically identify and process various types of data without relying on the data source structure, improving processing efficiency and accuracy, and reducing misjudgment rate and system power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124106B_ABST
    Figure CN120124106B_ABST
Patent Text Reader

Abstract

The embodiments of this specification disclose an incremental desensitization method and system for heterogeneous data sources. Among them, an incremental desensitization method for heterogeneous data sources includes obtaining the data to be desensitized from the heterogeneous data source; generating a structured feature vector containing different types of features for the data to be desensitized from the heterogeneous data source through streaming sampling and dynamic feature analysis; dynamically selecting a numerical channel or a fingerprint channel based on the structured feature vector to generate a composite feature fingerprint corresponding to the data to be desensitized; mapping the composite feature fingerprint to a probabilistic bit mapping matrix, and performing verification through a bit operation instruction set, marking the records that are not mapped to the probabilistic bit mapping matrix as incremental data, and performing desensitization on the incremental data. The embodiments of this specification significantly reduce the misjudgment rate while improving the desensitization accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of data security technology, and more particularly to an incremental desensitization method and system for heterogeneous data sources. Background Art

[0002] With the rapid development of information technology, data security and privacy protection have become crucial issues. Among many data security technologies, data desensitization technology has attracted much attention because it can meet the data usage requirements while protecting data privacy. Data desensitization refers to the processing of sensitive data so that it can still be used in scenarios such as analysis, testing, and sharing without revealing privacy. However, existing data desensitization technologies face many challenges when dealing with heterogeneous data sources, especially in incremental desensitization. For example, most existing data desensitization technologies rely on the primary key or unique key constraint of the database to identify and process data, and this dependence limits the scope of application of the desensitization technology. For data tables without a clear primary key or unique key, it is difficult to guarantee the desensitization effect. Another example is that most existing desensitization technologies are designed for specific types of data sources (such as relational databases), and their adaptability to heterogeneous data sources (such as non-relational databases, file systems, big data platforms, etc.) is poor. Existing desensitization technologies often face performance bottlenecks when dealing with large-scale data, and the misjudgment rate increases. Moreover, when dealing with complex data structures, the calculation efficiency is low, affecting the overall performance of the system. Therefore, there is an urgent need for an incremental desensitization method for heterogeneous data sources that can adapt to heterogeneous data sources and efficiently process incremental data to meet the growing needs of data security and privacy protection. Summary of the Invention

[0003] Embodiments of this specification provide an incremental desensitization method and system for heterogeneous data sources, and the technical solutions are as follows:

[0004] In a first aspect, embodiments of this specification provide an incremental desensitization method for heterogeneous data sources, including: obtaining the data to be desensitized from a heterogeneous data source; generating a structured feature vector containing different types of features for the data to be desensitized from the heterogeneous data source through streaming sampling and dynamic feature analysis; dynamically selecting a numerical channel or a fingerprint channel based on the structured feature vector to generate a composite feature fingerprint corresponding to the data to be desensitized, where the numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, and the fingerprint channel fuses non-encrypted hash and semantic weights; mapping the composite feature fingerprint to a probabilistic bit mapping matrix, and performing verification through a bit operation instruction set, marking the records that are not mapped to the probabilistic bit mapping matrix as incremental data, and performing desensitization on the incremental data.

[0005] Second aspect, embodiments of this specification provide an incremental desensitization system for heterogeneous data sources, including: a data access module for obtaining data to be desensitized from heterogeneous data sources; a vector generation module for generating structured feature vectors containing different types of features for the data to be desensitized from heterogeneous data sources through streaming sampling and dynamic feature analysis; a channel selection module for dynamically selecting a numerical channel or a fingerprint channel based on the structured feature vectors to generate a composite feature fingerprint corresponding to the data to be desensitized, where the numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, and the fingerprint channel fuses non-encrypted hashing and semantic weights; a matrix mapping module for mapping the composite feature fingerprint to a probabilistic bit mapping matrix and performing verification through a bit operation instruction set, marking records that are not mapped to the probabilistic bit mapping matrix as incremental data, and performing desensitization on the incremental data.

[0006] The beneficial effects brought by the technical solutions provided by some embodiments of this specification at least include:

[0007] Embodiments of this specification can obtain data to be desensitized from heterogeneous data sources, then generate structured feature vectors containing different types of features for the data to be desensitized from heterogeneous data sources through streaming sampling and dynamic feature analysis, then dynamically select a numerical channel or a fingerprint channel based on the structured feature vectors to generate a composite feature fingerprint corresponding to the data to be desensitized, then map it into a probabilistic bit mapping matrix, and perform verification through a bit operation instruction set, marking records that are not mapped to the probabilistic bit mapping matrix as incremental data, and performing desensitization on the incremental data. Embodiments of this specification provide a data desensitization technology that can adapt to heterogeneous data sources, efficiently process incremental data, and ensure data integrity. Without relying on the data source structure, it can automatically identify and process various types of data without the need to know the data table structure in advance or rely on primary key / unique key constraints; in addition, embodiments of this specification have efficient incremental desensitization capabilities, can accurately identify the differences between new data and processed data, and only desensitize new data, improving processing efficiency; embodiments of this specification have a low misjudgment rate. The numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, and the fingerprint channel fuses non-encrypted hashing and semantic weights, which can not only significantly reduce the misjudgment rate, improve desensitization accuracy, but also improve data processing speed and reduce system power consumption at the same time. Description of the Drawings

[0008] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0009] Figure 1It is a schematic diagram of the application scenario of an incremental desensitization method for heterogeneous data sources provided in this specification.

[0010] Figure 2 It is a schematic diagram of the process of an incremental desensitization method for heterogeneous data sources provided in this specification.

[0011] Figure 3 It is a schematic diagram of the process of generating a structured feature vector containing different types of features provided in this specification.

[0012] Figure 4 It is a schematic diagram of the process of generating a composite feature fingerprint provided in this specification.

[0013] Figure 5 It is a schematic diagram of the process of mapping the composite feature fingerprint to a probabilistic bit mapping matrix provided in this specification.

[0014] Figure 6 It is a schematic diagram of the structure of an incremental desensitization system for heterogeneous data sources provided in this specification.

[0015] Figure 7 It is a schematic diagram of the structure of an electronic device provided in this specification. Specific embodiments

[0016] Next, the technical solutions in the embodiments of this specification will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this specification.

[0017] Terms such as "first", "second", etc. in the specification, claims and the above-mentioned drawings of this specification are used to distinguish different objects, rather than to describe a specific order. In addition, the term "including" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0018] An incremental desensitization method for heterogeneous data sources provided in multiple embodiments of this specification, and the execution subject of the incremental desensitization method for heterogeneous data sources can be an incremental desensitization system for heterogeneous data sources provided in the embodiments of the present invention.

[0019] Before this specification elaborates on an incremental desensitization method for heterogeneous data sources in combination with one or more embodiments in detail, the application scenarios of the incremental desensitization method for heterogeneous data sources will be introduced first.

[0020] Please refer to Figure 1 , Figure 1The figure is a schematic diagram of the application scenario of an incremental desensitization method for heterogeneous data sources provided by an embodiment of the present invention. In this embodiment, an incremental desensitization system 100 for heterogeneous data sources may include a server 110 and a number of terminals 120, and the number of terminals 120 are respectively communicatively connected to the server 110.

[0021] In this embodiment of the specification, the terminal 120 may be a device such as a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC). The terminal 120 may send data from different types of data sources (such as non-relational databases, file systems, big data platforms, etc.) to the server 110.

[0022] The server 110 in this embodiment of the specification may be a single server or a server cluster composed of multiple servers, and a method for incremental desensitization of heterogeneous data sources of the present application is implemented by multiple servers.

[0023] The server 110 in this embodiment of the specification may include a data access module, a vector generation module, a channel selection module, a matrix mapping module, etc. The server 110 in this embodiment of the specification may first obtain the data to be desensitized of the heterogeneous data source; then, through streaming sampling and dynamic feature analysis, generate a structured feature vector containing different types of features for the data to be desensitized of the heterogeneous data source; then dynamically select a numerical channel or a fingerprint channel based on the structured feature vector to generate a composite feature fingerprint corresponding to the data to be desensitized. The numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, and the fingerprint channel fuses non-encrypted hashing and semantic weights; then map the composite feature fingerprint to a probabilistic bit mapping matrix, and perform verification through a bit operation instruction set, mark the records not mapped to the probabilistic bit mapping matrix as incremental data, and perform desensitization on the incremental data.

[0024] It should be noted that Figure 1 The schematic diagram of the scenario of an incremental desensitization system 100 for heterogeneous data sources shown is only an example. The incremental desensitization system and scenario for heterogeneous data sources described in the embodiments of the present invention are for more clearly explaining the technical solutions of the embodiments of the present invention, and do not constitute a limitation to the technical solutions provided by the embodiments of the present invention. Those skilled in the art know that with the evolution of an incremental desensitization system for heterogeneous data sources and the emergence of new scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0025] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an incremental desensitization method for heterogeneous data sources provided by an embodiment of the present invention. This incremental desensitization method for heterogeneous data sources may be performed by Figure 1An incremental desensitization system 100 for heterogeneous data sources as shown is executed. The incremental desensitization method for heterogeneous data sources may at least include the following steps:

[0026] 200. Obtain the data to be desensitized from the heterogeneous data source.

[0027] In this embodiment, the heterogeneous data source may be data sources of different types, such as non-relational databases, file systems, big data platforms, etc. The data to be desensitized from the heterogeneous data source may be data from different data sources, with different formats, structures, and characteristics. These data may differ in aspects such as data format, data structure, data source, and data characteristics.

[0028] For example, regarding the differences in data format, the data may be structured (such as tabular data in a relational database), semi-structured (such as JSON, XML files), or unstructured (such as text files, pictures, audio, etc.). For another example, regarding the differences in data structure, even if the data is structured, the table structures of different data sources may be different (for example, a table in one data source may contain fields such as user ID, name, and age, while a table in another data source may contain fields such as user ID, email, and phone number). For another example, regarding the differences in data source, the data may come from different systems, platforms, or organizations (for example, the data of an enterprise may come from multiple different systems such as its customer relationship management system (CRM), financial system, website logs, etc.). For another example, regarding the differences in data characteristics, the characteristics of the data (such as the distribution, entropy value, pattern, etc.) may be different (for example, some data may contain a large amount of text information, while other data may mainly contain numerical information).

[0029] In the embodiment of this specification, the data to be desensitized from the heterogeneous data source may be data of different types of data sources (such as non-relational databases, file systems, big data platforms, etc.) sent by the terminal 120 to the server 110, and these data are the data to be desensitized by the server 110.

[0030] 210. Through streaming sampling and dynamic feature analysis, generate a structured feature vector containing different types of features for the data to be desensitized from the heterogeneous data source.

[0031] In this embodiment, streaming sampling can be a sampling method for processing data streams in real-time or near real-time. Instead of waiting for the full dataset to be loaded, samples are dynamically extracted in the order of data arrival, and the representativeness of the sampling results is ensured. Streaming sampling can include incremental processing, dynamic adjustment, and lightweight processing, etc. Incremental processing can adapt to high-throughput data streams (such as logs and sensor data), avoiding the memory pressure caused by full-scale loading. Dynamic adjustment can automatically adjust the sampling rate according to changes in data distribution (such as new fields and data skew) (for example, downsampling high-frequency fields). Lightweight processing can sample in combination with time windows (such as sampling 100 records every 5 seconds).

[0032] In this embodiment, dynamic feature analysis can identify feature patterns (such as data types, distribution rules, and sensitivity levels) in data in real-time and dynamically generate structured feature vectors. Dynamic feature analysis can include feature recognition, sensitivity assessment, and vectorization representation, etc. Feature recognition can automatically detect field types (such as text, numeric, and date) and semantics (such as name, address, and ID number). Sensitivity assessment can use machine learning models to label sensitive fields (for example, mobile phone numbers need to be desensitized, while digital IDs can be retained). Vectorization representation can encode features as structured vectors (for example: [field name: "phone", type: "text", sensitivity: 0.9, desensitization method: "mask"]).

[0033] Embodiments of this specification can generate structured feature vectors containing different types of features for the data to be desensitized from heterogeneous data sources through streaming sampling and dynamic feature analysis. Different types of features can include numeric features, text features, etc. Numeric features can include distribution features, range features, incremental features, and density features, etc.; text features include entropy value features, pattern features, semantic features, and length features, etc.

[0034] In some embodiments, please refer to Figure 3 , Figure 3 is a schematic flowchart of the process for generating structured feature vectors containing different types of features provided by an embodiment of the present invention. Through streaming sampling and dynamic feature analysis, structured feature vectors containing different types of features are generated for the data to be desensitized from heterogeneous data sources, including:[[]]

[0035] 2100. Determine the field data type of the data to be desensitized;

[0036] 2102. Determine the standard deviation fluctuation value or entropy change rate corresponding to the data to be desensitized;

[0037] 2104. Based on the field data type, adjust the sampling depth according to the standard deviation fluctuation value or entropy change rate;

[0038] 2106. Generate a structured feature vector containing different types of features for the data to be desensitized from heterogeneous data sources based on the sampling depth.

[0039] In this embodiment, the field data types of the data to be desensitized may include numerical fields or text fields, etc. The sampling depth may be the scale of data samples extracted for analyzing field features, reflecting the dynamic adjustment ability of the sampling volume. Embodiments of this specification can generate a structured feature vector containing different types of features for the data to be desensitized from heterogeneous data sources based on the sampling depth.

[0040] In some embodiments, based on the field data type, adjust the sampling depth according to the standard deviation fluctuation value or the entropy value change rate, including: when the field data type is a numerical field and the standard deviation fluctuation value is greater than the preset fluctuation threshold, trigger incremental sampling; when the field data type is a text field and the entropy value change rate is greater than the preset change threshold, trigger supplementary sampling.

[0041] In this embodiment, the standard deviation fluctuation value may be when the field data type of the data to be desensitized is a numerical field, based on a time window, extract multiple sets of continuous data from the data to be desensitized to calculate statistical features (such as mean, standard deviation, etc.), obtain a series of standard deviations corresponding to the multiple sets of continuous data, and then determine the standard deviation fluctuation value according to the series of standard deviations. If the calculated standard deviation fluctuation value is greater than the preset fluctuation threshold, it indicates that the data distribution is unstable, and the sampling depth needs to be adjusted to improve the analysis reliability by increasing the sampling volume.

[0042] In this embodiment, the entropy value change rate may be when the field data type of the data to be desensitized is a text field, based on a time window, extract multiple sets of continuous data from the data to be desensitized to calculate the entropy value features corresponding to the text field, that is, calculate the information entropy value based on the Shannon entropy model, obtain a series of entropy values corresponding to the multiple sets of continuous data, and then determine the entropy value change rate according to the series of entropy values. When the entropy value change rate is greater than the preset change threshold, it indicates that the text information distribution has mutated (such as from a standardized address to free text), and supplementary sampling is required to cover the new pattern. Then, embodiments of this specification can generate a structured feature vector containing different types of features for the data to be desensitized from heterogeneous data sources based on the sampling depth.

[0043] Embodiments of this specification provide a dynamic streaming sampling mechanism for adaptive sample volume control, which can automatically adjust the sampling depth according to the complexity of the field data type. For example, regarding numerical fields: the basic sampling volume is 500 records, and incremental sampling is triggered when the standard deviation fluctuation value > 15%; another example, regarding text fields: the basic sampling volume is 1000 records, and supplementary sampling is started when the entropy value change rate > 10%.

[0044] In some embodiments, based on the sampling depth, a structured feature vector containing different types of features is generated for the data to be desensitized from heterogeneous data sources, including: when the field data type is a numerical field, obtaining the numerical features corresponding to the data to be desensitized, where the numerical features include distribution features, range features, incremental features, and density features; when the field data type is a text field, obtaining the text features corresponding to the data to be desensitized, where the text features include entropy value features, pattern features, semantic features, and length features; generating a structured feature vector containing different types of features according to the numerical features or text features corresponding to the data to be desensitized.

[0045] The embodiments of this specification design a multi-dimensional feature analysis system. From the analysis dimension of numerical fields, the embodiments of this specification can obtain the numerical features corresponding to the data to be desensitized, and the numerical features include distribution features, range features, incremental features, density features, etc. Among them, distribution features: calculate the mean, standard deviation, skewness, and kurtosis to construct the data distribution profile; range features: dynamically track the minimum value (Min), maximum value (Max), and value range span (Range); incremental features: analyze the numerical change trend of adjacent records and calculate the incremental stability index; density features: count the data distribution density within the unit interval to identify the aggregation section. From the analysis dimension of text fields, the embodiments of this specification can obtain the text features corresponding to the data to be desensitized, and the text features include entropy value features, pattern features, semantic features, and length features, etc. Among them, entropy value features: calculate the information entropy value based on the Shannon entropy model to quantify the data chaos degree. Pattern features: identify preset patterns such as dates, IP addresses, and ID numbers through a regular expression matching engine. Semantic features: apply a lightweight word embedding model to extract the semantic vector representation of the text field. Length features: count the mean, extreme values, and distribution rules of the string length.

[0046] 220. Dynamically select a numerical channel or a fingerprint channel based on the structured feature vector to generate a composite feature fingerprint corresponding to the data to be desensitized. The numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, and the fingerprint channel fuses non-encrypted hashing and semantic weights.

[0047] The embodiments of this specification dynamically select a numerical channel or a fingerprint channel based on the structured feature vector, which can match the best desensitization method according to the field features and reduce information loss. For structured numerical data (such as amounts, IDs), the embodiments of this specification can generate a composite feature fingerprint corresponding to the data to be desensitized through the numerical channel; for unstructured or high-entropy data (such as text, hash values, transaction records), the embodiments of this specification can generate a unique identifier (such as Murmur hash) to retain comparability but hide the original value.

[0048] In some embodiments, please refer to Figure 4 , Figure 4It is a schematic flowchart of generating a composite feature fingerprint provided by an embodiment of the present invention. Based on the structured feature vector, a numerical channel or a fingerprint channel is dynamically selected to generate a composite feature fingerprint corresponding to the data to be desensitized, including:

[0049] 2200. Determine channel selection parameters according to the structured feature vector, where the channel selection parameters include numerical continuity parameter, incremental density standard deviation, text entropy value, and semantic dispersion;

[0050] 2210. When the numerical continuity parameter is greater than the continuous parameter threshold and the incremental density standard deviation is less than the standard deviation threshold, select the numerical channel to generate a composite feature fingerprint corresponding to the data to be desensitized;

[0051] 2220. When the text entropy value is greater than the entropy threshold and the semantic dispersion is greater than the dispersion threshold, select the fingerprint channel to generate a composite feature fingerprint corresponding to the data to be desensitized.

[0052] In this embodiment, the composite feature fingerprint can be a unique identifier generated by fusing multiple feature dimensions. By generating the composite feature fingerprint, this embodiment can help the system distinguish new and old data and achieve incremental desensitization; by fusing multiple features, the misjudgment rate can be significantly reduced and the data processing efficiency can be improved.

[0053] In this embodiment, the numerical continuity parameter is used to measure whether the changes between adjacent records of a numerical field show stable and predictable progressive characteristics, and is applicable to ordered data such as timestamps and auto-incrementing IDs. For a sequence of field values, embodiments of this specification can determine the numerical continuity parameter by calculating the absolute difference or relative change rate of adjacent data.

[0054] In this embodiment, the incremental density standard deviation is used to illustrate the degree of concentration of the distribution of numerical increments (differences between adjacent records) and reflect the stability of data changes. When obtaining the incremental density standard deviation in embodiments of this specification, an incremental sequence can be calculated first according to the structured feature vector, then the incremental range is evenly divided into several intervals, the frequency of each interval is counted, then the normalized density is calculated according to the frequency of each interval, and then the standard deviation is calculated according to the normalized density, that is, the incremental density standard deviation is obtained.

[0055] In this embodiment, the text entropy value can quantify the degree of information chaos of a text field and is used to detect unstructured text (such as comments and logs). When obtaining the text entropy value in embodiments of this specification, the occurrence probability of each character in the text can be counted first, and then the entropy value can be calculated according to the occurrence probability of each character in the text.

[0056] In this embodiment, semantic dispersion is used to measure the degree of dispersion of semantic units (such as words, phrases) in a text field. When obtaining the semantic dispersion in the embodiments of this specification, a lightweight word segmentation tool (such as Trie tree word segmentation) can be used to extract a word list, and then the word frequency and inverse document frequency are calculated, and the semantic dispersion is calculated based on the word frequency and inverse document frequency.

[0057] The embodiments of this specification provide a dynamic routing decision to implement the channel selection logic. For example, the trigger condition for the numerical channel can be set as: IF (numerical continuity parameter > 0.7) AND (standard deviation of incremental density < 0.15) THEN route to the numerical channel; the trigger condition for the fingerprint channel can be set as: IF (text entropy value > 3.5) OR (semantic dispersion > 0.6) THEN route to the fingerprint channel.

[0058] In some embodiments, when dynamically selecting a numerical channel or a fingerprint channel based on a structured feature vector and generating a composite feature fingerprint corresponding to the data to be desensitized, dynamic weight allocation can also be performed, that is, the load ratio of the dual channels is adjusted according to the proportion of field types in the structured feature vector. For example, for a numerically dominant table: 70% of the traffic is allocated to the numerical channel and 30% is allocated to the fingerprint channel; another example is a textually dominant table: 30% of the traffic is allocated to the numerical channel and 70% is allocated to the fingerprint channel.

[0059] In some embodiments, the numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, including: initializing the optimization control window of the numerical channel; obtaining the data density and the slope of the incremental trend within the window; when the data density within the window is greater than the density threshold and the slope of the incremental trend is greater than the slope threshold, the optimization control window is enlarged to obtain a new optimization control window; based on the new optimization control window, the incremental density is determined according to the structured feature vector, and a composite key is generated through the incremental density; the composite key is inserted into the self-balancing binary search tree, and conflict detection queries are performed through the self-balancing binary search tree, and the allowed time deviation threshold is dynamically adjusted according to the variance of the timestamps.

[0060] In this embodiment, the data density within the window measures the filling rate of valid data points within the current sliding window, and is used to determine whether to expand the window to capture more valid information. The data density within the window can be the ratio of the number of valid data points (non-empty / non-outlier) to the current window capacity. The slope of the incremental trend can analyze the incremental (first-order difference) trend of the data within the window through linear regression, and is used to determine whether the data shows a significant growth / decline pattern.

[0061] In the embodiments of this specification, the incremental density can be determined based on the structured feature vector, and the composite key can be generated through the incremental density. The key value of the composite key can be generated by combining the mean value, standard deviation, and skewness of the incremental density of the feature vector. For example, for the structured feature vector V = [v1, v2,..., v n , the incremental sequence ΔV = [v2 - v1, v3 - v2,..., v n - v n-1 is calculated. Then, the incremental mean value, standard deviation, and skewness corresponding to the incremental sequence are calculated, and the composite key Key = (Avg(ΔV), StdDev(ΔV), Skewness(ΔV)) is determined from the incremental mean value, standard deviation, and skewness corresponding to the incremental sequence. Avg(ΔV) represents the incremental mean value, StdDev(ΔV) represents the standard deviation, and Skewness(ΔV) represents the skewness. In the embodiments of this specification, fast search and conflict determination can be achieved based on the optimization effect of the red - black tree index. Fast search: Utilize the insertion / query efficiency of the red - black tree, which is more adaptable to dynamic data streams than the hash table. Conflict determination: If the composite key distance (such as the Euclidean distance) between two feature vectors is less than the composite key distance threshold, it is determined as a potential conflict.

[0062] In this embodiment, initializing the optimization control window of the numerical channel can include the initial window size, and the window size = max(100, (Max - Min) / Std), where max() represents taking the maximum value in the parentheses, and Max, Min, and Std are respectively from the numerical statistical features in the structured feature vector: the maximum value, the minimum value, and the standard deviation. In the optimization control of the numerical channel in the embodiments of this specification, dynamic window adjustment can be achieved by setting the window expansion rule. For example, the window expansion rule: IF (the data density within the window > 90%) AND (the incremental trend slope > 0.5) THEN the window is expanded by 20%. In addition, in the optimization control of the numerical channel in the embodiments of this specification, conflict detection optimization is also performed through the red - black tree index key and the out - of - order tolerance. The red - black tree is a self - balancing binary search tree. The red - black tree index key: Generate a composite key based on the incremental density in the structured feature vector to accelerate the search for duplicate values. The out - of - order tolerance: Dynamically adjust the allowed time deviation threshold according to the variance of the timestamps. Specifically, the red - black tree index key is a composite index key structure based on the red - black tree (self - balancing binary search tree) for efficiently detecting potential conflicts (duplicate values) of the structured feature vector. In the embodiments of this specification, the key value can be generated through the incremental density feature, integrating the statistical characteristics of the data distribution into the index structure.

[0063] In some embodiments, the fingerprint channel fuses non - encrypted hashing with semantic weights, including: the fingerprint channel generates a composite feature fingerprint corresponding to the data to be desensitized from the structured feature vector based on hash algorithm selection, single - instruction multiple - data (SIMD) parallel optimization, and semantic fusion strategy; wherein, the hash algorithm selection includes: for the structured fields in the data to be desensitized, using the pattern encoding in the structured fields to select an optimized hash seed; for the unstructured fields in the data to be desensitized, determining the semantic hash corresponding to the unstructured field and generating a salt value according to the semantic hash; the SIMD parallel optimization includes: determining the average field length of the structured feature vector; dynamically adjusting the SIMD processing granularity according to the average field length; the semantic fusion strategy includes: determining the text entropy value of the structured feature vector and adjusting the proportion of semantic features in the composite feature fingerprint according to the text entropy value.

[0064] The embodiments of this specification can enhance the algorithm of the fingerprint channel by combining hash algorithm selection, single - instruction multiple - data (SIMD) parallel optimization, semantic fusion strategy, etc. In terms of hash algorithm selection, for structured fields (such as dates, IDs), the embodiments of this specification can use the pattern encoding in the structured feature vector to select an optimized hash seed; for unstructured fields (such as free text), the embodiments of this specification can generate a salt value based on the semantic hash to enhance hash discreteness. In terms of SIMD parallel optimization, the embodiments of this specification can adopt a block strategy, that is, dynamically adjusting the SIMD processing granularity according to the average field length in the structured feature vector. For example, IF (avg_length < 64B) THEN block size = 256B (4 64 - bit blocks) ELSE block size = 512B (8 64 - bit blocks). In terms of semantic fusion strategy, the embodiments of this specification can adjust the proportion of semantic features in the fingerprint based on the text entropy value in the feature vector to determine the semantic weight. For example, semantic weight = min(1.0, entropy / 8.0), where min() represents taking the minimum value in the parentheses and entropy represents the text entropy value.

[0065] The embodiments of this specification also provide a conflict resolution mechanism for predicting the conflict probability. For the numerical channel, the embodiments of this specification can calculate the expected conflict rate through the value range span and data density in the feature vector. The expected conflict rate = (Max - Min) / (window size * data density), where Max and Min are the numerical statistical features in the structured feature vector: the maximum value and the minimum value respectively. For the fingerprint channel, the embodiments of this specification can dynamically adjust the salt injection frequency based on the semantic dispersion and hash collision history. Hash collision means that different input data obtain the same output value after being processed by the hash function. The hash collision history records the frequency and pattern of collisions occurring during the hash process. The embodiments of this specification can understand which data is more likely to collide by analyzing the hash collision history, so as to optimize the hash algorithm. The salt value is a kind of random data used to enhance the security of the hash. The salt injection frequency can be the frequency of adding the salt value during the hash process.

[0066] 230. Map the composite feature fingerprint to the probability-based bit mapping matrix, perform verification through the bit operation instruction set, mark the records not mapped to the probability-based bit mapping matrix as incremental data, and desensitize the incremental data.

[0067] The embodiments of this specification can generate the composite feature fingerprint of the data using the structured feature vector, then map it into the probability-based bit mapping matrix, and perform the disk writing operation through the bit compression algorithm to prevent data loss caused by processing interruption. When a processing task ends and the second desensitization task starts, incremental desensitization can be performed through the above process. When the result calculated from the new data does not exist in the mapping matrix, it is new data, and this new data is desensitized, so as to realize the incremental desensitization function.

[0068] In some embodiments, please refer to Figure 5 , Figure 5 is the schematic flowchart of mapping the composite feature fingerprint to the probability-based bit mapping matrix provided by the embodiments of the present invention. Map the composite feature fingerprint to the probability-based bit mapping matrix, perform verification through the bit operation instruction set, mark the records not mapped to the probability-based bit mapping matrix as incremental data, and desensitize the incremental data, including:

[0069] 2300. Obtain the initial probability-based bit mapping matrix. The dimension D of the initial probability-based bit mapping matrix = 2 n ,n = ceil(log 2 (1.2 × estimated number of records)), ceil() represents rounding up, n represents the calculated minimum number of bits required, and the estimated number of records comes from the real-time feedback of the fingerprint generation rate in the second layer;

[0070] 2310. When mapping the composite feature fingerprint to the probabilistic bit mapping matrix, determine the matrix filling rate corresponding to the initial probabilistic bit mapping matrix;

[0071] 2320. When the matrix filling rate is greater than the filling rate threshold, perform dimension expansion processing on the initial probabilistic bit mapping matrix to obtain the expanded probabilistic bit mapping matrix;

[0072] 2330. Use the double - hash remapping algorithm to migrate the data in the initial probabilistic bit mapping matrix to the expanded probabilistic bit mapping matrix.

[0073] In this embodiment, the probabilistic bit mapping matrix is a two - dimensional bit array. The data is mapped to specific positions in the matrix through a hash function. By introducing a probability mechanism, the probabilistic bit mapping matrix allows a certain degree of misjudgment (i.e., false positives), thus significantly reducing the storage space. The probabilistic bit mapping matrix is a structure for efficient data storage and query, combining probability methods and bit mapping techniques to achieve fast data processing and a low misjudgment rate. In the data desensitization scenario, the probabilistic bit mapping matrix is used to quickly determine whether the data has been desensitized. By mapping the data into the matrix, the system can quickly query whether the data exists, and thus decide whether desensitization processing is required. Through the efficient query and storage mechanism of the probabilistic bit mapping matrix in this specification embodiment, the performance and efficiency of the system can be significantly improved.

[0074] In this embodiment, the second - layer fingerprint is to further process the composite feature fingerprint after preliminary feature extraction to generate more complex feature identifiers. These feature identifiers can be used to optimize the hash algorithm, reduce the misjudgment rate, and can also be used to dynamically adjust the storage and processing resources. The estimated record number in the embodiment of this specification can come from the real - time feedback of the second - layer fingerprint generation rate, that is, the system can dynamically adjust the estimated record number by obtaining the generation rate of the second - layer fingerprint, so as to optimize the allocation of storage and processing resources.

[0075] In this embodiment, the matrix filling rate can be the ratio of the number of filled elements to the total number of elements in the matrix. This specification embodiment sets a dynamic bit - width expansion mechanism, that is, when the matrix filling rate is greater than the filling rate threshold (such as matrix filling rate > 75%), dimension doubling is triggered, and the existing data needs to be migrated to a new and larger matrix: create a 2D×2D bit matrix; then use the double - hash remapping algorithm to migrate the data. Using the double - hash remapping algorithm can efficiently remap the data to the new matrix, ensure the uniform distribution of the data, and at the same time maintain the uniqueness and accuracy of the data; the atomic replacement of the old matrix expansion takes less than 50 ms, that is, the entire migration process can be completed in a short time, ensuring the stability of the system and the integrity of the data.

[0076] In this embodiment, the double-hashing remapping algorithm can be a technology used to solve hash conflicts and optimize data storage. The double-hashing remapping algorithm remaps data by using two hash functions, thereby reducing conflicts and improving the efficiency of data storage. The double-hashing remapping algorithm first uses two hash functions to calculate the storage location of the data; when the first hash function generates a conflict, the second hash function can be used to calculate a new storage location.

[0077] In the process of using the double-hashing remapping algorithm to map data in the embodiments of this specification, the first hash function h1(x) = CFC[0:63] mod D, where x represents the data in the composite feature fingerprint, CFC[0:63] represents the substring from the 0th bit to the 63rd bit of the composite feature fingerprint CFC, and D is the size of the hash table, which also represents the dimension of the probabilistic bit-mapping matrix. The first hash function maps the data to a location in the hash table through a modulo operation. The second hash function h2(x) = (CFC[64:127] ⊕ CFC[0:63]) mod D, where CFC[64:127] represents the substring from the 64th bit to the 127th bit of the composite feature fingerprint CFC, and ⊕ represents a bitwise exclusive OR operation. The second hash function generates an incremental value for conflict resolution by combining different parts of the composite feature fingerprint CFC and performing a bitwise exclusive OR operation.

[0078] In the hashing process of the embodiments of this specification, in order to further reduce the conflict probability, the semantic weight value of the second-layer fingerprint can also be dynamically injected as a salt value. The salt value is a kind of random or pseudo-random data used to enhance the security and uniqueness of the hash. By injecting the semantic weight value as a salt value into the hash function, the randomness of the hash value can be increased, thereby reducing the conflict probability.

[0079] In terms of controlling the collision probability in the embodiments of this specification, through the double-hashing remapping algorithm and dynamic salt value injection, the collision probability can be controlled at an extremely low level, such that in the processing of a large amount of data to be desensitized, the probability of hash conflicts is significantly reduced, thereby improving the efficiency and reliability of the system.

[0080] In some embodiments, the composite feature fingerprint is mapped to a probabilistic bit mapping matrix and verified through a bit operation instruction set. Records that are not mapped to the probabilistic bit mapping matrix are marked as incremental data, and desensitization is performed on the incremental data, including: during the process of mapping data using the double-hash remapping algorithm, the double-hash remapping algorithm includes using two hash functions to calculate the storage location of the data when the composite feature fingerprint is mapped to the probabilistic bit mapping matrix; when the first hash function generates a collision, the second hash function is used to calculate a new storage location; and the semantic weight value of the second-layer fingerprint is injected as a salt value into the second hash function; the first hash function h1(x) = CFC[0:63] mod D, where x represents the data in the composite feature fingerprint, CFC[0:63] represents the substring from the 0th bit to the 63rd bit of the composite feature fingerprint CFC, and D is the size of the hash table, which also represents the dimension of the probabilistic bit mapping matrix. The first hash function maps the data to a location in the hash table through a modulo operation; the second hash function h2(x) = (CFC[64:127] ⊕ CFC[0:63]) mod D, where CFC[64:127] represents the substring from the 64th bit to the 127th bit of the composite feature fingerprint CFC, and ⊕ represents a bitwise exclusive OR operation. The second hash function generates an incremental value for collision resolution by combining different parts of the composite feature fingerprint CFC and performing a bitwise exclusive OR operation; when continuous collisions are detected and the number of continuous collisions exceeds a threshold, extended fingerprint remapping is triggered; an incremental snapshot is generated every first preset duration using a breakpoint recovery mechanism, and a full snapshot is compressed and stored every second preset duration, and the operation log is appended in real time; when the system needs to be restored after a crash or abnormal termination, the latest snapshot is preferentially loaded and the memory matrix is rebuilt.

[0081] The embodiments of this specification adopt a double-hash remapping algorithm for collision control of the probabilistic bit mapping matrix, and when continuous collisions are detected and the number of continuous collisions exceeds a threshold, extended fingerprint remapping is triggered. That is, when the system detects continuous hash collisions, a longer composite feature fingerprint can be used to remap the data, so that the data can be accurately located and the misjudgment rate can be reduced.

[0082] The embodiments of this specification can significantly reduce hash collisions by using two hash functions for data remapping; moreover, when the matrix is extended, the double-hash remapping algorithm can quickly remap the data to the new matrix, reducing the migration time; at the same time, the entire migration process can be completed in a short time, ensuring the stability of the system and the integrity of the data. The embodiments of this specification play an important role in dynamic data processing and storage optimization through the double-hash remapping algorithm.

[0083] In the embodiments of this specification, a breakpoint recovery mechanism can be adopted to generate an incremental snapshot every first preset duration (e.g., 5 s), record the changes in the memory bit states since the last snapshot, and at the same time use the CRC64 algorithm to verify the snapshot to ensure the integrity and consistency of the snapshot. CRC64 is an efficient cyclic redundancy check algorithm that can detect whether errors occur during data transmission or storage. In the embodiments of this specification, a full snapshot can also be generated every second preset duration (e.g., 30 s), and the LZ4-HC (High Compression) algorithm is used to compress and store the full snapshot. LZ4-HC is an efficient lossless compression algorithm that can provide a high compression ratio while ensuring the compression speed, saving storage space. In the embodiments of this specification, real-time append of operation logs (WAL (Write-Ahead Logging) logs) can also be performed. The WAL log is a log recording mechanism that can record all modification operations on the memory bit states in real time. In the embodiments of this specification, the operation logs are written to the disk in an append manner in real time to ensure that all operations are recorded. During recovery, the memory bit states can be reconstructed through the operation logs.

[0084] When the system needs to be recovered after a crash or abnormal termination in the embodiments of this specification, the latest snapshot is preferentially loaded and the memory matrix is reconstructed. That is, during system recovery, the latest full snapshot or incremental snapshot is first loaded to restore the basic part of the memory bit states; then the memory matrix is reconstructed, and according to the snapshot content, the bit matrix in the memory is reconstructed to restore to the state at the time of the snapshot.

[0085] The embodiments of this specification adopt a breakpoint recovery guarantee mechanism through a three-level persistence strategy (incremental snapshot, full snapshot, and operation log) and a fast recovery process to ensure that the system can quickly recover to the nearest state when a failure occurs, guarantee the integrity and consistency of the data, and improve the reliability and fault tolerance of the system.

[0086] In the embodiments of this specification, the output bit width of the hash function can also be adjusted in real time according to the filling rate of the probabilistic bit mapping matrix. For example, when the filling rate of the probabilistic bit mapping matrix is less than 40%, 64-bit hashing is adopted, that is, when the filling rate is small, the data volume is small, and using a shorter hash bit width can save computing resources and storage space; for another example, when the filling rate of the probabilistic bit mapping matrix is between 40% and 75%, the data volume is moderate, and using 128-bit hashing can provide better discrimination and maintain high efficiency; for another example, when the filling rate of the probabilistic bit mapping matrix is greater than 75%, the data volume is large, and using a longer hash bit width (such as 256 bits) can significantly reduce the probability of hash collisions. At the same time, a matrix expansion operation is triggered to adapt to more data.

[0087] The embodiments of this specification can also set up a semantic awareness storage optimization mechanism. For data with a relatively high semantic weight (greater than 0.7) (such as text data with a high entropy value), these data have a relatively high importance in the system. Therefore, double copies are retained in the memory matrix to improve the reliability and access speed of the data. For data with a relatively low semantic weight (not greater than 0.7), the importance of these data is relatively low. Therefore, only a single copy is retained in the memory matrix, and the additional copy is stored on the disk to save memory space.

[0088] The embodiments of this specification can dynamically adjust the hash bit width according to the filling rate of the matrix to ensure efficient hash performance under different data volumes, while reducing the probability of hash collisions. The embodiments of this specification can also adjust the storage strategy according to the semantic weight value of the data to ensure the high availability and reliability of important data, while optimizing the use of storage resources. The embodiments of this specification can also detect the filling rate of the probabilistic bit mapping matrix in real time and dynamically adjust the hash bit width and storage strategy according to the current state to ensure the efficient operation of the system.

[0089] The embodiments of this specification provide a data desensitization technology that can adapt to heterogeneous data sources, efficiently process incremental data, and ensure data integrity. Without relying on the data source structure, it can automatically identify and process various types of data without the need to know the data table structure in advance or rely on primary key / unique key constraints. In addition, the embodiments of this specification have efficient incremental desensitization capabilities, can accurately identify the differences between new data and processed data, and only desensitize the new data to improve processing efficiency. The embodiments of this specification have a low misjudgment rate. The numerical channel adopts a self-balancing binary search tree index strategy with dynamic window adjustment, and the fingerprint channel fuses non-encrypted hashing and semantic weights, which can not only significantly reduce the misjudgment rate, improve desensitization accuracy, but also improve the data processing speed and reduce the system power consumption.

[0090] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0091] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an incremental desensitization system for heterogeneous data sources provided by the embodiments of this specification.

[0092] As Figure 6As shown in the figure, the incremental desensitization system for heterogeneous data sources can at least include a data access module 600, a vector generation module 610, a channel selection module 620, and a matrix mapping module 630, where:

[0093] The data access module 600 is used to obtain the data to be desensitized from heterogeneous data sources;

[0094] The vector generation module 610 is used to generate structured feature vectors containing different types of features for the data to be desensitized from heterogeneous data sources through streaming sampling and dynamic feature analysis;

[0095] The channel selection module 620 is used to dynamically select a numerical channel or a fingerprint channel based on the structured feature vector to generate a composite feature fingerprint corresponding to the data to be desensitized. The numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, and the fingerprint channel fuses non-encrypted hashing and semantic weights;

[0096] The matrix mapping module 630 is used to map the composite feature fingerprint to a probabilistic bit mapping matrix, and perform verification through a bit operation instruction set, mark the records that are not mapped to the probabilistic bit mapping matrix as incremental data, and perform desensitization on the incremental data.

[0097] In some embodiments, the vector generation module 610 includes a vector generation sub-module, and the vector generation sub-module is used to: determine the field data type of the data to be desensitized; determine the standard deviation fluctuation value or entropy value change rate corresponding to the data to be desensitized; adjust the sampling depth based on the field data type according to the standard deviation fluctuation value or entropy value change rate; generate structured feature vectors containing different types of features for the data to be desensitized from heterogeneous data sources based on the sampling depth.

[0098] In some embodiments, the vector generation sub-module includes a sampling module, and the sampling module is used to: trigger incremental sampling when the field data type is a numerical field and the standard deviation fluctuation value is greater than a preset fluctuation threshold; trigger supplementary sampling when the field data type is a text field and the entropy value change rate is greater than a preset change threshold.

[0099] In some embodiments, the vector generation sub-module includes a feature acquisition module, and the feature acquisition module is used to: when the field data type is a numerical field, acquire the numerical features corresponding to the data to be desensitized, and the numerical features include distribution features, range features, incremental features, and density features; when the field data type is a text field, acquire the text features corresponding to the data to be desensitized, and the text features include entropy value features, pattern features, semantic features, and length features; generate structured feature vectors containing different types of features according to the numerical features or text features corresponding to the data to be desensitized.

[0100] In some embodiments, the channel selection module 620 includes a determination module. The determination module is configured to: determine channel selection parameters according to the structured feature vector, where the channel selection parameters include a numerical continuity parameter, an incremental density standard deviation, a text entropy value, and a semantic dispersion degree; when the numerical continuity parameter is greater than the continuous parameter threshold and the incremental density standard deviation is less than the standard deviation threshold, select a numerical channel to generate a composite feature fingerprint corresponding to the data to be desensitized; when the text entropy value is greater than the entropy threshold and the semantic dispersion degree is greater than the dispersion threshold, select a fingerprint channel to generate a composite feature fingerprint corresponding to the data to be desensitized.

[0101] In some embodiments, the channel selection module 620 includes a numerical channel optimization module. The numerical channel optimization module is configured to: initialize an optimization control window for the numerical channel; obtain the data density and the incremental trend slope within the window; when the data density within the window is greater than the density threshold and the incremental trend slope is greater than the slope threshold, perform an expansion process on the optimization control window to obtain a new optimization control window; based on the new optimization control window, determine the incremental density according to the structured feature vector, and generate a composite key through the incremental density; insert the composite key into a self-balancing binary search tree, perform a conflict detection query through the self-balancing binary search tree, and dynamically adjust the allowed time deviation threshold according to the timestamp variance.

[0102] In some embodiments, the channel selection module 620 includes a fingerprint channel optimization module. The fingerprint channel optimization module is configured to: generate a composite feature fingerprint corresponding to the data to be desensitized from the structured feature vector based on a hash algorithm selection, single instruction multiple data parallel optimization, and a semantic fusion strategy; wherein, the hash algorithm selection includes: for the structured fields in the data to be desensitized, select an optimized hash seed using the pattern encoding in the structured fields; for the unstructured fields in the data to be desensitized, determine the semantic hash corresponding to the unstructured fields, and generate a salt value according to the semantic hash; the single instruction multiple data parallel optimization includes: determining the average field length of the structured feature vector; dynamically adjusting the single instruction multiple data processing granularity according to the average field length; the semantic fusion strategy includes: determining the text entropy value of the structured feature vector, and adjusting the proportion of semantic features in the composite feature fingerprint according to the text entropy value.

[0103] In some embodiments, the matrix mapping module 630 includes a mapping sub-module. The mapping sub-module is configured to: obtain an initial probabilistic bit mapping matrix, where the dimension D of the initial probabilistic bit mapping matrix is 2 n , n = ceil(log 2(1.2 × estimated number of records)), ceil() represents rounding up, n represents the calculated minimum number of digits required, and the estimated number of records comes from the real-time feedback of the fingerprint generation rate in the second layer; when mapping the composite feature fingerprint to the probabilistic bit mapping matrix, determine the matrix filling rate corresponding to the initial probabilistic bit mapping matrix; when the matrix filling rate is greater than the filling rate threshold, perform dimension expansion processing on the initial probabilistic bit mapping matrix to obtain the expanded probabilistic bit mapping matrix; use the double-hash remapping algorithm to migrate the data in the initial probabilistic bit mapping matrix to the expanded probabilistic bit mapping matrix.

[0104] In some embodiments, the matrix mapping module 630 includes a conflict detection module, and the conflict detection module is used for: during the process of mapping data using the double-hash remapping algorithm, the double-hash remapping algorithm includes using two hash functions to calculate the storage location of the data when the composite feature fingerprint is mapped to the probabilistic bit mapping matrix; when the first hash function generates a conflict, the second hash function is used to calculate a new storage location; and inject the semantic weight value of the second-layer fingerprint as a salt value into the second hash function; the first hash function h1(x) = CFC[0:63] mod D, where CFC[0:63] represents the substring from the 0th bit to the 63rd bit of the composite feature fingerprint CFC, and D is the size of the hash table, which also represents the dimension of the probabilistic bit mapping matrix. The first hash function maps the data to a location in the hash table through a modulo operation; the second hash function h2(x) = (CFC[64:127] ⊕ CFC[0:63]) mod D, where CFC[64:127] represents the substring from the 64th bit to the 127th bit of the composite feature fingerprint CFC, and ⊕ represents a bitwise exclusive OR operation. The second hash function generates an increment value for conflict resolution by combining different parts of the composite feature fingerprint CFC and performing a bitwise exclusive OR operation; when it is detected that continuous collisions occur and the number of continuous collisions exceeds the number threshold, trigger extended fingerprint remapping; use a breakpoint recovery mechanism to generate incremental snapshots every first preset duration and compress and store full snapshots every second preset duration, and append operation logs in real time; when the system needs to be restored after a crash or abnormal termination, preferentially load the latest snapshot and reconstruct the memory matrix.

[0105] The embodiments of this specification provide a data desensitization technology that can adapt to heterogeneous data sources, efficiently process incremental data, and ensure data integrity. Without relying on the data source structure, it can automatically identify and process various types of data without the need to know the data table structure in advance or rely on primary key / unique key constraints. Additionally, the embodiments of this specification have efficient incremental desensitization capabilities, can accurately identify the differences between new data and processed data, and only desensitize new data to improve processing efficiency. The embodiments of this specification have a low false positive rate. The numerical channel adopts a self-balancing binary search tree index strategy with dynamic window adjustment, and the fingerprint channel integrates non-encrypted hashing and semantic weights, which can not only significantly reduce the false positive rate, improve desensitization accuracy, but also increase the data processing speed while reducing system power consumption.

[0106] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for an embodiment of an incremental desensitization system for heterogeneous data sources, since it is basically similar to an embodiment of an incremental desensitization method for heterogeneous data sources, the description is relatively simple, and reference can be made to the partial description of the method embodiment for related parts.

[0107] Please refer to Figure 7 The structural schematic diagram of an electronic device of a server in an incremental desensitization system for heterogeneous data sources provided by the embodiments of this specification is shown.

[0108] As Figure 7 As shown, the electronic device 700 may include: at least one processor 710, at least one network interface 740, a user interface 730, a memory 750, and at least one communication bus 720.

[0109] Among them, the communication bus 720 can be used to realize the connection and communication of the above-mentioned various components.

[0110] Among them, the user interface 730 may include buttons, and the optional user interface may further include a standard wired interface and a wireless interface.

[0111] Among them, the network interface 740 may include, but is not limited to, a Bluetooth module, an NFC module, a ZigBee module, and a UWB module, etc.

[0112] Among them, the processor 710 may include one or more processing cores. The processor 710 connects various parts within the entire electronic device 700 through various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 750, and by calling the data stored in the memory 750, it executes various functions of the electronic device 700 and processes data. Optionally, the processor 710 may be implemented in at least one hardware form of DSP, FPGA, or PLA. The processor 710 may integrate one or a combination of several of CPU and GPU, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen.

[0113] Among them, the memory 750 may include RAM and may also include ROM. Optionally, the memory 750 includes a non-transitory computer-readable medium. The memory 750 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 750 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area can store the data involved in the above-mentioned method embodiments. Optionally, the memory 750 may also be at least one storage device located far from the aforementioned processor 710. The memory 750, as a computer storage medium, may include an operating system, a communication module, a user interface module, and an incremental desensitization application program for heterogeneous data sources. The processor 710 may be used to call an incremental desensitization application program for heterogeneous data sources stored in the memory 750 and execute the steps in an incremental desensitization method for heterogeneous data sources mentioned in the foregoing embodiments.

[0114] The embodiments of this specification also provide a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When they run on a computer or a processor, they cause the computer or the processor to execute one or more of the steps in the Figures 2 to 5 embodiments shown above. If the various component modules of the above electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0115] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this specification are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a Digital Versatile Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.

[0116] Those of ordinary skill in the art can understand that all or part of the processes in the above embodiments of the method can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above embodiments of the various methods. The aforementioned storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes. Without conflict, the technical features in this embodiment and the implementation scheme can be combined arbitrarily.

[0117] The above-described embodiments are merely described as the preferred embodiments of this specification, and do not limit the scope of this specification. Without departing from the design spirit of this specification, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of this specification shall fall within the protection scope determined by the claims of this specification.

Claims

1. Incremental desensitization method for heterogeneous data sources, characterized in that including: obtaining the data to be desensitized from heterogeneous data sources; generating a structured feature vector containing different types of features for the data to be desensitized from the heterogeneous data sources through streaming sampling and dynamic feature analysis; dynamically selecting a numerical channel or a fingerprint channel based on the structured feature vector to generate a composite feature fingerprint corresponding to the data to be desensitized, where the numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, and the fingerprint channel fuses non-encrypted hashing and semantic weights; mapping the composite feature fingerprint to a probabilistic bit mapping matrix, and performing verification through a bit operation instruction set, marking the records that are not mapped to the probabilistic bit mapping matrix as incremental data, and performing desensitization on the incremental data; the step of mapping the composite feature fingerprint to a probabilistic bit mapping matrix, and performing verification through a bit operation instruction set, marking the records that are not mapped to the probabilistic bit mapping matrix as incremental data, and performing desensitization on the incremental data includes: Obtain an initial probabilistic bit mapping matrix, where the dimension D of the initial probabilistic bit mapping matrix is 2 n , n = ceil(log 2 (1.2 × estimated number of records)), ceil() represents rounding up, n represents the calculated minimum number of bits required, and the estimated number of records comes from the real-time feedback of the fingerprint generation rate in the second layer; when mapping the composite feature fingerprint to a probabilistic bit mapping matrix, determining the matrix filling rate corresponding to the initial probabilistic bit mapping matrix; when the matrix filling rate is greater than a filling rate threshold, performing dimension expansion processing on the initial probabilistic bit mapping matrix to obtain an expanded probabilistic bit mapping matrix; using a double-hashing remapping algorithm to migrate the data in the initial probabilistic bit mapping matrix to the expanded probabilistic bit mapping matrix.

2. The incremental desensitization method for heterogeneous data sources according to claim 1, characterized in that the step of generating a structured feature vector containing different types of features for the data to be desensitized from the heterogeneous data sources through streaming sampling and dynamic feature analysis includes: determining the field data type of the data to be desensitized; determining the standard deviation fluctuation value or entropy value change rate corresponding to the data to be desensitized; adjusting the sampling depth based on the field data type according to the standard deviation fluctuation value or entropy value change rate; generating a structured feature vector containing different types of features for the data to be desensitized from the heterogeneous data sources based on the sampling depth.

3. The incremental desensitization method for heterogeneous data sources according to claim 2, wherein the step of adjusting the sampling depth based on the field data type according to the standard deviation fluctuation value or entropy value change rate includes: when the field data type is a numerical field and the standard deviation fluctuation value is greater than a preset fluctuation threshold, triggering incremental sampling; when the field data type is a text field and the entropy value change rate is greater than a preset change threshold, triggering supplementary sampling.

4. The incremental desensitization method for heterogeneous data sources according to claim 2, wherein the step of generating a structured feature vector containing different types of features for the data to be desensitized from the heterogeneous data sources based on the sampling depth includes: when the field data type is a numerical field, obtaining the numerical features corresponding to the data to be desensitized, where the numerical features include distribution features, range features, incremental features, and density features; when the field data type is a text field, obtaining the text features corresponding to the data to be desensitized, where the text features include entropy value features, pattern features, semantic features, and length features; generating a structured feature vector containing different types of features according to the numerical features or text features corresponding to the data to be desensitized.

5. The incremental desensitization method for heterogeneous data sources according to claim 1, wherein Dynamically selecting a numerical channel or a fingerprint channel based on the structured feature vector to generate a composite feature fingerprint corresponding to the data to be desensitized, including: Determining channel selection parameters according to the structured feature vector, where the channel selection parameters include a numerical continuity parameter, an incremental density standard deviation, a text entropy value, and a semantic dispersion; When the numerical continuity parameter is greater than a continuous parameter threshold and the incremental density standard deviation is less than a standard deviation threshold, selecting the numerical channel to generate a composite feature fingerprint corresponding to the data to be desensitized; When the text entropy value is greater than an entropy threshold and the semantic dispersion is greater than a dispersion threshold, selecting the fingerprint channel to generate a composite feature fingerprint corresponding to the data to be desensitized.

6. The incremental desensitization method for heterogeneous data sources according to claim 1, wherein The numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, including: Initializing an optimization control window of the numerical channel; Obtaining the data density and the incremental trend slope within the window; When the data density within the window is greater than a density threshold and the incremental trend slope is greater than a slope threshold, performing an expansion process on the optimization control window to obtain a new optimization control window; Based on the new optimization control window, determining an incremental density according to the structured feature vector and generating a composite key through the incremental density; Inserting the composite key into the self-balancing binary search tree, performing conflict detection and query through the self-balancing binary search tree, and dynamically adjusting an allowed time deviation threshold according to a timestamp variance.

7. The incremental desensitization method for heterogeneous data sources according to claim 1, wherein The fingerprint channel fuses non-encrypted hashing and semantic weights, including: The fingerprint channel generates a composite feature fingerprint corresponding to the data to be desensitized based on the structured feature vector by using a hashing algorithm selection, a single instruction multiple data parallel optimization, and a semantic fusion strategy; where, The hashing algorithm selection includes: regarding the structured fields in the data to be desensitized, selecting an optimized hashing seed by using the pattern encoding in the structured fields; regarding the unstructured fields in the data to be desensitized, determining a semantic hash corresponding to the unstructured fields and generating a salt value according to the semantic hash; The single instruction multiple data parallel optimization includes: determining an average field length of the structured feature vector; dynamically adjusting a single instruction multiple data processing granularity according to the average field length; The semantic fusion strategy includes: determining a text entropy value of the structured feature vector and adjusting a semantic feature proportion in the composite feature fingerprint according to the text entropy value.

8. The incremental desensitization method for heterogeneous data sources according to claim 1, wherein Mapping the composite feature fingerprint to a probabilistic bit mapping matrix and performing verification through a bit operation instruction set, marking records that are not mapped to the probabilistic bit mapping matrix as incremental data, and performing desensitization on the incremental data, including: During the process of mapping data using a double hashing remapping algorithm, the double hashing remapping algorithm includes using two hashing functions to calculate the storage location of data when the composite feature fingerprint is mapped to the probabilistic bit mapping matrix; when the first hashing function generates a conflict, the second hashing function is used to calculate a new storage location; and injecting the semantic weight value of the second layer fingerprint as a salt value into the second hashing function; The first hash function h1(x) = CFC[0:63] mod D, where x represents the data in the composite feature fingerprint, CFC[0:63] represents the substring from the 0th bit to the 63rd bit of the composite feature fingerprint CFC; D is the size of the hash table and also represents the dimension of the probabilistic bit mapping matrix; the first hash function maps the data to a position in the hash table through a modulo operation; the second hash function h2(x) = (CFC[64:127] ⊕ CFC[0:63]) mod D, where CFC[64:127] represents the substring from the 64th bit to the 127th bit of the composite feature fingerprint CFC, and ⊕ represents a bitwise exclusive OR operation. The second hash function generates an increment value for conflict resolution by combining different parts of the composite feature fingerprint CFC and performing a bitwise exclusive OR operation. When it is detected that consecutive collisions occur and the number of consecutive collisions exceeds the number threshold, extended fingerprint remapping is triggered. An incremental snapshot is generated every first preset duration using a breakpoint recovery mechanism, and a full snapshot is compressed and stored every second preset duration, with operation logs appended in real time. When the system needs to be restored after a crash or abnormal termination, the latest snapshot is preferentially loaded and the memory matrix is rebuilt.

9. Incremental desensitization system for heterogeneous data sources, characterized in that, It includes: A data access module for obtaining the data to be desensitized from heterogeneous data sources. A vector generation module for generating a structured feature vector containing different types of features for the data to be desensitized from the heterogeneous data source through streaming sampling and dynamic feature analysis. A channel selection module for dynamically selecting a numerical channel or a fingerprint channel based on the structured feature vector to generate a composite feature fingerprint corresponding to the data to be desensitized. The numerical channel adopts a self-balancing binary search tree indexing strategy with dynamic window adjustment, and the fingerprint channel fuses non-encrypted hashing and semantic weights. A matrix mapping module for mapping the composite feature fingerprint to a probabilistic bit mapping matrix, performing verification through a bit operation instruction set, marking the records not mapped to the probabilistic bit mapping matrix as incremental data, and performing desensitization on the incremental data. The mapping of the composite feature fingerprint to the probabilistic bit mapping matrix, performing verification through a bit operation instruction set, marking the records not mapped to the probabilistic bit mapping matrix as incremental data, and performing desensitization on the incremental data includes: Obtain an initial probabilistic bit mapping matrix, where the dimension D of the initial probabilistic bit mapping matrix is 2 n , n = ceil(log 2 (1.2 × estimated number of records)), ceil() represents rounding up, n represents the calculated minimum number of bits required, and the estimated number of records is from the real-time feedback of the fingerprint generation rate in the second layer; When mapping the composite feature fingerprint to the probabilistic bit mapping matrix, determining the matrix filling rate corresponding to the initial probabilistic bit mapping matrix. When the matrix filling rate is greater than the filling rate threshold, performing dimension amplification processing on the initial probabilistic bit mapping matrix to obtain an amplified probabilistic bit mapping matrix. Using a double-hash remapping algorithm to migrate the data in the initial probabilistic bit mapping matrix to the amplified probabilistic bit mapping matrix.

Citation Information

Patent Citations

  • Adverse drug reaction trace management method and system

    CN118280602A

  • Sensitive data protection system and method based on pseudonymous encryption

    CN119906583A