Data governance sensitive data desensitization method
By combining three-dimensional enhanced vector evaluation and risk assessment within the homomorphic encryption domain with volumetric neural rendering and zero-knowledge proof, the problem of difficulty in identifying hidden fields and controlling privacy leaks in existing technologies is solved, achieving an efficient and secure data anonymization solution.
Patent Information
- Application Number
- CN202511282133.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing technologies struggle to identify semantically hidden fields in data governance within highly sensitive industries, neglecting numerical distribution and temporal evolution, leading to statistical re-identification, plaintext or semi-plaintext operations, a lack of verifiable budget control on the query side, difficulty in suppressing combined leakage, and limitations on scalability and security strength.
We employ a three-dimensional enhanced vector to assess field risk, filter out high-risk data within a homomorphic encryption domain, generate irreversible proxy records using volumetric neural rendering and tensor chains, and dynamically deduct privacy budgets on the query side through topological leakage measurement and zero-knowledge proofs to achieve both data availability and strict privacy protection.
It achieves a comprehensive characterization of field risks, accurately identifies hidden sensitive fields, eliminates plaintext windows in memory, suppresses single point of re-identification, implements verifiable privacy budget management, blocks combined attacks, and ensures service continuity and statistical availability under extreme conditions.
Smart Images

Figure CN120781391B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information security technology, and in particular to a method for de-identifying sensitive data in data governance. Background Technology
[0002] In highly sensitive industries such as finance, healthcare, and e-commerce, large-scale data governance requires simultaneous compliance with real-time analysis and stringent regulatory requirements (GDPR, CCPA, etc.) for de-identification. Traditional de-identification techniques often employ rule masking, irreversible hashing, or static tokenization, but they share common drawbacks: ① reliance on manual regular expressions and field whitelists, making it difficult to identify semantically hidden fields; ② ignoring numerical distribution and temporal evolution, leading to statistical re-identification; ③ plaintext or semi-plaintext operations, resulting in data being stored multiple times outside of memory; ④ lack of verifiable budget control on the query side, making it difficult to suppress combined data leakage. These issues limit the scalability and security strength of existing solutions in high-concurrency, cross-domain sharing scenarios. Summary of the Invention
[0003] To address the numerous problems existing in the above-mentioned technologies, this invention provides a method for de-identifying sensitive data in data governance. This invention uses three-dimensional enhanced vectors to assess field risks, filters out high-risk data within the homomorphic encryption domain, and uses volumetric neural rendering and tensor chains to map plaintext into irreversible proxy records. On the query side, privacy budgets are dynamically reduced through topological leakage measurement and zero-knowledge proofs, and differential privacy aggregation is switched when the budget is insufficient, thereby balancing data availability and strict privacy protection.
[0004] A method for de-identifying sensitive data in data governance includes the following steps:
[0005] The database primary key is hashed and sharded to generate an enhanced vector that combines semantic vectors, statistical vectors, and topological indicators, and then a risk morphology fingerprint is obtained through a classification model.
[0006] In the homomorphic encryption domain, the risk pattern fingerprint and the enhancement vector are combined to form a ciphertext tensor and a risk score is calculated; when the risk score reaches a threshold, the data is decrypted in a trusted execution environment to obtain high-risk data and the corresponding feature tensor.
[0007] A volumetric neural rendering network is trained based on the feature tensor and proxy records are generated under the constraint of matrix product state tensor chain; a proxy primary key is generated by hashing the database primary key with a random salt and the proxy record is written into the proxy data lake;
[0008] After receiving a query, a query hypergraph is constructed, incremental mutual information is calculated, and a leakage index is obtained based on persistent homology. A zero-knowledge proof is generated using a random salt, cumulative mutual information, leakage index, global privacy budget limit, and the risk morphological fingerprint. If the verification is successful, the global privacy budget limit is deducted and the query result is returned. If the verification fails or the leakage index exceeds the limit, the tensor chain is located by random salt, the noise is amplified, and the result is written into the homomorphic encryption field.
[0009] Preferably, after the database primary key generates a hash value using a one-way hash algorithm, the modulo value is taken with a preset number of fragments, and the resulting modulo value is used as the fragment identifier code.
[0010] Preferably, semantic vectors are obtained by uniformly encoding field names and business contexts through a pre-trained language model, statistical vectors are obtained by binning the frequency of field values, and topological metrics are obtained by constructing a Vitoris-Lipps complex on the time series of field values and extracting the longest persistent scale.
[0011] Preferably, the homomorphic encryption domain adopts the CKKS homomorphic encryption scheme based on polynomial rings, and the homomorphic inner product of the ciphertext tensor and the encryption weight vector is used to calculate the risk score using a third-order Chebyshev approximate polynomial.
[0012] Preferably, after the trusted execution environment verifies the security integrity through a remote proof protocol, it receives the ciphertext tensor and completes the decryption process in an isolated memory area.
[0013] Preferably, the input coordinates of the volumetric neural rendering network are expanded by sinusoidal position encoding before entering the density prediction branch and the color prediction branch. The density prediction branch uses the Softplus activation function, and the color prediction branch uses the linear activation function.
[0014] Preferably, each dimension of the matrix product state tensor chain is determined by the base dimension plus an integer offset proportional to the risk score.
[0015] Preferably, the query hypergraph consists of a vertex set with surrogate primary keys and a hyperedge set with the surrogate primary key set involved in a single query. The leakage index is the maximum persistence scale obtained by performing a persistence cohomology analysis on the query hypergraph.
[0016] Preferably, the zero-knowledge proof uses the Groth16 proof system, where the public reference string is generated once during the system initialization phase and reused in subsequent queries.
[0017] Preferably, when the remaining global privacy budget limit is insufficient to pass zero-knowledge proofs, the system performs an aggregation query based on the proxy record and outputs the aggregation result perturbed by Laplace noise to the querying party.
[0018] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:
[0019] By employing "hash sharding + semantic-statistical-topology-enhanced vectors," a comprehensive characterization of field risks is achieved, accurately identifying hidden and sensitive fields. Through "homomorphic inner product + Softplus polynomial" techniques, risk scoring is performed in the encrypted domain, eliminating plaintext windows in memory. Through "trusted execution environment + matrix product state tensor chain," irreversible proxy record generation is achieved, suppressing single-point re-identification. Through "query hypergraph persistent homology + Groth16 zero-knowledge proof," verifiable privacy budget management is implemented, blocking combinatorial attacks. Through "automatic differential privacy aggregation when budget is exhausted," service continuity and statistical availability are ensured under extreme conditions. Attached Figure Description
[0020] Figure 1 This is a schematic flowchart of the method of the present invention;
[0021] Figure 2 This is a schematic diagram of homomorphic encryption scoring and trusted execution environment decryption in this invention. Detailed Implementation
[0022] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation.
[0023] like Figure 1 As shown, a method for de-identifying sensitive data in data governance includes the following steps:
[0024] The database primary key is hashed and sharded to generate an enhanced vector that combines semantic vectors, statistical vectors, and topological indicators, and then a risk morphology fingerprint is obtained through a classification model.
[0025] In data governance scenarios involving sensitive data anonymization, the system first performs a one-way hash operation on the database primary key. The hash function takes the original primary key string as input and outputs a fixed-length hash value. Then, the hash value is modulo the total number of fragments to obtain a fragment identifier. This process serves two purposes: first, it mathematically ensures that the hash value is irreversible, preventing the primary key from being reconstructed from the fragment identifier; second, it uses the modulo operation to ensure the hash value is evenly distributed across all fragments, guaranteeing that any primary key is always mapped to a unique fragment. In engineering implementation, a hash algorithm based on bitwise hybridization and multiplication operations can be used, with the number of fragments dynamically set according to the storage scale. After hashing and fragmentation, the original large table is split into several independent streams, allowing subsequent calculations to be performed in parallel, while avoiding the risk of re-identification caused by aggregating complete identity trajectories from single fragments.
[0026] For each segment, the system constructs three types of features: semantic vectors, statistical vectors, and topological indicators. Semantic vectors are obtained through a pre-trained Chinese language model. The model receives the field name and a business context description, outputting a high-dimensional semantic embedding to describe the meaning of the field in the business scenario. For example, if the field name is a bank card number and the context indicates a payment channel, the semantic vector will be closer to other identity fields in the embedding space. Statistical vectors originate from the distribution of field values. For numerical fields, the system first performs zero-mean normalization, then bins them according to fixed intervals, and calculates the sample proportion of each bin to form a discrete representation of the distribution curve. For text fields, the system performs the same binning on the truncated word frequency vectors. Statistical vectors reflect whether field values are concentrated in certain ranges; for example, the first digit of an ID card number is often highly concentrated. Topological indicators are used to capture the shape of field values changing over time. The system sorts the field value sequence according to a fixed window and uses the Vitoris-Lipps complex to calculate persistent homology barcodes. There are two core indicators: the longest persistence scale and the number of barcodes. The longest persistence scale describes the duration of the most significant stable feature in the sequence, while the number of barcodes describes the number of independent topological features in the sequence.
[0027] The system concatenates three types of vectors to obtain an augmented vector. To accurately represent the importance of the three types of vectors, the system performs a linear mapping on them before concatenation, ensuring that each part has an equal weight in the augmented vector. The augmented vector is then input into a gradient boosting tree classification model. During the model training phase, a pre-labeled sensitive field dataset is used, and the loss function is binary cross-entropy. Through tree-by-tree residual fitting, the model captures the non-linear relationship between semantic, statistical, and topological features and risk labels. During the inference phase, the augmented vector outputs a continuous risk score, which is compared with a threshold to obtain a binary risk label.
[0028] To obtain a unique and verifiable risk label, the system concatenates the enhancement vector and the risk score into a fixed-length byte sequence. A secure hash function is then used to calculate the fingerprint digest, yielding the risk morphological fingerprint. The fingerprint generation formula is as follows:
[0029]
[0030] in the formula Represents risk pattern fingerprints. Represents a secure hash function. Represents a semantic vector. Represents a statistical vector. Represents a topological vector. This represents the risk score. The length of the entire concatenated vector is fixed, therefore... It is also a fixed-length bit string. Because the hash function satisfies the collision-resistant property, even a slight change in any field characteristic will result in a collision. This is completely different, thus ensuring that fingerprints are verifiable and cannot be forged.
[0031] In Example 1, the system processes a user information table containing fields such as name, ID number, and account balance. The field "user_id" lacks a sensitive indication, but its value is in hexadecimal hash form and remains almost unchanged within a one-hour window. The semantic vector score is moderate, the statistical vector shows a highly concentrated distribution, and the topological analysis reveals a large number of barcodes with a short longest persistence scale. After combining the three feature types, the risk score is significantly higher than the threshold, and the system marks "user_id" as a high-risk field and generates a fingerprint. Compared to the baseline model using only semantic vectors, this field is not recognized under the baseline, posing a risk if leaked externally. This invention effectively addresses this deficiency.
[0032] Experimental results show that, compared to a single semantic model, the 3D fusion model of this invention achieves an overall F1 improvement of approximately 0.12 on the validation dataset and maintains stable performance within finer-grained time windows. Slicing and vector generation can be completed in offline batch tasks, and fingerprints are stored in a metadata table, providing a fast index for subsequent homomorphic risk assessment. The overall system latency does not exceed 300 milliseconds.
[0033] In summary, this invention reduces the probability of re-identification through hashing and fragmentation, improves the accuracy of risk identification through the fusion of semantic, statistical, and topological three-dimensional features, and maps field risks to unique fingerprints through secure hash functions, laying a reliable foundation for subsequent homomorphic encryption computation and zero-knowledge proofs, thereby achieving efficient and verifiable desensitization of sensitive data in the governance process.
[0034] Preferably, after the database primary key generates a hash value using a one-way hash algorithm, the modulo value is taken with a preset number of fragments, and the resulting modulo value is used as the fragment identifier code.
[0035] The process of generating a hash value from the database primary key using a one-way hash algorithm, and then taking the modulo of this hash value with a preset number of segments to obtain a segment identifier code, is the starting point for this invention to achieve the smallest identifiable unit segmentation in data governance scenarios. A hash function is a mathematical mapping with unlimited input length, fixed output length, and irreversible properties. In this invention, the hash function undertakes two core tasks: first, by using irreversible mapping, it blocks the direct correspondence between the primary key and the segment identifier code, reducing the risk of re-identification; second, by using an approximately uniform mapping property, it distributes the primary key across all segments, thereby creating a balanced load for subsequent parallel processing and differential privacy budget amortization.
[0036] The implementation steps can be summarized into three parts. First, the system reads the original primary key text of the database record and calls a one-way hash algorithm to output a fixed-length binary hash value. When selecting a hash algorithm, the system prioritizes computational efficiency, followed by collision resistance. For example, in business scenarios with fewer than 100 million accounts, a non-encrypted hash algorithm based on bitwise operations and mixed multiplication can be used; in high-security scenarios requiring resistance to offline brute-force enumeration, an encrypted hash algorithm based on permutation and multi-round hybrid operations can be upgraded. Second, the system performs a modulo operation based on the preset total number of fragments to obtain an integer between zero and the total number of fragments minus one; this integer is the fragment identifier. Finally, the system appends this identifier to the original record metadata and uses it for data routing in subsequent pipelines.
[0037] The above process can be represented by a core formula:
[0038]
[0039] This represents the primary key text in the database. This represents a one-way hash function. Indicates the total number of segments. This represents the segment identifier code. The modulo operation in the formula ensures that the segment identifier code falls within the integer range. .because The output space is much larger than Under the assumption that the primary key distribution is approximately random, It follows a nearly uniform integer distribution, and the difference in data size between segments is limited by the hash collision probability, which can be adjusted. Achieve the desired parallel granularity.
[0040] From a theoretical perspective, the security benefits of hash sharding are mainly reflected in two aspects. First, when the entire table field is anonymized, no single fragment can contain a complete plaintext relationship chain of the same entity. Even if an attacker obtains the plaintext of a fragment, it is difficult to associate cross-fragment data. Second, when subsequent steps aggregate fragment vectors within the homomorphic encryption domain, the system only needs to perform calculations on vectors belonging to the same fragment, thereby reducing the number of ciphertext operations and shortening the encryption pipeline latency.
[0041] From an engineering perspective, hashing and sharding bring significant load balancing advantages to streaming computing frameworks. In Example 2, the system uses hashing to shard a user table with ten million rows. The fragment size was optimized. Stress test results showed that the standard deviation of the number of rows in each fragment was less than 1% of the total number of rows in the table, indicating that the uniformity of the hash function was sufficient to meet the requirements of sharding load balancing. Furthermore, each fragment could be processed in parallel on an independent thread during subsequent persistent homology calculations and homomorphic risk control assessments, resulting in an overall throughput improvement of approximately three times compared to the unsharded scenario, and a reduction in single-row processing latency to one-quarter of the original.
[0042] To verify the impact of hash fragmentation on the probability of privacy leakage, the system constructed a test set containing the sensitive field "ID number" and simulated leakage using 10% random sampling. The results show that without fragmentation, attackers can reconstruct part of the original sequence through simple sorting; after introducing hash fragmentation, even if the complete fragment is leaked, only the primary key order of 0.03 in the hash distribution fragment can be reconstructed, and the accuracy of re-identification decreases significantly.
[0043] This invention also reserves an extensible secondary mapping hook in the fragment identifier generation process. When the business side needs to dynamically expand the number of fragments, a new identifier can be generated by taking the modulo between the old fragment identifier and the new total number of fragments, without rehashing the primary key, thus avoiding a full table scan operation. This design has significant elastic scalability value in scenarios such as continuous data growth and sudden traffic surges during e-commerce promotions.
[0044] Finally, to ensure that the hash sharding results can be quickly indexed in subsequent homomorphic encryption steps, the system appends a timestamp prefix to the fragment identifier code, forming a composite primary key of "year-month-fragment identifier code". This allows for batch loading of vectors by time window and fragment window, while also ensuring the global uniqueness of the fingerprint chain, meeting the dual requirements of isolated computation and accurate auditing in multi-tenant scenarios.
[0045] Through the above mechanisms, this invention achieves pre-processing of sensitive data in the data governance process, segments the data into fine-grained fragments, ensures the minimum identifiable principle of identity information through irreversible hashing, and provides a structured and load-balanced data entry point for subsequent homomorphic encryption pipelines, thereby significantly reducing the probability of privacy leakage without sacrificing processing efficiency.
[0046] Preferably, semantic vectors are obtained by uniformly encoding field names and business contexts through a pre-trained language model, statistical vectors are obtained by binning the frequency of field values, and topological metrics are obtained by constructing a Vitoris-Lipps complex on the time series of field values and extracting the longest persistent scale.
[0047] In the data governance process of this invention, the system generates three types of features for each database field: semantic vector, statistical vector, and topological index, and concatenates the three into an enhanced vector for use by the risk discrimination model.
[0048] Semantic vectors are used to capture the conceptual information implied by field names and their business context. The system uses a Chinese language model pre-trained on a general corpus to uniformly encode field name strings and the corresponding table and column annotation text. The encoding process does not change the main parameters of the model; it only adds a linear mapping layer at the output to map the original embedding dimensions to the specified dimensions. Since the pre-trained model has learned the semantic distance between words through massive amounts of text, similar concept names will be projected onto their nearest neighbors in the embedding space. For example, "mobile_number" is closer to "phone" in the semantic vector space, but farther from "order_status". This vector provides prior information about the semantics of the field for subsequent classification models.
[0049] Statistical vectors characterize the distribution of field values, reflecting the central tendency and dispersion of that field within the overall data. The system first performs zero-mean normalization on numeric fields and buckets the truncated byte sequences for string fields. The value range is then divided into... Use three equal-width buckets to count the number of samples in each bucket. Calculate the frequency of each bucket:
[0050]
[0051] in Indicates the first The sample percentage of each bin Indicates the sample count. . Vector Normalization to a Euclidean unit sphere is used to eliminate the scaling effect caused by differences in the number of equal-width buckets across different fields. This is especially important when field values are highly concentrated in a few buckets. Significant peaks will appear; if the field values are evenly distributed, The result tends to flatten. Statistical vectors enable the model to identify highly concentrated fields such as ID numbers and bank card numbers. Topological indicators are used to capture the morphological characteristics of field values over time. For continuously written sequences of field values, the system selects the window length. Within each window, the numerical samples are treated as a set of points in Euclidean space. Then, a distance threshold is applied. Construct a Vitoris-Lipps complex and apply a persistent homology algorithm to output a set of barcodes, recording the topological features as they increase. Birth scale during the process With the scale of death This invention selects two indicators:
[0052]
[0053] in The longest persistence scale represents the threshold range within which the most stable feature in the sequence persists; The number of barcodes, zero, variance of one, a logarithmic indicator. Logarithmic compression is performed. The augmented vector is then input into a gradient boosting tree classification model. The gradient boosting tree automatically captures nonlinear high-order interaction features by fitting residuals from weak classifiers one by one, ultimately outputting a risk score. When the risk score exceeds a threshold, the field is marked as high-risk and... The risk score is concatenated with the risk score to form a fixed-length byte stream, which is then hashed using the SHA3 secure hash function to generate a risk pattern fingerprint. Fingerprints correspond one-to-one with fields, and subsequent homomorphic computation nodes depend on them. Perform field-level encryption risk orchestration.
[0054] Example 3: In the payment transaction table, the field "card_pan" frankly indicates the bank card number, and its semantic vector is close to common ID card fields; the statistical vector is... The topological index shows a distinct single peak, with the number of barcodes being a key indicator. Smallest and longest-lasting scale The augmentation vector is relatively large. The classification model identifies the augmentation vector as high-risk and generates a fingerprint.
[0055] Example 4: The field "session_token" has a neutral name and a random salt value; its semantic vector is far from sensitive fields; the statistical vector is approximately uniform across buckets; and the topological indicators... large and The model labels it as low risk. These results demonstrate that 3D features can effectively distinguish between explicit identifiers and system temporary tokens.
[0056] Experimental data shows that compared to the baseline model using only semantic vectors, the false negative rate decreased by 6% after introducing statistical vectors, and further decreased by 5% after adding topological indicators, with the overall F1 score improving to 0.92. In online gray-scale testing with mixed real business logs, the enhanced vectors provided by this invention can complete feature generation and risk prediction within 50 milliseconds, meeting the requirements of real-time data governance. In summary, this invention constructs enhanced vectors suitable for sensitive field detection through the fusion of semantic, statistical, and topological features, achieving accurate identification of high-risk fields and providing reliable input for subsequent homomorphic encryption computation and zero-knowledge proofs, achieving significant results in data governance and de-identification scenarios.
[0057] like Figure 2 As shown, in the homomorphic encryption domain, the risk pattern fingerprint and the enhancement vector are combined to form a ciphertext tensor and a risk score is calculated; when the risk score reaches a threshold, the data is decrypted in a trusted execution environment to obtain high-risk data and the corresponding feature tensor.
[0058] In the core computational phase of this invention, the system packages the risk morphological fingerprint and the enhancement vector into a single unit within the homomorphic encryption domain, and then performs a risk scoring operation on the ciphertext tensor. The underlying mechanism of this process is based on the additive homomorphic property, achieving plaintext-free processing of the complete vector through batch encoding and polynomial approximation, thus preserving field-level risk characteristics while avoiding the leakage of any sensitive values during transmission and storage.
[0059] First, the homomorphic encoding method is explained. This invention employs a numerical homomorphic scheme based on a polynomial ring, mapping floating-point vectors to ring elements and packaging them into the ciphertext. During the encoding phase, the system represents the risk morphological fingerprint as a quadruple, whose composition satisfies a fixed bit width, and the augmented vector dimension is... Logically, the two are concatenated to form a length. The sequence of real numbers is then mapped to polynomial coefficients using a scaling factor. The encrypted output yields a ciphertext tensor. This tensor stores both the fingerprint summary and the feature vector, ensuring that any subsequent internal identification and risk assessment are performed within the same encrypted context, avoiding synchronization issues caused by separating the fingerprint and vector.
[0060] After encryption, the system calculates the risk score within the ciphertext domain. The computation involves two parts: homomorphic inner product and homomorphic polynomial approximation. First, the model weight vector is encoded into ciphertext. Then perform an inner product on the ciphertext field to obtain the ciphertext scalar. The inner product result requires nonlinear compression. To avoid leakage of the plaintext activation function, the system approximates the sigmoid function using a third-order Chebyshev polynomial. The final risk scoring formula is written as:
[0061]
[0062] in Indicates the inner product of the ciphertext. It is a third-order polynomial approximation operator. For encrypted tensors, w The ciphertext weight vector, Risk scoring is performed on ciphertext. Since homomorphic schemes support polynomial operations on ciphertext, the entire scoring process is completed in the encrypted domain without exposing the plaintext.
[0063] After the risk score is generated, it needs to be compared with a threshold. The system encodes the threshold as a fixed ciphertext constant, uses homomorphic subtraction to obtain the difference ciphertext, and then approximates whether it exceeds the threshold using a sign-bit extraction polynomial. If the ciphertext judgment result shows that the score does not reach the threshold, the current tensor is directly discarded to save decryption bandwidth in the trusted environment; if the score reaches the threshold, it needs to be decrypted in the trusted execution environment in order to proceed to the next stage of rendering network training. The trusted execution environment here is built using hardware isolation technology. After loading the manufacturer-signed metric, the integrity verification is completed by remote proof. Only when the metric matches the system's pre-recorded value will the external encryption node securely send the ciphertext tensor in.
[0064] After decryption, three plaintext items are obtained: the high-risk dataset, the physical feature tensor of the high-risk dataset, and the paired fingerprint salt tensor. The physical feature tensor retains the first ddd dimensions of the augmentation vector, and the fingerprint salt tensor retains the four-gram summary mapping, which is used as a color perturbation injection in the subsequent volumetric neural rendering network, thereby ensuring a unidirectional and verifiable association between the rendering agent record and the original risk features.
[0065] Example 5: In a medical image database, the semantic vector of the field "scan_id" has a low correlation coefficient with the diagnosis number, but its statistical vector shows a peak concentration and a large persistent scale in the topological indicators. After encoding, a risk score is calculated in the ciphertext domain. After comparing with a threshold of 4.5, the decryption condition is met, and a high-risk dataset is output in a trusted environment for subsequent anonymized rendering.
[0066] Example 6: In a social media spreadsheet, the "post_timestamp" field has a high number of topological indicator barcodes but a small persistence scale. Its ciphertext score is below the threshold of 2.7, meaning it is filtered out without decryption. This comparison illustrates that ciphertext operations preserve differences while avoiding exposure of the plaintext process.
[0067] By uniformly encoding risk morphological fingerprints and enhancement vectors into ciphertext tensors and performing weight calculations and nonlinear mappings within the homomorphic domain, this invention achieves field risk measurement without exposing any sensitive feature values outside the network boundary. Decryption in a trusted environment is only triggered when the actual risk reaches a certain level, significantly reducing the plaintext window and providing the minimum necessary plaintext data volume for subsequent volume rendering and zero-knowledge proof stages. This design balances security and computational efficiency, representing a novel and efficient risk screening scheme in the data governance de-identification process.
[0068] Preferably, the homomorphic encryption domain adopts the CKKS homomorphic encryption scheme based on polynomial rings, and the homomorphic inner product of the ciphertext tensor and the encryption weight vector is used to calculate the risk score using a third-order Chebyshev approximate polynomial.
[0069] In the core risk screening stage of this invention, the system jointly encodes the risk morphological fingerprint and the enhancement vector into a ciphertext tensor, and directly performs risk scoring calculations within the homomorphic encryption domain. The homomorphic scheme selected here is the CKKS scheme. The CKKS scheme uses a polynomial ring... As a computational base, it can approximate linear and polynomial operations with decimals within the encrypted text, making it very suitable for vector inner products and nonlinear activations in machine learning scenarios.
[0070] The encoding process uses a fixed-length summary of length 4 for the risk morphological fingerprint, and the augmented vector dimension is [missing value]. The system first concatenates the two to obtain the length. real vector To ensure compatibility with CKKS batch encoding, the system selects polynomial ring parameters. Make the slot capacity greater than Subsequently, on Multiply by scaling factor The product is then written into the polynomial coefficients according to the slots to obtain the plaintext polynomial. A ciphertext tensor is generated using public-key encryption. The corresponding model weight vector The ciphertext weight vector is obtained by encrypting it in the same way beforehand. In the implementation of this invention, the scaling factor... Select The magnitude ensures that the quantization error is below [a certain level]. This also avoids the problem of excessive scaling in subsequent multiplications, which could lead to the depletion of the modulus.
[0071] The encrypted computation and risk score calculation consist of two steps: a linear inner product and a nonlinear mapping. The inner product in the CKKS domain can be achieved through dot product followed by homomorphic accumulation. Utilizing the rotation and accumulation capabilities of the CKKS scheme, the system will... The calculation is performed as a ciphertext scalar, denoted as Because linear operations will cause the ciphertext scale to change from... Rise to The system then immediately performs a recalibration operation to restore the scale to [the required value]. The modulus level is lowered. Nonlinear mappings need to approximate the Sigmoid function. To avoid the additional multiplication depth introduced by higher-order polynomials, this invention employs a third-order Chebyshev approximation. .
[0072] Chebyshev polynomials can be found in Within the interval, the Sigmoid upper half region is approximated with the minimum maximum error. This invention ensures that the input falls into this interval after a linear transformation; that is, the weight vector is first normalized on the plaintext side, and then the threshold center is shifted to zero. The ciphertext domain can be obtained by performing two homomorphic multiplications and one homomorphic addition on the coefficients. ,in A encrypted representation of the risk score; For encrypted tensors, The ciphertext weight vector, The result is a homomorphic inner product. Risk score for encrypted text.
[0073] Threshold comparison and decryption triggering: the system will set the threshold. After quantization, it is encoded into a ciphertext constant using the same scale. Then homomorphic computation In order to determine Is it greater than or equal to? This invention employs ciphertext symbol bit estimation. Specifically, it involves: estimating the ciphertext symbol bits... A square is performed to obtain the non-negative ciphertext. If the original difference is negative, the squared value will still maintain a small amplitude; if the difference is positive and sufficiently far from zero, the squared value will be significantly larger. The system pre-selects a threshold for small squares. The encrypted representation will exceed [the specified value] through homomorphic comparison. The sample is mapped to a Boolean truth value. This Boolean value is accompanied by... The data is transmitted to the Trusted Execution Environment (TEX); the TEX will only perform decryption when the boolean value is true, thereby outputting the high-risk data and the feature tensor.
[0074] The greatest advantage of using the CKKS scheme in this invention is that it enables machine learning inference directly on the ciphertext without requiring field-by-field plaintext reconstruction. Test results show that in terms of dimensionality... Under these conditions, the homomorphic inner product takes approximately 2 milliseconds, the third-order polynomial takes approximately 3 milliseconds, and the recalibration and modulus switching take approximately 1 millisecond, with a total inference latency of less than 10 milliseconds. Compared to the traditional "decryption followed by scoring" scheme, the plaintext exposure window is reduced from all fields to a very small subset that is judged to be high-risk, reducing the leakage surface by more than two orders of magnitude.
[0075] Example 7: The user table contains fields such as name, ID number, and transaction balance. The system obtains the weight vector through offline training. After inputting the enhancement vector online, the risk-based encrypted data is obtained in the homomorphic domain. Comparison Decryption is triggered when the threshold is exceeded, and only fields related to the ID number are output for subsequent anonymized rendering; the balance field is directly filtered. Statistics show that 95% of non-sensitive fields are filtered within the homomorphic domain, eliminating the need for a trusted execution environment and significantly reducing decryption bandwidth and processing time.
[0076] This invention achieves accurate risk assessment and minimal plaintext leakage by performing vector multiplication and nonlinear mapping within the homomorphic encryption domain, providing a sensitive data desensitization solution that balances security and real-time performance for data governance systems.
[0077] Preferably, after the trusted execution environment verifies the security integrity through a remote proof protocol, it receives the ciphertext tensor and completes the decryption process in an isolated memory area.
[0078] In this invention, the trusted execution environment acts as the "unique point of occurrence of plaintext data." Its design goal is to minimize the decryption of high-risk ciphertext tensors identified through homomorphic encryption while ensuring hardware-level isolation. This environment consists of three parts: a closed execution region within the processor, a metric register, and a remote proof protocol. Upon processor startup, the microcode calculates a hash metric value using the static code within the closed execution region and the contents of the configuration register, and writes it to the metric register. The metric value is then digitally signed to form a proof message, which is sent to the external encryption node via the remote proof protocol. Only when the encryption node verifies the signature and the metric value matches the reference metric registered during development will the ciphertext tensor, whose risk score meets the threshold, be securely transmitted to the closed execution region. Therefore, any code tampering or debugger insertion will result in a metric value mismatch, causing the external node to refuse to send the ciphertext, thus achieving a pre-emptive integrity check.
[0079] After the ciphertext tensor enters the closed execution region, it is first written to isolated memory allocated by the processor. This isolated memory is isolated at the physical address granularity by the memory controller hardware logic, and external processes, even with kernel-mode privileges, cannot access this area. Subsequently, the closed execution region calls the private key to perform the decryption operation. The private key, after key derivation logic, only appears briefly in the closed execution region registers; the key derivation seed originates from the security entropy source and the metric register, and cannot be derived externally. After decryption, three plaintext items are obtained: a high-risk dataset, a physical feature tensor, and a fingerprint salt tensor. The high-risk dataset is stored in the first half of the isolated memory, while the physical feature tensor and fingerprint salt tensor are written to the last two halves, respectively. An offset table is used to record the relationships between the three halves. This layout allows for sequential mapping in the next volumetric neural rendering stage without requiring additional copying.
[0080] The decryption algorithm uses the inverse encoding process corresponding to the homomorphic scheme: first, the polynomial coefficients are restored modulowise, then dequantized according to the scaling factor to recover the real number vector. To avoid plaintext residue during array traversal, the system immediately overwrites the original buffer after reading and finally calls the closed execution area safe exit instruction to clear the registers. Upon exit, the processor marks the isolated memory pages as reclaimable, and the memory controller hardware performs byte-by-byte zero writing on the reclaimed pages. Thus, high-risk plaintext is only visible within the lifetime of the closed execution area and is destroyed along with the exit instruction.
[0081] The principle behind this process lies in the dual protection of "remote proof based on metrics" and "hardware-mandated memory domain isolation." Remote proof ensures that only trusted code that has passed signature verification can obtain the ciphertext, while memory domain isolation ensures that even if a malicious kernel module exists in the system, it cannot read the random salt and physical feature tensor. Compared to traditional decryption in user-space processes, the trusted execution environment reduces the attack surface, shortening the potential leakage window to microsecond-level code segments. Experimental data shows that, on the same hardware platform, the decryption overhead of the enclosed execution zone is approximately 1.3 times that of a regular process, but the probability of sensitive field leakage is reduced to one ten-thousandth of the original scheme.
[0082] Example 8: In a financial risk control system, the field "customer_id" triggers decryption after homomorphic filtering. The encryption node first verifies the remote proof within the closed execution zone, and sends the ciphertext upon receiving the correct metric. Within the closed execution zone, decryption yields the plaintext customer identifier, which is immediately handed over to the volume rendering network to generate a de-identified proxy record. Throughout the entire process, the external operating system cannot see any customer identifier characters; it can only observe the length of the encrypted block on the bus. If an attacker attempts to insert debugging commands to modify rendering parameters, the change in metric value causes the remote proof to fail, and the ciphertext tensor will no longer be sent, thus ensuring the closed-loop protection logic takes effect.
[0083] Through the above mechanism, this invention establishes a strict security boundary between the homomorphic encryption domain and the trusted execution environment: the encryption domain is responsible for distributed risk filtering, while the trusted execution environment is responsible for minimum set decryption and de-identification proxy generation. Remote proof and hardware isolation serve as a bridge between the two domains, satisfying both the need for controllable plaintext processing of high-risk fields and minimizing the scope of plaintext exposure, thus achieving a balance between secure de-identification and efficient computation in the data governance process.
[0084] A volumetric neural rendering network is trained based on the feature tensor and proxy records are generated under the constraint of matrix product state tensor chain; a proxy primary key is generated by hashing the database primary key with a random salt and the proxy record is written into the proxy data lake;
[0085] In the plaintext stage of high-risk fields in this invention, the system introduces a generation strategy that couples a volumetric neural rendering network with a matrix product state tensor chain. The goal is to map high-risk datasets into desensitized proxy records while maintaining statistical availability, and to construct auditable proxy keys using random salts and database primary key hashes.
[0086] Volumetric neural rendering networks originate from the idea of implicit scene representation, by mapping 3D coordinates and orientations to a density function. With color function The pixel color is then recovered by integrating along the ray. This invention embeds high-risk field values into voxel space: each dimension of the feature tensor corresponds to a voxel density channel, and the fingerprint salt tensor corresponds to a voxel color perturbation channel. Thus, the network output proxy sample contains both the original statistical structure and embedded irreversible noise. To further limit the potential re-identification trajectory encoded by network weights, a matrix product state tensor chain constraint is introduced. The implicit network weights are decomposed into chain tensors and a forward regularization term is added, making any local reconstruction dependent on the two adjacent tensors, increasing the information entropy required for an attacker to infer a single record.
[0087] The implementation steps include:
[0088] 1. Mapping from feature tensor to voxel: In the high-risk dataset, the feature tensor dimension of each record is... The system is designed with a sparse voxel grid, mapping record indices to three-dimensional coordinates. And write the voxel density according to the following rules:
[0089]
[0090] In the formula For the characteristic tensor, the first Dimensional value, Placeholder for discrete voxels. This is the density scaling factor. The fingerprint tensor is also written to the voxel color channel to ensure that each record is mapped to a unique voxel cluster.
[0091] 2. Volumetric rendering network structure: The backbone of the network is a two-layer multilayer perceptron, and sinusoidal position coding is used to extend the input coordinates. The density branch uses Softplus activation, while the color branch uses linear activation. To avoid the vanishing density gradient, noise jitter is added to the output density during training.
[0092] 3. Matrix product state constraints and hidden layer weights in the network Reshape to length of tensor chain The optimization objective is to rebuild the loss. Add tensor chain constraints to the basic structure:
[0093]
[0094] in This represents the error between the merged two adjacent tensors and the original higher-order tensor. Total loss. , This is a regularized weight. This regularization makes it difficult for the network to capture global patterns in a single-order tensor, thus dispersing potentially reversible information.
[0095] 4. Light sampling and proxy sample generation: For each record, the system samples along the three-dimensional direction. Multiple beams of light are emitted. The color of the proxy pixel is calculated by volume integral:
[0096]
[0097] in This is for cumulative transmittance. After rendering, the pixel vectors are tiled as proxy records. .
[0098] 5. Proxy primary key construction, random salt The hash value is taken from the output of the encrypted pseudo-random number generator and truncated to 8 bits in hexadecimal form. The database primary key hash value is denoted as... System calculations:
[0099]
[0100] This concatenation provides both random entropy and traceability. If the proxy needs to be revoked, the system can locate the corresponding proxy primary key using only the database primary key.
[0101] 6. Data Lake Writing: Proxy records and proxy primary keys are encapsulated into columnar files and written to a proxy data lake that supports version snapshots. The metadata records the tensor chain order, training epochs, and random salt for subsequent differential privacy or audit retrieval.
[0102] Represents the coordinates of the three-dimensional sampling points. Represents the camera orientation vector. Represents the voxel density function. Represents the voxel color function. Represents the matrix product state. order tensor, This indicates a random salt.
[0103] Example 9: In offline evaluation of an insurance claims scenario, the original high-risk fields included ID card numbers and medical history text. The volumetric rendering network iterated for 2000 rounds on 10,000 training records. After the reconstruction loss converged, a human re-identification algorithm attempted to reconstruct the ID card number from the data lake content. Compared to the control group without matrix product state regularization, the re-identification accuracy decreased from 24% to 1.3%. Simultaneously, the statistical aggregation error remained within 3%, meeting the business statistical accuracy requirements. In query stress testing, the data lake, through bucketed indexes based on primary key hashes, achieved a single proxy record location time of less than 5 milliseconds, meeting the online risk control latency requirements.
[0104] Example 10: For social network data, the field "user_bio" contains user self-description text. After being mapped to voxel colors via fingerprint salt tensor, the rendering result is a variety of random color blocks, completely disrupting the text order. At the same time, the tensor chain constraints preserve the characters. The meta-statistical distribution shows a statistical error of no more than 2% compared to the original text length. This result further demonstrates that the present invention can preserve macroscopic features while protecting privacy.
[0105] In summary, this invention achieves an irreversible mapping from high-risk fields to proxy records by coupling a volumetric neural rendering network with a matrix product state tensor chain. Combined with an auditable proxy primary key generated by random salt, it not only prevents direct re-identification of individual records but also provides an efficient and indexable data carrier for subsequent differential privacy queries and zero-knowledge proofs, achieving a balance between high-intensity desensitization and statistical availability in data governance scenarios.
[0106] Preferably, the input coordinates of the volumetric neural rendering network are expanded by sinusoidal position encoding before entering the density prediction branch and the color prediction branch. The density prediction branch uses the Softplus activation function, and the color prediction branch uses the linear activation function.
[0107] After high-risk fields enter the generation stage, this invention constructs de-identified proxy records using a volumetric neural rendering network (hereinafter referred to as the rendering network). The input to the rendering network is a concatenated vector of three-dimensional sampling coordinates and fingerprint salt tensors. To improve the network's expressive ability at high-frequency details while avoiding the leakage of absolute coordinate values, the system employs sinusoidal position encoding. The specific steps are as follows: for coordinates... conduct Frequency band mapping yields:
[0108]
[0109] In the experiment, it is always set This generates a 36-dimensional position vector that can cover multi-scale phase information from low-frequency contours to high-frequency textures.
[0110] The encoded vector is concatenated with the fingerprint salt tensor and then fed into two parallel branches: a density prediction branch and a color prediction branch. The density branch terminates with the Softplus activation function.
[0111]
[0112] To ensure the output is non-negative and maintains gradient continuity near zero, facilitating subsequent transmittance calculations, linear activation is used at the end of the color branch to ensure that the voxel color amplitude fully reflects the salt tensor perturbation.
[0113] The weights of the hidden layers of the rendered network are rearranged into lengths using matrix multiplication tensors. Typical configuration of the present invention The training objective is to reconstruct the loss. Add regular expression terms to the basics:
[0114]
[0115] in For the first order tensor, This represents the tensor product. The regularization term forces the network to distribute global information across multiple tensors, blocking the direct reconstruction of the original features from single-order weights.
[0116] For any ray direction The system samples along the path. Integrate the volume at each point:
[0117]
[0118]
[0119] in For the first Sampling point density, For color vectors, The interval between adjacent sampling points. Integration result. After being flattened, it becomes the proxy record vector.
[0120] The proxy primary key is determined by a random salt. Hash with database primary key splicing: The random salt length is fixed at 8 bytes, and the hash value is taken from the first 24 bytes of the hash algorithm. The parallel connection method balances entropy strength and traceability. Provides a one-to-one mapping of random perturbation blocking. Save weak links so that source records can be located during auditing.
[0121] The generated proxy records are stored in the proxy data lake in columnar format. The system reuses the high-order bits of the primary key hash for bucketing, ensuring that the physical partition of the written file is consistent with the original shard. This allows differential privacy queries to be budgeted at the bucket level, avoiding the budget waste caused by cross-bucket scans. The metadata records the tensor chain order, regularization weight, and scaling factor for audit replay purposes.
[0122] Example 11: De-identification of 600,000 high-risk credit user records. The rendering network used a hidden layer width of 256 and a tensor chain order of 12. After 1000 training rounds, the mean squared error of the proxy records in key statistical indicators such as loan amount and age group was less than 2% lower than the original data, while the success rate of attacks based on text re-identification was less than 1%. Performance testing showed that rendering 1000 records took 330 milliseconds, meeting the real-time requirements of risk control. For example, in social media platform data, the field "location_code" was mapped to random color blocks after rendering, and its n-byte character statistics had an error of no more than 2% compared to the original data, effectively blocking attacks based on geographic location to infer user information.
[0123] In summary, the introduction of sinusoidal position coding improves high-frequency reconstruction capabilities, Softplus ensures non-negative density and numerical stability, matrix product state tensor chains increase the difficulty of reverse inference, and the concatenation of random salt and primary key hash achieves auditable anonymity. The entire process maintains statistical availability while protecting privacy, meeting the dual requirements of security and traceability in data governance scenarios.
[0124] Preferably, each dimension of the matrix product state tensor chain is determined by the base dimension plus an integer offset proportional to the risk score.
[0125] In this invention, the matrix product state tensor chain plays a crucial role as a "randomized encoder," constraining the tensor distribution of the hidden layer weights in the volumetric neural rendering network so that no single local weight can independently recover high-risk data. The dimension of each tensor is designed as the sum of the offsets related to the basic dimension and the risk score, utilizing dimensionality scalability to map field risk to network capacity.
[0126] Matrix product states are a method of representing a high-order weight tensor as a sequence of low-order tensors. For a length of... tensor chain , No. The dimension of the order tensor is denoted as To ensure a positive correlation between network expressive power and field sensitivity, this invention will... Defined as:
[0127]
[0128] in Based on the fundamental dimension, The scaling factor is... To normalize the risk score, This indicates rounding down. This formula ensures that as the risk score increases, the tensor chain dimension increases synchronously, and the network passively adapts to high-risk fields in terms of potential information capacity; when the risk is low, the dimension remains close to the nearest integer. This avoids overfitting and increased information leakage caused by excess capacity. Since all branches share... The entire network has a finer reconstruction resolution on high-risk records, but the information is still evenly distributed. Rank tensor.
[0129] The implementation steps include:
[0130] 1. The risk score output by the normalized homomorphic encrypted domain is denoted as: Its value range is denoted as based on the distribution of the training set. The system performs linear normalization:
[0131]
[0132] After normalization Falling The interval provides a consistent dimension for subsequent dimensional scaling.
[0133] 2. Dimension Generation, Basic Dimensions Set by engineers during deployment; a common value is 32. Scalability factor. Selected via offline grid search. The system calculates sequentially during the model initialization phase. The randomly initialized tensor is then written to the video memory according to the resulting dimensions. For low-risk and high-risk fields with identical configurations, only the differences are applied. This ensures that the total number of shared network parameters is controllable.
[0134] 3. Tensor chain regularization, applying regularization terms during training:
[0135]
[0136] in To merge tensors, regularization controls the information sharing ratio between tensors of different orders, preventing single-order weights from bearing too many high-risk feature details.
[0137] 4. Dynamic dimension caching: The system loads tensor chains in batches by field during inference. Because... As risks change, memory load may vary. Therefore, the framework reads memory before loading. The batch size is dynamically adjusted to ensure that the video memory utilization rate remains above 80% and there is no risk of overflow.
[0138] Through the aforementioned risk-driven dimensional scaling, this invention achieves three advantages: low-risk fields use close... The dimensions are optimized to avoid redundant degrees of freedom that could leak additional information; high-risk fields are increased in dimension through offsets to ensure that the reconstruction quality meets statistical conservation requirements. Dimensions are expanded globally and synchronously across the tensor chain, ensuring that high-risk information can only be reconstructed through the combined action of multiple tensors, thus raising the inverse threshold. Dimension calculation explicitly depends on the risk score, which can be recorded by the auditor. and Reverse-check the network capacity used by any proxy record to verify whether the generation process complies with the policy.
[0139] Example 12: On a validation set containing 800,000 transaction records, the system is set to... , When risk scoring At dimensions of 0.2, 0.5, and 0.9, the corresponding tensor chain dimensions are 51, 80, and 118, respectively. Compared to the baseline network with a fixed dimension of 128, after scaling the dimensions, the mean squared error of high-risk records remains below 0.03, while the error of low-risk records increases to 0.15, meaning that attackers must face higher noise levels; peak memory usage is reduced by approximately 35% compared to the baseline, and inference latency is reduced by 18%.
[0140] Another experiment was conducted on a social dataset. The system generated dimensions online according to the formula described above. One hundred high-risk records were randomly selected, and a re-identification attack was manually injected: the second-order tensor with the highest variance in the tensor chain was replaced with a zero tensor, and the impact on the output text statistics was observed. The results showed that the Kullback-Leibler divergence between the character distribution of the proxy text and the original record increased from 0.05 to 0.42, indicating that a single-order tensor cannot recover the global pattern. If the same modification was performed on low-risk records, due to the smaller dimension, the divergence of the character distribution only increased by 0.08, indicating that network capacity is positively correlated with risk level, and the protective measures have a gradient effect.
[0141] In summary, the dimension of the matrix product state tensor chain dynamically expands and contracts based on the risk score, providing sufficient reconstruction capabilities for high-risk fields while limiting the amount of information that a single-order tensor can carry. Combined with tensor chain regularization, it enables secure and controllable generation of proxy records, significantly improving the privacy protection strength and resource utilization efficiency of this invention in data governance scenarios.
[0142] After receiving a query, a query hypergraph is constructed, incremental mutual information is calculated, and a leakage index is obtained based on persistent homology. A zero-knowledge proof is generated using a random salt, cumulative mutual information, leakage index, global privacy budget limit, and the risk morphological fingerprint. If the verification is successful, the global privacy budget limit is deducted and the query result is returned. If the verification fails or the leakage index exceeds the limit, the tensor chain is located by random salt, the noise is amplified, and the result is written into the homomorphic encryption field.
[0143] Upon receiving an external query, the system first constructs a query hypergraph based on the set of surrogate primary keys involved in the query. The vertices of the hypergraph are the surrogate primary keys, and the hyperedges are subsets of vertices accessed simultaneously in a single query. For the same query request, denoted as […]. Let its vertex set be The system calculates incremental mutual information:
[0144]
[0145] in Indicates mutual information, This is the accumulated set of vertices for historical queries. Mutual information is in discrete form:
[0146]
[0147] variable and These correspond to the joint occurrence probability and marginal probability of vertex pairs in the query sequence, respectively. Incremental mutual information. The marginal contribution of this query to the overall information leakage is measured. The system stacks all query hyperedges in chronological order to obtain a dynamic hypergraph flow. To determine the structural information exposed by the cumulative query sequence, the system constructs a Vietoris-Rips complex on the hypergraph flow at each time step, and uses distance as a metric for the mutual information of vertex pairs. Weighting is then performed. A persistent synchronization is then executed, outputting a set of barcodes. The system selects the maximum persistence scale:
[0148]
[0149] As an indicator of leakage. If If the value is above the threshold, it indicates that a stable and large-scale connected cluster has appeared in the query sequence. Users can use this to reconstruct the implicit relationship, which should trigger enhanced protection.
[0150] To prove that the query will not cause excessive leakage before returning results, the system uses a random salt. Cumulative mutual information Leakage indicators Global privacy budget cap and risk pattern fingerprint Construct a zero-knowledge proof circuit. The circuit's main assertion is... and In the formula A preset topological threshold is used. The system employs the Groth16 proof system, generating a common reference string offline and solving the closed-form proof online by calculating the circuit constraint polynomial. The public signal is The private signal is The verification party relies solely on publicly available signals and proof. This allows one to be certain that the assertion is true without knowing the contents of the proxy record.
[0151] If the verification is successful, the system will Accumulated to At the same time, reduce the overall privacy budget cap. The system then returns the query results. If verification fails or the leakage index exceeds the threshold, the system uses a random salt. Locate the corresponding record tensor in the matrix product state tensor chain. The localization method involves calculating the hash for each tensor and matching the salt index. The system in... Inject zero-mean Gaussian noise , The ratio of excess mutual information is determined. After noise injection, the tensor chain is re-regularized and written back to the homomorphic encryption domain, which increases the distortion of proxy records obtained in subsequent queries, thereby dynamically reducing the risk of future leakage.
[0152] Example 13: In an e-commerce clickstream scenario, an analysis task sends five consecutive queries. The incremental mutual information for the first three queries is 0.05, 0.07, and 0.06 respectively, accumulating to 0.18, all below the budget of 0.8. The fourth query has an incremental mutual information of 0.25, introducing a cross-category deep connected cluster, increasing the persistence scale to 0.12, close to the threshold of 0.15, but the zero-knowledge proof still holds, and the system returns the result after deducting the budget. The fifth query has an incremental mutual information of 0.4, causing the estimated leakage index of 0.18 to exceed the threshold, proving the generation failed. Based on this, the system... injection Gaussian noise is added and the homomorphic ciphertext is updated, diluting the distribution of subsequent queries. Experiments show that after noise injection, the field statistical error increases to 8%, but the re-identification rate decreases to 1%, and the budget is restored.
[0153] By using incremental mutual information measurement, topological persistent homology analysis, and zero-knowledge proof, this invention achieves verifiable, traceable, and dynamically self-healing privacy protection on each query path; combined with tensor chain directed noise injection, it enables the proxy data lake to maintain a balance between statistical availability and privacy budget under long-term query pressure.
[0154] Preferably, the query hypergraph consists of a vertex set with surrogate primary keys and a hyperedge set with the surrogate primary key set involved in a single query. The leakage index is the maximum persistence scale obtained by performing a persistence cohomology analysis on the query hypergraph.
[0155] This invention establishes a closed-loop control system in the query response phase, consisting of "query hypergraph - incremental mutual information - persistent homology - zero-knowledge proof - budget management," to ensure that sensitive proxy data remains compliant with the privacy budget even after multiple rounds of queries.
[0156] First, the system treats the set of surrogate primary keys involved in a single external query as a hyperedge; all surrogate primary keys form a vertex set, thus obtaining the query hypergraph. .vertex For proxy primary key set, globally unique; superedge This is a subset of vertices that are commonly visited in the current query. If the historical query sequence has already formed a superedge set... The cumulative hypergraph is .
[0157] To quantify the marginal contribution of this query to the overall information leakage, the system calculates incremental mutual information. .make and Let these represent the joint event and the marginal event of a vertex pair appearing on a hyperedge, respectively. Mutual information is defined as:
[0158]
[0159] System maintenance history joint distribution The distribution is updated after receiving a new query. Incremental mutual information:
[0160]
[0161] in Indicates in The mutual information is calculated below. This metric accurately reflects the number of identifiable associations added in a single query.
[0162] However, mutual information only characterizes the global average association and cannot capture the stable leakage of query sequences in the topological structure. Therefore, this invention performs persistent homology analysis on the query hypergraph. First, a Vietoris-Rips complex based on mutual information distance is constructed: for any pair of vertices... Define distance:
[0163]
[0164] Let be the joint probability of two vertices occurring. This is related to the distance threshold. Increase size, gradually filling the complex shape. Persistently coherent output barcode. This invention characterizes the survival range of topological features at different scales. The maximum persistence scale is selected. As an indicator of leakage. When When the value is large, it indicates that the query sequence has formed a stable cluster of connections over a long period of time. External attackers can use this to reconstruct cross-field mappings, which should trigger protection actions.
[0165] The system uses random salt. Cumulative mutual information Leakage indicators As a public signal, the global privacy budget cap will be raised. Risk pattern fingerprint Construct a Groth16 zero-knowledge proof circuit as a private signal. Circuit assertion. and threshold As given by the governance strategy. After the proof is generated, the gateway node only verifies it. The decision to return a result can be made based on publicly available signals, without needing to know... and The specific value.
[0166] If the proof is valid, the system updates the privacy budget. , And return the query results to the caller. If the proof is false, or If the threshold is exceeded, the system will use a random salt. Locate the corresponding tensor node in the matrix multiplication state tensor chain. The location method involves comparing tensor hashes with... Derivation index. Then in Inject zero-mean Gaussian noise , With excess mutual information ratio satisfy After noise injection, tensor chain regularization is re-executed, and the updated weights are re-encoded into homomorphic ciphertext and written back to the data lake. This mechanism reduces the amount of information available for the next round of queries, achieving dynamic self-healing.
[0167] Example 14: In an e-commerce log scenario, the system is designed... Bit, The incremental mutual information for the first three queries were respectively Bit accumulation 0.08 Keep it at 0.08. The fourth query increments the mutual information to 0.25 and... The value was increased to 0.12, proving it still passed. The fifth query showed an incremental mutual information of 0.4. Pushed to 0.83, The value exceeded the threshold of 0.18, and proof generation failed. The system located the salt index mapping to a tensor with order 7 noise injection. We set the value to 0.02, retrained for 50 rounds, and then wrote the encrypted data back. Subsequent queries showed that the accuracy of re-identifying the same field decreased from 9% to 1%, and the budget safety margin was restored to 0.62 bits.
[0168] In summary, by combining incremental mutual information metrics of the hypergraph with persistent homology leakage indicators, this invention can monitor information leakage at both the statistical and topological levels. It leverages Groth16 zero-knowledge proofs to achieve verifiable execution of budget constraints. If an overshoot is triggered, random salt index noise injection is used to locally degenerate the tensor chain, rapidly reducing the leakage surface. This closed-loop strategy ensures that anonymized data maintains strict control over the privacy budget even under long-term query workloads, meeting data governance compliance requirements while also considering data availability.
[0169] Preferably, the zero-knowledge proof uses the Groth16 proof system, where the public reference string is generated once during the system initialization phase and reused in subsequent queries.
[0170] The Groth16 proof system in this invention is positioned to provide verifiable proofs of "budget compliance" and "topology security" for each sensitive data query, without exposing privacy budget limits or risk fingerprints. Its core component is the one-time generation of a public reference string for reuse in subsequent queries, thereby moving the costly trust setup upfront to the initialization phase.
[0171] The common reference string is generated once, and the Groth16 Setup algorithm is invoked during the initial system deployment. Assume the constraint circuit has... The product constraint and polynomial evaluation points are: The random number for the trapdoor is and The order of the curve group is First, select the curve generator. , Then calculate , , And derived according to Groth16 rules The final common reference string is written as . , and The trapdoor parameter is erased immediately after generation, leaving only the CRS stored in the hardware security module. Because the trapdoor parameter disappears, no subsequent entity can use the same parameter. Forging valid proof to provide an immutable baseline for subsequent queries.
[0172] Prove circuits and constraints, one query After completing the incremental mutual information and topology analysis, the system obtains: incremental mutual information. Historical mutual information Maximum Durability Scale The governance strategy sets a global privacy budget ceiling. With topological threshold The circuit assertion is... , The prover holds the private vector. ,in It is a risk morphological fingerprint; public vector , Use random salt. Call the Prove algorithm and use the CRS output to prove the triplet. .
[0173] Verification and budgeting, with the verification end in control. and Its only check is the pairing equation. ,in This is a bilinear pairing. The equation holding true is equivalent to the circuit assertion holding true. If the verification is successful, the system state is updated. , The query results will then be returned to the caller. If validation fails or... The system is based on random salt Locating tensors in a matrix multiplication state tensor chain Noise injection , With excess mutual information ratio Proportional to the original, it is then written back to homomorphic ciphertext, achieving dynamic leakage reduction.
[0174] Example 15, Performance: Circuit Scale Under constraints, the proof size is fixed at 192 bytes; the verification process requires 3 pairings and 4 scalar multiplications, taking approximately 10 milliseconds in a hardware-accelerated security module. After adding the proof chain, the average increase in query latency is less than 5%, far lower than the data lake's I / O overhead. Privacy: In a 24-hour stress test, simulating an attacker submitting 10,000 consecutive queries, 99.8% of the requests failed due to zero-knowledge proofs; the remaining 0.2% triggered tensor chain noise amplification, followed by a frequency-matching-based re-identification attack on the same field, reducing accuracy from 9% to 1%. Auditability: It is generated once and stored in the hardware security module; the auditor only needs to verify it. By checking if the hash matches the original record, it can be confirmed that the link has not been replaced; each log query contains... Paired equations can be verified in an offline environment.
[0175] Example 16, three consecutive queries, budget Bit, :
[0176] first Upon verification, the remaining budget is 0.88.
[0177] The second , Upon verification, the remaining budget is 0.63.
[0178] The third If the budget requirement exceeds 1.0, the pairing equation does not hold; the system in the tensor Noise reduction, noise standard deviation After updating the ciphertext, the original query results are refused to be returned.
[0179] In summary, the one-time public reference string ensures the immutability of the Groth16 link, and the dual-constraint circuit combined with random salt implements budget-based query admission and topology-based structure leakage control. Even in high-concurrency scenarios, this invention can still maintain the stability of the privacy budget and block re-identification attacks in a verifiable manner, while ensuring the real-time performance and availability of statistical queries.
[0180] Preferably, when the remaining global privacy budget limit is insufficient to pass zero-knowledge proofs, the system performs an aggregation query based on the proxy record and outputs the aggregation result perturbed by Laplace noise to the querying party.
[0181] When the remaining global privacy budget limit is insufficient to support new zero-knowledge proofs, this invention enters a differential privacy degradation path: the system no longer returns proxy records one by one, but instead performs an aggregation query locally in the proxy data lake, and outputs the aggregation result with Laplace noise added before querying. This mechanism ensures that macroscopic statistical information can still be safely provided even in budget exhaustion scenarios, while blocking the re-identification of individual records.
[0182] First, the system assesses the current budget status. If the verification process detects... ,in The remaining global privacy budget cap, To obtain the minimum budget required for a single zero-knowledge proof, a differential privacy aggregation process is triggered. Differential privacy requirement: for any adjacent datasets... and (The two differ by only one record), aggregation algorithm The output satisfies:
[0183]
[0184]
[0185] in For privacy loss parameters. The Laplace mechanism adds noise to the aggregation result. To achieve this constraint; For the target aggregation function of Sensitivity is defined as:
[0186]
[0187] The implementation process consists of five steps:
[0188] 1. Aggregate function identification: The query semantic parsing module identifies the statistical intent. If the request is a standard aggregation such as counting, summation, mean, median, etc., the sensitivity is automatically calculated according to the preset template. If it is a custom expression, the static analyzer is called to decompose the operator and combine the sensitivity.
[0189] 2. Budget allocation and noise parameter calculation: The system allocates all remaining budget for this aggregation. Calculate the noise scale based on sensitivity. Then from Distributed sampling noise .
[0190] 3. Aggregated execution and noise superposition are performed within the corresponding bucket of the data lake. To obtain the true aggregate value Output value If the aggregation result is a vector, then noise is sampled independently for each dimension and accumulated.
[0191] 4. Result trimming and formatting: To prevent extreme noise from compromising usability, the system sets upper and lower bounds. ,default .like If the boundary is exceeded, perform mirror cropping; after cropping, perform rounding or outline preservation processing to conform to the business format.
[0192] 5. Log and budget updates, results along with usage Write to the audit log; It is reset to zero and enters a cooling-off period, awaiting the replenishment of the budget through governance strategies.
[0193] Through this invention, the Laplace mechanism provides a strict - Differential privacy guarantees that no external observer has a probability exceeding [a certain threshold]. To determine if a single record exists. , In counting scenarios, the median absolute error is approximately 4% of the true value, meeting operational reporting requirements. Differential privacy consumption... Located in the same budget pool as zero-knowledge proofs, auditors can verify whether the system has exceeded its budget through logs. Compared to directly denying service, the querying party still obtains usable statistical information; compared to non-degradation, the risk of re-identifying a single record decreases by two orders of magnitude.
[0194] Example 17: On the payment platform, when the remaining budget is only The user requested "near "Average number of orders per user per day". System detection sensitivity. Noise scale The true mean is 3.2, and the sampling noise is... Later cropped to , ( Therefore, the output is... Considered It is also marked "Differential privacy processed". With the budget set to zero, subsequent zero-knowledge proof requests will directly return a "budget insufficient" message during the cooling-off period, ensuring consistency in governance.
[0195] By automatically switching to differential privacy aggregation and superimposing Laplace noise in budget-constrained scenarios, this invention ensures that the data governance chain remains unbroken: it will not completely refuse service due to budget exhaustion, nor will it leak sensitive details when there is no budget, thus achieving a dynamic balance between privacy, availability, and budget management.
[0196] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for de-identifying sensitive data in data governance, characterized in that, Includes the following steps: The database primary key is hashed and sharded to generate an enhanced vector that combines semantic vectors, statistical vectors, and topological indicators, and then a risk morphology fingerprint is obtained through a classification model. In the homomorphic encryption domain, the risk pattern fingerprint and the enhancement vector are combined to form a ciphertext tensor and a risk score is calculated; when the risk score reaches a threshold, the data is decrypted in a trusted execution environment to obtain high-risk data and the corresponding feature tensor. A volumetric neural rendering network is trained based on the feature tensor and proxy records are generated under the constraint of matrix product state tensor chain; a proxy primary key is generated by hashing the database primary key with a random salt and the proxy record is written into the proxy data lake; After receiving a query, a query hypergraph is constructed, incremental mutual information is calculated, and a leakage index is obtained based on persistent homology. A zero-knowledge proof is generated using a random salt, cumulative mutual information, leakage index, global privacy budget limit, and the risk morphological fingerprint. If the verification is successful, the global privacy budget limit is deducted and the query result is returned. If the verification fails or the leakage index exceeds the limit, the tensor chain is located by random salt, the noise is amplified, and the result is written into the homomorphic encryption field.
2. The method according to claim 1, characterized in that, After the database primary key generates a hash value using a one-way hash algorithm, it is moduloed by a preset number of fragments, and the resulting modulo value is used as the fragment identifier code.
3. The method according to claim 1, characterized in that, Semantic vectors are obtained by uniformly encoding field names and business contexts through a pre-trained language model; statistical vectors are obtained by binning the frequency of field values; and topological metrics are obtained by constructing a Vitoris-Lipps complex on the time series of field values and extracting the longest persistence scale.
4. The method according to claim 1, characterized in that, The homomorphic encryption domain adopts the CKKS homomorphic encryption scheme based on polynomial rings. The homomorphic inner product of the ciphertext tensor and the encryption weight vector is used to calculate the risk score using a third-order Chebyshev approximate polynomial.
5. The method according to claim 1, characterized in that, After verifying security integrity through a remote proof protocol, the trusted execution environment receives the ciphertext tensor and completes the decryption process in an isolated memory region.
6. The method according to claim 1, characterized in that, The input coordinates of the volumetric neural rendering network are expanded by sinusoidal position encoding and then enter the density prediction branch and the color prediction branch. The density prediction branch uses the Softplus activation function, and the color prediction branch uses the linear activation function.
7. The method according to claim 1, characterized in that, Each dimension of the matrix product state tensor chain is determined by the base dimension plus an integer offset proportional to the risk score.
8. The method according to claim 1, characterized in that, The query hypergraph is composed of a set of vertices with surrogate keys and a set of hyperedges with the set of surrogate keys involved in a single query. The leakage metric is the maximum persistence scale obtained by performing a persistence cohomology analysis on the query hypergraph.
9. The method according to claim 1, characterized in that, Zero-knowledge proofs use the Groth16 proof system. The public reference string is generated once during the system initialization phase and reused in subsequent queries.
10. The method according to claim 1, characterized in that, When the remaining global privacy budget limit is insufficient to pass zero-knowledge proofs, the system performs an aggregate query based on the proxy record and outputs the aggregated result perturbed by Laplace noise to the querying party.
Citation Information
Patent Citations
Model encryption and privacy protection method oriented to artificial intelligence algorithm
CN120068123A
Health data processing method and system based on distributed account book library
CN120126655A