Cross-institutional user numerical data processing method based on privacy protection computing
By using dimensional metadata and global quantile summary commitments to achieve standardization and distribution alignment of cross-institutional data in the dense domain, the problem of privacy leakage and data inconsistency in cross-institutional data collaborative processing is solved, ensuring data comparability and privacy protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- KAIENTAI (NANJING) TECH CO LTD
- Filing Date
- 2026-03-03
- Publication Date
- 2026-05-29
AI Technical Summary
When processing data collaboratively across institutions, existing technologies pose privacy risks under plaintext alignment and centralized normalization methods, and cannot achieve dimensional standardization and distribution alignment in the dense domain, resulting in inconsistent numerical scales, distorted statistical features, and deviations in the task model.
By determining dimensional metadata, constructing a multi-objective matching graph and solving the matching relationship, generating reversible scaling transformation parameters and global quantile summary commitments, dense-state encoding and monotonic quantile mapping of user numerical data are realized, and data standardization and distribution alignment within the dense-state domain are completed.
It achieves data standardization and distribution alignment under heterogeneous units, time granularity, and statistical caliber while maintaining privacy and security, ensuring consistency in accuracy and protection of privacy.
Smart Images

Figure CN121786887B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method for cross-organizational user numerical data processing based on privacy-preserving computation. Background Technology
[0002] Cross-institutional data collaborative processing technologies, when dealing with user numerical data, generally rely on plaintext alignment and centralized normalization. Due to differences in data units, statistical standards, and time granularity among different institutions, traditional methods often require unified preprocessing before data aggregation, leading to high privacy risks and high processing costs. In scenarios employing encrypted computing, existing technologies typically cannot achieve dimensional standardization and distribution alignment in the encrypted domain, resulting in inconsistencies in numerical scales, statistical distortion, and task model deviations between data from different institutions. Furthermore, when multiple recipients exist, traditional alignment strategies struggle to simultaneously accommodate the distribution characteristics of multiple terminals, easily leading to results biased towards a single institution or a decrease in overall accuracy, making it difficult to achieve fair fusion and stable modeling effects across institutions.
[0003] To address the above issues, this application proposes a cross-institutional user numerical data processing method based on privacy-preserving computation. Summary of the Invention
[0004] The technical problem this application aims to solve is to address the shortcomings of existing technologies by providing a cross-institutional user numerical data processing method based on privacy-preserving computation. This method determines dimensional metadata corresponding to a first data terminal and multiple second data terminals. The auditing end constructs a multi-objective matching graph based on the dimensional metadata and generates matching relationships through weighted matching and minimum inconsistency closure. Based on this, reversible scaling parameters and global quantile summary commitments are calculated. The first data terminal performs dense-state encoding on the user numerical data according to the scaling parameters to generate standardized dense-state features. The second terminals perform monotone quantile mapping and calibration processing in the dense-state domain based on the global quantile summary commitments to obtain aligned dense-state user numerical data. This method achieves privacy-preserving data standardization and distribution alignment under heterogeneous units, time granularities, and statistical calibers, balancing accuracy consistency and privacy protection.
[0005] To achieve the above objectives, this application provides the following technical solution:
[0006] A cross-institutional user numerical data processing method based on privacy-preserving computation is applied to the data terminal of a cross-institutional privacy-preserving computation system, which also includes an auditing terminal. The method includes:
[0007] Determine the dimensional metadata corresponding to the first mechanism end to send user numerical data and the second mechanism end to receive user numerical data, wherein the second mechanism end includes at least one data terminal;
[0008] The dimensional metadata is sent to the auditing end to obtain the reversible scaling parameters and the corresponding parameter commitments and global quantile summary commitments.
[0009] The first mechanism performs dense-state encoding on the user numerical data according to the reversible scaling transformation parameters to obtain standardized dense-state features, and then sends the standardized dense-state features to the second mechanism.
[0010] The second agency performs monotonic quantile mapping on the standardized dense-state features in the dense-state domain based on the global quantile summary commitment, and obtains aligned dense-state user numerical data through truncation processing.
[0011] The dimensional metadata corresponding to the first mechanism end for determining the user numerical data to be sent and the second mechanism end for determining the user numerical data to be received includes:
[0012] Based on the attribute information of the user numerical data to be sent stored in the data terminal of the first institution, a first dimension metadata is generated, wherein the attribute information includes data unit, currency type, time granularity and statistical caliber.
[0013] For each data terminal in the second agency, obtain the corresponding attribute description information and generate a second-dimensional metadata set.
[0014] Based on the reversible scaling transformation parameters, dense-state encoding is performed on the user numerical data to obtain standardized dense-state features, including:
[0015] The zero-knowledge range verification is performed on the scaling coefficient, translation coefficient, and time normalization factor in the reversible scaling parameters based on the preset public key verification signature.
[0016] The piecewise linear scaling operator is calculated based on the verified scaling and translation coefficients, and the time resampling kernel is calculated based on the verified time normalization factor.
[0017] For each user's numerical data, the plaintext numerical data is quantized at fixed points according to the piecewise linear scaling operator and the time resampling kernel, the numerical granularity is made consistent, and the quantized and consistent user numerical data is compressed in intervals and truncated at bit width by using the quantization step size and truncation bit width that match the parameter commitment, so as to obtain a standardized fixed-point vector.
[0018] The standardized fixed-point vector is encoded in a secret state according to the secret state mask that corresponds one-to-one with the standardized fixed-point vector to generate a standardized secret state feature. The secret state mask is obtained by secret sharing decomposition of the integrity identifier and abnormal missing identifier of the user numerical data and adding random noise perturbation.
[0019] The calculation method for the dense-state mask includes:
[0020] Extract the integrity identifier corresponding to the user numerical data, wherein the integrity identifier is used to characterize the validity status of the user numerical data during the acquisition, transmission and storage stages;
[0021] Based on the integrity identifier and the preset anomaly detection rules, the consistency of user numerical data is determined and missing data is identified to obtain an anomaly missing identifier.
[0022] The abnormal missing identifier is used as input, and the user's numerical data is decomposed by the threshold secret sharing algorithm to obtain multiple secret shares that meet the preset recovery threshold. Each secret share includes a random salt value and a corresponding timestamp. The random salt value is calculated by the threshold secret sharing algorithm.
[0023] For each secret share, a perturbation share corresponding to the secret share is calculated based on a preset privacy budget and perturbation distribution function, and the perturbation share is encapsulated into a secret mask component through homomorphic encryption;
[0024] By combining all the dense state mask components through weighted aggregation, a dense state mask corresponding one-to-one with the normalized fixed-point vector is obtained.
[0025] Based on the global quantile summary commitment, the normalized dense-state features are monotonically mapped in the dense-state domain, and aligned dense-state user numerical data is obtained through truncation, including:
[0026] For each data terminal in the second agency, the standardized encrypted features transmitted by the first agency and the global quantile digest commitment transmitted by the auditing agency are received.
[0027] The global quantile digest commitment is signed and verified based on the public key and commitment issuance information corresponding to the data terminal, and the quantile cut-off sequence and corresponding representative value sequence of the global quantile digest commitment in the data terminal are extracted.
[0028] Based on the quantile tangent sequence, piecewise linear interpolation is performed on the normalized dense state features in the dense state domain to map the normalized dense state features to the initial quantile mapping result;
[0029] Based on the representative value sequence, the initial quantile mapping result is segmented in the dense state domain, the difference in the dense state mean of each segment is calculated to obtain residual data, and the residual correction function is calculated based on the residual data.
[0030] The calibration function corresponding to the residual correction function is generated by the order-preserving regression constraint, wherein the calibration function is a monotonically non-decreasing function, and the constraint conditions of the order-preserving regression constraint correspond to the upper and lower threshold information in the global quantile summary commitment;
[0031] The initial quantile mapping result is corrected and truncated using the calibration function to obtain the aligned dense quantile mapping result.
[0032] Extracting the global quantile summary commitment into the data terminal the quantile cut-off sequence and the corresponding representative value sequence, including:
[0033] Based on the global quantile index mapping table in the global quantile digest commitment, and in conjunction with the service domain identifier of the data terminal, the quantile index range corresponding to the data terminal is determined;
[0034] The quantile digest structure in the global quantile digest commitment is indexed and located by a dense state retrieval algorithm, and the dense state quantile cut-off point value corresponding to the quantile index range is extracted to generate a quantile cut-off point sequence.
[0035] Based on the quantile cut-off sequence, the corresponding dense state representative values are extracted from the representative value set in the global quantile summary commitment to obtain a representative value sequence that corresponds one-to-one with the quantile cut-off sequence.
[0036] The step of generating the calibration function corresponding to the residual correction function through ordinal-preserving regression constraints includes:
[0037] Calculate the piecewise residual sequence corresponding to the residual correction function in the dense state domain, wherein the piecewise residual sequence is used to characterize the deviation trend of the initial quantile mapping result in the quantile interval;
[0038] Based on the upper and lower threshold information in the global quantile summary commitment, constraints are constructed for the residual correction function. These constraints include non-subtractive relation constraints on quantile segment residuals, residual magnitude limitation constraints, and total change constraints.
[0039] Based on the constraints, the segmented residual sequence is subjected to order-preserving regression calculation. The residuals that do not satisfy the monotonic relationship are merged by dense-state weighted average and dense-state comparison operation to obtain a dense-state calibration sequence that is adapted to the monotonic non-decreasing characteristics.
[0040] Based on the dense-state calibration sequence, a piecewise linear expression of the calibration function is constructed in the dense-state domain, and the calibration slope and offset parameters of each segment are determined to generate the calibration function.
[0041] A cross-institutional user numerical data processing method based on privacy-preserving computation is applied to the audit end of a cross-institutional privacy-preserving computation system. The cross-institutional privacy-preserving computation system further includes data terminals: a first institutional end for transmitting user numerical data, and a second institutional end for receiving user numerical data. The second institutional end includes at least one data terminal. The method includes:
[0042] Receive dimensional metadata from the first institution and a set of dimensional metadata from the second institution;
[0043] Based on a preset public key, the dimensional metadata and the dimensional metadata set are signed and their integrity verified, and then standardized and encoded to generate a first dimensional structure signature and a second dimensional structure signature.
[0044] Using the first dimensional structure signature as the source node and the second dimensional structure signature as the target node, a multi-target matching graph is constructed, wherein the mapping edges of the multi-target matching graph are established through the unit interchangeability, time granularity scalability and statistical caliber compatibility between dimensional structure signatures.
[0045] The matching relationships are obtained by performing weighted matching and minimum inconsistency closure on the multi-objective matching graph.
[0046] The reversible scale transformation parameters of the first mechanism end are calculated based on the matching relationship, and the corresponding parameter commitments are generated. The reversible scale transformation parameters include a scaling factor, a translation factor, and a time normalization factor.
[0047] Based on the matching cost corresponding to the matching relationship, the multi-target matching graph is updated through differential privacy protection to obtain a matching distribution graph, and a corresponding global quantile summary commitment is generated based on the matching distribution graph.
[0048] Weighted matching and minimum inconsistency closure are performed on the multi-objective matching graph to obtain the matching relationships, including:
[0049] In the multi-target matching graph, conflicting mapping edges are identified. For the detected inconsistent mapping edges, the conflict priority is marked according to the weight value and reversibility constraint. The conflict includes one-to-many mapping conflict, many-to-one mapping conflict and same-dimensional parameter conflict.
[0050] The marked conflict mapping edges are subjected to closure operation, and the conflict mapping edges are eliminated by constraint optimization algorithm to obtain the preliminary matching relationship.
[0051] Based on the mapping edge corresponding to the preliminary matching relationship, a hierarchical adjudication is performed according to the preset adjudication priority rules. The preliminary matching relationship is then modified based on the adjudication results to obtain a matching relationship. The adjudication priority rules include regulatory priority, compliance priority, business priority, and historical version priority.
[0052] Based on the matching cost corresponding to the matching relationship, the multi-target matching graph is updated using differential privacy protection to obtain a matching distribution graph, including:
[0053] For each mapping edge in the matching relationship, the matching cost is calculated based on the unit conversion error, time granularity scaling error and statistical caliber difference between the dimensional structure signatures, and a matching cost matrix is generated.
[0054] Based on the preset privacy budget parameters, differential privacy perturbation is performed on each matching cost value in the matching cost matrix; the edge weights in the multi-target matching graph are updated according to the perturbed matching cost values, the matching probability distribution is calculated, and a matching distribution graph is generated.
[0055] Compared with the prior art, the beneficial effects of this application are:
[0056] This application achieves unified scale standardization and distribution alignment of numerical data from multiple institutions in the dense domain by introducing reversible scaling parameters and global quantile summary commitments. This avoids the risks of data conversion and feature leakage in plaintext. It can automatically match and calibrate differences in units, time granularity and statistical caliber while maintaining the data privacy of each institution, making multi-source data comparable and fusionable in the dense space. Attached Figure Description
[0057] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0058] Figure 1 A schematic diagram illustrating the problem principle provided in the embodiments of this application;
[0059] Figure 2 An exemplary application scenario diagram provided for an embodiment of this application;
[0060] Figure 3 A flowchart illustrating a cross-organizational user numerical data processing method based on privacy-preserving computation, provided for an embodiment of this application;
[0061] Figure 4 This is a flowchart illustrating another cross-organizational user numerical data processing method based on privacy-preserving computation provided in this application embodiment. Detailed Implementation
[0062] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0063] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0064] This application addresses a common collaborative scenario:
[0065] Multiple business units, acting as institutional clients, have accumulated user data with the same name but different definitions over a long period. Regulatory and compliance requirements explicitly prohibit the sharing of this data in plaintext. Theoretically, all institutional clients should perform joint modeling or evaluation within a privacy-preserving computation framework. However, in practice, even the same feature table can differ in currency, unit, time granularity, and even whether refunds are included or extreme value processing is applied. Once in the dense training or federated inference phase, these differences are amplified into spurious correlations, tail drift, and threshold mismatches. Differential privacy noise further erodes the limited effective signal.
[0066] Traditional approaches include: Method 1: relying on trusted hardware at the central side to perform intermediate plaintext preprocessing; Method 2: forcibly using a one-size-fits-all global quantile mapping to cover all participants.
[0067] The former is difficult to implement in terms of trust and auditing, while the latter is costly in terms of accuracy and fairness, and quickly becomes ineffective when faced with seasonal changes or changes in business structure.
[0068] Based on the aforementioned contradictions, this application does not assume that there are directly usable field dictionaries or fixed mapping relationships between different institutions, but instead delegates the responsibility for alignment to the encrypted side of each data terminal:
[0069] The auditing end only generates commitments to the reversible scaling parameters and global quantile digest commitments obtained under differential privacy protection. Any visible information is limited to the commitment and proof level. The organization that needs to transmit data uses this to densify its local numerical features and complete scale unification before sending it. The organization that needs to receive data receives the standardized densified features without decryption, first completes the basic monotonic mapping according to the global quantile digest, and then performs order-preserving calibration and secure truncation on the mapping residuals based on its own business profile, ultimately forming a densified result that is both aligned with the global distribution and retains local interpretability. Without exposing the plaintext distribution, the embodiments of this application can stably maintain the ordering relationship, control the correction magnitude, and allow zero-knowledge proofs to self-prove the link of correct source, bounded process, and budget compliance.
[0070] The application scenarios of this application are not limited to finance, but finance is the most representative.
[0071] For example, in joint anti-fraud and pre-credit assessment, banks and payment clearing networks often face mixed data with multiple currencies, different reconciliation cycles, and different refund criteria. Similarly, cross-operator credit scoring, cross-insurance company claims fraud prevention, and cross-media attribution in privacy advertising measurement all suffer from numerical features with the same name but different meanings and distribution misalignment caused by asynchronous sampling.
[0072] The commonality of application scenarios is not whether or not they are willing to share, but how to achieve verifiable consistency when sharing is not possible.
[0073] The core processing logic of this embodiment is precisely designed to solve this dilemma:
[0074] By allowing the knowledge required for alignment to flow in the form of commitments and proofs, and by ensuring that the computation of alignment itself only occurs within their respective dense state boundaries, the fragile links that originally required a trust center are decomposed into verifiable, rollbackable, and auditable endogenous capabilities of the terminal through the connection of reversible scaling, global quantile, and order-preserving calibration. This allows for the acquisition of reusable cross-terminal numerical alignment paths without altering the internal governance and IT boundaries of each organization.
[0075] Understandably, in application scenarios, the institutions receiving data are often not just one, but multiple collaborative nodes from different business systems and institutional environments. For example, in typical privacy computing tasks such as cross-bank credit granting, cross-payment channel risk control, cross-provincial tax collection, or cross-hospital diagnostic analysis, a source data terminal often needs to synchronize or share encrypted indicator data with multiple data terminals.
[0076] In traditional methods, the problem of cross-domain data transmission and alignment from a single source to multiple endpoints is almost impossible to solve consistently under encrypted conditions. While methods might suggest that a central node or trusted execution environment uniformly performs scaling and distribution alignment in decrypted mode, then encrypts and returns the results separately, or that each data terminal independently establishes a calibration model to standardize the encrypted data based on its own standards and units, these approaches superficially achieve cross-institutional data matching. However, from both engineering and compliance perspectives, these approaches have significant shortcomings. The former disrupts the privacy loop of end-to-end encrypted computation; if the central node is attacked or accessed without authorization, the data privacy of all participants will be exposed. The latter, due to the lack of a unified statistical reference, leads to inconsistent mapping relationships between terminals, further amplifying distribution offsets and task accuracy losses.
[0077] What's even more challenging is that the differences in the dimensions of numerical data across different data terminals are not simply a matter of proportion or unit, but involve multi-dimensional differences such as statistical caliber, time window, and business logic.
[0078] For example, when calculating transaction amounts, one terminal might use daily averages in RMB, while another might use weekly cumulative amounts in USD. Some terminals might also include additional items such as refunds, compensation, and transaction fees in their statistics. These differences can be resolved through rule preprocessing or manual mapping in the plaintext stage. However, in a encrypted computing environment, because the data cannot be directly compared, traditional methods struggle to establish reliable correspondences without decryption.
[0079] refer to Figure 1 , Figure 1 This is a schematic diagram illustrating the problem principle provided in the embodiments of this application.
[0080] Figure 1 The aforementioned methods one and two are shown, wherein:
[0081] Method 1 is shown as follows Figure 1 The upper section. Before sending user numerical data to multiple second data terminals, the first data terminal typically performs plaintext conversion and numerical unification operations through a trusted third party.
[0082] Understandably, while Method 1 can standardize units, currencies, or definitions before transmission, the entire process relies on a central node processing data in plaintext, posing a significant risk of privacy breaches. If access control or key management by a trusted third party fails, the data content of all participating terminals could be reverse-engineered or illegally obtained. Furthermore, in scenarios with dynamic multi-terminal access, ensuring consistency in updating and synchronizing unified definitions is difficult.
[0083] Method 2 is shown as follows Figure 1The lower section. To avoid plaintext exposure, some systems adopt a decentralized multi-terminal direct collaboration mode. The first data terminal directly sends data to multiple second data terminals, and each second data terminal performs independent secondary standardization processing according to its own caliber and unit.
[0084] Understandably, this approach eliminates the reliance on a central node, but due to the lack of a global scale reference, numerical mappings between different second data terminals can be biased. For example, one terminal might use RMB while another uses USD, or the time granularity and statistical caliber might differ, causing the scale relationship of the same indicator to be inconsistent across terminals. Ultimately, in privacy-preserving computation or encrypted modeling, the offset in data distribution across terminals can lead to problems such as incorrect model parameter estimation, imbalance in joint computation, and decreased comparability of results.
[0085] It is understood that, in this application, the data terminal and the institution end are logically defined to be the same, and both can be regarded as business nodes that perform data acquisition, encryption processing, transmission, and distribution alignment. The difference lies in that, in order to more clearly express the situation where there may be multiple independent computing nodes under one institution in a cross-domain collaborative computing scenario, this application further refines the institution end into one or more data terminals.
[0086] In other words, each institution represents a data holder, and the multiple data terminals under each institution correspond to computing nodes of different business systems, statistical methods, or geographical branches. For example, in a cross-bank credit reporting scenario, a bank institution may contain three data terminals for processing credit card, loan, and wealth management data respectively; in a cross-medical collaborative modeling scenario, a hospital institution may contain three types of data terminals: outpatient, imaging, and laboratory. This division allows the technical solution of this application to address more precisely the differences in statistical methods, time windows, or data structures within the same institution.
[0087] Based on this, the following description of the second agency terminal in this application embodiment refers to the inclusion of multiple data terminals:
[0088] When performing cross-domain privacy computation, a receiving organization may have multiple data terminal nodes with independent processing logic. Each node needs to independently complete quantization mapping and correction in the secret domain according to its own business characteristics and definition.
[0089] refer to Figure 2 , Figure 2 This is an exemplary application scenario diagram provided for an embodiment of this application.
[0090] Figure 2 The overall structure and data interaction process of a cross-data terminal user numerical data processing system based on privacy-preserving computing are shown.
[0091] like Figure 2 As shown, the system in this embodiment mainly includes an audit terminal, a first institution terminal, and a second institution terminal. Figure 2 In embodiments not shown, the second mechanism includes at least one data terminal. The functional modules of each part and their interaction relationships are as follows:
[0092] The first mechanism is used to perform encrypted encoding and transmission of data, including:
[0093] The dimensional metadata data transmission module is used to send dimensional metadata related to the numerical data of the user to be processed to the audit end;
[0094] The dense-state encoding module is used to perform dense-state encoding on local user numerical data based on the reversible scaling transformation parameters issued by the audit end, obtain standardized dense-state features, and send them to the second agency end.
[0095] The second mechanism is used to receive dense-state encoded data and perform local distribution correction, including:
[0096] The dimension metadata data transmission module is used to send its own dimension metadata to the auditing end so that the auditing end can establish cross-terminal matching relationships;
[0097] The signature verification module is used to receive and verify the global quantile digest commitment and parameter commitment transmitted by the audit end, ensuring the integrity of the commitment source and content;
[0098] The residual correction module is used to perform monotonic quantile mapping in the dense domain based on the global quantile summary commitment, and to perform order-preserving calibration and truncation processing in combination with local distributed residual information to generate aligned dense-state user numerical data.
[0099] The auditing end acts as a coordination and verification center, used to ensure data scale consistency across all terminals without decryption, including:
[0100] The encoding module is used to standardize and encode the received terminal dimension metadata.
[0101] The matching graph construction module establishes a multi-target matching graph based on the dimensional metadata and defines the dimensional correspondence between different terminals.
[0102] The closure solving module is used to perform weighted matching and minimum inconsistency closure operations on the matching graph to obtain globally consistent matching relationships.
[0103] The matching graph update module is used to update the weights and perturb the matching relationships according to the differential privacy protection mechanism, and generate reversible scaling transformation parameter commitments and global quantile summary commitments.
[0104] Next, with reference to the accompanying drawings, a cross-organizational user numerical data processing method based on privacy-preserving computation provided by an embodiment of this application will be described. Figure 3 The method shown is applied to a data terminal of a cross-agency privacy computing system, which also includes an auditing terminal. The method includes:
[0105] S1: Determine the dimensional metadata corresponding to the first mechanism end of the user numerical data to be sent and the second mechanism end of the user numerical data to be received;
[0106] The second agency includes at least one data terminal.
[0107] In this embodiment, the first institution can be a business-side node holding the original statistical data, such as a bank credit scoring platform, payment channel gateway, or regional tax terminal, while the second institution can be a collaborative node that needs to receive and participate in joint calculations. The first institution extracts attribute parameters, including data unit, currency type, time granularity, and statistical caliber, based on its stored user numerical data attributes to generate dimensional metadata. The second institution then generates a corresponding set of dimensional metadata according to its own indicator system and business domain rules.
[0108] Furthermore, this application defines dimensional attributes at the metadata layer rather than directly relying on numerical features, which avoids plaintext alignment operations before dense-state computation and effectively reduces the risk of information exposure.
[0109] Those skilled in the art will understand that the specific field settings of dimensional metadata can be flexibly adjusted according to the application scenario, as long as they can represent the convertible relationship between data across terminals. This application does not impose any further limitations.
[0110] S2: Send the dimensional metadata to the auditing end to obtain the reversible scaling parameters and the parameter commitments corresponding to the reversible scaling parameters, as well as the global quantile summary commitments;
[0111] In this embodiment, the first agency and each of the second agencies submit dimensional metadata to the auditing agency. The auditing agency verifies the integrity and source of the dimensional metadata through a preset signature verification mechanism, and then establishes a multi-objective matching graph in the dense domain. The multi-objective matching graph uses the first dimensional structure signature as the source node and each of the second dimensional structure signatures as the target nodes, and establishes mapping edges through dimensions such as unit interchangeability, time granularity scalability, and statistical caliber compatibility.
[0112] Understandably, the auditing end only outputs commitments and summaries, without directly participating in data decryption and transmission, thus not violating the privacy constraints of the entire encrypted state. This ensures both global consistency of parameters and maintains the compliance and verifiability of the computation process.
[0113] S3: The first mechanism performs dense-state encoding on the user numerical data according to the reversible scaling transformation parameters to obtain standardized dense-state features, and sends the standardized dense-state features to the second mechanism.
[0114] In this embodiment, the first institution first adjusts the user's numerical data by scaling and resampling it over time based on the reversible scaling parameters provided by the auditing institution. The scaling factor is used for currency or unit conversion, the shift factor is used to align the statistical interval benchmark, and the time normalization factor is used to eliminate periodic granularity differences. Subsequently, the first institution performs fixed-point quantization, converting continuous numerical values into structured vectors and compressing them according to the quantization step size and truncation bit width specified in the parameter commitment to obtain a standardized fixed-point representation.
[0115] Furthermore, the first mechanism performs homomorphic encoding based on the dense state mask that corresponds one-to-one with the fixed-point vector, generating standardized dense state features. The dense state mask is obtained by decomposing the data integrity identifier and the abnormal missing identifier through threshold secret sharing and then adding random noise perturbation.
[0116] Those skilled in the art will understand that the fixed-point quantization method and the encrypted coding algorithm can be adjusted according to the computational framework (such as homomorphic encryption or secret sharing protocol), the core of which is to ensure the reversibility of the transformation and the consistency of the encrypted state.
[0117] S4: The second agency performs monotonic quantile mapping on the standardized dense state features in the dense state domain according to the global quantile summary commitment, and obtains aligned dense state user numerical data through truncation processing.
[0118] In this embodiment, the second agency first verifies the signature validity of the global quantile digest commitment and extracts the quantile cut-off sequence and representative value sequence corresponding to its own business domain. Subsequently, it performs piecewise linear interpolation mapping on the received standardized dense features in the dense domain to obtain the initial quantile mapping result.
[0119] Furthermore, to address the mapping bias, the second mechanism calculates the dense-state residuals for each segment and generates a residual correction function. Subsequently, a monotonically non-decreasing calibration function is established in the dense-state domain using order-preserving regression constraints to correct the cumulative offset of the residuals for each segment. After dense-state correction and truncation of the initial mapping results using this calibration function, aligned dense-state user numerical data is obtained.
[0120] Before delving into the specific technical details of the steps, it is necessary to further explain the computational logic and privacy protection framework applicable to the embodiments of this application.
[0121] Understandably, in privacy-preserving computation tasks involving multiple institutions and data terminals, user numerical data are often not directly comparable indicators from the same source, but rather heterogeneous variables with independent statistical definitions, sampling periods, and data granularities. Traditional methods typically standardize the data during the data aggregation phase using unified rules or centralized algorithms, but these methods all require the plaintext exposure of the original data or intermediate variables, making them unacceptable in scenarios with strict privacy constraints. This embodiment does not attempt to optimize the encryption algorithm itself at the computational level, but rather starts from the mappability of the data space, reconstructing a set of verifiable numerical equivalence relations within the encrypted domain to achieve distribution alignment and consistency of definitions among multiple terminals.
[0122] Specifically, the core of the method in this application is not to directly compare the data values themselves, but to establish numerical relationships based on dimensional metadata.
[0123] In other words, each data terminal, when participating in the calculation, does not provide raw numerical values, but instead provides a dimensional signature describing its own statistical rules and unit system. Without decrypting any data, the auditing end constructs a multi-objective matching graph based on the logical compatibility between the signatures, thereby determining the scale mapping path between different terminals.
[0124] Compared to traditional centralized decryption methods, the method in this application is closer to a structural-level matching reasoning, and the judgment is based not on the data content, but on the formal consistency and reversible transformation constraints between dimensions.
[0125] Furthermore, after obtaining the matching relationship, this embodiment also includes two types of verifiable structures: parameter commitment and quantile digest commitment. The former constrains the scaling ratio, translation, and time normalization parameters of the scaling transformation through cryptographic signatures, so that all subsequent calculations can be verified in the dense domain; the latter summarizes the distribution characteristics of each terminal and constructs a global statistical framework through the ordered combination of quantile cut points, so that even if the original distribution patterns of each terminal are significantly different, the reconstruction of the mapping relationship and consistency verification can be completed without exposing the specific values.
[0126] Furthermore, to address the residual offset issue that may arise during multi-terminal distribution alignment, this embodiment introduces a monotonic constraint quantile correction logic in the dense domain. By reweighting the residual gradient under order-preserving conditions, the corrected mapping function retains its non-decreasing property in a statistical sense. Mathematically, this is equivalent to implementing a constrained order-preserving regression under encrypted conditions, except that the input features are encrypted statistics rather than plaintext samples. This ensures that the correction process does not compromise the original privacy and security structure, while also providing verifiable stability in the distribution consistency of the data output from each terminal.
[0127] Next, we will further elaborate on the technical content of the dimensional metadata in this application.
[0128] It is understandable that in cross-domain data collaboration scenarios, although the user numerical data held by different institutions or data terminals may represent the same indicator semantically, there are often differences in physical quantity, time scale and statistical logic.
[0129] In one example, the first agency automatically extracts and encodes the attribute information of the numerical data of the user to be sent, generating first-dimensional metadata. The attribute information can be obtained from data tags, field annotations, or database metadata, and can be further expanded to include dimensional parameters such as unit system, currency type, time granularity, statistical caliber, data periodicity, and anomaly correction strategy.
[0130] In this embodiment, each parameter is recorded as a parsable meta-field, enabling it to participate in subsequent matching graph construction and also serve as a structural signature input auditing end in a closed environment. In this way, the first agency can achieve standardized expression at the data structure level without exposing any real values.
[0131] In another example, for multiple data terminals that may exist in the second agency, this embodiment uses a parallel acquisition method to generate a second-dimensional metadata set. Each data terminal extracts corresponding attribute description information and performs metadata encoding according to its own business domain's statistical rules and data caliber definitions. The attribute description information can be obtained by automatically crawling the terminal database's meta-schema, calling the indicator definition table returned by the business interface, or by the terminal security proxy module transmitting it in a secure channel.
[0132] Those skilled in the art will understand that, during implementation, a centralized or distributed data collection method can be selected based on the actual network topology and the number of terminals, as long as the dimensional metadata set can cover the attribute space of all target terminals.
[0133] It is important to note that the dimensional metadata and dimensional metadata set described in this embodiment are not only static descriptions of units and definitions, but also include logical constraint information formed by each data terminal during the actual sampling and statistical process. For example, for the same indicator, different data terminals may use different time windows, anomaly removal strategies, or aggregation algorithms. These differences are often the root cause of inconsistencies in numerical definitions. Therefore, when generating dimensional metadata, both the first and second institutions encode their respective statistical logic in a structured form through the statistical strategy identifier field. This allows the auditing end to automatically determine the logical compatibility between different terminals without accessing plaintext data during subsequent matching and solving processes.
[0134] Next, we will further elaborate on the technical content of the standardized dense state characteristics of the method in this application.
[0135] It is understandable that standardized dense-state features specifically refer to a class of cryptographic representation vectors formed in the dense-state domain through scale uniformity and cryptographic encoding. These vectors are computable, verifiable, but cannot be used to reverse-engineer the original values. This not only preserves the structural relationships of the original user numerical data but also ensures metric consistency across institutions or data terminals through reversible scaling parameters, giving them equivalent statistical significance in subsequent dense-state distribution alignment and joint computation.
[0136] It is worth noting that the standardized encrypted features in this embodiment are not only an encrypted data structure, but also a verifiable feature representation. They maintain the operability of multiplication and addition within the homomorphic encryption domain, enabling the second mechanism to perform operations such as quantile mapping, residual correction, and distribution truncation without decryption. Furthermore, since the generation process involves parameter commitment and signature verification, any tampering or deviation of intermediate results can be detected through auditing, thus ensuring the reliability of the calculation and the consistency of the results.
[0137] In one example, dense-state encoding of user numerical data is performed based on the reversible scaling parameters to obtain standardized dense-state features, including:
[0138] S3.1: Based on the preset public key verification signature, perform zero-knowledge range verification on the scaling coefficient, translation coefficient and time normalization factor in the reversible scaling parameters;
[0139] Specifically, before performing a scaling transformation, it is necessary to ensure that the reversible scaling transformation parameters are valid in both numerical range and logical relationship. Otherwise, problems such as quantization overflow and time mapping misalignment may occur during encrypted computation. Traditional methods often directly detect the parameter value range through plaintext verification, which compromises the closed nature of privacy computation and leads to the risk of parameter leakage. Therefore, this embodiment adopts a range verification method based on zero-knowledge proofs, using a public key verification signature mechanism to achieve encrypted verification of parameter validity without revealing the specific parameter values.
[0140] In this embodiment, the scaling factor, translation factor, and time normalization factor are all generated by the auditing end in the encrypted domain, along with a set of public key signature information. Upon receiving these parameters, the first institution verifies whether each parameter falls within a preset security range through cryptographic calculation. During the verification process, the parameter values are not directly read; instead, a result flag is obtained through homomorphic verification calculation of the signature function and the range constraint function. This result only returns a Boolean state indicating whether the parameter is valid or invalid, without exposing any plaintext parameter information.
[0141] Understandably, if the verification result of the zero-knowledge scope check is valid, the first agency can proceed to the subsequent scaling operator and time resampling kernel calculation steps; if the verification result is invalid, the parameter regeneration or correction process needs to be triggered.
[0142] Specifically, if any of the scaling factor, translation factor, or time normalization factor fails the verification, the first agency will refuse to execute subsequent coding operations and send a parameter anomaly feedback flag to the auditing agency. Upon receiving this flag, the auditing agency will recalculate the corresponding parameter based on a preset parameter correction strategy. Correction methods may include resetting the safety range, dynamically scaling the abnormal scaling factor, or adjusting the time normalization granularity. Once the parameter passes the re-signature verification, the auditing agency will regenerate the parameter commitment and issue an updated result to ensure that subsequent operations are performed in a secure and reliable parameter environment.
[0143] Those skilled in the art will understand that there are various implementation paths for performing secure interval correction and logical relationship verification on reversible scaling transformation parameters. For example, interval remapping based on homomorphic addition constraints, parameter rollback based on signature hashing, or regenerating the parameter set through a trusted module can all be used. The core purpose of the above process is to ensure that the scaling parameters participating in the encrypted computation meet the requirements of secure numerical range and logical consistency, thereby avoiding problems such as computational anomalies or scaling mismatches in subsequent encrypted computations. This application will not elaborate further on these points.
[0144] S3.2: Calculate the piecewise linear scaling operator based on the verified scaling and translation coefficients, and calculate the time resampling kernel based on the verified time normalization factor;
[0145] Specifically, user numerical data exhibits orders of magnitude and time-to-base differences across different institutions. A single scaling transformation cannot guarantee linear consistency in alignment, especially in the presence of outliers or non-uniform time sampling. Therefore, this embodiment designs the scaling mapping as a piecewise linear function to ensure adaptive scalability across different numerical ranges. Simultaneously, to address the data granularity mismatch caused by different sampling periods, a resampling kernel function based on a time normalization factor is introduced to achieve data alignment across time scales.
[0146] In this embodiment, the original numerical interval is segmented according to the quantiles of the data distribution, with each segment corresponding to a pair of scaling coefficients and translation biases. By executing the segmented mapping function in the dense domain, the data in the high-value interval and the low-value interval can obtain relatively independent scaling transformations, thereby maintaining the linear consistency of the overall distribution pattern and obtaining the corresponding segmented linear scaling operator.
[0147] Furthermore, non-uniform resampling is performed on the original time series signal by using a time normalization factor, so that data from different statistical periods are mapped to a unified standard time reference in the dense state domain, and the corresponding time resampling kernel is obtained based on the non-uniform resampling results.
[0148] S3.3: For each user's numerical data, the plaintext numerical data is quantized at fixed points according to the piecewise linear scaling operator and the time resampling kernel, the numerical granularity is made consistent, and the quantized and consistent user numerical data is compressed in intervals and truncated at bit width by using the quantization step size and truncation bit width that match the parameter commitment, so as to obtain a standardized fixed-point vector.
[0149] Specifically, after achieving scale and time alignment, to ensure the data is computable in the dense domain, continuous floating-point values need to be converted into fixed-point quantized representations. Traditional fixed-point methods often rely on plaintext mapping, which suffers from precision loss and uncontrollable truncation. This embodiment uses parameter commitment to ensure that the quantization process is executed under safe and verifiable conditions, guaranteeing that the quantization step size, bit width, and compression interval all meet global consistency requirements.
[0150] In this embodiment, the first agency performs fixed-point quantization on each user value based on the result calculated by the piecewise linear scaling operator. The quantization step size is determined by the auditing end when generating the parameter commitment, and is usually dynamically adjusted based on the standard deviation or quantile spacing of the global distribution; the truncation bit width is determined based on the target precision and the bit width supported by the encryption algorithm. Each value is mapped to integer form after quantization, and an interval compression operation is performed in the plaintext domain to truncate values exceeding the set range according to a symmetric truncation rule. Subsequently, the first agency truncates the compressed integer values, retaining only the agreed number of valid bits, thereby generating a standardized fixed-point vector.
[0151] S3.4: Perform dense-state encoding on the standardized fixed-point vector according to the dense-state mask that corresponds one-to-one with the standardized fixed-point vector to generate standardized dense-state features, wherein the dense-state mask is obtained by secret sharing decomposition of the integrity identifier and abnormal missing identifier of the user numerical data and adding random noise perturbation.
[0152] Specifically, in encrypted computing scenarios, if the data distributions of multiple terminals exhibit statistical correlations, attackers may be able to infer hidden patterns through side-channel attacks. To overcome this potential correlation, this embodiment introduces an encrypted masking mechanism before standardized fixed-point vector coding. The masking is designed not only to prevent numerical leakage but also to enhance the independence of encrypted inputs, ensuring semantic security during encrypted computing across multiple terminals.
[0153] In one example, the calculation method of the dense state mask includes:
[0154] Extract the integrity identifier corresponding to the user numerical data, wherein the integrity identifier is used to characterize the validity status of the user numerical data during the acquisition, transmission and storage stages;
[0155] Based on the integrity identifier and the preset anomaly detection rules, the consistency of user numerical data is determined and missing data is identified to obtain an anomaly missing identifier.
[0156] The abnormal missing identifier is used as input, and the user's numerical data is decomposed by the threshold secret sharing algorithm to obtain multiple secret shares that meet the preset recovery threshold. Each secret share includes a random salt value and a corresponding timestamp. The random salt value is calculated by the threshold secret sharing algorithm.
[0157] For each secret share, a perturbation share corresponding to the secret share is calculated based on a preset privacy budget and perturbation distribution function, and the perturbation share is encapsulated into a secret mask component through homomorphic encryption;
[0158] By combining all the dense state mask components through weighted aggregation, a dense state mask corresponding one-to-one with the normalized fixed-point vector is obtained.
[0159] In this embodiment, the generation of the secret state mask follows the design principle of being recoverable and verifiable but unpredictable. Its core logic lies in using a two-layer structure of secret sharing and random perturbation to inject independent randomness into each data point in the secret state domain, thereby completely breaking the data dependency relationship across terminals without affecting the overall statistical regularity.
[0160] Specifically, the original data is first divided into a usable area and an uncertain area based on integrity and abnormal missing data identifiers. The data in the uncertain area is then decomposed using a threshold secret sharing algorithm. This algorithm splits a single plaintext value into multiple random shares, and the original data can only be reconstructed when a recovery threshold is reached. Simultaneously, each share embeds an independent random salt value and a timestamp, ensuring that the same data point exhibits non-repeatability across different computation rounds. This structure effectively prevents attackers from using cross-task statistical correlations or homogeneous distribution features for reverse engineering, theoretically guaranteeing the semantic independence of the encrypted input.
[0161] Furthermore, to eliminate potential residual correlations between shares at the statistical level, this embodiment introduces a perturbation distribution function after secret sharing decomposition, independently adding random noise that follows a differential privacy mechanism to each share. The perturbation distribution function can be a Gaussian or Laplace model, and the noise amplitude is dynamically controlled according to a set privacy budget, ensuring that the value of each share still has sufficient uncertainty without distorting the overall distribution. Subsequently, each perturbation share is homomorphically encrypted and encapsulated to form a cryptographic mask component. All components are aggregated according to share weights to obtain the final cryptographic mask. This cryptographic mask corresponds one-to-one with a standardized fixed-point vector and is used to obfuscate the original values in subsequent homomorphic encoding processes. After this processing, even if an attacker obtains the complete ciphertext sequence, they cannot infer any original information through correlation or temporal patterns, achieving global semantic security and privacy isolation under multi-terminal interaction.
[0162] Next, we will further elaborate on the technical content of the second-machine processing in this application.
[0163] Understandably, when the second agency receives the standardized encrypted features forwarded by the first agency, the core challenge it faces is not simply receiving encrypted data, but rather how to achieve a consistent mapping and verifiable correction of the encrypted data distribution based on the global quantile digest commitment generated by the auditing end without decryption.
[0164] It is understandable that due to systematic deviations in the statistical caliber and distribution characteristics of data from different institutions, performing arithmetic operations directly in the dense domain would cause the model to shift during the joint computation phase. This embodiment introduces a dense quantile mapping and order-preserving calibration mechanism at the second institution, enabling the data to achieve statistical alignment while maintaining encryption, that is, re-embedding standardized dense features into the statistical domain of each terminal.
[0165] In one example, the normalized dense-state features are monotonically mapped in the dense-state domain according to the global quantile summary commitment, and aligned dense-state user numerical data is obtained through truncation, including:
[0166] S4.1: For each data terminal in the second agency, receive the standardized encrypted features transmitted by the first agency and the global quantile digest commitment transmitted by the auditing end;
[0167] S4.2: Verify the signature of the global quantile digest commitment based on the public key and commitment issuance information corresponding to the data terminal, and extract the quantile cut-off sequence and the corresponding representative value sequence of the global quantile digest commitment in the data terminal;
[0168] Specifically, to prevent the commitment from being tampered with or replaced during transmission, each data terminal in the second agency needs to perform dual verification of the commitment's source and content within the encrypted domain, and then selectively extract statistical scales that match its own business domain. The quantile cutoff sequence can be understood as a set of interval boundaries arranged according to a probability scale, used to divide the target distribution into equally probable or predetermined probability intervals; the corresponding representative value sequence can be understood as a set of target values used to represent the output scale within each interval, often given by the auditing end according to a global summary strategy.
[0169] In one example, extracting the global quantile summary commitment into the data terminal includes the quantile cut-off sequence and the corresponding representative value sequence, including:
[0170] Based on the global quantile index mapping table in the global quantile digest commitment, and in conjunction with the service domain identifier of the data terminal, the quantile index range corresponding to the data terminal is determined;
[0171] The quantile digest structure in the global quantile digest commitment is indexed and located by a dense state retrieval algorithm, and the dense state quantile cut-off point value corresponding to the quantile index range is extracted to generate a quantile cut-off point sequence.
[0172] Based on the quantile cut-off sequence, the corresponding dense state representative values are extracted from the representative value set in the global quantile summary commitment to obtain a representative value sequence that corresponds one-to-one with the quantile cut-off sequence.
[0173] In this embodiment, the extraction of the quantile cutoff sequence and the representative value sequence follows the principles of index-first, dense-state localization, and commitment consistency:
[0174] The global quantile digest commitment pre-defines a quantile index mapping table to indicate the correspondence between different business domains, indicator categories, and target quantile intervals. The data terminal first determines the available quantile index range within the commitment based on its own business domain and indicator identifiers. This process does not read any plaintext quantiles; instead, it only verifies and matches the index with the version number to ensure consistency between subsequent retrieval and the statistical scale of the local task. Subsequently, the encrypted retrieval algorithm is invoked to locate the index within the digest structure: for each target index, the position is confirmed using the commitment issuance information and encrypted comparison primitives, and the corresponding encrypted quantile cutoff values are securely retrieved and assembled into a quantile cutoff sequence according to the index order. The entire retrieval process is constrained by commitment hash and version tag verification, ensuring that the retrieved cutoff values and commitment content are consistent in structure and order, and can be independently verified afterward.
[0175] Furthermore, the acquisition of the representative value sequence maintains a one-to-one correspondence with the cut-off point sequence. Based on the existing quantile cut-off point sequence, the data terminal uses the same indexing method to retrieve the corresponding entries from the representative value set, forming a dense representative value sequence. The representative value can be the target scale for each quantile interval, and its selection strategy is fixed in the commitment using a digest. To prevent cross-batch or cross-version misuse, the extraction process simultaneously records the version identifier, index trajectory, and issuance verification digest, and binds them to the local mapping context. If a version inconsistency or issuance verification failure is subsequently discovered, use is immediately stopped and a commitment re-retrieval is triggered.
[0176] S4.3: Based on the quantile tangent sequence, perform piecewise linear interpolation calculation on the normalized dense state features in the dense state domain to map the normalized dense state features to the initial quantile mapping result;
[0177] Specifically, the quantile tangent sequence provides the interval skeleton of the target distribution. Using the interval skeleton, monotonic and combinable intra-segment mappings can be performed on standardized dense features without decryption. Piecewise linear interpolation has the characteristics of simple implementation, homomorphism friendliness, and controllable boundaries, and can complete the smooth transition of the value range while ensuring that the ordering relationship is not destroyed.
[0178] In this embodiment, the data terminal binds standardized dense-state features with tangent sequence and completes interval localization through homomorphic comparison and mask selection: for each dense-state data point, the dense-state comparison primitive is used to determine the tangent interval it falls into, and an interval selection mask is generated; then, based on the relative positions of adjacent tangent points, the homomorphic linear combination primitive is called to calculate the interpolation result within the segment, and the initial quantile mapping result is obtained. The entire process uses only linear and comparison-based dense-state operations, avoiding computational amplification and noise accumulation caused by complex nonlinear kernels.
[0179] S4.4: Based on the representative value sequence, the initial quantile mapping result is segmented in the dense state domain, the difference of the dense state mean value of each segment is calculated to obtain residual data, and the residual correction function is calculated based on the residual data;
[0180] Specifically, the initial quantile mapping projects the standardized dense-state features onto the global scale. However, due to differences in the local structure and sampling rules of each data terminal, systematic shifts may still occur within segments. The representative value sequence provides the target scale for each interval. Using the representative value sequence, the shift direction and magnitude of each interval can be measured in the dense-state domain, and an interpretable residual correction function can be constructed accordingly.
[0181] In this embodiment, the data terminal performs segmented aggregation on the initial quantile mapping results according to the interval number obtained in S4.3. It uses homomorphic addition and homomorphic counting primitives to calculate the dense mean or weighted statistic of each segment, and performs homomorphic difference with the corresponding representative value to obtain the difference of the dense mean of each segment as residual data. Subsequently, based on the order structure and amplitude constraints of the residual data, a set of segment-level correction coefficients is generated, and the intra-segment offset and inter-segment transition strategy of the residual correction function are defined so that the function can be structurally verified and can be connected with subsequent order-preserving constraints.
[0182] S4.5: Generate a calibration function corresponding to the residual correction function through the order-preserving regression constraint, wherein the calibration function is a monotonically non-decreasing function, and the constraint conditions of the order-preserving regression constraint correspond to the upper and lower threshold information in the global quantile summary commitment;
[0183] Specifically, directly superimposing residual correction functions may disrupt monotonicity, leading to local order reversal after mapping. To ensure that the output remains statistically ordered, this step applies order-preserving regression constraints to the residual correction functions in the dense domain, ensuring that each correction segment is globally monotonically non-decreasing, and the correction magnitude is limited by a committed threshold to prevent overcorrection.
[0184] In one example, generating the calibration function corresponding to the residual correction function through ordinal-preserving regression constraints includes:
[0185] S4.5.1: Calculate the piecewise residual sequence corresponding to the residual correction function in the dense domain, wherein the piecewise residual sequence is used to characterize the deviation trend of the initial quantile mapping result in the quantile interval;
[0186] Specifically, the initial quantile mapping only completes the intra-segment linear projection based on the global tangent point. The sampling calibers and anomaly handling strategies of different data terminals will create systematic offsets within the segment. Without segment-level offset measurement, the subsequent calibration function lacks a constrained reference trajectory. To establish a verifiable offset characterization in the dense domain, this step binds each data point to its corresponding quantile interval, calculates the dense state statistics on an interval-by-interval basis, and differs them with the target representative scale of that interval to obtain a residual chain with an ordered structure. The residual chain does not directly correct a single point but rather abstracts the offset trend at the interval level, facilitating subsequent global consistency processing under order-preserving constraints.
[0187] In this embodiment, the data terminal uses the interval index and relative position within the segment recorded in the preceding steps to perform homomorphic aggregation on the initial quantile mapping results by interval, obtaining the dense mean or dense weighted statistical value of each interval; then, it calls the homomorphic difference primitive to differ with the corresponding interval representative value in the commitment, generating interval residuals. To reduce the impact of fluctuations in a single batch of samples, segmented statistics can enable a dense sliding window, incorporating interval statistics from adjacent batches into the calculation in a decaying manner, thereby obtaining a more stable residual estimate without exposing the plaintext.
[0188] S4.5.2: Based on the upper and lower threshold information in the global quantile summary commitment, construct the constraints of the residual correction function. The constraints include the non-subtractive relation constraint of the quantile segment residuals, the residual magnitude limit constraint, and the total change constraint.
[0189] Specifically, if the residual sequence is not constrained by boundary conditions, it is prone to inter-segment crossover or amplitude overshoot during calibration, thereby compromising the orderliness and numerical stability of the output sequence. Therefore, a set of constraints that can be homomorphically verified needs to be established for the residual sequence in the dense domain. Non-subtractive relations are used to limit the global order of the sequence, amplitude constraints are used to suppress excessive corrections in single segments, and total change constraints are used to control the overall correction strength, ensuring that the calibrated mapping conforms to the statistical scale while avoiding excessive interference with the original ordering.
[0190] In this embodiment, the data terminal reads the upper and lower threshold labels bound to the local task from the global quantile summary commitment, forming the boundary source of the constraint set. Subsequently, a constraint description is established for each interval residual. Non-subtractive relations are formed into a set of adjacent constraint pairs using dense-state comparison primitives. Amplitude constraints are formed into a set of amplitude masks using the homomorphic difference between the threshold labels and the interval residuals. The total change is formed into a total amount mask by dense-state aggregation of the absolute corrections of the entire sequence. These masks are used as selection signals in subsequent solution steps to participate in coefficient updates, allowing each update step to be replayed and independently verified via the commitment information.
[0191] S4.5.3: Based on the constraints, perform order-preserving regression calculation on the segmented residual sequence, and merge the residuals that do not satisfy the monotonic relationship through dense-state weighted average and dense-state comparison operations to obtain a dense-state calibration sequence that is adapted to the monotonic non-decreasing characteristics.
[0192] Specifically, the essence of order-preserving regression is to merge adjacent segments that violate the order constraint and re-estimate them so that the output sequence satisfies a monotonic relationship globally, while approximating the original residual sequence as closely as possible. Since both the residuals and constraints are in the dense state domain, merging and updating need to be achieved by comparing homomorphic weighted averages with dense state data, and the constraint mask is used as the update trigger condition, so that each merging and backtracking has a clear dense state evidence chain that can be audited.
[0193] In this embodiment, the data terminal scans the segmented residual sequence and uses dense-state comparison to identify adjacent intervals that violate the non-decreasing constraint. After a violation is found, adjacent intervals are merged into new interval nodes using homomorphic weighted averaging, with the weights taken from the dense-state combination of the aforementioned interval weight labels and the segment count. After merging, dense-state comparison is performed again on the new node and its adjacent nodes. If the non-decreasing relationship is still violated, the merging continues to expand to both sides until the constraint is satisfied or the sequence boundary is reached. To avoid frequent jitter, the amplitude limit mask and the total change mask are used during the merging process to reduce the update step size, maintaining the smoothness of the sequence.
[0194] S4.5.4: Based on the dense-state calibration sequence, construct a piecewise linear expression for the calibration function in the dense-state domain, determine the calibration slope and offset parameters for each segment, and generate the calibration function;
[0195] Specifically, the dense-state calibration sequence provides an interval-level correction objective that satisfies the order-preserving constraint, but a functional form that can be directly applied to dense-state samples still needs to be constructed. Piecewise linear expressions combine computability and verifiability, enabling smooth transitions within segments with minimal computational complexity, while preserving the traceable structure of inter-segment boundaries, providing clear parameterized entry points for subsequent truncation, combination, and version replay.
[0196] In this embodiment, the data terminal generates a pair of segment-level parameters for each interval based on the dense-state calibration sequence. The parameters consist of the dense-state slope and the dense-state offset. The parameters are obtained using a homomorphic linear regression approximation, and parameter estimation is completed with the intra-segment statistical summary and calibration target as inputs without decryption. A transition band strategy is adopted between segments, so that adjacent segments gradually switch in a homomorphically weighted manner within the transition band, avoiding the amplification of subsequent computational noise caused by boundary discontinuities. The function parameters, along with the interval index, transition bandwidth, upper bound of amplitude, and version tag, are written into the calibration function description structure, becoming an auditable dense-state object.
[0197] S4.6: Perform dense state correction and dense state truncation on the initial quantile mapping result according to the calibration function to obtain the aligned dense state quantile mapping result;
[0198] Specifically, the calibration function, after order preservation, possesses global monotonicity and amplitude-controlled properties, and can serve as a terminal corrector for the initial quantile mapping results. To suppress the impact of outliers on subsequent joint calculations, dense-state truncation is also required based on the upper and lower thresholds in the global commitment, ensuring that the output falls within a verifiable safe range.
[0199] In this embodiment, the data terminal applies a calibration function to the initial quantile mapping result in a homomorphic manner, performs segment-level correction on each data point, and employs a transition strategy between segments to achieve smooth connection. Subsequently, based on the upper and lower thresholds given by the commitment, the dense state comparison and selection primitives are invoked to saturate the out-of-bounds data, resulting in the aligned dense state quantile mapping result. This result, along with the version identifier and commitment citation digest in the mapping context, is encapsulated to form a replayable output unit, supporting direct consumption in subsequent joint training or inference processes.
[0200] For example, the following numerical example is provided to illustrate how to convert to consistent reversible scale parameters (scale factor and translation factor) and the calculation process of matching relationship closure. This example is only used to explain the operation link and dimensional relationship. The selected parameters and values are illustrative values and do not represent the actual calibration results or engineering recommended values.
[0201] Suppose there are three organizations, denoted as A, B, and C, all reporting the same type of user data, but with inconsistent definitions. Organization A's "Consumption Amount" is aggregated in "Yuan / Day"; Organization B's "Transaction Amount" is aggregated in "Minutes / Hour"; and Organization C's "Transaction History" is aggregated in "Yuan / Week". To achieve structural alignment without accessing plaintext, the auditing end only obtains the dimensional signatures and statistical summary commitments (e.g., quantile summaries, mean range proofs, etc.) for each field, and constructs candidate matching relationships based on these. In this multi-objective matching graph, nodes represent "field definitions," edges represent candidate mappings where "a certain field in A may correspond to a certain field in B," and edge weights represent the comprehensive inconsistency cost of this mapping under multiple objectives; the lower the cost, the more credible the mapping.
[0202] In this example, the audit side calculates three types of costs for each candidate edge and sums them up with weights: The dimension cost is used to measure whether the unit conversion is self - consistent, the time - granularity cost is used to measure whether the aggregated period after conversion falls within an acceptable range, and the distribution - form cost is used to measure the consistency of the quantile summary after scale normalization. The weight values are for illustration only. Let the dimension weight be 0.5, the time weight be 0.3, and the form weight be 0.2. First, look at a candidate edge from A to B: A is "yuan / day" and B is "cents / hour". For unit conversion, 1 yuan equals 100 cents; for time conversion, 1 day equals 24 hours. If they are indeed the same semantic quantity, then the value of B multiplied by 0.01 can be converted to "yuan / hour", and then multiplied by 24 can be converted to "yuan / day", and the combined overall scale factor is 0.24. The audit side uses the summary commitment for consistency verification: Assume that under their respective calibers, the median commitment on the A side is 48 (yuan / day), and the median commitment on the B side is 200 (cents / hour). First, convert the median on the B side to "yuan / day" according to 0.24, getting 200 multiplied by 0.24 which equals 48, consistent with the median on the A side. Therefore, the form cost can take a very small illustrative value of 0.05; at the same time, the dimension and time conversions are completely closed, the dimension cost is 0, and the time cost is 0.02. The comprehensive cost is summed up by weights as 0.5 times 0 plus 0.3 times 0.02 plus 0.2 times 0.05, getting 0.016. This value is very small, indicating that "A's consumption amount B's transaction amount" is a highly credible candidate edge. Then, look at a candidate edge from B to C: B is "cents / hour" and C is "yuan / week". If they correspond, the scale factor from B to "yuan / week" should be first multiplied by 0.01 to convert to yuan, and then multiplied by 168 to convert to weeks (7 days times 24 hours), getting an overall scale factor of 1.68. Assume that the median commitment on the C side is 336 (yuan / week), and the median on the B side is still 200 (cents / hour). Then the median after conversion on the B side is 200 multiplied by 1.68 which equals 336, consistent with the C side. Therefore, the form cost also takes the illustrative value of 0.05; the dimension cost is 0, the time cost is 0.02, and the comprehensive cost is also 0.016. At this time, "B's transaction amount ↔ C's turnover" also seems credible.
[0203] Furthermore, the problem appears in another candidate edge "A's consumption amount C's turnover". A is "yuan / day" and C is "yuan / week". According to time conversion, the scale factor from C to "yuan / day" should be divided by 7, equal to approximately 0.142857; the reverse factor from A to "yuan / week" should be multiplied by 7. If all three are completely self - consistent, then the implied scale from A to C derived from the two edges "A B" and "B C" should be the same as the "A The direct candidate edge "C" is consistent. Now, deduce from the first two edges: from B to A is divided by 0.24, which is approximately 4.166666; from C to B is divided by 1.68, which is approximately 0.595238; the combined deduced factor from C to A is 0.595238 multiplied by 4.166666, approximately equal to 2.480158. This deduced factor means that "the value of yuan / week for C multiplied by 2.480158 should be close to the value of yuan / day for A". However, according to common sense of time, converting "yuan / week" to "yuan / day" should be divided by 7, approximately 0.142857, and the two differ by an order of magnitude, indicating that if "A C" is also forced to be the same semantic quantity, an obvious contradiction will occur in the scale link. To quantify the contradiction, the audit side still gives a hint using the summary commitment: the median on the C side is 336 (yuan / week), and when converted to "yuan / day" according to the deduced factor 2.480158, it will become approximately 833.33 (yuan / day), while the median on the A side is 48 (yuan / day), and the two deviate greatly, so the form cost is raised, for example, taking 0.95; at the same time, "yuan / week → yuan / day" does not hold semantically in terms of time (it should be divided by 7 but is pushed by the link to multiply by 2.48), and the time cost can be taken as 0.9; although the dimension is the same "yuan", the dimension cost can still be taken as 0.1 to reflect the penalty for "seemingly the same unit but conflicting statistical calibers". The comprehensive cost of this edge is 0.5 multiplied by 0.1 plus 0.3 multiplied by 0.9 plus 0.2 multiplied by 0.95, resulting in 0.05 plus 0.27 plus 0.19, equal to 0.51, which is significantly higher than the aforementioned 0.016.
[0204] At this point, a three-node closed loop appears in the multi-objective matching graph: A B, B C, A C. The idea of the minimum inconsistent closure is to identify and eliminate the minimum set of edges that cause the closed loop to be inconsistent while keeping "as many low-cost consistent relationships as possible" valid, so as to obtain a non-contradictory global matching subgraph. For this closed loop, if the three edges are retained, there will inevitably be a scale deduction contradiction, so at least one edge must be eliminated. The audit side compares the comprehensive costs of the three edges and finds that the cost of A B is 0.016, the cost of B C is 0.016, while the cost of A C is 0.51. So the closure operation chooses to eliminate the high-cost edge "A C", and retains the first two low-cost edges, finally obtaining a consistent matching result: the A field corresponds to the B field, the B field corresponds to the C field, and A and C do not directly establish a one-hop mapping to avoid introducing unexplainable scale conflicts.
[0205] After the closure is completed, the audit end can solidify the "executable conversion parameters" into a unified set of scaling parameters and use it as input for subsequent close-state alignment. Continuing with the above standard, taking "yuan / day" as the unified standard for A, the scaling factor from B to A is 0.24, meaning that multiplying the value on the B side by 0.24 converts it to the A standard. The scaling factor from C to A is obtained by multiplying the value of C to B and then to A. The value of C is first multiplied by 0.595238 to convert to the B standard, then multiplied by 0.24 to convert to the A standard, and combined to obtain approximately 0.142857, which is exactly equal to dividing by 7, satisfying the time semantics.
[0206] refer to Figure 4 , Figure 4 This is a flowchart illustrating another cross-organizational user numerical data processing method based on privacy-preserving computation provided in this application embodiment.
[0207] Figure 4 The method shown is applied to the audit end of a cross-organizational privacy computing system. The cross-organizational privacy computing system also includes data terminals. A data terminal for transmitting user numerical data is a first organizational end, and a data terminal for receiving user numerical data is a second organizational end. The second organizational end includes at least one data terminal. The method includes:
[0208] A1: Receive dimensional metadata from the first institution and a set of dimensional metadata from the second institution;
[0209] Specifically, cross-domain participants often have inconsistent units, currencies, time granularities, and statistical standards for indicators with the same name. If direct comparison or aggregation is performed during the matching stage, structural biases will be introduced due to semantic misalignment. Therefore, it is necessary to establish a comparable dimensional description layer on the source side to receive and solidify the structural information of each participant regarding the indicator measurement method before conducting subsequent matching inference and parameter calculation.
[0210] In this embodiment, the auditing end is configured with a dimensional metadata access gateway, employing a dual-channel reception method:
[0211] One channel receives first-dimensional metadata from the first institution; another channel receives second-dimensional metadata sets from data terminals in multiple second institutions. Each message contains metadata fields such as indicator identifier, unit code, currency code, time granularity code, statistical caliber label, anomaly handling strategy identifier, sampling window identifier, version number, and timestamp. After reception, the data is placed in a buffer, and an arrival order hash is generated to support subsequent audit traceability.
[0212] A2: Based on the preset public key, the dimensional metadata and the dimensional metadata set are signed and their integrity verified, and standardized encoding is performed to generate a first dimensional structure signature and a second dimensional structure signature.
[0213] Specifically, dimensional metadata is at risk of being replaced or partially tampered with during cross-domain transmission. If source and integrity verification is not completed before entering the matching stage, subsequent mapping inference will be based on unreliable evidence. After verification, the heterogeneous metadata fields from multiple sources need to be folded into comparable structural fingerprints, enabling the matching stage to perform structural-level comparisons without reading plaintext business rules.
[0214] In this embodiment, the auditing end performs two levels of verification on each dimensional metadata record:
[0215] First, the digital signature and certificate chain are verified using the institution's public key. Then, the payload is hashed and compared with the commitment digest in the message field. If they match, the integrity is deemed passed. The verified data is sent to the encoder. The encoder normalizes and encodes the unit, currency, time granularity, caliber label, anomaly policy, window identifier, etc., according to a unified field dictionary and weight table. The dependencies between fields are added to the fingerprint graph as directed attribute edges. Finally, the structured signature is obtained by hash folding. The first-dimensional structured signature is output to the first institution, and the corresponding second-dimensional structured signature is output to each second institution's data terminal.
[0216] A3: Construct a multi-target matching graph using the first dimension structure signature as the source node and the second dimension structure signature as the target node;
[0217] The mapping edges of the multi-target matching graph are established through unit interchangeability, temporal granularity scalability and statistical caliber compatibility between dimensional structure signatures;
[0218] Specifically, alignment from one source to multiple ends is not the same as a simple one-to-one matching. It requires expressing all potential feasible mappings at the structural level and using costs and constraints to represent their credibility and invertibility boundaries. The graph structure can accommodate three types of evidence—unit conversion, time scaling, and caliber compatibility—within the same framework, and uses edge attributes to characterize the upper bound of errors and the source of constraints, providing a unified representation for subsequent global solutions.
[0219] In this embodiment, the auditing end uses the first-dimensional signature structure as the source node and compares it one by one with the second-dimensional signature structure set. The establishment of the mapping edges adopts the collaboration of the rule engine and the metric engine:
[0220] The rule engine determines whether an edge is interchangeable, scalable, or compatible based on a dictionary and regulatory rules. The metric engine calculates three cost components for candidate edges selected through rule filtering: unit conversion error bound (combining the validity period and confidence level of the exchange rate / unit table), time scaling loss (estimated by the approximate error upper bound of the aggregation / interpolation strategy), and caliber gap penalty (given based on differences in statistical range and anomaly strategies). These three components, along with the invertibility flag (whether an inverse path exists and the error is controllable), together form the edge attributes.
[0221] A4: Perform weighted matching and minimum inconsistency closure on the multi-objective matching graph to obtain the matching relationship;
[0222] Specifically, multi-objective matching needs to simultaneously satisfy two objectives: minimizing global cost and ensuring invertibility. However, these two objectives often conflict when mutually exclusive edges and shared constraints exist. If only locally optimal choices are made using a greedy approach, contradictory mappings across terminals will occur, leading to the inability to uniformly collect subsequent scale parameters. Introducing a minimum inconsistency closure can minimize the set of contradictions in the graph while maintaining an acceptable overall cost, making the output relationship logically consistent and invertible.
[0223] In one example, weighted matching and minimum inconsistency closure are solved on the multi-objective matching graph to obtain matching relationships, including:
[0224] In the multi-target matching graph, conflicting mapping edges are identified. For the detected inconsistent mapping edges, the conflict priority is marked according to the weight value and reversibility constraint. The conflict includes one-to-many mapping conflict, many-to-one mapping conflict and same-dimensional parameter conflict.
[0225] The marked conflict mapping edges are subjected to closure operation, and the conflict mapping edges are eliminated by constraint optimization algorithm to obtain the preliminary matching relationship.
[0226] Based on the mapping edge corresponding to the preliminary matching relationship, a hierarchical adjudication is performed according to the preset adjudication priority rules. The preliminary matching relationship is then modified based on the adjudication results to obtain a matching relationship. The adjudication priority rules include regulatory priority, compliance priority, business priority, and historical version priority.
[0227] A5: Calculate the reversible scaling parameters of the first mechanism end based on the matching relationship, and generate the corresponding parameter commitment;
[0228] The reversible scaling parameters include a scaling factor, a translation factor, and a time normalization factor.
[0229] Specifically, once the matching relationship is determined, the mapping evidence of the structural layer needs to be translated into executable scale parameters and released externally in the form of commitments. This allows the endpoints to perform unified scale and time alignment within the dense domain. The parameters must be reversible and carry error and validity information to facilitate boundary control and version verification on the endpoints during execution.
[0230] In this embodiment, the auditing end calculates a parameter triple for each selected mapping edge: the proportional coefficient is generated by unit / currency mapping and bound to the exchange rate source and validity period; the translation coefficient is given by the benchmark correction amount converted from the difference in caliber (whether it includes tax, refund, or handling fees, etc.); and the time normalization factor is derived from the time granularity difference and the target reference window. Subsequently, the parameters and their evidence chain (dictionary version, exchange rate source, rule number, closure trajectory, and validity period) are written into a parameter commitment. The commitment is fixed by a signature and timestamp and includes a zero-knowledge range proof to prove that the parameters fall within the safe interval and reversible domain. The commitment uses a dual index of version number and digest hash for easy reference by the auditing end and for audit sampling.
[0231] A6: Based on the matching cost corresponding to the matching relationship, the multi-target matching graph is updated through differential privacy protection to obtain a matching distribution graph, and a corresponding global quantile summary commitment is generated based on the matching distribution graph;
[0232] Specifically, the global quantile summary needs to reflect the contribution weight of each target terminal in the statistical fusion, but directly using the original matching cost or terminal density will leak local distribution information. In order to generate a usable global quantile skeleton without exposing the plaintext statistics of terminals, differential privacy perturbations need to be introduced to form a stable probability distribution on the graph, and then the construction of cut points and representative values is organized accordingly.
[0233] In one example, based on the matching cost corresponding to the matching relationship, the multi-objective matching graph is updated using differential privacy protection to obtain a matching distribution graph, including:
[0234] For each mapping edge in the matching relationship, the matching cost is calculated based on the unit conversion error, time granularity scaling error and statistical caliber difference between the dimensional structure signatures, and a matching cost matrix is generated.
[0235] Based on the preset privacy budget parameters, differential privacy perturbation is performed on each matching cost value in the matching cost matrix; the edge weights in the multi-target matching graph are updated according to the perturbed matching cost values, the matching probability distribution is calculated, and a matching distribution graph is generated.
[0236] In this embodiment, the generation of the matching distribution graph is based on the graph weight update mechanism under differential privacy. Its core principle is to transform the structural difference of each mapping edge into a quantization cost, and then achieve the smoothing of the global probability distribution and privacy protection through controlled perturbation.
[0237] Specifically, the auditing end first calculates the matching cost based on the characteristic differences of each mapping edge in the matching relationship. The cost calculation comprehensively considers three types of deviations: unit conversion error, time granularity scaling error, and statistical caliber difference. The unit conversion error is obtained by comparing the convertibility matrix of the unit system and the currency; the time granularity scaling error is derived from the ratio of the sampling window to the target window; and the statistical caliber difference is mapped from the differences in the statistical range, anomaly exclusion rules, and aggregation methods defined by each terminal. The three types of indicators are standardized and weighted to synthesize a cost value, forming a cost matrix. Each row of the matrix corresponds to the indicator of the first institution, and each column corresponds to the candidate mapping path of the data terminal of the second institution.
[0238] Furthermore, after obtaining the cost matrix, the auditing end introduces a differential privacy mechanism to prevent the inference of the internal statistical characteristics of individual terminals from the cost distribution. Specifically, the auditing end applies a random perturbation to each cost value in the matrix according to preset privacy budget parameters. The perturbation amount is generated by a Laplace or Gaussian distribution, and the noise amplitude is dynamically adjusted based on the sensitivity of the cost value. The perturbed matrix is then renormalized to ensure that the weights of all mapping edges still satisfy the probability constraints. Subsequently, the edge weights in the multi-objective matching graph are updated, and the probability of each edge being adopted is calculated based on the perturbed weights, forming a matching probability distribution. The matching probability reflects the credibility of the mapping under global constraints while retaining a certain degree of randomness, preventing external observers from inferring individual mapping relationships through statistical patterns. Finally, the weight distribution protected by differential privacy constitutes the matching distribution graph, providing a robust and secure statistical foundation for the subsequent generation of global quantile summary commitments, theoretically realizing cross-institutional dimensional alignment and statistical weight fusion under privacy constraints.
[0239] It is easy to understand that the specific method for generating the corresponding global quantile summary commitment based on the matching distribution map can be achieved through various feasible statistical fusion methods, such as:
[0240] A quantile fusion algorithm based on matching probability weighting can be used to weight and aggregate the local quantile summaries reported by each data terminal under differential privacy protection to obtain the global quantile cut-off points and representative values. Alternatively, a dense distribution fitting can be performed on the matching distribution graph, and the weighted statistics corresponding to each mapping edge can be generated into a quantile sketch after dense-state weighted averaging and truncation smoothing, which is then solidified into a global summary structure using hash commitment. Those skilled in the art will understand that both of the above methods and their equivalent variations can achieve the generation of global quantile summary commitments, as long as the generated summary structure can be referenced by each data terminal in the dense domain and supports subsequent quantile mapping and order-preserving calibration calculations. This application will not elaborate further on this.
[0241] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A cross-institutional user numerical data processing method based on privacy-preserving computation, applied to the data terminal of a cross-institutional privacy-preserving computation system, characterized in that: The cross-agency privacy computing system also includes an auditing arm, and the method includes: Determine the dimensional metadata corresponding to the first mechanism end to send user numerical data and the second mechanism end to receive user numerical data, wherein the second mechanism end includes at least one data terminal; The dimensional metadata is sent to the auditing end to obtain the reversible scaling parameters and the corresponding parameter commitments and global quantile summary commitments. The first mechanism performs dense-state encoding on the user numerical data according to the reversible scaling transformation parameters to obtain standardized dense-state features, and then sends the standardized dense-state features to the second mechanism. The second agency performs monotonic quantile mapping on the standardized dense state features in the dense state domain based on the global quantile summary commitment, and obtains aligned dense state user numerical data through truncation processing. Based on the reversible scaling transformation parameters, dense-state encoding is performed on the user numerical data to obtain standardized dense-state features, including: The zero-knowledge range verification is performed on the scaling coefficient, translation coefficient, and time normalization factor in the reversible scaling parameters based on the preset public key verification signature. The piecewise linear scaling operator is calculated based on the verified scaling and translation coefficients, and the time resampling kernel is calculated based on the verified time normalization factor. For each user's numerical data, the plaintext numerical data is quantized at fixed points according to the piecewise linear scaling operator and the time resampling kernel, the numerical granularity is made consistent, and the quantized and consistent user numerical data is compressed in intervals and truncated at bit width by using the quantization step size and truncation bit width that match the parameter commitment, so as to obtain a standardized fixed-point vector. The standardized fixed-point vector is encoded in a dense state according to the dense state mask that corresponds one-to-one with the standardized fixed-point vector to generate a standardized dense state feature. The dense state mask is obtained by secret sharing decomposition of the integrity identifier and abnormal missing identifier of the user numerical data and adding random noise perturbation. Based on the global quantile summary commitment, the normalized dense-state features are monotonically mapped in the dense-state domain, and aligned dense-state user numerical data is obtained through truncation, including: For each data terminal in the second agency, the standardized encrypted features transmitted by the first agency and the global quantile digest commitment transmitted by the auditing agency are received. The global quantile digest commitment is signed and verified based on the public key and commitment issuance information corresponding to the data terminal, and the quantile cut-off sequence and corresponding representative value sequence of the global quantile digest commitment in the data terminal are extracted. Based on the quantile tangent sequence, piecewise linear interpolation is performed on the normalized dense state features in the dense state domain to map the normalized dense state features to the initial quantile mapping result; Based on the representative value sequence, the initial quantile mapping result is segmented in the dense state domain, the difference in the dense state mean of each segment is calculated to obtain residual data, and the residual correction function is calculated based on the residual data. The calibration function corresponding to the residual correction function is generated by the order-preserving regression constraint, wherein the calibration function is a monotonically non-decreasing function, and the constraint conditions of the order-preserving regression constraint correspond to the upper and lower threshold information in the global quantile summary commitment; The initial quantile mapping result is corrected and truncated using the calibration function to obtain the aligned dense quantile mapping result.
2. The cross-institutional user numerical data processing method based on privacy-preserving computation according to claim 1, characterized in that, The dimensional metadata corresponding to the first mechanism end for determining the user numerical data to be sent and the second mechanism end for determining the user numerical data to be received includes: Based on the attribute information of the user numerical data to be sent stored in the data terminal of the first institution, a first dimension metadata is generated, wherein the attribute information includes data unit, currency type, time granularity and statistical caliber. For each data terminal in the second agency, obtain the corresponding attribute description information and generate a second-dimensional metadata set.
3. The cross-institutional user numerical data processing method based on privacy-preserving computation according to claim 1, characterized in that, The calculation method for the dense-state mask includes: Extract the integrity identifier corresponding to the user numerical data, wherein the integrity identifier is used to characterize the validity status of the user numerical data during the acquisition, transmission and storage stages; Based on the integrity identifier and the preset anomaly detection rules, the consistency of user numerical data is determined and missing data is identified to obtain an anomaly missing identifier. The abnormal missing identifier is used as input, and the user's numerical data is decomposed by the threshold secret sharing algorithm to obtain multiple secret shares that meet the preset recovery threshold. Each secret share includes a random salt value and a corresponding timestamp. The random salt value is calculated by the threshold secret sharing algorithm. For each secret share, a perturbation share corresponding to the secret share is calculated based on a preset privacy budget and perturbation distribution function, and the perturbation share is encapsulated into a secret mask component through homomorphic encryption; By combining all the dense state mask components through weighted aggregation, a dense state mask corresponding one-to-one with the normalized fixed-point vector is obtained.
4. The cross-institutional user numerical data processing method based on privacy-preserving computation according to claim 1, characterized in that, Extracting the global quantile summary commitment into the data terminal the quantile cut-off sequence and the corresponding representative value sequence, including: Based on the global quantile index mapping table in the global quantile digest commitment, and in conjunction with the service domain identifier of the data terminal, the quantile index range corresponding to the data terminal is determined; The quantile digest structure in the global quantile digest commitment is indexed and located by a dense state retrieval algorithm, and the dense state quantile cut-off point value corresponding to the quantile index range is extracted to generate a quantile cut-off point sequence. Based on the quantile cut-off sequence, the corresponding dense state representative values are extracted from the representative value set in the global quantile summary commitment to obtain a representative value sequence that corresponds one-to-one with the quantile cut-off sequence.
5. The cross-institutional user numerical data processing method based on privacy-preserving computation according to claim 1, characterized in that, The step of generating the calibration function corresponding to the residual correction function through ordinal-preserving regression constraints includes: Calculate the piecewise residual sequence corresponding to the residual correction function in the dense state domain, wherein the piecewise residual sequence is used to characterize the deviation trend of the initial quantile mapping result in the quantile interval; Based on the upper and lower threshold information in the global quantile summary commitment, constraints are constructed for the residual correction function. These constraints include non-subtractive relation constraints on quantile segment residuals, residual magnitude limitation constraints, and total change constraints. Based on the constraints, the segmented residual sequence is subjected to order-preserving regression calculation. The residuals that do not satisfy the monotonic relationship are merged by dense-state weighted average and dense-state comparison operation to obtain a dense-state calibration sequence that is adapted to the monotonic non-decreasing characteristics. Based on the dense-state calibration sequence, a piecewise linear expression of the calibration function is constructed in the dense-state domain, and the calibration slope and offset parameters of each segment are determined to generate the calibration function.
6. A cross-organizational user numerical data processing method based on privacy-preserving computation, applied to the audit end of a cross-organizational privacy-preserving computation system, characterized in that... The cross-agency privacy computing system further includes data terminals, wherein the data terminal for transmitting user numerical data is a first agency end, and the data terminal for receiving user numerical data is a second agency end, wherein the second agency end includes at least one data terminal, and the method includes: Receive dimensional metadata from the first institution and a set of dimensional metadata from the second institution; Based on a preset public key, the dimensional metadata and the dimensional metadata set are signed and their integrity verified, and then standardized and encoded to generate a first dimensional structure signature and a second dimensional structure signature. Using the first dimensional structure signature as the source node and the second dimensional structure signature as the target node, a multi-target matching graph is constructed, wherein the mapping edges of the multi-target matching graph are established through the unit interchangeability, time granularity scalability and statistical caliber compatibility between dimensional structure signatures. The matching relationships are obtained by performing weighted matching and minimum inconsistency closure on the multi-objective matching graph. The reversible scale transformation parameters of the first mechanism end are calculated based on the matching relationship, and the corresponding parameter commitments are generated. The reversible scale transformation parameters include a scaling factor, a translation factor, and a time normalization factor. Based on the matching cost corresponding to the matching relationship, the multi-target matching graph is updated through differential privacy protection to obtain a matching distribution graph, and a corresponding global quantile summary commitment is generated based on the matching distribution graph.
7. The cross-institutional user numerical data processing method based on privacy-preserving computation according to claim 6, characterized in that, Weighted matching and minimum inconsistency closure are performed on the multi-objective matching graph to obtain the matching relationships, including: In the multi-target matching graph, conflicting mapping edges are identified. For the detected inconsistent mapping edges, the conflict priority is marked according to the weight value and reversibility constraint. The conflict includes one-to-many mapping conflict, many-to-one mapping conflict and same-dimensional parameter conflict. The marked conflict mapping edges are subjected to closure operation, and the conflict mapping edges are eliminated by constraint optimization algorithm to obtain the preliminary matching relationship. Based on the mapping edge corresponding to the preliminary matching relationship, a hierarchical adjudication is performed according to the preset adjudication priority rules. The preliminary matching relationship is then modified based on the adjudication results to obtain a matching relationship. The adjudication priority rules include regulatory priority, compliance priority, business priority, and historical version priority.
8. The cross-institutional user numerical data processing method based on privacy-preserving computation according to claim 6, characterized in that, Based on the matching cost corresponding to the matching relationship, the multi-target matching graph is updated using differential privacy protection to obtain a matching distribution graph, including: For each mapping edge in the matching relationship, the matching cost is calculated based on the unit conversion error, time granularity scaling error and statistical caliber difference between the dimensional structure signatures, and a matching cost matrix is generated. Based on the preset privacy budget parameters, differential privacy perturbation is performed on each matching cost value in the matching cost matrix; the edge weights in the multi-target matching graph are updated according to the perturbed matching cost values, the matching probability distribution is calculated, and a matching distribution graph is generated.