A Privacy-Preserving Collaborative Modeling and Intelligent Generation Method for Multi-Institutional Medical Data

By constructing a cross-institutional feature semantic mapping graph and a hierarchical privacy-preserving hybrid coding data structure, the heterogeneity and privacy protection issues of medical data are solved, enabling semantic alignment and efficient collaborative modeling of cross-institutional data, and improving model performance and privacy security.

CN121167788BActive Publication Date: 2026-03-13NEWLINK TECH INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the heterogeneity of feature spaces and differences in semantic representation among multi-institutional medical data, impacting the effective integration of cross-institutional data and model performance. Furthermore, traditional privacy protection methods fail to strike a balance between protecting privacy and maintaining data availability.

Method used

We construct a cross-institutional feature semantic mapping graph that integrates a contrastive learning mechanism, achieve feature semantic alignment through a feature semantic dual-space encoding structure, generate a hybrid encoding data structure with hierarchical protection strength, combine a homomorphic encryption mechanism for privacy protection, and use feature-level privacy sensitivity vectors to allocate differentiated privacy budgets for features with different sensitivities.

Benefits of technology

It achieves semantic alignment of medical data across institutions, improves the effectiveness and accuracy of collaborative modeling, reduces the risk of privacy leakage by 30%, maintains data availability, and enhances the generalization ability and robustness of generative models, improving the accuracy of multiple clinical prediction tasks by 15%-25%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167788B_ABST
    Figure CN121167788B_ABST
Patent Text Reader

Abstract

This invention provides a privacy-preserving method for collaborative modeling and intelligent generation of multi-institutional medical data, relating to the field of privacy protection technology. It includes addressing feature heterogeneity through cross-institutional feature semantic mapping graphs, constructing a hybrid coding data structure for layered privacy protection, securely fusing multiple local feature representations using homomorphic encryption to generate a global representation, and allocating differentiated privacy budgets based on feature sensitivity. This invention achieves secure collaborative modeling of multi-institutional data, improving model performance while ensuring data privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to privacy protection technology, and more particularly to a method for collaborative modeling and intelligent generation of multi-institutional medical data based on privacy protection. Background Technology

[0002] With the rapid development of medical informatization, medical institutions have accumulated a large amount of electronic health records, medical images, and diagnostic data. This heterogeneous medical data has enormous research and application value, particularly in areas such as intelligent medical decision support, precision diagnosis, and personalized treatment planning. However, due to the high privacy sensitivity of medical data and the data silos between different medical institutions, these valuable data resources are difficult to fully integrate and utilize. Cross-institutional collaborative analysis and modeling of medical data has become a key technical means to solve problems such as the dispersion of medical resources and limited diagnostic accuracy.

[0003] In recent years, collaborative modeling of multi-institutional medical data has gradually attracted widespread attention from academia and industry. Existing research mainly focuses on technologies such as federated learning, differential privacy, and homomorphic encryption, aiming to achieve multi-party collaborative analysis without sharing the original data. Traditional collaborative modeling methods typically employ centralized data processing before model training, or utilize simple federated learning frameworks for decentralized training. While these methods improve model performance, they also bring challenges in data privacy protection and handling data heterogeneity among institutions.

[0004] Existing technologies struggle to effectively address the heterogeneity of feature spaces and differences in semantic representation among multi-institutional medical data. Differences in equipment, recording standards, and treatment protocols among different medical institutions lead to significant variations in feature representation and semantic interpretation of collected medical data. This severely impacts the effective integration of cross-institutional data and model performance.

[0005] Traditional privacy protection methods typically apply the same level of protection to all features, ignoring the significantly different privacy sensitivities of different features in medical data. This "one-size-fits-all" approach leads to either overprotection of low-sensitivity features, reducing data usability, or insufficient protection of high-sensitivity features, resulting in the risk of privacy breaches.

[0006] Existing collaborative modeling frameworks struggle to balance privacy protection with model performance. Most methods experience a significant decline in model accuracy and generalization ability after enhancing privacy protection; conversely, methods prioritizing model performance often suffer from insufficient privacy protection, failing to achieve a good balance between the security and effective use of medical data. Summary of the Invention

[0007] This invention provides a privacy-preserving method for collaborative modeling and intelligent generation of multi-institutional medical data, which can solve the problems in the prior art.

[0008] A first aspect of this invention provides a method for collaborative modeling and intelligent generation of multi-institutional medical data based on privacy protection, comprising:

[0009] Local medical datasets are obtained from multiple data holding institutions. Based on the heterogeneity of the feature space and the differences in semantic expression of the local medical datasets, a cross-institutional feature semantic mapping map with a contrastive learning mechanism is constructed. The cross-institutional feature semantic mapping map captures and quantifies the semantic similarity of features between different institutions through the contrastive learning mechanism.

[0010] Heterogeneous features are mapped to a unified semantic space to obtain a semantically aligned dataset. A privacy risk assessment is performed on the semantically aligned dataset to generate a feature-level privacy sensitivity vector. Based on this, a hybrid coding data structure with hierarchical protection strength is constructed.

[0011] Based on the hybrid coding data structure, the local feature representations of each data holding institution are extracted through the contrastive learning mechanism, and the local feature representations are encrypted using the homomorphic encryption mechanism. The multiple encrypted local feature representations are then fused to generate a global feature representation.

[0012] The global feature representation is used to train and generate a model. During the training process, a differentiated privacy budget is allocated to features with different sensitivity based on the feature-level privacy sensitivity vector. The model parameters are then updated iteratively and protectively through the homomorphic encryption mechanism.

[0013] To address the heterogeneity of the feature space and the differences in semantic representation of the local medical dataset, a cross-institutional feature semantic mapping map fused with a contrastive learning mechanism is constructed. This cross-institutional feature semantic mapping map captures and quantifies the semantic similarity of features between different institutions through the contrastive learning mechanism, including:

[0014] To address the heterogeneity issue of the local medical dataset, a feature-semantic dual-space encoding structure is constructed. This structure preserves the original feature expression through the institution's private semantic space and establishes a mapping relationship between features through a cross-institutional shared semantic space. This achieves semantic interoperability of features while protecting the uniqueness of institutional data, resulting in a dual-space feature embedding representation.

[0015] A contrastive learning framework is constructed based on the dual-space feature embedding representation. The contrastive learning framework identifies anchor features from data of various institutions through semantic stability analysis, constructs positive sample pairs by embedding the anchor features in the cross-institutional shared semantic space, and selects negative sample pairs by feature correlation constraints. The mapping parameters of the feature semantic dual-space encoding structure are optimized by contrastive learning so that features with similar semantics have similarity in the cross-institutional shared semantic space, resulting in the optimized dual-space feature embedding representation.

[0016] The optimized dual-space feature embedding representation is used to calculate the semantic similarity matrix of cross-institutional features. The semantic similarity matrix is ​​verified by introducing semantic transitivity consistency constraints. A cross-institutional feature semantic mapping map is constructed by incorporating the contrastive learning mechanism. Based on the cross-institutional feature semantic mapping map, the heterogeneous features of each institution are semantically unified to achieve semantic alignment of cross-institutional data.

[0017] Heterogeneous features are mapped to a unified semantic space to obtain a semantically aligned dataset. A privacy risk assessment is then performed on the semantically aligned dataset to generate a feature-level privacy sensitivity vector. Based on this vector, a hybrid coding data structure with hierarchical protection strength is constructed, including:

[0018] A multidimensional privacy risk analysis framework is constructed, and a privacy assessment is performed on the semantically aligned local medical dataset through the multidimensional privacy risk analysis framework to obtain multidimensional privacy risk indicators.

[0019] Membership inference and attribute inference attacks are performed on the semantically aligned local medical dataset. Multidimensional risk aggregation weights are obtained by analyzing the impact of the multidimensional privacy risk indicators on the attack success rate. The multidimensional risk aggregation weights are then comprehensively calculated and integrated with the privacy risk indicators of the corresponding dimensions to generate a feature-level privacy sensitivity vector that quantifies the privacy protection requirements of features.

[0020] Based on the feature-level privacy sensitivity vector, features are protected hierarchically. According to the sensitivity level, the features are divided into different sensitive feature sets by an adaptive threshold. For feature sets with different sensitivity levels, dense feature representations and perturbation feature representations are generated respectively. The dense feature representations and perturbation feature representations are recombined according to the structural relationship between the original features to construct a hybrid coding data structure with hierarchical protection strength.

[0021] Membership inference and attribute inference attacks are performed on the semantically aligned local medical dataset. The multidimensional risk aggregation weights are obtained by analyzing the impact of the multidimensional privacy risk indicators on the attack success rate, including:

[0022] The target record set is extracted from the semantically aligned local medical dataset, and a background knowledge set simulating attacker auxiliary information is constructed. Based on the known values ​​of the background knowledge set, a member inference attack executor is constructed. The member inference attack executor analyzes the association features between the target record and the original dataset, identifies and determines the dataset affiliation relationship of each record in the target record set, and obtains a member determination result with multi-dimensional verification.

[0023] Based on the member determination results of the multi-dimensional verification, member records with deterministic confidence are obtained. An attribute inference attack executor is constructed by combining the background knowledge set. The attribute inference attack executor is used to mine the implicit correlation patterns between features in the confidence member records. Undisclosed sensitive feature values ​​are inferred from multiple angles to form a systematic attribute inference result.

[0024] A multi-layered risk assessment system is constructed using the member inference attack executor and the attribute inference attack executor. Based on the risk assessment system, attack tests of different intensities and strategies are performed on the semantically aligned local medical dataset to explore the data vulnerability boundary and obtain the attack success rate under each dimension. Based on the attack success rate, the propagation characteristics and impact of privacy risks in each dimension are analyzed and evaluated, and a multi-dimensional risk aggregation weight that can dynamically reflect the importance of the dimensions is generated.

[0025] Based on the hybrid coding data structure, local feature representations of each data holding institution are extracted through the contrastive learning mechanism, and homomorphic encryption is used to encrypt the local feature representations. The fusion of multiple encrypted local feature representations to generate a global feature representation includes:

[0026] Based on the aforementioned hybrid coding data structure, a contrastive learning module is designed. The contrastive learning module analyzes the multi-level semantic associations and dynamic structural features in the hybrid coding data structures of various data holding institutions, and establishes an adaptive cross-institutional data similarity measurement system.

[0027] Based on the cross-institutional data similarity measurement system, local feature representations are generated for the hybrid coded data structures of each data holding institution; the local feature representations are enhanced with security by a homomorphic encryption module to generate encrypted feature representations that are tamper-proof and maintain the original structure and computational properties; the encrypted feature representations are aggregated to a central server according to the data holding institutions, and the central server forms aggregated encrypted features in an adaptive weighted manner within the encryption domain, and after decryption, obtains a global feature representation that integrates differentiated knowledge from multiple parties.

[0028] The generative model is trained using the global feature representation. During training, differentiated privacy budgets are allocated to features with different sensitivity based on the feature-level privacy sensitivity vector, including:

[0029] The training mechanism for constructing a generative model is constructed using the global feature representation. The global feature representation is deconstructed in multiple dimensions through the training mechanism to establish a dimensional index system that reflects the essential attributes of the features.

[0030] Based on the dimensional indexing system, the feature-level privacy sensitivity vector is associated, the sensitivity of each dimension feature is quantitatively evaluated and a sensitivity mapping relationship is formed. A privacy budget allocation function is designed according to the sensitivity mapping relationship. An adaptive mapping mechanism between the sensitivity quantification value and the budget amount is constructed according to the privacy budget allocation function. The budget is allocated differently for each dimension feature through the adaptive mapping mechanism to form a feature-level privacy budget allocation scheme.

[0031] During the parameter optimization process of the generative model, the gradient components of each dimension feature are dynamically calibrated according to the feature-level privacy budget allocation scheme to generate a gradient update direction with privacy protection effect; the network parameters of the generative model are adjusted based on the gradient update direction until the model converges and the training process is completed.

[0032] Based on the sensitivity mapping relationship, a privacy budget allocation function is designed. Based on the privacy budget allocation function, an adaptive mapping mechanism is constructed between sensitivity metrics and budget amounts, including:

[0033] Based on the sensitivity mapping relationship, the sensitivity metrics are deconstructed in multiple dimensions to construct a feature characterization system that represents the sensitivity metrics. Based on the feature characterization system, the function form of the privacy budget allocation function is designed. Based on the function form of the privacy budget allocation function, a mapping calculation unit is constructed. The sensitivity metrics of each feature dimension in the sensitivity mapping relationship are evaluated through the mapping calculation unit, and differential transformation operations are performed to generate the initial budget amount corresponding to each feature dimension.

[0034] A budget calibration mechanism is constructed based on the initial budget amount. The hierarchical relationship and competitive game characteristics between budget amounts in each dimension are explored according to the budget calibration mechanism to obtain a calibrated budget amount that considers the constraint coupling between dimensions.

[0035] A bidirectional mapping architecture is constructed between the calibration budget and the sensitivity quantification. The structural deviation between the calibration budget and the sensitivity quantification is measured through the bidirectional mapping architecture, and an adjustment signal is constructed. The function shape of the privacy budget allocation function is dynamically optimized based on the adjustment signal, thus forming an adaptive mapping mechanism.

[0036] A second aspect of the present invention provides an electronic device, comprising:

[0037] processor;

[0038] Memory used to store processor-executable instructions;

[0039] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0040] A third aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0041] The beneficial effects of this application are as follows:

[0042] By constructing a cross-institutional feature semantic mapping map with a fusion contrastive learning mechanism, the problem of feature space heterogeneity and semantic expression differences of medical data in a multi-institutional environment is solved, feature semantic alignment between different data sources is achieved, and the effectiveness and accuracy of collaborative modeling are improved.

[0043] By employing a hybrid coding data structure based on feature-level privacy sensitivity vectors, we have achieved hierarchical and refined privacy protection for medical data. Compared with traditional methods, this reduces the risk of privacy leakage by approximately 30% while maintaining data availability. It protects patient privacy without compromising the analytical value of medical data.

[0044] By combining homomorphic encryption with a global feature fusion mechanism based on differentiated privacy budgets, institutions can collaborate on modeling without directly sharing raw data. This solves the problem of data silos among medical institutions and improves the generalization ability and robustness of the generated model. Compared with single-institution models, the accuracy is improved by 15%-25% in multiple clinical prediction tasks. Attached Figure Description

[0045] Figure 1 This is a flowchart illustrating the privacy-preserving multi-institutional medical data collaborative modeling and intelligent generation method according to an embodiment of the present invention.

[0046] Figure 2 This is a flowchart illustrating the privacy protection process of distributed institutional feature fusion and knowledge integration in an embodiment of the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0049] Figure 1This is a flowchart illustrating the privacy-preserving multi-institutional medical data collaborative modeling and intelligent generation method according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0050] Local medical datasets are obtained from multiple data holding institutions. Based on the heterogeneity of the feature space and the differences in semantic expression of the local medical datasets, a cross-institutional feature semantic mapping map with a contrastive learning mechanism is constructed. The cross-institutional feature semantic mapping map captures and quantifies the semantic similarity of features between different institutions through the contrastive learning mechanism.

[0051] Heterogeneous features are mapped to a unified semantic space to obtain a semantically aligned dataset. A privacy risk assessment is performed on the semantically aligned dataset to generate a feature-level privacy sensitivity vector. Based on this, a hybrid coding data structure with hierarchical protection strength is constructed.

[0052] Based on the hybrid coding data structure, the local feature representations of each data holding institution are extracted through the contrastive learning mechanism, and the local feature representations are encrypted using the homomorphic encryption mechanism. The multiple encrypted local feature representations are then fused to generate a global feature representation.

[0053] The global feature representation is used to train and generate a model. During the training process, a differentiated privacy budget is allocated to features with different sensitivity based on the feature-level privacy sensitivity vector. The model parameters are then updated iteratively and protectively through the homomorphic encryption mechanism.

[0054] In one optional implementation, considering the feature space heterogeneity and semantic expression differences of the local medical dataset, a cross-institutional feature semantic mapping map integrating a contrastive learning mechanism is constructed. This cross-institutional feature semantic mapping map captures and quantifies the semantic similarity of features between different institutions through the contrastive learning mechanism, including:

[0055] To address the heterogeneity issue of the local medical dataset, a feature-semantic dual-space encoding structure is constructed. This structure preserves the original feature expression through the institution's private semantic space and establishes a mapping relationship between features through a cross-institutional shared semantic space. This achieves semantic interoperability of features while protecting the uniqueness of institutional data, resulting in a dual-space feature embedding representation.

[0056] A contrastive learning framework is constructed based on the dual-space feature embedding representation. The contrastive learning framework identifies anchor features from data of various institutions through semantic stability analysis, constructs positive sample pairs by embedding the anchor features in the cross-institutional shared semantic space, and selects negative sample pairs by feature correlation constraints. The mapping parameters of the feature semantic dual-space encoding structure are optimized by contrastive learning so that features with similar semantics have similarity in the cross-institutional shared semantic space, resulting in the optimized dual-space feature embedding representation.

[0057] The optimized dual-space feature embedding representation is used to calculate the semantic similarity matrix of cross-institutional features. The semantic similarity matrix is ​​verified by introducing semantic transitivity consistency constraints. A cross-institutional feature semantic mapping map is constructed by incorporating the contrastive learning mechanism. Based on the cross-institutional feature semantic mapping map, the heterogeneous features of each institution are semantically unified to achieve semantic alignment of cross-institutional data.

[0058] To address the feature space heterogeneity issue in multi-institutional medical datasets, a dual-space feature-semantic encoding structure employs a hierarchical mapping architecture to achieve unified data representation. This encoding structure consists of two independent but related representation layers: an institution-specific semantic space and a cross-institutional shared semantic space. The institution-specific semantic space maintains the semantic integrity of the original features, with a 512-dimensional vector space. Densely connected layers perform nonlinear transformations on the input medical features, and the Swish activation function is used to maintain the smoothness of gradient flow. The cross-institutional shared semantic space is designed as a 256-dimensional common representation space. A learnable linear projection layer maps the features from the private space to the shared space. The projection matrix is ​​initialized using a Xavier normal distribution, with a learning rate of 0.001 and a weight decay coefficient of 1e-4. The dual-space encoder ensures the effectiveness of information transmission through a residual connection mechanism, sets the dropout probability to 0.1 to prevent overfitting, and applies batch normalization before each transformation layer to stabilize the training process.

[0059] In constructing the institution's private semantic space, each data-holding institution's local medical data first undergoes feature standardization. The Z-score standardization method maps continuous features to a zero-mean, unit-variance distribution, while discrete features are converted into binary vector representations through one-hot encoding. The feature embedding layer uses a trainable embedding matrix with a dimension equal to the number of features multiplied by the embedding dimension. The embedding dimension is dynamically adjusted based on feature complexity, ranging from 64 to 128 dimensions. The private encoder network structure consists of three fully connected layers with 1024, 768, and 512 hidden neurons, respectively. ReLU activation and batch normalization are applied between each layer, ultimately outputting a 512-dimensional private feature representation vector.

[0060] A cross-institutional shared semantic space is achieved through an alignment loss function, which maps the semantic features of different institutions. The alignment loss uses cosine similarity as a metric, aiming to maximize the similarity of semantically similar features within the shared space. The shared space mapping network employs a two-layer fully connected architecture. The first layer compresses the 512-dimensional private features into a 384-dimensional intermediate representation, and the second layer further maps them to a 256-dimensional shared representation. Orthogonality constraints are introduced during the mapping process to ensure the independence of features across different dimensions. The orthogonality loss weight is set to 0.01, and the orthogonality of the mapping matrix is ​​maintained through a Gram-Schmidt orthogonalization process.

[0061] The contrastive learning framework constructs a positive-negative sample pair selection mechanism based on dual-space feature embedding representations. Semantic stability analysis determines anchor features by calculating the variance stability of features within a time window. The time window is set to 30 consecutive data batches, and the variance threshold is set to 0.05; features below this threshold are identified as semantically stable anchor features. The anchor feature recognition algorithm traverses all feature dimensions, calculates the moving variance of each dimension within the time window, and smooths the variance calculation results using an exponential moving average method with a smoothing factor set to 0.9. The embedding representations of the identified anchor features in the cross-institutional shared semantic space constitute a candidate set of positive sample pairs.

[0062] During the construction of positive sample pairs, anchor features from different institutions but with similar semantics are paired using cosine similarity calculation. A similarity threshold of 0.8 is set; feature pairs exceeding this threshold are selected as positive samples. Feature correlation constraints quantify the linear relationship between features using the Pearson correlation coefficient, with a correlation threshold of 0.3. Feature pairs below this threshold are selected as negative samples. The negative sample sampling strategy employs a hard negative sample mining method, prioritizing feature pairs with similarity between 0.4 and 0.6 as hard negative samples to improve the training effect of contrastive learning. The ratio of positive to negative samples in each training batch is maintained at 1:2, and the batch size is set to 128 sample pairs.

[0063] The contrastive learning optimization process employs the InfoNCE loss function, with a temperature parameter set to 0.07. This parameter controls the sharpness of the similarity distribution; a lower temperature value helps learn more refined feature distinctions. The optimizer used is AdamW, with a cosine annealing scheduling strategy for the learning rate. The initial learning rate is set to 3e-4, the minimum learning rate to 1e-6, and the training epochs are 200. The gradient clipping threshold is set to 1.0 to prevent gradient explosion, and the learning rate decays every 50 epochs with a decay factor of 0.8.

[0064] The optimized dual-space feature embedding representation is used to calculate the semantic similarity matrix for cross-institutional features. Similarity calculation employs a standardized dot product operation, with matrix element values ​​ranging from -1 to 1. Semantic transitivity consistency constraints are implemented through a triplet consistency check. For a feature triplet (A, B, C), if A is similar to B and B is similar to C, then the similarity between A and C must satisfy the transitivity constraint. The transitivity threshold is set to 0.1, meaning that consistency constraints are considered satisfied when the similarity difference does not exceed this threshold. The consistency check algorithm iterates through all feature triplets, calculating the proportion of transitivity violations. When the violation proportion exceeds 5%, the similarity matrix is ​​recalculated.

[0065] In the construction of the cross-institutional feature semantic mapping graph, the similarity matrix is ​​modeled using a graph neural network structure. Nodes in the graph represent features of different institutions, and edge weights correspond to semantic similarity values. The graph convolutional layers adopt the GraphSAGE architecture, using a mean aggregator as the aggregation function. The neighborhood sampling number is set to 15 nodes, and the number of layers is set to 3. The graph update adopts an asynchronous update strategy. Each institution independently updates its local graph structure, and the global graph parameters are synchronized using a federated averaging method, with a synchronization frequency set to once every 10 training batches.

[0066] Semantic unification processing is based on a constructed cross-institutional feature semantic mapping graph to align and transform heterogeneous features. The alignment algorithm employs the Wasserstein distance minimization method from optimal transfer theory, achieving feature distribution alignment by solving for the optimal transfer matrix. The transfer matrix is ​​solved using the Sinkhorn iterative algorithm, with 100 iterations, a convergence threshold of 1e-4, and a regularization parameter of 0.1 to control the smoothness of the transfer. The aligned features retain the statistical properties of the original features. The alignment effect is verified using KL divergence, with a KL divergence threshold set to 0.05; values ​​below this threshold indicate successful alignment.

[0067] In the implementation case, three medical institutions provided electronic health record data for cardiovascular disease, diabetes, and oncology, respectively, with feature dimensions of 1024, 768, and 896 dimensions. After dual-space encoding, the private space dimension was unified to 512 dimensions, and the shared space dimension to 256 dimensions. Anchor feature recognition identified 156, 142, and 168 stable features in the three institutions, respectively, with 2847 positive sample pairs and 5694 negative sample pairs. After 150 epochs of comparative learning training, the average similarity of the cross-institutional feature similarity matrix improved from the initial 0.23 to 0.67, and the semantic transitivity consistency violation rate decreased from the initial 12.3% to 2.8%. After semantic unification, the KL divergences of the feature distributions of the three institutions were 0.032, 0.041, and 0.038, respectively, all meeting the alignment success threshold requirements, achieving effective semantic alignment of heterogeneous medical data.

[0068] In one optional implementation, heterogeneous features are mapped to a unified semantic space to obtain a semantically aligned dataset. A privacy risk assessment is then performed on the semantically aligned dataset to generate a feature-level privacy sensitivity vector. Based on this vector, a hybrid encoded data structure with hierarchical protection strength is constructed, including:

[0069] A multidimensional privacy risk analysis framework is constructed, and a privacy assessment is performed on the semantically aligned local medical dataset through the multidimensional privacy risk analysis framework to obtain multidimensional privacy risk indicators.

[0070] Membership inference and attribute inference attacks are performed on the semantically aligned local medical dataset. Multidimensional risk aggregation weights are obtained by analyzing the impact of the multidimensional privacy risk indicators on the attack success rate. The multidimensional risk aggregation weights are then comprehensively calculated and integrated with the privacy risk indicators of the corresponding dimensions to generate a feature-level privacy sensitivity vector that quantifies the privacy protection requirements of features.

[0071] Based on the feature-level privacy sensitivity vector, features are protected hierarchically. According to the sensitivity level, the features are divided into different sensitive feature sets by an adaptive threshold. For feature sets with different sensitivity levels, dense feature representations and perturbation feature representations are generated respectively. The dense feature representations and perturbation feature representations are recombined according to the structural relationship between the original features to construct a hybrid coding data structure with hierarchical protection strength.

[0072] The multidimensional privacy risk analysis framework employs five independent risk assessment dimensions to conduct a comprehensive privacy assessment of the semantically aligned local medical dataset, including uniqueness risk, association risk, sensitivity risk, inference risk, and background knowledge risk. The uniqueness risk assessment module quantifies the recognition potential of features by calculating the frequency and sparsity of feature values ​​in the dataset, using Shannon entropy to calculate the information content of features, with entropy values ​​ranging from 0 to logarithm N, where N is the number of samples in the dataset. The association risk assessment quantifies the dependencies between features using mutual information theory, calculating the joint probability distribution through a kernel density estimation algorithm, and adaptively setting the bandwidth parameter using the Scott criterion, with values ​​ranging from 0.1 to 2.0. The sensitivity risk assessment constructs a sensitivity dictionary based on expert knowledge in the medical field, including highly sensitive categories such as disease diagnosis, genetic information, and mental health, with each category assigned a sensitivity weight between 0 and 1.

[0073] Inference risk assessment measures the predictability of target features by constructing a decision tree classifier. The maximum depth of the decision tree is set to 8 layers, the minimum number of split samples is 50, and the minimum number of leaf node samples is 20. Features with a classification accuracy exceeding 0.7 are marked as high inference risk, features with an accuracy between 0.5 and 0.7 are marked as medium inference risk, and features with an accuracy below 0.5 are marked as low inference risk. Background knowledge risk is assessed through matching with an external knowledge base containing publicly available medical statistics, census information, and medical literature data. The matching degree is calculated using cosine similarity, with a threshold set at 0.6. Features exceeding the threshold are considered to have a risk of background knowledge leakage.

[0074] In the calculation of privacy risk indicators, each dimension generates a normalized risk score between 0 and 1. The uniqueness risk score is normalized by dividing the entropy value by the maximum entropy value, and the correlation risk score is normalized by dividing the mutual information value by the maximum mutual information value. The sensitivity risk score directly uses predefined sensitivity weights, the inference risk score equals the accuracy of the decision tree classifier, and the background knowledge risk score equals the matching similarity of the external knowledge base. The five dimensions of risk indicators constitute a multidimensional privacy risk vector with a dimension of 5, and each element takes a value between 0 and 1.

[0075] The member inference attack process employs a shadow model approach. The training dataset is divided into member and non-member data in an 8:2 ratio. The shadow model uses the same neural network architecture as the target model, containing three fully connected layers with 512, 256, and 128 hidden neurons, respectively. The attack model uses a logistic regression classifier, with input features being the target model's predicted probability vector for the samples and the loss value. The attack success rate is quantified by the area under the receiver operating characteristic (ROC) curve, with values ​​ranging from 0.5 to 1.0. 0.5 represents a random guess level, and 1.0 represents a perfect attack.

[0076] The attribute inference attack targets specific sensitive attributes. The attacker's model employs a gradient boosting decision tree algorithm with 100 trees, a learning rate of 0.1, and a maximum depth of 6 layers. Attack features include all observable features other than the target attribute. The attack success rate is comprehensively evaluated using precision, recall, and a combined evaluation score. The attack experiments use a 10-fold cross-validation method, with 90% of the data used for training and 10% for testing in each experiment. The attack success rate is the average of the ten experimental results.

[0077] The calculation of multidimensional risk aggregation weights is based on the sensitivity analysis of attack success rate to each risk dimension, and uses partial correlation analysis to quantify the independent contribution of each risk dimension to the attack success rate. The influence of other risk dimensions is controlled during the calculation of the partial correlation coefficient to obtain the net influence strength of each dimension. The absolute value of the partial correlation coefficient is used as the aggregation weight for that dimension. Weight normalization ensures that the sum of all dimension weights equals 1. The normalization method uses first-order norm standardization, i.e., each weight is divided by the sum of the absolute values ​​of all weights.

[0078] Feature-level privacy sensitivity vectors are generated through weighted linear combination calculations. The sensitivity value equals the sum of the products of each risk dimension score and its corresponding aggregation weight. Vectorization is employed in the calculation process to improve efficiency, supporting batch processing of sensitivity calculations for multiple features. The dimension of the sensitivity vector equals the number of features, with each element corresponding to the privacy sensitivity score of a single feature, ranging from 0 to 1, where 0 indicates no privacy risk and 1 indicates extremely high privacy risk.

[0079] The adaptive threshold determination employs the K-means clustering algorithm to divide features into three sensitive feature sets based on their sensitivity scores. The number of clusters, K, is set to 3, corresponding to the high-sensitivity, medium-sensitivity, and low-sensitivity feature sets, respectively. Cluster initialization uses the K-means enhanced initialization method, with a maximum of 300 iterations and a convergence threshold of 0.01%. The clustering results are evaluated using the silhouette coefficient; a silhouette coefficient exceeding 0.5 is considered a good clustering effect, while a coefficient below 0.3 requires readjustment of clustering parameters or an increase in the number of clusters.

[0080] The highly sensitive feature set is represented using homomorphic encryption to generate a encrypted feature representation. The encryption scheme employs a homomorphic encryption algorithm with a polynomial degree of 16384, a modular chain length of 5 layers, and a security parameter of 128 bits. The encrypted feature representation maintains the homomorphic operation capability under encryption, supporting addition and multiplication operations. The ciphertext size is approximately 50 to 100 times that of the plaintext. The encryption process uses batch encoding technology to improve efficiency, processing 2048 plaintext elements in a single encoding operation, achieving an encoding density of over 50%.

[0081] The sensitive feature set is represented by perturbations using a differential privacy mechanism, with a Laplace's algorithm used for noise generation. The privacy budget parameter is set to 1.0. Sensitivity is calculated based on the global sensitivity of the feature values, taking the difference between the maximum and minimum feature values. Noise generation employs an inverse transform sampling method, and the pseudo-random number generator uses the Mason tween algorithm. The seed value is generated based on the current timestamp and process identifier to ensure randomness. During noise injection, noise is added independently to each feature value, with the noise amplitude dynamically adjusted according to the privacy budget and sensitivity.

[0082] The low-sensitivity feature set retains its original numerical representation, undergoing only data type conversion and format standardization. Numerical features are represented using 32-bit floating-point numbers, categorical features use integer encoding, and text features are stored using a common character encoding. Feature value range checks ensure that values ​​are within a reasonable range; outliers exceeding the range are detected and handled using box plot methods, and are replaced with quantile boundary values.

[0083] The hybrid coding data structure reconstruction process maintains the structural relationships between the original features and uses a sparse matrix storage format to reduce storage overhead. The data structure includes a feature index mapping table, a sensitivity label array, and a coding type identifier. The index mapping table records the correspondence between the original feature positions and their encoded positions, the sensitivity labels identify the sensitivity level of each feature, and the coding type identifier indicates the protection method used for the feature.

[0084] The structural reorganization algorithm employs topological sorting to maintain dependencies between features. The dependency graph is constructed based on feature correlation analysis results, with a correlation threshold set to 0.3. Directed edges are established between feature pairs exceeding this threshold. The sorting results determine the storage order of features within the hybrid encoding structure, ensuring that dependent features are processed before those they depend on. Data structure serialization uses a protocol buffer format, supporting cross-platform data exchange and version compatibility.

[0085] In the implementation case, a cardiovascular disease dataset containing 15,000 patient records, after semantic alignment, had 128 features. The multidimensional privacy risk analysis framework identified 32 features with high uniqueness risk, 28 features with high association risk, and 45 features with high sensitivity risk. The area under the curve for membership inference attacks on the original data was 0.78, and the comprehensive evaluation score for attribute inference attacks was 0.65. The multidimensional risk aggregation weight calculation results showed that the sensitivity risk weight was the highest at 0.35, the uniqueness risk weight was 0.28, the association risk weight was 0.22, the inference risk weight was 0.10, and the background knowledge risk weight was 0.05. After generating the feature-level privacy sensitivity vector, K-means clustering divided the features into 41 high-sensitivity features, 52 medium-sensitivity features, and 35 low-sensitivity features. After homomorphic encryption, the ciphertext size of highly sensitive features increased by 68 times. After adding Laplace noise to medium-sensitive features, the average signal-to-noise ratio was 15.2 dB. The storage size of the hybrid encoded data structure was 2.3 times that of the original data, and the average feature query response time was 8.5 milliseconds.

[0086] In one optional implementation, membership inference and attribute inference attacks are performed on the semantically aligned local medical dataset. The multidimensional risk aggregation weights are obtained by analyzing the impact of the multidimensional privacy risk indicators on the attack success rate, including:

[0087] The target record set is extracted from the semantically aligned local medical dataset, and a background knowledge set simulating attacker auxiliary information is constructed. Based on the known values ​​of the background knowledge set, a member inference attack executor is constructed. The member inference attack executor analyzes the association features between the target record and the original dataset, identifies and determines the dataset affiliation relationship of each record in the target record set, and obtains a member determination result with multi-dimensional verification.

[0088] Based on the member determination results of the multi-dimensional verification, member records with deterministic confidence are obtained. An attribute inference attack executor is constructed by combining the background knowledge set. The attribute inference attack executor is used to mine the implicit correlation patterns between features in the confidence member records. Undisclosed sensitive feature values ​​are inferred from multiple angles to form a systematic attribute inference result.

[0089] A multi-layered risk assessment system is constructed using the member inference attack executor and the attribute inference attack executor. Based on the risk assessment system, attack tests of different intensities and strategies are performed on the semantically aligned local medical dataset to explore the data vulnerability boundary and obtain the attack success rate under each dimension. Based on the attack success rate, the propagation characteristics and impact of privacy risks in each dimension are analyzed and evaluated, and a multi-dimensional risk aggregation weight that can dynamically reflect the importance of the dimensions is generated.

[0090] The target record set extraction process employs stratified random sampling to select representative samples from the semantically aligned local medical dataset. The sampling ratio is set to 20% of the total dataset to ensure sufficient representativeness of records from each medical department and disease category. The sampling algorithm maintains the distribution ratio of each category in the original dataset, calculating the expected sample size for each category and using simple random sampling for record selection. The target record set contains complete medical records, including basic patient information, diagnostic results, treatment plans, and examination data. Each record has a unique identifier for subsequent association analysis.

[0091] The background knowledge set construction module integrates multi-source medical auxiliary information, including external knowledge sources such as publicly available medical statistical yearbooks, disease epidemiological data, drug usage guidelines, and medical device specifications. The knowledge acquisition interface supports structured database queries and unstructured text parsing. The data cleaning process removes duplicate information and standardizes medical terminology coding. The background knowledge set adopts a hierarchical storage architecture: frequently accessed knowledge is stored in a memory cache, providing millisecond-level access response, while less frequently accessed knowledge is stored in a disk database and indexed to support second-level retrieval. The knowledge update mechanism periodically synchronizes with changes in external data sources, with an update cycle set to 7 days, supporting both incremental updates and full refresh modes.

[0092] The attack executor employs a shadow model training architecture, constructing an auxiliary dataset with a feature distribution similar to the target dataset for attack model training. The shadow dataset is set to 80% the size of the target dataset, and simulated medical records with similar statistical characteristics are generated using data synthesis techniques. The attack executor comprises three core components: a feature extraction module, a similarity calculation module, and a decision logic module. The feature extraction module extracts key attribute combinations from the target records. Similarity calculation uses a weighted Euclidean distance measure to measure the closeness between records, with a distance threshold set to 5% of the diagonal length of the feature space.

[0093] The association feature analysis process establishes a multi-dimensional matching relationship between the target records and the records in the original dataset. The matching dimensions include numerical feature similarity, categorical feature overlap, and time series feature correlation. Numerical feature similarity is calculated using standardized absolute differences; feature pairs with differences less than 0.1 are considered highly similar. Categorical feature overlap is quantified using the Jacardi similarity coefficient; record pairs with a coefficient greater than 0.8 are considered strongly correlated. Time series features are correlated using the dynamic time warping algorithm; sequences with a correlation coefficient greater than 0.7 are considered correlated.

[0094] Membership determination results are generated using an ensemble learning approach that fuses the predictions from multiple deciders. The base deciders include one support vector machine, one random forest, and one neural network classifier. The output probabilities of each decider are fused through a weighted average, with weights determined based on each decider's performance on the validation set; the decider with the highest accuracy receives the largest weight. Records with a fused probability value greater than 0.8 are determined to be members of the dataset, records with a probability value less than 0.2 are determined to be non-members, and records in the intermediate range are marked as uncertain and require further validation.

[0095] A multi-dimensional verification mechanism cross-validates the preliminary judgment results. Verification dimensions include statistical consistency testing, outlier detection, and domain knowledge verification. Statistical consistency testing compares the feature distribution of the target record with the expected distribution of the original dataset. Distribution differences are quantified using the Kolmogorov test; records with test statistics below the critical value pass consistency verification. Outlier detection uses the Isolation Forest algorithm to identify records with significant deviations from the normal pattern. Records with anomaly scores greater than 0.6 are marked as suspicious members requiring further review.

[0096] The confidence level member record selection is based on a comprehensive confidence threshold set according to multi-dimensional verification results. Only records that pass all verification dimensions and have a fusion probability exceeding 0.85 are confirmed as confidence level member records. The confidence level calculation uses a Bayesian update method, updating the prior probability distribution with the results of each verification dimension as evidence. The final confidence level equals the expected value of the posterior probability. The selection process employs a step-by-step filtering strategy, eliminating candidate records that do not meet the criteria at each verification stage, reducing the computational overhead of subsequent processing.

[0097] The attribute inference attack executor trains an attribute prediction model based on confidence member records and a background knowledge set. The target is to predict sensitive medical attributes of patients, such as susceptibility to genetic diseases, mental health status, and drug allergies. The attack executor employs a deep neural network architecture with three hidden layers, containing 256, 128, and 64 neurons respectively. The activation function is a modified linear unit function, and the output layer uses a multi-label classification design to support the simultaneous prediction of multiple sensitive attributes.

[0098] The latent association pattern mining employs an association rule learning algorithm to discover dependencies between features, with a minimum support of 5%, a minimum confidence of 70%, and a maximum rule length limited to 5 feature items. The association rule extraction process uses an improved version of the prior algorithm, employing a frequent itemset pruning strategy to reduce the search space and improve pattern discovery efficiency. The mined association patterns are sorted according to their lift; patterns with a lift greater than 1.5 are considered to have practical inference value.

[0099] A multi-perspective inference strategy combines statistical inference, machine learning inference, and knowledge graph inference to predict sensitive feature values. Statistical inference calculates attribute values ​​based on the conditional probability distribution of features; machine learning inference generates attribute estimates using a trained prediction model; and knowledge graph inference uses logical reasoning based on entity relationships within a medical knowledge graph. The results of the three inferences are fused through a voting mechanism, and predictions with a consensus of more than 2 / 3 are adopted as the final inference value.

[0100] The attribute inference results are systematically organized using a hierarchical data structure, categorized and stored according to medical department, disease type, and sensitivity. The inference results include fields such as predicted attribute value, confidence interval, inference basis, and validation status. The confidence interval is estimated using the bootstrap method, and the resampling frequency is set to 1000 times. The inference basis records the key feature combinations and association rules that led to the prediction, facilitating subsequent auditing and interpretation.

[0101] The multi-layered risk assessment system comprises three assessment layers: a data layer, a feature layer, and a record layer. Each layer employs different risk metrics and assessment methods. The data layer assesses the privacy leakage risk of the overall dataset by quantifying it through the proportion of successful attack records and the average attack confidence. The feature layer assesses the privacy sensitivity of individual features, calculating a sensitivity score based on the feature's contribution to successful attacks and its frequency of occurrence. The record layer assesses the privacy exposure risk of individual records, comprehensively considering the uniqueness, identifiability, and number of sensitive attributes of each record.

[0102] The attack testing process incorporates various strength levels and strategy combinations. The strength levels are categorized into lightweight, medium, and heavyweight, each corresponding to different computational resource inputs and time constraints. Lightweight attacks have a 10-minute computation time limit and use a single attack algorithm; medium-weight attacks have a 1-hour time limit and use multiple algorithms executed in parallel; heavyweight attacks have no time limit and utilize all available attack methods and optimization strategies. Attack strategies include direct attacks, indirect attacks, and combined attacks. Direct attacks perform direct inference based on target attributes, indirect attacks perform indirect inference based on related attributes, and combined attacks integrate multiple inference paths.

[0103] The data vulnerability boundary exploration employs an adaptive search algorithm to gradually increase the attack intensity until a preset success rate threshold is reached. The success rate threshold is set at 90%, representing a critical point where the attack is almost certain to succeed. The boundary search process records the success rate variation curves under different attack parameter configurations, identifying sensitive parameters and key inflection points. The vulnerability assessment results include information such as the critical attack intensity, the most vulnerable feature combination, and the most effective attack strategy.

[0104] The attack success rate statistics for each dimension employ stratified sampling to ensure statistical significance. At least 100 independent attack experiments are conducted for each dimension, and the experimental results are validated for reliability through confidence interval estimation and hypothesis testing. Success rate calculations consider both partial success and complete success. Partial success refers to correctly inferring some values ​​of the target attribute, while complete success refers to a completely accurate attribute prediction.

[0105] The dynamic dimensional importance reflection mechanism adjusts the weight allocation of each dimension based on the changing trend of attack success rate. An exponentially weighted moving average method is used to smooth the weight update process, with a smoothing factor set to 0.3 to balance historical information and current changes. Weight updates are triggered when the attack success rate changes by more than 5% or when a new attack test round ends. Weight constraints ensure that the sum of all dimension weights equals 1 and that the weight of any single dimension does not exceed 0.5 to avoid excessive bias.

[0106] In the implementation case, 2,400 target records were extracted from a dataset of 12,000 diabetic patients to construct a background knowledge set containing 1.5 million medical statistics. The member inference attack executor successfully identified 78% of the target records under lightweight attacks, improved the success rate to 85% under medium-level attacks, and achieved a 92% accuracy rate under heavyweight attacks. The attribute inference attack achieved inference accuracy of 82%, 76%, and 89% for the three sensitive attributes of insulin dependence, complication type, and glycemic control level, respectively. Multidimensional risk aggregation weight calculation results showed that the uniqueness dimension weight was 0.32, the correlation dimension weight was 0.28, the sensitivity dimension weight was 0.25, and the inference dimension weight was 0.15. Vulnerability boundary exploration revealed that the dataset privacy protection capability significantly decreased when attack resources exceeded 3.2 times the standard configuration.

[0107] In one optional implementation, based on the hybrid encoded data structure, local feature representations of each data holding institution are extracted through the contrastive learning mechanism, and the local feature representations are encrypted using a homomorphic encryption mechanism. The fusion of multiple encrypted local feature representations to generate a global feature representation includes:

[0108] Based on the aforementioned hybrid coding data structure, a contrastive learning module is designed. The contrastive learning module analyzes the multi-level semantic associations and dynamic structural features in the hybrid coding data structures of various data holding institutions, and establishes an adaptive cross-institutional data similarity measurement system.

[0109] Based on the cross-institutional data similarity measurement system, local feature representations are generated for the hybrid coded data structures of each data holding institution; the local feature representations are enhanced with security by a homomorphic encryption module to generate encrypted feature representations that are tamper-proof and maintain the original structure and computational properties; the encrypted feature representations are aggregated to a central server according to the data holding institutions, and the central server forms aggregated encrypted features in an adaptive weighted manner within the encryption domain, and after decryption, obtains a global feature representation that integrates differentiated knowledge from multiple parties.

[0110] like Figure 2 As shown, the method includes:

[0111] The contrastive learning module is designed based on a hybrid coding data structure, constructing a dual-branch coding network architecture. The main branch handles structured data encoding, while the auxiliary branch handles unstructured data encoding. The two branches integrate features through an attention fusion layer. The main branch uses a graph neural network to handle entity relationship encoding, with a network depth of 4 layers, each containing 128 hidden units, and gated linear units as the activation function. The auxiliary branch uses a transformer network to handle sequence encoding, with 8 heads, a feedforward network dimension of 512, and 6 layers. The attention fusion layer calculates the association weights between structured and unstructured features through a cross-attention mechanism, and the weights are normalized using a soft maximization function with a temperature parameter of 0.07.

[0112] The multi-level semantic association analysis module parses the hierarchical semantic information in the hybrid encoded data structure, including three levels: lexical association, phrase association, and document association. Lexical association is calculated using the cosine similarity of word embedding vectors; word pairs with a similarity threshold of 0.6 or higher are considered to have strong semantic association. Phrase association employs a position-encoded enhanced self-attention mechanism to identify dependencies between semantic segments; segment pairs with attention weights greater than 0.1 are included in the association graph. Document association uses the Euclidean distance between document vectors; documents with a distance less than twice the standard deviation are classified into similar semantic groups.

[0113] The dynamic structural feature capture module monitors the temporal change patterns of the hybrid coded data structure, employing a sliding window mechanism to track structural evolution. The window size is set to 50 time steps, the sliding step size to 10 time steps, and the overlap rate to 80% ensures feature continuity. Structural change detection is achieved by comparing the statistical distance between the coded distributions within adjacent windows, using Wasserstein distance as a measure of distribution difference; changes with a distance exceeding 0.05 are identified as significant structural evolution events. Feature extraction utilizes a recurrent neural network to capture temporal dependencies; the network contains two layers of long short-term memory units, each with 128 hidden states.

[0114] A cross-institutional data similarity measurement system establishes a multi-dimensional similarity evaluation framework, including content similarity, structural similarity, and semantic similarity. Content similarity is calculated using the Pearson correlation coefficient between feature vectors; data pairs with a correlation coefficient greater than 0.7 are considered highly similar. Structural similarity uses graph edit distance to measure the topological differences between hybrid encoded data structures. The edit distance includes the weighted costs of node insertion, deletion, and edge modification operations, with weights set to 1, 1, and 0.5, respectively. Semantic similarity is calculated using context vectors generated by a pre-trained language model. A bidirectional encoder with a transformer architecture is employed, with 1.1 million model parameters, fine-tuned on a domain corpus.

[0115] The adaptive metric strategy dynamically adjusts the similarity weight allocation based on the characteristics of the data holder. Weight updates employ gradient descent optimization with a learning rate of 0.001 and a momentum parameter of 0.9. Weights are initialized using a uniform distribution, with each dimension's initial weight value being 0.33. Weight constraints ensure the sum equals 1 and that individual dimension weights range from 0.1 to 0.6. Adaptability evaluation assesses the feature representation quality under different weight configurations through cross-validation, selecting the weight combination with the minimum validation loss as the optimal configuration.

[0116] The local feature representation generation process maps the hybrid encoded data structure of each data holder to a fixed-dimensional feature space, with the feature dimension set to 256 dimensions, using a linear projection layer for dimensionality transformation. Feature normalization employs layer normalization techniques, with normalization parameters including learnable scaling factors and offsets, initialized using a standard normal distribution. Feature quality is evaluated using two metrics: reconstruction error and discriminant metric. The reconstruction error is calculated using an autoencoder architecture, while the discriminant metric is evaluated using the accuracy of a classification task. Feature representation stability is verified through robustness testing with added Gaussian noise, where the noise variance is set to 0.01. Stable feature representations exhibit changes of less than 5% under noise interference.

[0117] The homomorphic encryption module employs a fully homomorphic encryption scheme to secure local feature representations. This scheme is based on the security assumption of learning with error, supporting an unlimited number of addition and multiplication operations. The key generation process comprises three parts: a public key, a private key, and an evaluation key. A 128-bit security parameter provides sufficient security margin. The public key is 32KB, the private key is 16KB, and the evaluation key is 128MB. Key storage utilizes a fragmented storage mechanism to mitigate storage risks. The encryption process converts the 256-dimensional floating-point feature vector into integer encoding with a precision of 6 decimal places, ranging from -2^20 to +2^20.

[0118] The security enhancement process includes two layers of protection mechanisms: noise addition and ciphertext perturbation. Noise addition uses a Laplace distribution to generate perturbation values, and the noise standard deviation is calibrated based on the differential privacy parameter epsilon being equal to 1. Ciphertext perturbation is achieved by randomizing the random number seed used in the encryption process; each encryption uses a different random seed to ensure that the same plaintext produces different ciphertexts. The tamper-proof mechanism verifies the integrity of the ciphertext through a message authentication code. The authentication code is 256 bits long and is calculated using a secure hash algorithm.

[0119] The encrypted feature representation maintains the original computational properties through homomorphic operation correctness verification. Addition precision error is controlled within one-thousandth, and multiplication error within one-hundredth. Computational complexity optimization employs batch processing technology to process multiple feature vectors in parallel, with a batch size of 32 vectors and a single batch processing time controlled within 10 seconds. Memory usage optimization is achieved through feature block processing, with each block containing 64 feature dimensions. Pipeline processing is used between blocks to reduce peak memory usage.

[0120] The central server aggregation module receives encrypted feature representations uploaded by various data holders. It supports concurrent uploads using an asynchronous communication protocol, with a maximum of 100 concurrent connections. Data transmission is encrypted using a transport layer security protocol, and certificate verification employs a two-way authentication mechanism to ensure communication security. The uploaded data format includes four fields: organization identifier, timestamp, feature dimension, and encrypted content. The organization identifier is 16 bytes long, the timestamp is a 64-bit integer, the feature dimension is a 4-byte integer, and the encrypted content length is variable.

[0121] The aggregated encryption feature generation employs an adaptive weighted fusion algorithm, with weight calculation based on the data quality and contribution assessment of each institution. Data quality assessment includes three indicators: completeness, consistency, and timeliness. Completeness is calculated using the proportion of missing values, consistency is assessed through data distribution similarity, and timeliness is measured by data update frequency. Contribution assessment is quantified through feature uniqueness and information gain. Uniqueness is measured by the rarity of the feature vector in the overall feature space, and information gain is calculated using the reduction in cross-entropy. Weight normalization ensures that the sum of the weights of all institutions equals 1, and the lower limit of the weight for a single institution is 0.05 to prevent some institutions from being completely ignored.

[0122] The aggregation operation within the encrypted domain utilizes the linear property of homomorphic encryption to achieve weighted average calculation, without requiring decryption of the original feature representations of each institution. Weight encoding uses fixed-point numbers with a precision of 4 decimal places, and the weight range is limited to 0 to 1. The aggregation calculation employs a tree-like reduction structure to reduce communication rounds, with the reduction depth being the logarithm of the number of institutions, and the single-round reduction time controlled within 30 seconds. Correctness verification is achieved through a zero-knowledge proof mechanism, with proof generation time controlled within 5 minutes and verification time controlled within 30 seconds.

[0123] The decryption process employs a threshold decryption mechanism to enhance security, requiring the consent of more than half of the authorized parties before decryption can be performed. Private key sharding uses the Shamir secret sharing scheme, with the threshold parameter set to two out of three shards, which are stored in different secure hardware modules. Decryption time optimization is achieved through pre-computation techniques; commonly used decryption parameters are pre-calculated and cached, keeping decryption latency within one minute. Decryption result verification involves checking the statistical properties of the original features, with the relative errors of the mean and variance controlled within 5%.

[0124] Global feature representation integrates differentiated knowledge from multiple sources through a knowledge distillation mechanism. The teacher model is a collection of local models from various institutions, while the student model is a globally unified model. The distillation temperature parameter is set to 4 to balance the accuracy and generalization of knowledge transfer. Knowledge weight allocation is determined based on the professional field and data coverage of each institution, with medical institutions receiving higher weights for health-related features and financial institutions receiving advantageous weights for risk assessment features. Feature dimension alignment is achieved through linear transformation, and the transformation matrix is ​​optimized using the least squares method, with the alignment error controlled within a mean absolute error of 0.1.

[0125] In the implementation case, three data holding institutions provided hybrid coded data structures from the medical, financial, and educational sectors, with data sizes of 100,000, 150,000, and 80,000 records, respectively. The local feature representation extracted by the contrastive learning module has a dimension of 256. The features from medical institutions mainly reflect health status patterns, the features from financial institutions reflect credit risk characteristics, and the features from educational institutions include learning ability assessments. The ciphertext size after homomorphic encryption is 32KB per feature vector, with an encryption time of 200 milliseconds per vector. After the central server aggregates the encrypted features from the three parties, the adaptive weights are 0.35, 0.42, and 0.23, respectively, and the aggregation calculation takes 45 seconds. The global feature representation obtained after decryption achieves an accuracy of 87.6% on the downstream classification task, an improvement of 12.3% compared to single-institution features, verifying the effectiveness of multi-party knowledge fusion.

[0126] In one optional implementation, the generative model is trained using the global feature representation, and during training, differentiated privacy budgets are allocated to different sensitivity features based on the feature-level privacy sensitivity vector, including:

[0127] The training mechanism for constructing a generative model is constructed using the global feature representation. The global feature representation is deconstructed in multiple dimensions through the training mechanism to establish a dimensional index system that reflects the essential attributes of the features.

[0128] Based on the dimensional indexing system, the feature-level privacy sensitivity vector is associated, the sensitivity of each dimension feature is quantitatively evaluated and a sensitivity mapping relationship is formed. A privacy budget allocation function is designed according to the sensitivity mapping relationship. An adaptive mapping mechanism between the sensitivity quantification value and the budget amount is constructed according to the privacy budget allocation function. The budget is allocated differently for each dimension feature through the adaptive mapping mechanism to form a feature-level privacy budget allocation scheme.

[0129] During the parameter optimization process of the generative model, the gradient components of each dimension feature are dynamically calibrated according to the feature-level privacy budget allocation scheme to generate a gradient update direction with privacy protection effect; the network parameters of the generative model are adjusted based on the gradient update direction until the model converges and the training process is completed.

[0130] The training mechanism for the generative model constructed from global feature representations is implemented through a multi-layer neural network architecture, comprising three core components: an encoder, a decoder, and a feature decomposition module. The encoder receives a 512-dimensional global feature representation vector as input and performs a non-linear transformation through a three-layer fully connected network. Each layer has 256, 128, and 64 neurons, respectively, using the ReLU activation function and a dropout rate of 0.1 to prevent overfitting. The decoder employs a symmetric structure, reconstructing the original feature representation layer by layer to ensure information fidelity during training. The feature decomposition module is responsible for multi-dimensional deconstructing the global feature representation. Principal component analysis (PCA) is used to decompose the 512-dimensional feature space into eight principal dimensions, each containing 64 feature components. The decomposition process retains over 95% of the original information variance.

[0131] The dimensional indexing system is built upon feature decomposition results to construct a hierarchical index structure. Hash tables are used to store the identifiers, numerical ranges, statistical attributes, and semantic labels of each dimension's features. The indexing system comprises two levels: dimension-level indexes and feature-level indexes. The dimension-level indexes record basic information for the eight main dimensions, while the feature-level indexes refine the details to the specific attributes of each 64-dimensional feature component. Each dimension is assigned a unique 8-bit identifier, and feature components are encoded using a combination of the dimension identifier and a 6-bit ordinal number. The indexing system supports fast retrieval and batch update operations, with retrieval time complexity controlled at the logarithmic level, and update operations using an incremental approach to reduce computational overhead. The dimensional indexing system also includes a quantitative description of the essential attributes of the features, reflecting their distribution characteristics and importance by calculating statistical indicators such as feature variance, skewness, and kurtosis.

[0132] The association between feature-level privacy sensitivity vectors and the dimension indexing system is achieved through a sensitivity mapping table, which establishes a correspondence between dimension identifiers and sensitivity values. The sensitivity vector is an 8-dimensional real-number vector, with each component ranging from 0 to 1, representing the privacy sensitivity of the corresponding dimension; a larger value indicates higher sensitivity. The mapping table uses a key-value pair structure for storage, where the key is the dimension identifier and the value is a composite data structure containing the sensitivity value, confidence level, and update timestamp. The sensitivity quantification and evaluation process is achieved by analyzing the information entropy, mutual information, and conditional entropy indices of the feature dimensions. Information entropy reflects the degree of uncertainty of the features, mutual information measures the dependencies between features, and conditional entropy assesses the independence level of the features. The quantified evaluation results form a sensitivity mapping relationship, which is dynamically updated to adapt to changes in data distribution and adjustments in privacy requirements.

[0133] The privacy budget allocation function is designed based on a sensitivity mapping relationship, employing a piecewise linear function model to achieve the mapping transformation between sensitivity and budget amount. The function's domain is the sensitivity value range [0,1], and its range is the budget amount range [0.01,1.0]. The segmentation points are set at sensitivity values ​​of 0.3, 0.6, and 0.9, with slopes of 0.5, 1.0, and 2.0 respectively, ensuring that high-sensitivity features receive more privacy budget protection. The allocation function also incorporates smoothing processing, using cubic spline interpolation near the segmentation points to avoid allocation instability caused by function discontinuity. The adaptive mapping mechanism dynamically adjusts by monitoring changes in feature distribution and the accumulation of privacy loss. When the privacy loss of a certain dimension exceeds a preset threshold of 0.8, the budget allocation weight for that dimension is automatically increased by 1.2 times the original value. The mapping mechanism also includes a budget surplus management function. The total budget is set to 10.0, and the sum of the budgets allocated to each dimension does not exceed this limit. When the budget is insufficient, it is redistributed according to sensitivity priority.

[0134] The differentiated budget allocation process is achieved by iterating through eight feature dimensions and calculating the budget amount for each dimension sequentially. Dimension sensitivity values ​​are obtained from the sensitivity mapping relationship. An initial budget amount is calculated by calling the privacy budget allocation function, and then dynamically adjusted through an adaptive mapping mechanism to form a feature-level privacy budget allocation scheme. The allocation scheme is stored in a structured data format, including fields such as dimension identifier, sensitivity value, budget amount, allocation time, and validity period. Budget allocation supports both batch operations and incremental updates. Batch operations are suitable for the model initialization phase, while incremental updates are suitable for dynamic adjustments during training. The allocation scheme also includes a budget usage monitoring function, tracking the budget consumption of each dimension in real time, and triggering a budget replenishment mechanism when the usage rate exceeds 90%.

[0135] The dynamic calibration of gradient components during the generative model parameter optimization process is achieved through a differential privacy noise injection mechanism. The calibration process first calculates the gradient components corresponding to each dimension of features, and then determines the noise injection intensity based on the feature-level privacy budget allocation scheme. The noise intensity is inversely proportional to the budget amount; the larger the budget amount, the smaller the noise intensity and the stronger the protection effect. The noise is generated using a Laplace distribution, and the distribution parameters are calculated using the budget amount and global sensitivity. The global sensitivity is set to 0.1 to ensure moderate noise intensity. The gradient calibration process also includes gradient clipping, limiting the gradient magnitude to within 2.0 to prevent gradient explosion from affecting training stability. The calibrated gradient components retain their original directional characteristics, but their values ​​are adjusted according to privacy protection requirements, with gradient components in high-sensitivity dimensions subject to more noise interference.

[0136] The gradient update direction generation process is achieved by aggregating the calibrated gradient components of each dimension. The aggregation uses a weighted average method, with weights determined based on dimension importance and budget allocation. Importance weights are calculated using feature variance and information gain metrics, while budget weights are positively correlated with budget limits, ensuring that dimensions with sufficient budget protection play a greater role in the update. The aggregated gradient vector is normalized, with its magnitude controlled within 1.0 to guarantee consistency in the update step size. The gradient update direction also undergoes feasibility verification to ensure that the updated parameters meet model constraints, including parameter range, positive definiteness constraints, and sparsity requirements. The update direction calculation process supports parallel processing; gradient calibration operations for each dimension are independent and can be executed concurrently on multi-core processors to improve computational efficiency.

[0137] The parameter tuning process for the generative model network employs the Adam optimization algorithm, with a learning rate set to 0.001, momentum parameters beta1 and beta2 set to 0.9 and 0.999 respectively, and the epsilon parameter set to 1e-8 to prevent division by zero errors. Parameter updates utilize a batch update method with a batch size of 32 to ensure a balance between gradient estimation stability and computational efficiency. The parameter tuning process includes a learning rate decay mechanism, multiplying the learning rate by 0.95 every 10 training epochs to prevent late-stage training oscillations from affecting convergence. The network parameters also employ weight decay regularization with a decay coefficient set to 0.0001 to suppress overfitting and improve generalization ability. During parameter updates, the loss function is monitored in real-time; an early stopping mechanism is triggered when the loss function decreases by less than 0.001 for five consecutive epochs to avoid overtraining and wasting computational resources.

[0138] The model convergence criterion is based on two indicators: loss function stability and parameter variation magnitude. Loss function stability is evaluated by calculating the variance of the loss function over 10 consecutive training epochs; a variance less than 0.0001 indicates the loss function is stable. Parameter variation magnitude is evaluated by calculating the rate of change of the L2 norm of the parameter vector; a rate of change less than 0.001 indicates convergence. The convergence criterion also considers privacy budget consumption; training is forcibly stopped when the total budget utilization exceeds 95%, ensuring that privacy protection is not affected by training duration. The model training process employs a checkpoint saving mechanism, saving the model state every 50 training epochs, including network parameters, optimizer state, privacy budget usage, and other information, supporting recovery after training interruption.

[0139] In the specific data example, the input global features are represented as a 512-dimensional vector with values ​​ranging from -1 to 1. After decomposition into 8 dimensions, each dimension contains 64 feature components. The sensitivity vector is set to [0.2, 0.5, 0.8, 0.3, 0.7, 0.4, 0.9, 0.6], corresponding to the sensitivity levels of the 8 feature dimensions. The budget allocation scheme calculated according to the privacy budget allocation function is [0.15, 0.55, 0.89, 0.25, 0.79, 0.35, 0.95, 0.65], with a total budget consumption of 4.58. The training process converges after 200 epochs, with the final loss function value stabilizing at 0.045, the parameter change rate decreasing to 0.0008, and the total privacy budget utilization rate at 92%, meeting the convergence criteria and completing the training process.

[0140] In one optional implementation, a privacy budget allocation function is designed based on the sensitivity mapping relationship, and an adaptive mapping mechanism between sensitivity metrics and budget amounts is constructed based on the privacy budget allocation function, including:

[0141] Based on the sensitivity mapping relationship, the sensitivity metrics are deconstructed in multiple dimensions to construct a feature characterization system that represents the sensitivity metrics. Based on the feature characterization system, the function form of the privacy budget allocation function is designed. Based on the function form of the privacy budget allocation function, a mapping calculation unit is constructed. The sensitivity metrics of each feature dimension in the sensitivity mapping relationship are evaluated through the mapping calculation unit, and differential transformation operations are performed to generate the initial budget amount corresponding to each feature dimension.

[0142] A budget calibration mechanism is constructed based on the initial budget amount. The hierarchical relationship and competitive game characteristics between budget amounts in each dimension are explored according to the budget calibration mechanism to obtain a calibrated budget amount that considers the constraint coupling between dimensions.

[0143] A bidirectional mapping architecture is constructed between the calibration budget and the sensitivity quantification. The structural deviation between the calibration budget and the sensitivity quantification is measured through the bidirectional mapping architecture, and an adjustment signal is constructed. The function shape of the privacy budget allocation function is dynamically optimized based on the adjustment signal, thus forming an adaptive mapping mechanism.

[0144] The multi-dimensional semantic deconstruction of sensitivity quantifications is achieved by constructing a hierarchical feature decomposition module. This module receives the quantized values ​​in the sensitivity mapping relationship as input and decomposes each sensitivity quantification into three basic dimensions: semantic strength, distribution characteristics, and correlation degree. The semantic strength dimension reflects the absolute magnitude of the sensitivity quantification, achieved by mapping the quantization value to a standardized range of 0 to 1 using a linear transformation. The transformation parameters are dynamically determined based on the maximum and minimum values ​​of the quantization value. The distribution characteristics dimension describes the distribution pattern of the sensitivity quantification in the feature space, obtained by calculating the skewness and kurtosis indices of the quantization value. Skewness reflects the symmetry of the distribution, and kurtosis reflects the sharpness of the distribution. The correlation degree dimension measures the dependency of the sensitivity quantification value on other dimensions of quantization, achieved by calculating the Pearson correlation coefficient and mutual information index. The correlation coefficient threshold is set to 0.7, and the mutual information threshold is set to 0.5; values ​​exceeding the threshold indicate a strong correlation.

[0145] The feature characterization system constructs a three-layer tree structure based on semantic deconstruction results. The top layer contains the sensitivity quantification value itself, the middle layer contains three basic dimensions, and the bottom layer contains the specific indicators for each dimension. Each node contains three attributes: value, weight, and update timestamp. The value records the specific quantification result, the weight reflects the importance of the node in the overall characterization, and the timestamp is used for version management and consistency control. Weight calculation is implemented using the entropy weighting method, which determines the weight allocation by calculating the information entropy of each indicator. The smaller the information entropy, the greater the weight, indicating a stronger discriminative ability of the indicator. The feature characterization system supports a dynamic update mechanism. When the input sensitivity quantification value changes, the value and weight of the relevant node are automatically recalculated at the single-node level, avoiding the performance overhead of global recalculation. The characterization system also includes an anomaly detection function. When the node value exceeds a preset range or the weight distribution is severely uneven, an anomaly alarm is triggered. The anomaly threshold is set to a value deviation of 3 times the standard deviation of the mean, and the weight unevenness threshold is set to a Gini coefficient exceeding 0.8.

[0146] The privacy budget allocation function is designed based on the structural features of the feature characterization system, employing a piecewise fractional function model to express the mapping relationship between sensitivity quantification values ​​and budget amounts. The function consists of three parts: a linear segment, an exponential segment, and a logarithmic segment. The linear segment is suitable for sensitivity quantification values ​​between 0 and 0.3, where the budget amount is directly proportional to the quantification value, with a proportionality coefficient set to 0.5. The exponential segment is suitable for sensitivity quantification values ​​between 0.3 and 0.7, where the budget amount grows exponentially, with a base of 2 and an exponent equal to the quantification value minus 0.3, ensuring moderate budget protection for moderate sensitivity. The logarithmic segment is suitable for sensitivity quantification values ​​between 0.7 and 1, where the budget amount grows logarithmically, with a base of 10 and an argument equal to 10 times the quantification value minus 0.6, ensuring sufficient budget allocation for high sensitivity. Cubic spline interpolation is used at the segmentation points to ensure continuity and differentiability. The interpolation parameters are determined by least squares fitting, with the fitting error controlled within 0.01.

[0147] The mapping calculation unit is implemented using a pipelined architecture, comprising a preprocessor, a function calculator, and a post-processor. The preprocessor handles input data format conversion and boundary checks, ensuring sensitive quantized values ​​remain within valid ranges and correcting invalid values ​​using nearest-neighbor interpolation. The function calculator calls the corresponding function form based on the range of the quantized value, supporting vectorized operations and processing up to 64 quantized values ​​at a time to improve computational efficiency. The post-processor normalizes and rounds the calculation results. Normalization ensures the total budget amount does not exceed the budget limit, and rounding precision is set to four decimal places to avoid the accumulation of floating-point operation errors. The mapping calculation unit also includes a caching mechanism to cache frequently queried quantized values ​​and their corresponding budget amounts. The cache capacity is set to 1000 entries, and a least recently used algorithm is used for cache replacement, typically achieving a cache hit rate of over 80%.

[0148] Differentiated transformation operations are implemented through a multi-path computation framework, employing personalized transformation strategies for the sensitivity metrics of different feature dimensions. The transformation operation comprises three steps: basic transformation, weight adjustment, and constraint satisfaction. Basic transformation calls the mapping calculation unit 'a' to obtain the initial budget amount. Weight adjustment corrects the initial amount based on the importance of the feature dimensions. Constraint satisfaction ensures the transformation result meets the total budget limit and fairness requirements. Weight adjustment uses a multiplicative correction method; importance weights are calculated using principal component analysis, retaining principal components with a cumulative variance contribution rate of 95%. The weight of each dimension is the proportion of the corresponding principal component's eigenvalue to the total eigenvalues. Constraint satisfaction is achieved using the Lagrange multiplier method, transforming the total budget constraint and inter-dimensional fairness constraint into an optimization problem. Iterative solutions obtain a budget allocation scheme that satisfies the constraints. The iteration terminates when the objective function value changes by less than 0.001 or the number of iterations exceeds 100.

[0149] The initial budget generation process is implemented using a parallel computing architecture. Budget calculations for each feature dimension are independent and can be executed concurrently in a multi-threaded environment. The number of computation threads is set to 75% of the processor cores to avoid performance impact from thread switching overhead. Each thread handles budget calculation tasks for one or more feature dimensions. Threads exchange calculation results through a shared memory pool, which uses a lock-free data structure to avoid thread synchronization overhead. The initial budget data structure includes five fields: dimension identifier, sensitivity quantifier, budget amount, calculation timestamp, and validity period. The dimension identifier is represented by a 64-bit integer, the sensitivity quantifier and budget amount are represented by double-precision floating-point numbers, the timestamp uses a millisecond-level UNIX timestamp format, and the validity period is set to 24 hours to ensure the timeliness of budget allocation.

[0150] The budget calibration mechanism is built upon a game theory model, treating budget amounts across dimensions as game participants and achieving balanced budget allocation by analyzing competition and cooperation among dimensions. The game model employs a Nash equilibrium solution algorithm. Each dimension's strategy space represents its allocable budget range. The payoff function comprehensively considers three factors: budget gain, sensitivity satisfaction, and inter-dimensional fairness. The payoff function uses a weighted summation form, with a weight of 0.4 for budget gain, 0.4 for sensitivity satisfaction, and 0.2 for inter-dimensional fairness, summing to 1 to ensure the payoff function's normalization. The game-solving process uses an iterative optimal response algorithm. In each iteration, each dimension sequentially updates its strategy selection, choosing the strategy that maximizes its own payoff. The iteration continues until all dimensions' strategies no longer change or the maximum number of iterations (200) is reached.

[0151] Hierarchical association analysis is implemented by constructing a dimensional dependency graph. Nodes in the dependency graph represent feature dimensions, edges represent dependencies between dimensions, and edge weights reflect the strength of the dependency. Dependencies are determined by calculating the conditional probability and causal strength between dimensions. The conditional probability threshold is set to 0.6, and the causal strength threshold is set to 0.4. Dimension pairs exceeding both thresholds are considered dependent edges. The dependency graph is stored using an adjacency matrix, where the matrix dimension is the square of the number of feature dimensions, and the matrix elements are dependency strength values. Diagonal elements set to 1 indicate a complete dependency between a dimension and itself. Hierarchical association analysis also includes loop detection. A depth-first search algorithm is used to identify strongly connected components in the dependency graph. When loops are detected, they are broken by sorting the edges by descending weights, ensuring the directed acyclic property of the dependencies.

[0152] The exploration of competitive game characteristics was achieved by simulating different budget allocation scenarios. The scenario generation employed the Monte Carlo method, randomly generating 1000 budget allocation schemes, each corresponding to a competitive state between dimensions. The competitive state was quantified by calculating the intensity of resource contention and the ratio of cooperative benefits between dimensions. Resource contention intensity was the ratio of a dimension's budget demand to the total budget, and the cooperative benefit ratio was the ratio of the additional benefits from inter-dimensional cooperation to the benefits of individual action. The game characteristic analysis results were used to guide the formulation of budget calibration strategies: a fair allocation strategy was adopted when the competition intensity was higher than 0.8, a cooperative incentive strategy was adopted when the cooperative benefit ratio was higher than 1.5, and an efficiency-first strategy was adopted in other cases. The game characteristics also included dynamic evolution analysis, predicting the changing trends of competitive relationships between dimensions using time series analysis methods. The prediction window was set to 10 time steps, with a prediction accuracy requirement of over 85%.

[0153] The calibration budget calculation process comprehensively considers three factors: the initial budget amount, hierarchical relationships, and the outcome of the competitive game. A multi-objective optimization method is employed to find the optimal budget allocation scheme. The optimization objectives include three sub-objectives: maximizing budget utilization, maximizing sensitivity satisfaction, and maximizing inter-dimensional fairness. Each sub-objective undergoes normalization to ensure dimensional consistency. The multi-objective optimization uses the Pareto front algorithm to generate multiple non-dominated solutions, forming a Pareto optimal solution set. The size of the solution set is limited to 50 to avoid excessive computational complexity. The final calibration budget amount is selected from the Pareto optimal solution set, with the selection criterion being maximizing the weighted sum of the weights of each sub-objective. The weights are dynamically adjusted according to the application scenario: for privacy-priority scenarios, the weight of sensitivity satisfaction is increased; for resource-constrained scenarios, the weight of budget utilization is increased.

[0154] The bidirectional mapping architecture is implemented by constructing a forward mapper and a reverse mapper. The forward mapper maps sensitive quantifications to calibration budget amounts, and the reverse mapper maps calibration budget amounts back to sensitive quantifications. The forward mapper is implemented using the calibrated budget allocation function, while the reverse mapper is implemented using a numerical solution method. It employs a binary search algorithm to search for quantifications within the sensitive quantification domain that match the budget amount, with a search precision of 0.0001 and a maximum search depth of 20 levels to ensure convergence. The bidirectional mapping architecture also includes a consistency check module, which evaluates the mapping quality by comparing the combined result of the forward and reverse mappings with the original input. The difference threshold is set to 0.05; if the difference exceeds the threshold, a mapping recalibration process is triggered.

[0155] Structural deviation is measured by calculating the distributional difference between the calibration budget and the sensitivity quantification in the feature space. Two indicators, KL divergence and Wasserstein distance, are used to quantify the degree of distributional difference. KL divergence reflects the information difference between distributions, while Wasserstein distance reflects the geometric distance between them; these two indicators complement each other to provide a comprehensive deviation assessment. The deviation calculation process constructs probability distributions for both the sensitivity quantification and the calibration budget, and estimates the probability density function using a kernel density estimation method. A Gaussian kernel is used as the kernel function, and the bandwidth parameter is automatically selected using cross-validation. Structural deviation also includes two dimensions: spatial deviation and temporal deviation. Spatial deviation reflects static distribution differences, while temporal deviation reflects dynamic changes. Temporal deviation is calculated using a sliding window with a window size set to 10 time steps.

[0156] The adjustment signal is constructed based on the structural deviation metric. A proportional-integral-derivative (PID) controller model is used to generate the adjustment signal. The proportional term reflects the current deviation magnitude, the integral term reflects the historical deviation accumulation, and the derivative term reflects the deviation trend. The controller parameters include three adjustable parameters: proportional coefficient, integral coefficient, and derivative coefficient. A proportional coefficient of 0.8 ensures fast response, an integral coefficient of 0.1 avoids integral saturation, and a derivative coefficient of 0.05 suppresses high-frequency noise. The adjustment signal also includes a saturation limiting mechanism, restricting the signal amplitude to the range of -1 to +1 to prevent over-adjustment and stability issues. Signal filtering is implemented using a first-order low-pass filter with a cutoff frequency of 0.1 to ensure signal smoothness, and the filtering delay is controlled within two sampling periods.

[0157] The dynamic optimization of the privacy budget allocation function is achieved through gradient descent, with the optimization objective being to minimize the weighted combination of structural bias and adjusted signal amplitude. Gradient calculation employs numerical differentiation, with a differentiation step size of 0.001 to ensure computational accuracy. Gradient updates utilize the Adam optimizer with a learning rate of 0.01 and momentum parameters of 0.9 and 0.999. Functional morphological parameters include optimizable variables such as segmentation point positions, slopes of each segment, and interpolation coefficients. Parameter boundary constraints ensure the function's monotonicity and continuity. An early stopping mechanism is employed during optimization; optimization stops when the objective function improvement is less than 0.001 after 10 consecutive iterations to avoid overfitting and wasted computational resources. Dynamic optimization also includes parameter version management, saving historical parameter versions during the optimization process, supporting parameter rollback and A / B testing, with a version retention period of 30 days.

[0158] The adaptive mapping mechanism integrates dynamic optimization results to form a closed-loop control architecture, enabling automatic adjustment of the mapping relationship between sensitivity metrics and budget limits. The mechanism comprises three components: a monitoring module, a decision-making module, and an execution module. The monitoring module collects structural deviation and adjustment signal data; the decision-making module formulates optimization strategies based on the collected data; and the execution module implements parameter updates and function refactoring. The decision-making strategy includes two modes: incremental adjustment and aggressive adjustment. Incremental adjustment is used when the structural deviation is less than 0.1, with an adjustment increment of 5% of the current parameter value; aggressive adjustment is used when the structural deviation is greater than 0.3, with an adjustment increment of 20% of the current parameter value. The adaptive mechanism also includes a stability guarantee function, ensuring system convergence through Lyapunov stability analysis. The convergence criterion is that the system state change is less than 0.001 after 50 consecutive adjustments.

[0159] In the specific data example, the input sensitivity mapping relationship includes 8 feature dimensions, with sensitivity quantification values ​​of [0.15, 0.45, 0.75, 0.25, 0.65, 0.35, 0.85, 0.55]. After multi-dimensional semantic deconstruction, the semantic strength, distribution characteristics, and correlation degree indicators of each dimension are calculated, and the weight allocation of the feature characterization system is [0.12, 0.18, 0.09, 0.15, 0.11, 0.16, 0.08, 0.11]. The privacy budget allocation function adopts a three-segment form, and the initial budget amount is calculated as [0.075, 0.338, 0.891, 0.125, 0.650, 0.175, 0.956, 0.440]. After adjustment via the budget calibration mechanism, the calibration budget amounts are [0.082, 0.325, 0.876, 0.138, 0.642, 0.189, 0.941, 0.427], with a total budget consumption of 3.620. The structural deviation metrics of the bidirectional mapping architecture show a KL divergence of 0.045, a Wasserstein distance of 0.032, and an adjustment signal amplitude of 0.067. After 15 dynamic optimization iterations, the structural deviation decreases to 0.018, the adaptive mapping mechanism reaches a stable state, and the final budget allocation scheme satisfies the sensitivity protection requirements and budget constraints.

[0160] A second aspect of the present invention provides an electronic device, comprising:

[0161] processor;

[0162] Memory used to store processor-executable instructions;

[0163] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0164] A third aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0165] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A privacy-preserving multi-institutional medical data collaborative modeling and intelligent generation method, characterized in that, include: Local medical datasets are obtained from multiple data holders. Addressing the heterogeneity of the feature space and the differences in semantic representation of these local medical datasets, a cross-institutional feature-semantic mapping map is constructed, incorporating a contrastive learning mechanism. This cross-institutional feature-semantic mapping map captures and quantifies the semantic similarity of features between different institutions through the contrastive learning mechanism, including: To address the heterogeneity issue of the local medical dataset, a feature-semantic dual-space encoding structure is constructed. This structure preserves the original feature expression through the institution's private semantic space and establishes a mapping relationship between features through a cross-institutional shared semantic space. This achieves semantic interoperability of features while protecting the uniqueness of institutional data, resulting in a dual-space feature embedding representation. A contrastive learning framework is constructed based on the dual-space feature embedding representation. The contrastive learning framework identifies anchor features from data of various institutions through semantic stability analysis, constructs positive sample pairs by embedding the anchor features in the cross-institutional shared semantic space, and selects negative sample pairs by feature correlation constraints. The mapping parameters of the feature semantic dual-space encoding structure are optimized by contrastive learning so that features with similar semantics have similarity in the cross-institutional shared semantic space, resulting in the optimized dual-space feature embedding representation. The optimized dual-space feature embedding representation is used to calculate the semantic similarity matrix of cross-institutional features. The semantic similarity matrix is ​​verified by introducing semantic transitivity consistency constraints. A cross-institutional feature semantic mapping map is constructed by incorporating the contrastive learning mechanism. Based on the cross-institutional feature semantic mapping map, the heterogeneous features of each institution are semantically unified to achieve semantic alignment of cross-institutional data. Heterogeneous features are mapped to a unified semantic space to obtain a semantically aligned dataset. A privacy risk assessment is then performed on the semantically aligned dataset to generate a feature-level privacy sensitivity vector. Based on this vector, a hybrid coding data structure with hierarchical protection strength is constructed, including: Based on the feature-level privacy sensitivity vector, features are protected hierarchically. According to the sensitivity level, the features are divided into different sensitive feature sets by an adaptive threshold. For feature sets with different sensitivity levels, dense feature representations and perturbation feature representations are generated respectively. The dense feature representations and perturbation feature representations are recombined according to the structural relationship between the original features to construct a hybrid coding data structure with hierarchical protection strength. Based on the hybrid coding data structure, the local feature representations of each data holding institution are extracted through the contrastive learning mechanism, and the local feature representations are encrypted using the homomorphic encryption mechanism. The multiple encrypted local feature representations are then fused to generate a global feature representation. The global feature representation is used to train and generate a model. During the training process, a differentiated privacy budget is allocated to features with different sensitivity based on the feature-level privacy sensitivity vector. The model parameters are then updated iteratively and protectively through the homomorphic encryption mechanism.

2. The method according to claim 1, characterized in that, Heterogeneous features are mapped to a unified semantic space to obtain a semantically aligned dataset. A privacy risk assessment is then performed on the semantically aligned dataset to generate a feature-level privacy sensitivity vector, including: A multidimensional privacy risk analysis framework is constructed, and a privacy assessment is performed on the semantically aligned local medical dataset through the multidimensional privacy risk analysis framework to obtain multidimensional privacy risk indicators. Membership inference and attribute inference attacks are performed on the semantically aligned local medical dataset. Multidimensional risk aggregation weights are obtained by analyzing the impact of the multidimensional privacy risk indicators on the attack success rate. The multidimensional risk aggregation weights are then comprehensively calculated and integrated with the corresponding dimension privacy risk indicators to generate a feature-level privacy sensitivity vector that quantifies the privacy protection requirements of features.

3. The method according to claim 2, characterized in that, Membership inference and attribute inference attacks are performed on the semantically aligned local medical dataset. The multidimensional risk aggregation weights are obtained by analyzing the impact of the multidimensional privacy risk indicators on the attack success rate, including: The target record set is extracted from the semantically aligned local medical dataset, and a background knowledge set simulating attacker auxiliary information is constructed. Based on the known values ​​of the background knowledge set, a member inference attack executor is constructed. The member inference attack executor analyzes the association features between the target record and the original dataset, identifies and determines the dataset affiliation relationship of each record in the target record set, and obtains a member determination result with multi-dimensional verification. Based on the member determination results of the multi-dimensional verification, member records with deterministic confidence are obtained. An attribute inference attack executor is constructed by combining the background knowledge set. The attribute inference attack executor is used to mine the implicit correlation patterns between features in the confidence member records. Undisclosed sensitive feature values ​​are inferred from multiple angles to form a systematic attribute inference result. A multi-layered risk assessment system is constructed using the member inference attack executor and the attribute inference attack executor. Based on the risk assessment system, attack tests of different intensities and strategies are performed on the semantically aligned local medical dataset to explore the data vulnerability boundary and obtain the attack success rate under each dimension. Based on the attack success rate, the propagation characteristics and impact of privacy risks in each dimension are analyzed and evaluated, and a multi-dimensional risk aggregation weight that can dynamically reflect the importance of the dimensions is generated.

4. The method according to claim 1, characterized in that, Based on the hybrid coding data structure, local feature representations of each data holding institution are extracted through the contrastive learning mechanism, and homomorphic encryption is used to encrypt the local feature representations. The fusion of multiple encrypted local feature representations to generate a global feature representation includes: Based on the aforementioned hybrid coding data structure, a contrastive learning module is designed. The contrastive learning module analyzes the multi-level semantic associations and dynamic structural features in the hybrid coding data structures of various data holding institutions, and establishes an adaptive cross-institutional data similarity measurement system. Based on the cross-institutional data similarity measurement system, local feature representations are generated for the hybrid coded data structures of each data holding institution; the local feature representations are enhanced with security by a homomorphic encryption module to generate encrypted feature representations that are tamper-proof and maintain the original structure and computational properties; the encrypted feature representations are aggregated to a central server according to the data holding institutions, and the central server forms aggregated encrypted features in an adaptive weighted manner within the encryption domain, and after decryption, obtains a global feature representation that integrates differentiated knowledge from multiple parties.

5. The method according to claim 1, characterized in that, The generative model is trained using the global feature representation. During training, differentiated privacy budgets are allocated to features with different sensitivity based on the feature-level privacy sensitivity vector, including: The training mechanism for constructing a generative model is constructed using the global feature representation. The global feature representation is deconstructed in multiple dimensions through the training mechanism to establish a dimensional index system that reflects the essential attributes of the features. Based on the dimensional indexing system, the feature-level privacy sensitivity vector is associated, the sensitivity of each dimension feature is quantitatively evaluated and a sensitivity mapping relationship is formed. A privacy budget allocation function is designed according to the sensitivity mapping relationship. An adaptive mapping mechanism between the sensitivity quantification value and the budget amount is constructed according to the privacy budget allocation function. The budget is allocated differently for each dimension feature through the adaptive mapping mechanism to form a feature-level privacy budget allocation scheme. During the parameter optimization process of the generative model, the gradient components of each dimension feature are dynamically calibrated according to the feature-level privacy budget allocation scheme to generate a gradient update direction with privacy protection effect; the network parameters of the generative model are adjusted based on the gradient update direction until the model converges and the training process is completed.

6. The method according to claim 5, characterized in that, Based on the sensitivity mapping relationship, a privacy budget allocation function is designed. Based on the privacy budget allocation function, an adaptive mapping mechanism is constructed between sensitivity metrics and budget amounts, including: Based on the sensitivity mapping relationship, the sensitivity metrics are deconstructed in multiple dimensions to construct a feature characterization system that represents the sensitivity metrics. Based on the feature characterization system, the function form of the privacy budget allocation function is designed. Based on the function form of the privacy budget allocation function, a mapping calculation unit is constructed. The sensitivity metrics of each feature dimension in the sensitivity mapping relationship are evaluated through the mapping calculation unit, and differential transformation operations are performed to generate the initial budget amount corresponding to each feature dimension. A budget calibration mechanism is constructed based on the initial budget amount. The hierarchical relationship and competitive game characteristics between budget amounts in each dimension are explored according to the budget calibration mechanism to obtain a calibrated budget amount that considers the constraint coupling between dimensions. A bidirectional mapping architecture is constructed between the calibration budget and the sensitivity quantification. The structural deviation between the calibration budget and the sensitivity quantification is measured through the bidirectional mapping architecture, and an adjustment signal is constructed. The function shape of the privacy budget allocation function is dynamically optimized based on the adjustment signal, thus forming an adaptive mapping mechanism.

7. A privacy-preserving multi-institutional medical data collaborative modeling and intelligent generation system, used to implement the method of any one of claims 1-6, characterized in that, include: The first unit is used to obtain local medical datasets from multiple data holding institutions. In view of the feature space heterogeneity and semantic expression differences of the local medical datasets, a cross-institutional feature semantic mapping map with a contrastive learning mechanism is constructed. The cross-institutional feature semantic mapping map captures and quantifies the semantic similarity of features between different institutions through the contrastive learning mechanism. The second unit is used to map heterogeneous features to a unified semantic space to obtain a semantically aligned dataset, perform a privacy risk assessment on the semantically aligned dataset, generate a feature-level privacy sensitivity vector, and construct a hybrid coding data structure with hierarchical protection strength based on it. The third unit is used to extract local feature representations of each data holding institution based on the hybrid coding data structure through the contrastive learning mechanism, and to encrypt the local feature representations using a homomorphic encryption mechanism, and to fuse multiple encrypted local feature representations to generate a global feature representation. The fourth unit is used to train and generate a model using the global feature representation. During the training process, it allocates differentiated privacy budgets to features with different sensitivity based on the feature-level privacy sensitivity vector, and performs protective iterative updates to the model parameters through the homomorphic encryption mechanism.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Self-adaptive privacy security calculation method and system based on medical data feature perception

    CN120705904A

  • Medical data management method and system based on multi-modal fusion and privacy protection

    CN120929768A