A privacy weight adaptive heterogeneous data federated collaborative training method and system

By using a knowledge graph-guided contrastive learning model and dynamic privacy weights, the modal compatibility and privacy protection adaptability issues of heterogeneous data in federated learning are solved, enabling efficient collaborative training of heterogeneous data and improving the system's compatibility, flexibility, and efficiency.

CN120974543BActive Publication Date: 2026-04-24FUJIAN THINKWIN BIG DATA APPLICATION SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511484044.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-04-24
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing federated learning frameworks face challenges such as heterogeneous data modality compatibility issues, adaptability deficiencies of static privacy protection mechanisms, and high communication overhead when handling collaborative training with multi-source heterogeneous data, making it difficult to achieve improvements in compatibility, flexibility, and efficiency.

Method used

A heterogeneous data federated collaborative training method with privacy weights is adopted. Semantic alignment is performed through a contrastive learning model guided by a knowledge graph. Combined with dynamic privacy weights and differential privacy operations, a heterogeneous data encoding model is used to transform data of different modalities into structured vectors. Communication is optimized through compression and encrypted transmission, and participation strategies are dynamically adjusted to improve training efficiency.

Benefits of technology

It effectively solves the problem of heterogeneous data modality compatibility, realizes the adaptability of static privacy protection mechanism, significantly improves the compatibility, flexibility and efficiency of federated collaborative training, supports joint modeling of multimodal data, and improves model generalization ability and system robustness while ensuring privacy and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974543B_ABST
    Figure CN120974543B_ABST
Patent Text Reader

Abstract

The application provides a privacy weight adaptive heterogeneous data federal collaborative training method and system in the technical field of federal learning and privacy calculation, and the method comprises the following steps: S1, each client performs a differential privacy operation on a local data set based on a privacy weight to obtain a desensitized data set, and encodes the desensitized data set through a heterogeneous data encoding model; S2, through a contrast learning model, semantic alignment is performed on each encoding vector to obtain an aligned vector set; S3, the local model is trained through the aligned vector set to generate a local gradient, local model parameters are extracted, and the privacy weight, the local gradient and a local difference parameter are uploaded to a server; and S4, the server trains a global model based on the local difference parameter and a global gradient, extracts global model parameters and sends the global model parameters to each client for training. The application has the advantages that the compatibility, flexibility and efficiency of the heterogeneous data federal collaborative training are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of federated learning and privacy computing technology, and in particular to a method and system for federated collaborative training of heterogeneous data with adaptive privacy weights. Background Technology

[0002] With increasingly stringent data privacy protections, Federated Learning (FL) has emerged as a distributed machine learning paradigm. Federated Learning, through its core mechanism of "the model moves while the data remains stationary," allows participants such as medical institutions, financial companies, and smart terminals to collaboratively train a global model based on local datasets without exchanging raw data. The core objective of Federated Learning is to achieve joint modeling of data while protecting privacy and data security, thereby improving model performance and generalization capabilities.

[0003] Typical federated learning systems adopt a client-server architecture, and their operation process includes four technical stages: (1) the server initializes global model parameters and broadcasts them to the clients; (2) each client performs a preset number of SGD optimizations based on local data (such as 5-10 rounds of local iteration in the FedAvg algorithm); (3) the clients upload model updates (including gradient tensors, weight matrices, and other parameters) to the server; (4) the server generates a new global model through aggregation algorithms (such as weighted average in FedAvg and approximate optimization in FedProx). After multiple rounds of iteration (usually 50-200 rounds of communication), the model converges to a stable state that meets the preset accuracy. Although federated learning has inherent advantages in privacy protection, it still faces the following technical bottlenecks when processing multi-source heterogeneous data for collaborative training:

[0004] 1. Heterogeneous data modal compatibility issues:

[0005] In real-world scenarios, the data structures of participating parties exhibit significant heterogeneity. For example, medical institutions store 3D medical images in DICOM format (unstructured data), banking systems store customer transaction records (structured relational data), and IoT devices generate sensor logs in JSON format (semi-structured time-series data). Existing federated learning frameworks (such as TensorFlow Federated) are primarily designed for homogeneous data, employing a unified feature embedding layer to process multimodal data. This leads to: ① Difficulty in cross-modal feature alignment: Semantic gaps exist in the representations of different data structures in the latent space; for example, CT scan slices and financial transaction records differ significantly in feature dimensions and numerical distribution, and directly concatenating the inputs can easily cause gradient conflicts. ② Severe information loss: Forcibly converting unstructured data (such as images) into structured feature vectors results in the loss of crucial information such as spatial topological relationships.

[0006] 2. Adaptability limitations of static privacy protection mechanisms:

[0007] Current privacy protection schemes often employ fixed parameter configurations. For example, differential privacy (DP) sets a constant noise level (such as Gaussian noise σ=1.2). This rigid strategy leads to the following contradictions: ① Over-protection of low-sensitivity data (such as publicly available traffic flow data) results in a loss of model utility. ② Insufficient protection of highly sensitive data (such as gene sequences), making it difficult for existing schemes to dynamically adapt to the trust levels of different stakeholders.

[0008] 3. High communication overhead:

[0009] Heterogeneous data distribution leads to the following communication bottlenecks, which in turn affect the efficiency of collaborative training: ① Waste of bandwidth resources: Using full-connection communication, each client needs to upload all parameters (e.g., ResNet-50 contains 25 million parameters). ② Device heterogeneity constraints: The hardware performance of the participants varies significantly (e.g., the computing power of a mobile phone GPU is only 1 / 100 of that of a server-grade GPU), resulting in long synchronization waiting times.

[0010] Therefore, how to provide a heterogeneous data federated collaborative training method and system with privacy weights that can improve the compatibility, flexibility and efficiency of heterogeneous data federated collaborative training has become an urgent technical problem to be solved. Summary of the Invention

[0011] The technical problem to be solved by this invention is to provide a heterogeneous data federated collaborative training method and system with privacy weight adaptive, so as to improve the compatibility, flexibility and efficiency of heterogeneous data federated collaborative training.

[0012] In a first aspect, the present invention provides a privacy-weighted adaptive heterogeneous data federated collaborative training method, comprising the following steps:

[0013] Step S1: The server initializes a global model, deploys the global model to each client as a local model, creates and trains a heterogeneous data encoding model to deploy to each client, and creates a knowledge graph-guided contrastive learning model.

[0014] Step S2: Each client dynamically calculates the privacy weight of each data in the local dataset based on data sensitivity, participant trust score, and data distribution differences; the local dataset includes structured data, semi-structured data, or unstructured data.

[0015] Step S3: Each client performs differential privacy operation on the local dataset based on the privacy weight to obtain a de-identified dataset. Through the deployed heterogeneous data encoding model, an encoding operation is performed on the de-identified dataset to obtain an initial structured encoding vector, an initial semi-structured encoding vector, or an initial unstructured encoding vector.

[0016] Step S4: Each client performs a collaborative semantic alignment operation on the initial structured encoding vector, initial semi-structured encoding vector, or initial unstructured encoding vector through the contrastive learning model created by the server, to obtain aligned structured encoding vectors, aligned semi-structured encoding vectors, or aligned unstructured encoding vectors, and constructs an aligned vector set;

[0017] Step S5: Each client trains the local model using the alignment vector set, records the local training log, generates the local gradient of the local model, extracts the local model parameters of the trained local model, compares the local model parameters extracted from the two training rounds with the local training log to obtain the local difference parameters, compresses and encrypts the privacy weights, local gradients and local difference parameters into a local encrypted data packet, and uploads the local encrypted data packet to the server based on the preset elastic participation strategy.

[0018] Step S6: The server decrypts and decompresses the received local encrypted data packet to obtain privacy weights, local gradients, and local difference parameters. Based on each privacy weight, the server performs global aggregation on each local gradient to obtain the global gradient. Based on the received local difference parameters and global gradient, the server performs collaborative training on the global model. During the training process, the server optimizes the model using a preset composite loss function, extracts the global model parameters of the trained global model, compresses and encrypts the global model parameters into an online encrypted data packet, and sends the online encrypted data packet to each client.

[0019] Step S7: Each client decrypts and decompresses the received online encrypted data packet to obtain global model parameters, and trains the local model based on the global model parameters until the preset training rounds or early stopping conditions are met.

[0020] Secondly, this invention provides a privacy-weighted adaptive heterogeneous data federated collaborative training system, comprising the following modules:

[0021] The initialization module is used to initialize a global model on the server, deploy the global model to each client as a local model, create and train a heterogeneous data encoding model to deploy to each client, and create a knowledge graph-guided contrastive learning model.

[0022] The privacy weight calculation module is used by each client to dynamically calculate the privacy weight of each data in the local dataset based on data sensitivity, participant trust scores, and differences in data distribution; the local dataset includes structured data, semi-structured data, or unstructured data.

[0023] The de-identification encoding module is used by each client to perform differential privacy operations on the local dataset based on the privacy weight to obtain a de-identified dataset. Through the deployed heterogeneous data encoding model, the module performs encoding operations on the de-identified dataset to obtain an initial structured encoding vector, an initial semi-structured encoding vector, or an initial unstructured encoding vector.

[0024] The semantic alignment module is used by each client to perform a collaborative semantic alignment operation on the initial structured encoding vector, initial semi-structured encoding vector, or initial unstructured encoding vector through the contrastive learning model created by the server, to obtain aligned structured encoding vectors, aligned semi-structured encoding vectors, or aligned unstructured encoding vectors, and to construct an aligned vector set;

[0025] The local training module is used by each client to train the local model through the alignment vector set, record local training logs, generate local gradients of the local model, extract local model parameters of the trained local model, obtain local difference parameters by comparing the local model parameters extracted in the two training rounds with the local training logs, compress and encrypt the privacy weights, local gradients and local difference parameters into a local encrypted data packet, and upload the local encrypted data packet to the server based on a preset elastic participation strategy.

[0026] The online training module is used by the server to decrypt and decompress the received local encrypted data packets to obtain privacy weights, local gradients, and local difference parameters. Based on each privacy weight, the local gradients are globally aggregated to obtain the global gradient. Based on the received local difference parameters and global gradient, the global model is collaboratively trained. During the training process, optimization is performed through a preset composite loss function. The global model parameters of the trained global model are extracted, and the global model parameters are compressed and encrypted into an online encrypted data packet, which is then sent to each client.

[0027] The federated collaborative training module is used by each client to decrypt and decompress the received online encrypted data packets to obtain global model parameters, and to train the local model based on the global model parameters until the preset training rounds or early stopping conditions are met.

[0028] The advantages of this invention are:

[0029] 1. Initialize a global model on the server, then collaboratively deploy the global model to each client as a local model. Create and train a heterogeneous data encoding model, then collaboratively deploy it to each client. Also, create a knowledge graph-guided contrastive learning model. Each client dynamically calculates the privacy weight of each data point in its local dataset based on data sensitivity, participant trust scores, and data distribution differences. Based on these privacy weights, perform differential privacy operations on the local dataset to obtain a de-identified dataset. Then, use the heterogeneous data encoding model to encode the de-identified dataset, obtaining initial structured encoding vectors, initial semi-structured encoding vectors, or initial unstructured encoding vectors. Finally, each client uses the contrastive learning model to process the initial... Collaborative semantic alignment is performed on the structured encoded vector, the initial semi-structured encoded vector, or the initial unstructured encoded vector to obtain aligned structured encoded vectors, aligned semi-structured encoded vectors, or aligned unstructured encoded vectors, and an aligned vector set is constructed. Each client trains its local model using the aligned vector set, records local training logs, generates local gradients for the local model, extracts the local model parameters after training, and obtains local difference parameters by comparing the local model parameters extracted from two training rounds using the local training logs. The privacy weights, local gradients, and local difference parameters are compressed and encrypted into a local encrypted data packet, and the local encrypted data packet is then distributed based on a preset elastic participation strategy. The encrypted data packet is uploaded to the server. The server decrypts and decompresses the local encrypted data packet to obtain privacy weights, local gradients, and local difference parameters. Based on each privacy weight, the local gradients are globally aggregated to obtain the global gradient. Based on the local difference parameters and the global gradient, the global model is co-trained. During training, optimization is performed using a preset composite loss function. The global model parameters of the trained global model are extracted, compressed, and encrypted into an online encrypted data packet, which is then sent to each client. Each client decrypts and decompresses the online encrypted data packet to obtain the global model parameters. Based on the global model parameters, the local model is trained until a preset number of training rounds is met or early stopping occurs. The conditions are as follows: First, a knowledge graph-guided contrastive learning model performs collaborative semantic alignment on heterogeneous encoding vectors to overcome the traditional problem of heterogeneous data modality compatibility. Second, differential privacy operations are performed on the local dataset using privacy weights, and local gradients are globally aggregated using privacy weights to overcome the adaptability defects of traditional static privacy protection mechanisms. Third, local difference parameters are uploaded instead of the traditional full upload, and the transmitted data is compressed. Finally, a flexible participation strategy for federated collaborative training (i.e., selecting the round of federated collaborative training based on hardware performance to shorten synchronization waiting time) is used to significantly improve the compatibility, flexibility, and efficiency of heterogeneous data federated collaborative training.

[0030] 2. By dynamically calculating privacy weights across three dimensions—data sensitivity (data inherent attributes), participant trust scores (collaboration history assessment), and data distribution differences (data heterogeneity index)—it addresses the limitation of traditional federated learning's single privacy budget in adapting to complex scenarios. It supports differentiated processing of structured / semi-structured / unstructured data; for example, assigning higher privacy weights to highly sensitive data while reducing protection for low-risk data to improve data usability. Furthermore, it directly correlates privacy weights with the variance of Gaussian noise (σ²·W). i This achieves the optimal balance between high-strength protection for sensitive data and low interference with non-sensitive data.

[0031] 3. By combining a heterogeneous data encoding model with a knowledge graph-guided contrastive learning model, the limitation of traditional federated learning in handling only single data types is overcome: the heterogeneous data encoding model transforms different modal data into structured vectors, eliminating differences in data form; the knowledge graph provides semantic association constraints, and contrastive learning achieves semantic space alignment across clients, supporting scenarios that integrate multimodal data.

[0032] 4. By integrating federated learning, privacy-preserving computation, and heterogeneous data processing, a dynamic privacy weight mechanism is used to achieve multi-dimensional adaptive protection of data sensitivity, participant trust, and distribution differences, maximizing data utility while ensuring privacy and security. Utilizing heterogeneous data encoding models and knowledge graph-guided contrastive learning models, the semantic alignment challenge of cross-modal data (structured / semi-structured / unstructured) is overcome, constructing a general federated training framework. Combining compression, encrypted transmission, composite loss optimization, and difference parameter detection mechanisms, the model training efficiency and system robustness are significantly improved. It can effectively defend against malicious attacks and adapt to complex network environments, providing a secure and efficient collaborative learning solution.

[0033] 5. The global model and the local model are deployed collaboratively through the federated gateway, ensuring that the original data does not leave the local machine, while the distributed data effectively improves the generalization ability of the model; and it supports joint modeling of heterogeneous data (structured / semi-structured / unstructured) from multiple clients, breaking through the limitation of traditional federated learning that only supports a single data type.

[0034] 6. Dedicated encoders are designed for different data types: Graph Attention Network (GAT) is used for structured data to capture topological relationships between entities and enhance relational reasoning capabilities; Temporal Convolutional Network (TCN) + Transformer are used to jointly model temporal dependencies and long-range semantics for semi-structured data, which is suitable for dynamic data such as logs and sensor streams; 3D ResNet is used to efficiently extract spatiotemporal features of images / videos and avoid the gradient vanishing problem for unstructured data; Each encoder is optimized independently to avoid mutual interference between features of different modalities, while providing high-quality vector representations for subsequent cross-modal alignment.

[0035] 7. By matching the similarity between entity embedding vectors and multimodal encoding vectors (entity alignment unit), the semantic constraints of the knowledge graph are injected into the data representation to improve cross-modal semantic consistency; based on the knowledge graph relation path reasoning latent semantic relations (relation reasoning unit), the problem of implicit association mining in data sparse scenarios is solved; by using a multi-head attention mechanism to dynamically fuse entity association weights and semantic relation vectors, graph embedding vectors rich in semantic context are generated to enhance the targeting of knowledge guidance.

[0036] 8. Eliminate modal differences by unifying the contrast space to achieve semantic alignment of multimodal data; improve the model's ability to distinguish fine-grained differences by using dynamic negative sampling to filter difficult negative samples based on semantic similarity; dynamically control the discriminative power of positive / negative sample pairs by using a temperature coefficient to balance model convergence speed and robustness, avoiding overfitting or underfitting; use a gradient reversal mechanism to resolve the conflict between the optimization direction of knowledge graph guided loss and contrastive loss, ensuring multi-task collaborative optimization; dynamically adjust the loss weights according to the alignment difficulty of each modality (e.g., assign higher weights to unstructured data with higher alignment difficulty) to avoid a single modality dominating the training process.

[0037] 9. By providing prior semantic constraints through knowledge graphs, the dependence of contrastive learning on massive negative samples is reduced, thus lowering computational costs; by combining the federated framework with dynamic negative sampling, the problems of data heterogeneity and sample imbalance are effectively alleviated, improving the model's convergence speed and generalization performance.

[0038] 10. Knowledge graph-guided contrastive learning reduces reliance on massive negative samples, and dynamic negative sampling further filters high-value samples, reducing computational complexity; 3D ResNet accelerates feature extraction from unstructured data (such as videos) through residual connections, avoiding the gradient vanishing problem in deep networks; Temporal Convolutional Network (TCN) efficiently processes semi-structured data (such as time-series logs) with its causal convolutional structure, reducing training time compared to RNN.

[0039] 11. Ensuring data privacy and security through a federated learning framework, it deeply integrates the semantic priors of knowledge graphs with the heterogeneous encoding capabilities of multimodal data (structured, semi-structured, and unstructured). It achieves cross-modal semantic alignment through knowledge-guided contrastive learning, and optimizes model training efficiency and robustness by combining dynamic negative sampling, gradient coordinators, and adaptive weight allocation mechanisms. Its modular design supports flexible expansion, significantly improving model generalization, interpretability, and cross-scenario transfer capabilities while protecting user privacy, providing efficient, reliable, and low-energy intelligent solutions for complex business scenarios.

[0040] 12. By adopting a multi-level compression mechanism of CRC check + normalized quantization + spatiotemporal differential coding + sparse matrix block, it balances data integrity (CRC check code) and compression efficiency (spatiotemporal differential coding eliminates redundancy); by jointly applying Huffman coding (global dictionary) and adaptive arithmetic coding, it significantly reduces the overhead of repeated coding of sparse submatrices, which is suitable for federated transmission scenarios of high-dimensional vector data; by combining sparse matrix block (preset block size) with global dictionary index replacement, it achieves coordinated compression of data locality and global repetition patterns, which is especially suitable for sparse feature representation of unstructured data.

[0041] 13. By embedding CRC checksums during the compression stage and generating HMAC authentication tags (128 bits) during the encryption stage, the data is ensured to remain unchanged during transmission through double verification; an encryption seed is generated using SHA3-512 to provide anti-collision hash protection and prevent the encryption parameters from being maliciously replaced.

[0042] 14. By compressing data packets and encapsulating metadata such as normalization parameters and quantization step size, and encrypting data packets and binding key information such as Nonce values ​​and authentication tags, the consistency of processing rules between the client and the server is ensured, and data parsing errors caused by parameter mismatches are avoided.

[0043] 15. The server performs collaborative semantic alignment on the initial vector sets of multiple clients through the knowledge graph, which solves the semantic gap problem of heterogeneous data (structured / semi-structured / unstructured) in federated learning and improves the model aggregation effect. The client and the server transmit compressed and encrypted data packets through the federated gateway to avoid the exposure of the original data in plaintext, while reducing the transmission bandwidth consumption and adapting to low bandwidth scenarios such as edge computing.

[0044] 16. Update the dynamic character replacement table through smart contracts to ensure the transparency and immutability of rule changes and avoid the risk of single point of control by centralized servers; provide audit traceability capabilities by storing hash values ​​on the blockchain.

[0045] 17. By coordinating the design of compression rules (normalization, quantization, spatiotemporal differential coding, sparse matrix processing) and encryption rules (HMAC, Argon2id, NTRUEncrypt, AES-256-GCM, Base85, Feistel), the project balances transmission efficiency and security, and solves the problem of the "security-efficiency" trade-off in federated learning.

[0046] 18. By deeply integrating an efficient multi-level compression mechanism (CRC check, spatiotemporal differential coding, sparse matrix block division, and global dictionary replacement) with a quantum-resistant encryption system (NTRU post-quantum algorithm, dynamic blockchain key management, and Feistel network obfuscation), the bandwidth consumption of cross-client data transmission in federated learning is significantly reduced while ensuring data integrity and privacy. At the same time, relying on blockchain smart contracts to dynamically update encryption rules and knowledge graph-driven collaborative semantic alignment, the semantic gap problem in heterogeneous data federation scenarios is solved, achieving end-to-end secure transmission and efficient semantic aggregation. It also has high scalability (adapting to structured / semi-structured / unstructured data) and resistance to reverse attacks (dynamic replacement table, Argon2id key derivation).

[0047] 19. By calling hardware acceleration technologies (such as GPU / TPU / FPGA) to optimize the local model training process, the computational efficiency is significantly improved, the single training cycle is shortened, and the client resource consumption is reduced; the frequency of client participation in federated training is dynamically adjusted based on hardware performance scores to avoid low-performance devices becoming system bottlenecks, ensure the overall efficiency of federated collaborative training, achieve dynamic balance between resource allocation and task load, and improve system throughput and stability.

[0048] 20. Hardware acceleration technology significantly improves the efficiency of local model training, and flexible participation strategy dynamically adapts to the computing resources of clients with different performance levels to achieve overall load balancing in federated collaborative training. The integrated compression and encryption design (multi-layer processing of privacy weights, gradients and differential parameters) effectively reduces communication overhead and ensures data privacy and security. At the same time, the traceability and fault tolerance of the training process are enhanced through full-dimensional training log recording and differential parameter comparison mechanism. Finally, a highly efficient, secure, low-resource-consumption federated learning system that supports flexible access from heterogeneous devices is built.

[0049] 21. By employing a bidirectional data transmission mechanism combining encryption and compression, a dynamic gradient aggregation strategy based on privacy weights, and a collaborative optimization design of the Adam optimizer and composite loss function, a multi-faceted balance is achieved within the federated learning framework, encompassing data privacy and security (end-to-end encryption and local data processing), improved communication efficiency (data compression to reduce transmission load), and optimized model performance (adaptive gradient aggregation to accelerate convergence). Furthermore, leveraging the automated collaborative training process and modular expansion capabilities of the federated gateway, an efficient, secure, and scalable distributed model training solution is provided.

[0050] 22. Achieving a precise balance between privacy protection and data utility through dynamic privacy weights and adaptive noise injection, and combining multimodal coding systems (graph attention networks, spatiotemporal convolution, 3D residual networks) with knowledge graph-guided comparative learning, effectively solves the problem of semantic alignment of heterogeneous data; a high-security, low-bandwidth communication link is constructed using five-layer quantum-secure encryption (NTRU+AES-256-GCM) and intelligent compression algorithms, and the robustness of the system is improved through elastic participation strategies and composite loss functions. Furthermore, the innovative integration of blockchain dynamic trust mechanisms and edge computing adaptation architecture ensures the accuracy of cross-modal models while effectively improving hardware resource utilization and reducing training energy consumption, forming a secure, efficient, and scalable federated learning full-stack solution. Attached Figure Description

[0051] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0052] Figure 1 This is a flowchart of a privacy-weighted adaptive heterogeneous data federated collaborative training method according to the present invention.

[0053] Figure 2 This is a schematic diagram of the structure of a privacy-weighted adaptive heterogeneous data federated collaborative training system according to the present invention. Detailed Implementation

[0054] The overall approach of the technical solution in this application is as follows: A knowledge graph-guided contrastive learning model performs collaborative semantic alignment on heterogeneous encoding vectors to overcome the traditional problem of heterogeneous data modality compatibility; differential privacy operations are performed on the local dataset using privacy weights, and local gradients are globally aggregated using privacy weights to overcome the adaptability defects of traditional static privacy protection mechanisms; uploading local difference parameters replaces the traditional full upload, and the transmitted data is compressed, combined with a flexible participation strategy for federated collaborative training, thereby improving the compatibility, flexibility, and efficiency of heterogeneous data federated collaborative training.

[0055] Please refer to Figures 1 to 2 As shown, a preferred embodiment of the privacy-weight-adaptive heterogeneous data federated collaborative training method of the present invention includes the following steps:

[0056] Step S1: The server initializes a global model, deploys the global model to each client as a local model, creates and trains a heterogeneous data encoding model to deploy to each client, and creates a knowledge graph-guided contrastive learning model.

[0057] Step S2: Each client dynamically calculates the privacy weight of each data in the local dataset based on data sensitivity, participant trust score, and data distribution differences; the local dataset includes structured data, semi-structured data, or unstructured data.

[0058] Step S3: Each client performs differential privacy operation on the local dataset based on the privacy weight to obtain a de-identified dataset. Through the deployed heterogeneous data encoding model, an encoding operation is performed on the de-identified dataset to obtain an initial structured encoding vector, an initial semi-structured encoding vector, or an initial unstructured encoding vector.

[0059] Step S4: Each client performs a collaborative semantic alignment operation on the initial structured encoding vector, initial semi-structured encoding vector, or initial unstructured encoding vector through the contrastive learning model created by the server, to obtain aligned structured encoding vectors, aligned semi-structured encoding vectors, or aligned unstructured encoding vectors, and constructs an aligned vector set;

[0060] Step S5: Each client trains the local model using the alignment vector set, records the local training log, generates the local gradient of the local model, extracts the local model parameters of the trained local model, compares the local model parameters extracted from the two training rounds with the local training log to obtain the local difference parameters, compresses and encrypts the privacy weights, local gradients and local difference parameters into a local encrypted data packet, and uploads the local encrypted data packet to the server based on the preset elastic participation strategy.

[0061] Step S6: The server decrypts and decompresses the received local encrypted data packet to obtain privacy weights, local gradients, and local difference parameters. Based on each privacy weight, the server performs global aggregation on each local gradient to obtain the global gradient. Based on the received local difference parameters and global gradient, the server performs collaborative training on the global model. During the training process, the server optimizes the model using a preset composite loss function, extracts the global model parameters of the trained global model, compresses and encrypts the global model parameters into an online encrypted data packet, and sends the online encrypted data packet to each client.

[0062] Step S7: Each client decrypts and decompresses the received online encrypted data packet to obtain global model parameters, and trains the local model based on the global model parameters until the preset training rounds or early stopping conditions are met.

[0063] By combining a heterogeneous data encoding model with a knowledge graph-guided contrastive learning model, the limitation of traditional federated learning in handling only single data types is overcome: the heterogeneous data encoding model transforms different modalities into structured vectors, eliminating differences in data form; the knowledge graph provides semantic association constraints, and contrastive learning achieves semantic space alignment across clients, supporting scenarios that integrate multimodal data.

[0064] Step S1 specifically involves:

[0065] The server initializes a global model based on a neural network, and deploys the global model to each client as a local model through a pre-configured federated gateway. It also creates and trains a heterogeneous data encoding model, deploys the heterogeneous data encoding model to each client through the federated gateway, and creates a knowledge graph-guided contrastive learning model.

[0066] The federated gateway enables collaborative deployment of global and local models, ensuring that the original data does not leave the local machine, while effectively improving the model's generalization ability by utilizing distributed data. It also supports joint modeling of heterogeneous data (structured / semi-structured / unstructured) from multiple clients, breaking through the limitation of traditional federated learning that only supports a single data type.

[0067] The heterogeneous data encoding model is constructed based on a structured data encoder, a semi-structured data encoder, and an unstructured data encoder. The structured data encoder is used to encode structured data through a first graph attention network to obtain a structured encoding vector. The semi-structured data encoder is used to encode semi-structured data through a temporal convolutional network and a Transformer to obtain a semi-structured encoding vector. The unstructured data encoder is used to encode unstructured data through a three-dimensional residual network to obtain an unstructured encoding vector.

[0068] Dedicated encoders are designed for different data types: Graph Attention Network (GAT) is used for structured data to capture topological relationships between entities and enhance relational reasoning capabilities; Temporal Convolutional Network (TCN) + Transformer are used to jointly model temporal dependencies and long-range semantics for semi-structured data, which is suitable for dynamic data such as logs and sensor streams; 3D ResNet is used to efficiently extract spatiotemporal features of images / videos and avoid the gradient vanishing problem for unstructured data; Each encoder is optimized independently to avoid mutual interference between features of different modalities, while providing high-quality vector representations for subsequent cross-modal alignment.

[0069] The contrastive learning model is constructed based on a knowledge graph guidance module, a contrastive learning module, and a collaborative optimization module.

[0070] The knowledge graph guidance module is constructed based on an entity alignment unit, a relation reasoning unit, and a second graph attention network. The entity alignment unit is used to perform similarity matching between entity embedding vectors in the knowledge graph and structured encoding vectors, semi-structured encoding vectors, and unstructured encoding vectors, respectively, to generate entity association weights. The relation reasoning unit is used to reason about the potential semantic relationships between structured encoding vectors, semi-structured encoding vectors, and unstructured encoding vectors based on relation paths in the knowledge graph, to obtain semantic relation vectors. The second graph attention network is used to aggregate neighborhood node information in the knowledge graph based on the entity association weights and semantic relation vectors through a multi-head attention mechanism to generate graph embedding vectors.

[0071] By matching the similarity between entity embedding vectors and multimodal encoding vectors (entity alignment unit), the semantic constraints of the knowledge graph are injected into the data representation, improving cross-modal semantic consistency; based on the knowledge graph relation path reasoning latent semantic relations (relation reasoning unit), the problem of implicit association mining in data sparse scenarios is solved; by using a multi-head attention mechanism to dynamically fuse entity association weights and semantic relation vectors, graph embedding vectors rich in semantic context are generated, enhancing the targeting of knowledge guidance.

[0072] The contrastive learning module is constructed based on a cross-modal projection unit, a dynamic negative sampling unit, and a contrastive loss calculation unit. The cross-modal projection unit maps structured encoding vectors, semi-structured encoding vectors, unstructured encoding vectors, and corresponding graph embedding vectors to a unified contrastive space, performs cross-modal contrastive learning within the contrastive space, and constructs positive and negative sample pairs to perform semantic alignment operations. The dynamic negative sampling unit calculates the semantic similarity of each positive and negative sample pair, and dynamically selects difficult negative sample pairs from each negative sample pair based on the semantic similarity. The contrastive loss calculation unit calculates the contrastive loss between each positive and negative sample pair using the InfoNCE loss function, and adjusts the discriminative power between the positive and difficult negative sample pairs using the temperature coefficient in the InfoNCE loss function.

[0073] The collaborative optimization module is built based on a gradient coordinator and an adaptive weight allocator. The gradient coordinator is used to balance the optimization direction of the guidance loss of the knowledge graph guidance module and the contrastive loss of the contrastive learning module through a gradient reversal mechanism. The adaptive weight allocator is used to dynamically adjust the loss weights of the guidance loss and the contrastive loss according to the alignment difficulty of structured encoding vectors, semi-structured encoding vectors and unstructured encoding vectors.

[0074] The formula for the guiding loss is: L1=ɑ'*L contrastive +β'*L cross-entropy +γ'*L MSE;

[0075] Where L1 represents the value of the bootstrapping loss; L contrastive The entity alignment unit sub-loss is represented by the contrastive loss function; L cross-entropy The sub-loss of the relational reasoning unit is represented by the cross-entropy loss function; L MSE The second graph shows the sub-loss of the attention network, which uses the mean squared error loss function; α', β', and γ' all represent weight coefficients.

[0076] By providing prior semantic constraints through knowledge graphs, the dependence of contrastive learning on massive negative samples is reduced, thus lowering computational costs. By combining a federated framework with dynamic negative sampling, the problems of data heterogeneity and sample imbalance are effectively alleviated, improving the model's convergence speed and generalization performance.

[0077] By ensuring data privacy and security through a federated learning framework, it deeply integrates the semantic priors of knowledge graphs with the heterogeneous encoding capabilities of multimodal data (structured, semi-structured, and unstructured). It achieves cross-modal semantic alignment through knowledge-guided contrastive learning, and optimizes model training efficiency and robustness by combining dynamic negative sampling, gradient coordinators, and adaptive weight allocation mechanisms. Its modular design supports flexible expansion, and while protecting user privacy, it significantly improves model generalization, interpretability, and cross-scenario transfer capabilities, providing efficient, reliable, and low-energy intelligent solutions for complex business scenarios.

[0078] In step S2, the formula for calculating the privacy weight is: ; ;

[0079] in, This represents the privacy weight of the i-th data item; Indicates data sensitivity; This represents the trust score among participating parties, based on the reliability of historical collaborations. It represents the difference in data distribution. In practice, it can be calculated based on KL divergence, JS divergence, Wasserstein distance, etc. This represents the i-th data; This represents the smoothing coefficient used to prevent division by zero errors; and All represent adjustment coefficients; This indicates the number of personally identifiable information entries contained in the i-th data item; This represents the total amount of data for the i-th data item; This indicates the percentage of sensitive information. This represents the entropy value of the data distribution;

[0080] Entropy of data distribution is an indicator that measures the uncertainty of data distribution. It comes from the concept of entropy in information theory and is used to quantify the randomness and uncertainty of data. The higher the entropy value, the more uniform the data distribution and the greater the uncertainty; the lower the entropy value, the more concentrated the data distribution and the smaller the uncertainty.

[0081] In step S3, the differential privacy operation specifically involves:

[0082] The Gaussian noise added to the local dataset is dynamically adjusted based on the privacy weights. The formula for the Gaussian noise is: ;

[0083] in, represents the Gaussian noise function; 0 represents the expectation of the Gaussian noise, which takes the value of zero; The variance of Gaussian noise is represented; This represents the fundamental intensity of Gaussian noise; This represents the privacy weight of the i-th data point.

[0084] Step S4 specifically involves:

[0085] Each client compresses the initial vector set, including the initial structured encoding vector, the initial semi-structured encoding vector, or the initial unstructured encoding vector, into an initial vector compressed package using preset compression rules. The initial vector compressed package is then encrypted into an initial vector encrypted package using preset encryption rules. The initial vector encrypted package is then uploaded to the server through the federated gateway.

[0086] A Federated Gateway is a middleware for federated learning and distributed data processing, designed to solve the problem of data silos, enabling data sharing and collaborative modeling among different institutions while ensuring data privacy and security. Through federated learning algorithms, privacy computing, and other technologies, the Federated Gateway allows data to be jointly modeled and analyzed without leaving the local machine.

[0087] The server decrypts and decompresses the initial vector encryption packets sent by each client using preset encryption and compression rules to obtain an initial vector set. Then, using the contrastive learning model, it invokes a preset knowledge graph to perform a collaborative semantic alignment operation on the initial vector set, resulting in aligned structured encoded vectors, aligned semi-structured encoded vectors, or aligned unstructured encoded vectors, thereby constructing corresponding aligned vector sets for each client. The knowledge graph stores knowledge used for semantic alignment of heterogeneous encoded vectors, including structured data entities, semi-structured data entities, unstructured data entities, structured data relationships, semi-structured data relationships, unstructured data relationships, structured data attributes, semi-structured data attributes, unstructured data attributes, semantic embeddings (vector representations of entities and relationships), and contextual information.

[0088] The server compresses the alignment vector set into an alignment vector compressed package using the compression rules, encrypts the alignment vector compressed package into an alignment vector encrypted package using the encryption rules, and sends the alignment vector encrypted package to each client through the federated gateway.

[0089] Each client decrypts and decompresses the aligned vector encryption packet sent by the server using the encryption and compression rules to obtain the aligned vector set;

[0090] The compression rules are as follows:

[0091] The process involves: performing CRC calculation on the data to be compressed to obtain a CRC checksum; normalizing the data using preset normalization parameters to obtain normalized data; quantizing the normalized data using a preset quantization step size to obtain quantized data; performing spatiotemporal differential encoding on the quantized data to obtain differential coded data; constructing a sparse matrix based on the positions and values ​​of non-zero values ​​in the differential coded data; dividing the sparse matrix into blocks using a preset block size to obtain sparse sub-matrices; and replacing repeated sparse sub-matrices in the sparse matrix with corresponding short codes using a preset global dictionary to obtain a first-level compressed sparse matrix. The short codes are obtained by performing Huffman coding on dictionary indices, which are used to match sparse sub-matrices from the global dictionary.

[0092] Spatiotemporal differential coding is a differential coding technique that combines time and spatial dimensions, primarily used for the efficient processing and transmission of data with spatiotemporal correlations. It reduces data redundancy by calculating the difference between adjacent time steps or spatial locations, thereby achieving data compression and efficient transmission.

[0093] Run-length encoding is performed on the positions in the first-level compressed sparse matrix, and adaptive arithmetic encoding is performed on the values ​​to obtain the second-level compressed sparse matrix.

[0094] The secondary compressed sparse matrix, CRC check code, normalization parameter, quantization step size, and global dictionary are encapsulated into a compressed data packet.

[0095] The decompression process of the compressed data packet is as follows: Parsing the compressed data packet yields a secondary compressed sparse matrix, a CRC checksum, normalization parameters, a quantization step size, and a global dictionary; running-length decoding is performed on the positions in the secondary compressed sparse matrix, and adaptive arithmetic decoding is performed on the values ​​to obtain a primary compressed sparse matrix; Huffman decoding is performed on the short codes in the primary compressed sparse matrix to obtain a dictionary index, and sparse submatrices are matched and replaced from the global dictionary using the dictionary index to obtain a sparse matrix; parsing the sparse matrix yields the positions and values ​​of non-zero values ​​in the differential encoded data, and the differential encoded data is restored; spatiotemporal differential decoding is performed on the differential encoded data to obtain quantized data, and the quantized data is dequantized using the quantization step size to obtain normalized data, and the normalized data is denormalized using the normalization parameters to obtain the data to be compressed; the integrity of the data to be compressed is verified using the CRC checksum to complete the decompression.

[0096] The encryption rules are as follows:

[0097] Generate a master key and a random nonce value of 12 bytes, obtain the data to be encrypted, calculate a 128-bit authentication tag for the data to be encrypted, the master key and the random nonce value using the HMAC algorithm, and perform SHA3-512 calculation on the random nonce value as the encryption seed;

[0098] The master key is encrypted using the encryption seed via the Argon2id algorithm to obtain a first-level encryption key, and then the first-level encryption key is encrypted using the NTRUEncrypt quantum algorithm to obtain a second-level encryption key.

[0099] The master key is called using the AES-256-GCM algorithm to encrypt the data to be encrypted into first-level encrypted data. The first-level encrypted data is Base85 encoded to obtain Base85 encoded data. The characters in the Base85 encoded data are replaced by a dynamic character replacement table to obtain second-level encrypted data.

[0100] The secondary encrypted data, the secondary encryption key, the random Nonce value, and the 128-bit authentication tag are concatenated to obtain concatenated data. The concatenated data is then subjected to packet permutation encryption using a Feistel network structure to obtain an encrypted data packet.

[0101] The dynamic character replacement table is stored on the blockchain platform in the form of a smart contract and is updated every 5 minutes based on preset update rules.

[0102] The dynamic character replacement table is stored on the blockchain platform in the form of a smart contract. The update function of the smart contract is called periodically by an external script to update the dynamic character replacement table, and the hash value of the updated dynamic character replacement table is stored on the blockchain.

[0103] The decryption process of the encrypted data packet is as follows: The encrypted data packet is decrypted using a Feistel network structure with packet permutation to obtain concatenated data. The concatenated data is parsed to obtain secondary encrypted data, a secondary encryption key, a random nonce value, and a 128-bit authentication tag. Characters in the secondary encrypted data are replaced using the dynamic character replacement table to obtain Base85 encoded data. The Base85 encoded data is then Base85 decoded to obtain primary encrypted data. The random nonce value is calculated using SHA3-512 as an encryption seed. The secondary encryption key is decrypted using the NTRUEncrypt quantum algorithm to obtain the primary encryption key. The primary encryption key is then decrypted using the Argon2id algorithm with the encryption seed to obtain the master key. The primary encrypted data is then decrypted into data to be encrypted using the AES-256-GCM algorithm with the master key. Finally, the 128-bit authentication tag is used to authenticate the data to be encrypted, the master key, and the random nonce value to complete the decryption process.

[0104] By compressing data packets and encapsulating metadata such as normalization parameters and quantization step size, and encrypting data packets to bind key information such as Nonce values ​​and authentication tags, the consistency of processing rules between the client and the server is ensured, and data parsing errors caused by parameter mismatches are avoided.

[0105] By updating the dynamic character replacement table through smart contracts, the transparency and immutability of rule changes are ensured, avoiding the risk of single point of control by centralized servers; and the on-chain storage of hash values ​​provides audit traceability capabilities.

[0106] The server-side uses a knowledge graph to perform collaborative semantic alignment on the initial vector sets of multiple clients, solving the semantic gap problem of heterogeneous data (structured / semi-structured / unstructured) in federated learning and improving the model aggregation effect. The client and server transmit compressed and encrypted data packets through the federated gateway to avoid the exposure of the original data in plaintext, while reducing the transmission bandwidth consumption and adapting to low-bandwidth scenarios such as edge computing.

[0107] By coordinating compression rules (normalization, quantization, spatiotemporal differential coding, sparse matrix processing) and encryption rules (HMAC, Argon2id, NTRUEncrypt, AES-256-GCM, Base85, Feistel), the system balances transmission efficiency and security, solving the "security-efficiency" trade-off problem in federated learning.

[0108] Step S5 specifically involves:

[0109] Each client uses hardware acceleration technology to train its local model by calling the alignment vector set. During training, it records local training logs in real time, including at least model information, training data information, model structure, model parameters, training process, hyperparameter adjustment process, and anomaly information. It generates local gradients of the local model through backpropagation algorithm, extracts local model parameters of the trained local model, and obtains local difference parameters by comparing the local model parameters extracted from the local training logs with those from the previous and next training rounds. It compresses the privacy weights, local gradients, and local difference parameters into a local compressed data packet using preset compression rules, encrypts the local compressed data packet into a local encrypted data packet using preset encryption rules, and uploads the local encrypted data packet to the server through the federated gateway based on a preset elastic participation strategy.

[0110] The flexible participation strategy is as follows: the hardware performance of the client is scored to obtain a performance score, and the performance score is converted into the training frequency for participating in federated collaborative training; for example, if a client's performance score is 30 points, and the average performance score of all clients is 90 points, then the other clients will participate in 3 federated collaborative training sessions, while the client is only allowed to participate in 1 federated collaborative training session to ensure the synchronization of training.

[0111] Step S6 specifically involves:

[0112] The server decrypts and decompresses the local encrypted data packets sent by each client using preset encryption and compression rules to obtain privacy weights, local gradients, and local difference parameters. Based on the privacy weights, it globally aggregates the local gradients to obtain the global gradient and aggregates the local difference parameters uploaded by each client to obtain aggregated difference parameters. The server then uses the Adam optimizer to update the global model with the global gradients and aggregated difference parameters for collaborative training of the global model. During training, the server optimizes the model using a preset composite loss function and extracts the global model parameters after training. The server then compresses the global model parameters into an online compressed data packet using the compression rules and encrypts the online compressed data packet into an online encrypted data packet using the encryption rules. Finally, the server distributes the encrypted data packet to each client through the federated gateway.

[0113] The formula for the global gradient is: ;

[0114] in, This represents the global gradient; n represents the total number of participants. This represents the global privacy weight of the j-th client, achieved by... The average value is obtained. This represents the privacy weight of the i-th data item on the client side; This represents the gradient clipping operation; This represents the local gradient of the j-th client; Indicates the gradient clipping threshold; represents the Gaussian noise function; 0 represents the expectation of the Gaussian noise, which takes the value of zero; The variance of Gaussian noise is represented; This represents the fundamental intensity of Gaussian noise;

[0115] The formula for the composite loss function is: ;

[0116] in, This represents the loss value of the composite loss function; The task loss is represented by the cross-entropy loss function; Indicates the smoothing coefficient; Indicates privacy constraints;

[0117] Step S7 specifically involves:

[0118] Each client decrypts and decompresses the received online encrypted data packets using preset encryption and compression rules to obtain global model parameters. Each client then updates the global model parameters to its local model using the Adam optimizer and trains the local model using hardware acceleration technology until the preset training rounds or early stopping conditions are met.

[0119] A preferred embodiment of the privacy-weight-adaptive heterogeneous data federated collaborative training system of the present invention includes the following modules:

[0120] The initialization module is used to initialize a global model on the server, deploy the global model to each client as a local model, create and train a heterogeneous data encoding model to deploy to each client, and create a knowledge graph-guided contrastive learning model.

[0121] The privacy weight calculation module is used by each client to dynamically calculate the privacy weight of each data in the local dataset based on data sensitivity, participant trust scores, and differences in data distribution; the local dataset includes structured data, semi-structured data, or unstructured data.

[0122] The de-identification encoding module is used by each client to perform differential privacy operations on the local dataset based on the privacy weight to obtain a de-identified dataset. Through the deployed heterogeneous data encoding model, the module performs encoding operations on the de-identified dataset to obtain an initial structured encoding vector, an initial semi-structured encoding vector, or an initial unstructured encoding vector.

[0123] The semantic alignment module is used by each client to perform a collaborative semantic alignment operation on the initial structured encoding vector, initial semi-structured encoding vector, or initial unstructured encoding vector through the contrastive learning model created by the server, to obtain aligned structured encoding vectors, aligned semi-structured encoding vectors, or aligned unstructured encoding vectors, and to construct an aligned vector set;

[0124] The local training module is used by each client to train the local model through the alignment vector set, record local training logs, generate local gradients of the local model, extract local model parameters of the trained local model, obtain local difference parameters by comparing the local model parameters extracted in the two training rounds with the local training logs, compress and encrypt the privacy weights, local gradients and local difference parameters into a local encrypted data packet, and upload the local encrypted data packet to the server based on a preset elastic participation strategy.

[0125] The online training module is used by the server to decrypt and decompress the received local encrypted data packets to obtain privacy weights, local gradients, and local difference parameters. Based on each privacy weight, the local gradients are globally aggregated to obtain the global gradient. Based on the received local difference parameters and global gradient, the global model is collaboratively trained. During the training process, optimization is performed through a preset composite loss function. The global model parameters of the trained global model are extracted, and the global model parameters are compressed and encrypted into an online encrypted data packet, which is then sent to each client.

[0126] The federated collaborative training module is used by each client to decrypt and decompress the received online encrypted data packets to obtain global model parameters, and to train the local model based on the global model parameters until the preset training rounds or early stopping conditions are met.

[0127] By combining a heterogeneous data encoding model with a knowledge graph-guided contrastive learning model, the limitation of traditional federated learning in handling only single data types is overcome: the heterogeneous data encoding model transforms different modalities into structured vectors, eliminating differences in data form; the knowledge graph provides semantic association constraints, and contrastive learning achieves semantic space alignment across clients, supporting scenarios that integrate multimodal data.

[0128] The initialization module is specifically used for:

[0129] The server initializes a global model based on a neural network, and deploys the global model to each client as a local model through a pre-configured federated gateway. It also creates and trains a heterogeneous data encoding model, deploys the heterogeneous data encoding model to each client through the federated gateway, and creates a knowledge graph-guided contrastive learning model.

[0130] The federated gateway enables collaborative deployment of global and local models, ensuring that the original data does not leave the local machine, while effectively improving the model's generalization ability by utilizing distributed data. It also supports joint modeling of heterogeneous data (structured / semi-structured / unstructured) from multiple clients, breaking through the limitation of traditional federated learning that only supports a single data type.

[0131] The heterogeneous data encoding model is constructed based on a structured data encoder, a semi-structured data encoder, and an unstructured data encoder. The structured data encoder is used to encode structured data through a first graph attention network to obtain a structured encoding vector. The semi-structured data encoder is used to encode semi-structured data through a temporal convolutional network and a Transformer to obtain a semi-structured encoding vector. The unstructured data encoder is used to encode unstructured data through a three-dimensional residual network to obtain an unstructured encoding vector.

[0132] Dedicated encoders are designed for different data types: Graph Attention Network (GAT) is used for structured data to capture topological relationships between entities and enhance relational reasoning capabilities; Temporal Convolutional Network (TCN) + Transformer are used to jointly model temporal dependencies and long-range semantics for semi-structured data, which is suitable for dynamic data such as logs and sensor streams; 3D ResNet is used to efficiently extract spatiotemporal features of images / videos and avoid the gradient vanishing problem for unstructured data; Each encoder is optimized independently to avoid mutual interference between features of different modalities, while providing high-quality vector representations for subsequent cross-modal alignment.

[0133] The contrastive learning model is constructed based on a knowledge graph guidance module, a contrastive learning module, and a collaborative optimization module.

[0134] The knowledge graph guidance module is constructed based on an entity alignment unit, a relation reasoning unit, and a second graph attention network. The entity alignment unit is used to perform similarity matching between entity embedding vectors in the knowledge graph and structured encoding vectors, semi-structured encoding vectors, and unstructured encoding vectors, respectively, to generate entity association weights. The relation reasoning unit is used to reason about the potential semantic relationships between structured encoding vectors, semi-structured encoding vectors, and unstructured encoding vectors based on relation paths in the knowledge graph, to obtain semantic relation vectors. The second graph attention network is used to aggregate neighborhood node information in the knowledge graph based on the entity association weights and semantic relation vectors through a multi-head attention mechanism to generate graph embedding vectors.

[0135] By matching the similarity between entity embedding vectors and multimodal encoding vectors (entity alignment unit), the semantic constraints of the knowledge graph are injected into the data representation, improving cross-modal semantic consistency; based on the knowledge graph relation path reasoning latent semantic relations (relation reasoning unit), the problem of implicit association mining in data sparse scenarios is solved; by using a multi-head attention mechanism to dynamically fuse entity association weights and semantic relation vectors, graph embedding vectors rich in semantic context are generated, enhancing the targeting of knowledge guidance.

[0136] The contrastive learning module is constructed based on a cross-modal projection unit, a dynamic negative sampling unit, and a contrastive loss calculation unit. The cross-modal projection unit maps structured encoding vectors, semi-structured encoding vectors, unstructured encoding vectors, and corresponding graph embedding vectors to a unified contrastive space, performs cross-modal contrastive learning within the contrastive space, and constructs positive and negative sample pairs to perform semantic alignment operations. The dynamic negative sampling unit calculates the semantic similarity of each positive and negative sample pair, and dynamically selects difficult negative sample pairs from each negative sample pair based on the semantic similarity. The contrastive loss calculation unit calculates the contrastive loss between each positive and negative sample pair using the InfoNCE loss function, and adjusts the discriminative power between the positive and difficult negative sample pairs using the temperature coefficient in the InfoNCE loss function.

[0137] The collaborative optimization module is built based on a gradient coordinator and an adaptive weight allocator. The gradient coordinator is used to balance the optimization direction of the guidance loss of the knowledge graph guidance module and the contrastive loss of the contrastive learning module through a gradient reversal mechanism. The adaptive weight allocator is used to dynamically adjust the loss weights of the guidance loss and the contrastive loss according to the alignment difficulty of structured encoding vectors, semi-structured encoding vectors and unstructured encoding vectors.

[0138] The formula for the guiding loss is: L1=ɑ'*L contrastive +β'*L cross-entropy +γ'*L MSE ;

[0139] Where L1 represents the value of the bootstrapping loss; L contrastive The entity alignment unit sub-loss is represented by the contrastive loss function; L cross-entropy The sub-loss of the relational reasoning unit is represented by the cross-entropy loss function; L MSE The second graph shows the sub-loss of the attention network, which uses the mean squared error loss function; α', β', and γ' all represent weight coefficients.

[0140] By providing prior semantic constraints through knowledge graphs, the dependence of contrastive learning on massive negative samples is reduced, thus lowering computational costs. By combining a federated framework with dynamic negative sampling, the problems of data heterogeneity and sample imbalance are effectively alleviated, improving the model's convergence speed and generalization performance.

[0141] By ensuring data privacy and security through a federated learning framework, it deeply integrates the semantic priors of knowledge graphs with the heterogeneous encoding capabilities of multimodal data (structured, semi-structured, and unstructured). It achieves cross-modal semantic alignment through knowledge-guided contrastive learning, and optimizes model training efficiency and robustness by combining dynamic negative sampling, gradient coordinators, and adaptive weight allocation mechanisms. Its modular design supports flexible expansion, and while protecting user privacy, it significantly improves model generalization, interpretability, and cross-scenario transfer capabilities, providing efficient, reliable, and low-energy intelligent solutions for complex business scenarios.

[0142] In the privacy weight calculation module, the formula for calculating the privacy weight is: ; ;

[0143] in, This represents the privacy weight of the i-th data item; Indicates data sensitivity; This represents the trust score among participating parties, based on the reliability of historical collaborations. It represents the difference in data distribution. In practice, it can be calculated based on KL divergence, JS divergence, Wasserstein distance, etc. This represents the i-th data; This represents the smoothing coefficient used to prevent division by zero errors; and All represent adjustment coefficients; This indicates the number of personally identifiable information entries contained in the i-th data item; This represents the total amount of data for the i-th data item; This indicates the percentage of sensitive information. This represents the entropy value of the data distribution;

[0144] Entropy of data distribution is an indicator that measures the uncertainty of data distribution. It comes from the concept of entropy in information theory and is used to quantify the randomness and uncertainty of data. The higher the entropy value, the more uniform the data distribution and the greater the uncertainty; the lower the entropy value, the more concentrated the data distribution and the smaller the uncertainty.

[0145] In the de-identification encoding module, the differential privacy operation specifically includes:

[0146] The Gaussian noise added to the local dataset is dynamically adjusted based on the privacy weights. The formula for the Gaussian noise is: ;

[0147] in, represents the Gaussian noise function; 0 represents the expectation of the Gaussian noise, which takes the value of zero; The variance of Gaussian noise is represented; This represents the fundamental intensity of Gaussian noise; This represents the privacy weight of the i-th data point.

[0148] The semantic alignment module is specifically used for:

[0149] Each client compresses the initial vector set, including the initial structured encoding vector, the initial semi-structured encoding vector, or the initial unstructured encoding vector, into an initial vector compressed package using preset compression rules. The initial vector compressed package is then encrypted into an initial vector encrypted package using preset encryption rules. The initial vector encrypted package is then uploaded to the server through the federated gateway.

[0150] A Federated Gateway is a middleware for federated learning and distributed data processing, designed to solve the problem of data silos, enabling data sharing and collaborative modeling among different institutions while ensuring data privacy and security. Through federated learning algorithms, privacy computing, and other technologies, the Federated Gateway allows data to be jointly modeled and analyzed without leaving the local machine.

[0151] The server decrypts and decompresses the initial vector encryption packets sent by each client using preset encryption and compression rules to obtain an initial vector set. Then, using the contrastive learning model, it invokes a preset knowledge graph to perform a collaborative semantic alignment operation on the initial vector set, resulting in aligned structured encoded vectors, aligned semi-structured encoded vectors, or aligned unstructured encoded vectors, thereby constructing corresponding aligned vector sets for each client. The knowledge graph stores knowledge used for semantic alignment of heterogeneous encoded vectors, including structured data entities, semi-structured data entities, unstructured data entities, structured data relationships, semi-structured data relationships, unstructured data relationships, structured data attributes, semi-structured data attributes, unstructured data attributes, semantic embeddings (vector representations of entities and relationships), and contextual information.

[0152] The server compresses the alignment vector set into an alignment vector compressed package using the compression rules, encrypts the alignment vector compressed package into an alignment vector encrypted package using the encryption rules, and sends the alignment vector encrypted package to each client through the federated gateway.

[0153] Each client decrypts and decompresses the aligned vector encryption packet sent by the server using the encryption and compression rules to obtain the aligned vector set;

[0154] The compression rules are as follows:

[0155] The process involves: performing CRC calculation on the data to be compressed to obtain a CRC checksum; normalizing the data using preset normalization parameters to obtain normalized data; quantizing the normalized data using a preset quantization step size to obtain quantized data; performing spatiotemporal differential encoding on the quantized data to obtain differential coded data; constructing a sparse matrix based on the positions and values ​​of non-zero values ​​in the differential coded data; dividing the sparse matrix into blocks using a preset block size to obtain sparse sub-matrices; and replacing repeated sparse sub-matrices in the sparse matrix with corresponding short codes using a preset global dictionary to obtain a first-level compressed sparse matrix. The short codes are obtained by performing Huffman coding on dictionary indices, which are used to match sparse sub-matrices from the global dictionary.

[0156] Spatiotemporal differential coding is a differential coding technique that combines time and spatial dimensions, primarily used for the efficient processing and transmission of data with spatiotemporal correlations. It reduces data redundancy by calculating the difference between adjacent time steps or spatial locations, thereby achieving data compression and efficient transmission.

[0157] Run-length encoding is performed on the positions in the first-level compressed sparse matrix, and adaptive arithmetic encoding is performed on the values ​​to obtain the second-level compressed sparse matrix.

[0158] The secondary compressed sparse matrix, CRC check code, normalization parameter, quantization step size, and global dictionary are encapsulated into a compressed data packet.

[0159] The decompression process of the compressed data packet is as follows: Parsing the compressed data packet yields a secondary compressed sparse matrix, a CRC checksum, normalization parameters, a quantization step size, and a global dictionary; running-length decoding is performed on the positions in the secondary compressed sparse matrix, and adaptive arithmetic decoding is performed on the values ​​to obtain a primary compressed sparse matrix; Huffman decoding is performed on the short codes in the primary compressed sparse matrix to obtain a dictionary index, and sparse submatrices are matched and replaced from the global dictionary using the dictionary index to obtain a sparse matrix; parsing the sparse matrix yields the positions and values ​​of non-zero values ​​in the differential encoded data, and the differential encoded data is restored; spatiotemporal differential decoding is performed on the differential encoded data to obtain quantized data, and the quantized data is dequantized using the quantization step size to obtain normalized data, and the normalized data is denormalized using the normalization parameters to obtain the data to be compressed; the integrity of the data to be compressed is verified using the CRC checksum to complete the decompression.

[0160] The encryption rules are as follows:

[0161] Generate a master key and a random nonce value of 12 bytes, obtain the data to be encrypted, calculate a 128-bit authentication tag for the data to be encrypted, the master key and the random nonce value using the HMAC algorithm, and perform SHA3-512 calculation on the random nonce value as the encryption seed;

[0162] The master key is encrypted using the encryption seed via the Argon2id algorithm to obtain a first-level encryption key, and then the first-level encryption key is encrypted using the NTRUEncrypt quantum algorithm to obtain a second-level encryption key.

[0163] The master key is called using the AES-256-GCM algorithm to encrypt the data to be encrypted into first-level encrypted data. The first-level encrypted data is Base85 encoded to obtain Base85 encoded data. The characters in the Base85 encoded data are replaced by a dynamic character replacement table to obtain second-level encrypted data.

[0164] The secondary encrypted data, the secondary encryption key, the random Nonce value, and the 128-bit authentication tag are concatenated to obtain concatenated data. The concatenated data is then subjected to packet permutation encryption using a Feistel network structure to obtain an encrypted data packet.

[0165] The dynamic character replacement table is stored on the blockchain platform in the form of a smart contract and is updated every 5 minutes based on preset update rules.

[0166] The dynamic character replacement table is stored on the blockchain platform in the form of a smart contract. The update function of the smart contract is called periodically by an external script to update the dynamic character replacement table, and the hash value of the updated dynamic character replacement table is stored on the blockchain.

[0167] The decryption process of the encrypted data packet is as follows: The encrypted data packet is decrypted using a Feistel network structure with packet permutation to obtain concatenated data. The concatenated data is parsed to obtain secondary encrypted data, a secondary encryption key, a random nonce value, and a 128-bit authentication tag. Characters in the secondary encrypted data are replaced using the dynamic character replacement table to obtain Base85 encoded data. The Base85 encoded data is then Base85 decoded to obtain primary encrypted data. The random nonce value is calculated using SHA3-512 as an encryption seed. The secondary encryption key is decrypted using the NTRUEncrypt quantum algorithm to obtain the primary encryption key. The primary encryption key is then decrypted using the Argon2id algorithm with the encryption seed to obtain the master key. The primary encrypted data is then decrypted into data to be encrypted using the AES-256-GCM algorithm with the master key. Finally, the 128-bit authentication tag is used to authenticate the data to be encrypted, the master key, and the random nonce value to complete the decryption process.

[0168] By compressing data packets and encapsulating metadata such as normalization parameters and quantization step size, and encrypting data packets to bind key information such as Nonce values ​​and authentication tags, the consistency of processing rules between the client and the server is ensured, and data parsing errors caused by parameter mismatches are avoided.

[0169] By updating the dynamic character replacement table through smart contracts, the transparency and immutability of rule changes are ensured, avoiding the risk of single point of control by centralized servers; and the on-chain storage of hash values ​​provides audit traceability capabilities.

[0170] The server-side uses a knowledge graph to perform collaborative semantic alignment on the initial vector sets of multiple clients, solving the semantic gap problem of heterogeneous data (structured / semi-structured / unstructured) in federated learning and improving the model aggregation effect. The client and server transmit compressed and encrypted data packets through the federated gateway to avoid the exposure of the original data in plaintext, while reducing the transmission bandwidth consumption and adapting to low-bandwidth scenarios such as edge computing.

[0171] By coordinating compression rules (normalization, quantization, spatiotemporal differential coding, sparse matrix processing) and encryption rules (HMAC, Argon2id, NTRUEncrypt, AES-256-GCM, Base85, Feistel), the system balances transmission efficiency and security, solving the "security-efficiency" trade-off problem in federated learning.

[0172] The local training module is specifically used for:

[0173] Each client uses hardware acceleration technology to train its local model by calling the alignment vector set. During training, it records local training logs in real time, including at least model information, training data information, model structure, model parameters, training process, hyperparameter adjustment process, and anomaly information. It generates local gradients of the local model through backpropagation algorithm, extracts local model parameters of the trained local model, and obtains local difference parameters by comparing the local model parameters extracted from the local training logs with those from the previous and next training rounds. It compresses the privacy weights, local gradients, and local difference parameters into a local compressed data packet using preset compression rules, encrypts the local compressed data packet into a local encrypted data packet using preset encryption rules, and uploads the local encrypted data packet to the server through the federated gateway based on a preset elastic participation strategy.

[0174] The flexible participation strategy is as follows: the hardware performance of the client is scored to obtain a performance score, and the performance score is converted into the training frequency for participating in federated collaborative training; for example, if a client's performance score is 30 points, and the average performance score of all clients is 90 points, then the other clients will participate in 3 federated collaborative training sessions, while the client is only allowed to participate in 1 federated collaborative training session to ensure the synchronization of training.

[0175] The online training module is specifically used for:

[0176] The server decrypts and decompresses the local encrypted data packets sent by each client using preset encryption and compression rules to obtain privacy weights, local gradients, and local difference parameters. Based on the privacy weights, it globally aggregates the local gradients to obtain the global gradient and aggregates the local difference parameters uploaded by each client to obtain aggregated difference parameters. The server then uses the Adam optimizer to update the global model with the global gradients and aggregated difference parameters for collaborative training of the global model. During training, the server optimizes the model using a preset composite loss function and extracts the global model parameters after training. The server then compresses the global model parameters into an online compressed data packet using the compression rules and encrypts the online compressed data packet into an online encrypted data packet using the encryption rules. Finally, the server distributes the encrypted data packet to each client through the federated gateway.

[0177] The formula for the global gradient is: ;

[0178] in, This represents the global gradient; n represents the total number of participants. This represents the global privacy weight of the j-th client, achieved by... The average value is obtained. This represents the privacy weight of the i-th data item on the client side; This represents the gradient clipping operation; This represents the local gradient of the j-th client; Indicates the gradient clipping threshold; represents the Gaussian noise function; 0 represents the expectation of the Gaussian noise, which takes the value of zero; The variance of Gaussian noise is represented; This represents the fundamental intensity of Gaussian noise;

[0179] The formula for the composite loss function is: ;

[0180] in, This represents the loss value of the composite loss function; The task loss is represented by the cross-entropy loss function; Indicates the smoothing coefficient; Indicates privacy constraints;

[0181] The federated collaborative training module is specifically used for:

[0182] Each client decrypts and decompresses the received online encrypted data packets using preset encryption and compression rules to obtain global model parameters. Each client then updates the global model parameters to its local model using the Adam optimizer and trains the local model using hardware acceleration technology until the preset training rounds or early stopping conditions are met.

[0183] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A privacy-weight-adaptive heterogeneous data federated collaborative training method, characterized in that: Includes the following steps: Step S1: The server initializes a global model, deploys the global model to each client as a local model, creates and trains a heterogeneous data encoding model to deploy to each client, and creates a knowledge graph-guided contrastive learning model. Step S2: Each client dynamically calculates the privacy weight of each data in the local dataset based on data sensitivity, participant trust score, and data distribution differences; The local dataset includes structured data, semi-structured data, or unstructured data; Step S3: Each client performs differential privacy operation on the local dataset based on the privacy weight to obtain a de-identified dataset. Through the deployed heterogeneous data encoding model, an encoding operation is performed on the de-identified dataset to obtain an initial structured encoding vector, an initial semi-structured encoding vector, or an initial unstructured encoding vector. Step S4: Each client performs a collaborative semantic alignment operation on the initial structured encoding vector, initial semi-structured encoding vector, or initial unstructured encoding vector through the contrastive learning model created by the server, to obtain aligned structured encoding vectors, aligned semi-structured encoding vectors, or aligned unstructured encoding vectors, and constructs an aligned vector set; Step S5: Each client trains the local model using the alignment vector set, records the local training log, generates the local gradient of the local model, extracts the local model parameters of the trained local model, compares the local model parameters extracted from the two training rounds with the local training log to obtain the local difference parameters, compresses and encrypts the privacy weights, local gradients and local difference parameters into a local encrypted data packet, and uploads the local encrypted data packet to the server based on the preset elastic participation strategy. Step S6: The server decrypts and decompresses the received local encrypted data packet to obtain privacy weights, local gradients, and local difference parameters. Based on each privacy weight, the server performs global aggregation on each local gradient to obtain the global gradient. Based on the received local difference parameters and global gradient, the server performs collaborative training on the global model. During the training process, the server optimizes the model using a preset composite loss function, extracts the global model parameters of the trained global model, compresses and encrypts the global model parameters into an online encrypted data packet, and sends the online encrypted data packet to each client. Step S7: Each client decrypts and decompresses the received online encrypted data packet to obtain global model parameters, and trains the local model based on the global model parameters until the preset training rounds or early stopping conditions are met.

2. The privacy-weighted adaptive heterogeneous data federated collaborative training method as described in claim 1, characterized in that: Step S1 specifically involves: The server initializes a global model based on a neural network, and deploys the global model to each client as a local model through a pre-configured federated gateway. It also creates and trains a heterogeneous data encoding model, deploys the heterogeneous data encoding model to each client through the federated gateway, and creates a knowledge graph-guided contrastive learning model. The heterogeneous data encoding model is constructed based on a structured data encoder, a semi-structured data encoder, and an unstructured data encoder. The structured data encoder is used to encode structured data through a first graph attention network to obtain a structured encoding vector. The semi-structured data encoder is used to encode semi-structured data through a temporal convolutional network and a Transformer to obtain a semi-structured encoding vector. The unstructured data encoder is used to encode unstructured data through a three-dimensional residual network to obtain an unstructured encoding vector. The contrastive learning model is constructed based on a knowledge graph guidance module, a contrastive learning module, and a collaborative optimization module. The knowledge graph guidance module is constructed based on an entity alignment unit, a relation reasoning unit, and a second graph attention network. The entity alignment unit is used to perform similarity matching between entity embedding vectors in the knowledge graph and structured encoding vectors, semi-structured encoding vectors, and unstructured encoding vectors, respectively, to generate entity association weights. The relation reasoning unit is used to reason about the potential semantic relationships between structured encoding vectors, semi-structured encoding vectors, and unstructured encoding vectors based on relation paths in the knowledge graph, to obtain semantic relation vectors. The second graph attention network is used to aggregate neighborhood node information in the knowledge graph based on the entity association weights and semantic relation vectors through a multi-head attention mechanism to generate graph embedding vectors. The contrastive learning module is constructed based on a cross-modal projection unit, a dynamic negative sampling unit, and a contrastive loss calculation unit. The cross-modal projection unit maps structured encoding vectors, semi-structured encoding vectors, unstructured encoding vectors, and corresponding graph embedding vectors to a unified contrastive space, performs cross-modal contrastive learning within the contrastive space, and constructs positive and negative sample pairs to perform semantic alignment operations. The dynamic negative sampling unit calculates the semantic similarity of each positive and negative sample pair, and dynamically selects difficult negative sample pairs from each negative sample pair based on the semantic similarity. The contrastive loss calculation unit calculates the contrastive loss between each positive and negative sample pair using the InfoNCE loss function, and adjusts the discriminative power between the positive and difficult negative sample pairs using the temperature coefficient in the InfoNCE loss function. The collaborative optimization module is built based on a gradient coordinator and an adaptive weight allocator. The gradient coordinator is used to balance the optimization direction of the guidance loss of the knowledge graph guidance module and the contrastive loss of the contrastive learning module through a gradient reversal mechanism. The adaptive weight allocator is used to dynamically adjust the loss weights of the guidance loss and the contrastive loss according to the alignment difficulty of structured encoding vectors, semi-structured encoding vectors and unstructured encoding vectors.

3. The privacy-weighted adaptive heterogeneous data federated collaborative training method as described in claim 1, characterized in that: In step S2, the formula for calculating the privacy weight is: ; ; in, This represents the privacy weight of the i-th data item on the client side. Indicates data sensitivity; Indicates the trust score of the participants; Indicates differences in data distribution; This represents the i-th data; This represents the smoothing coefficient used to prevent division by zero errors; and All represent adjustment coefficients; This indicates the number of personally identifiable information entries contained in the i-th data item; This represents the total amount of data for the i-th data item; This indicates the percentage of sensitive information. This represents the entropy value of the data distribution; In step S3, the differential privacy operation specifically involves: The Gaussian noise added to the local dataset is dynamically adjusted based on the privacy weights. The formula for the Gaussian noise is: ; in, represents the Gaussian noise function; 0 represents the expectation of the Gaussian noise, which takes the value of zero; The variance of Gaussian noise is represented; This represents the fundamental intensity of Gaussian noise; This represents the privacy weight of the i-th data point.

4. The privacy-weighted adaptive heterogeneous data federated collaborative training method as described in claim 1, characterized in that: Step S4 specifically involves: Each client compresses the initial vector set, including the initial structured encoding vector, the initial semi-structured encoding vector, or the initial unstructured encoding vector, into an initial vector compressed package using preset compression rules. The initial vector compressed package is then encrypted into an initial vector encrypted package using preset encryption rules. The initial vector encrypted package is then uploaded to the server through the federated gateway. The server decrypts and decompresses the initial vector encryption packets sent by each client using the preset encryption and compression rules to obtain an initial vector set. The server then calls the preset knowledge graph through the contrastive learning model to perform a collaborative semantic alignment operation on the initial vector set to obtain aligned structured encoded vectors, aligned semi-structured encoded vectors, or aligned unstructured encoded vectors, thereby constructing corresponding aligned vector sets for each client. The server compresses the alignment vector set into an alignment vector compressed package using the compression rules, encrypts the alignment vector compressed package into an alignment vector encrypted package using the encryption rules, and sends the alignment vector encrypted package to each client through the federated gateway. Each client decrypts and decompresses the aligned vector encryption packet sent by the server using the encryption and compression rules to obtain the aligned vector set; The compression rules are specifically as follows: CRC calculation is performed on the data to be compressed to obtain a CRC checksum. The data to be compressed is normalized using a preset normalization parameter to obtain normalized data. The normalized data is quantized using a preset quantization step size to obtain quantized data. Spatiotemporal differential encoding is performed on the quantized data to obtain differential coded data. A sparse matrix is ​​constructed based on the position and value of non-zero values ​​in the differential coded data. The sparse matrix is ​​divided into blocks using a preset block size to obtain sparse sub-matrices. Using a preset global dictionary, repeated sparse sub-matrices in the sparse matrix are replaced with corresponding short codes to obtain a first-level compressed sparse matrix. The short code is obtained by performing Huffman coding on the dictionary index, which is used to match sparse submatrices from the global dictionary; Run-length encoding is performed on the positions in the first-level compressed sparse matrix, and adaptive arithmetic encoding is performed on the values ​​to obtain the second-level compressed sparse matrix. The secondary compressed sparse matrix, CRC check code, normalization parameter, quantization step size, and global dictionary are encapsulated into a compressed data packet. The encryption rules are as follows: Generate a master key and a random nonce value of 12 bytes, obtain the data to be encrypted, calculate a 128-bit authentication tag for the data to be encrypted, the master key and the random nonce value using the HMAC algorithm, and perform SHA3-512 calculation on the random nonce value as the encryption seed; The master key is encrypted using the encryption seed via the Argon2id algorithm to obtain a first-level encryption key, and then the first-level encryption key is encrypted using the NTRUEncrypt quantum algorithm to obtain a second-level encryption key. The master key is called using the AES-256-GCM algorithm to encrypt the data to be encrypted into first-level encrypted data. The first-level encrypted data is Base85 encoded to obtain Base85 encoded data. The characters in the Base85 encoded data are replaced by a dynamic character replacement table to obtain second-level encrypted data. The secondary encrypted data, the secondary encryption key, the random Nonce value, and the 128-bit authentication tag are concatenated to obtain concatenated data. The concatenated data is then subjected to packet permutation encryption using a Feistel network structure to obtain an encrypted data packet. The dynamic character replacement table is stored on the blockchain platform in the form of a smart contract and is updated every 5 minutes based on preset update rules. The dynamic character replacement table is stored on the blockchain platform in the form of a smart contract. The update function of the smart contract is called periodically by an external script to update the dynamic character replacement table, and the hash value of the updated dynamic character replacement table is stored on the blockchain.

5. The privacy-weighted adaptive heterogeneous data federated collaborative training method as described in claim 1, characterized in that: Step S5 specifically involves: Each client uses hardware acceleration technology to train its local model by calling the alignment vector set. During training, it records local training logs in real time, including at least model information, training data information, model structure, model parameters, training process, hyperparameter adjustment process, and anomaly information. It generates local gradients of the local model through backpropagation algorithm, extracts local model parameters of the trained local model, and obtains local difference parameters by comparing the local model parameters extracted from the local training logs with those from the previous and next training rounds. It compresses the privacy weights, local gradients, and local difference parameters into a local compressed data packet using preset compression rules, encrypts the local compressed data packet into a local encrypted data packet using preset encryption rules, and uploads the local encrypted data packet to the server through the federated gateway based on a preset elastic participation strategy. The elastic participation strategy specifically involves: scoring the hardware performance of the client to obtain a performance score, and converting the performance score into the training frequency for participating in federated collaborative training; Step S6 specifically involves: The server decrypts and decompresses the local encrypted data packets sent by each client using preset encryption and compression rules to obtain privacy weights, local gradients, and local difference parameters. Based on the privacy weights, it globally aggregates the local gradients to obtain the global gradient and aggregates the local difference parameters uploaded by each client to obtain aggregated difference parameters. The server then uses the Adam optimizer to update the global model with the global gradients and aggregated difference parameters for collaborative training of the global model. During training, the server optimizes the model using a preset composite loss function and extracts the global model parameters after training. The server then compresses the global model parameters into an online compressed data packet using the compression rules and encrypts the online compressed data packet into an online encrypted data packet using the encryption rules. Finally, the server distributes the encrypted data packet to each client through the federated gateway. The formula for the global gradient is: ; in, This represents the global gradient; n represents the total number of participants. This represents the global privacy weight of the j-th client, achieved by... The average value is obtained. This represents the privacy weight of the i-th data item on the client side. This represents the gradient clipping operation; This represents the local gradient of the j-th client; Indicates the gradient clipping threshold; represents the Gaussian noise function; 0 represents the expectation of the Gaussian noise, which takes the value of zero; The variance of Gaussian noise is represented; This represents the fundamental intensity of Gaussian noise; The formula for the composite loss function is: ; in, This represents the loss value of the composite loss function; The task loss is represented by the cross-entropy loss function; Indicates the smoothing coefficient; Indicates privacy constraints; Step S7 specifically involves: Each client decrypts and decompresses the received online encrypted data packets using preset encryption and compression rules to obtain global model parameters. Each client then updates the global model parameters to its local model using the Adam optimizer and trains the local model using hardware acceleration technology until the preset training rounds or early stopping conditions are met.

6. A privacy-weight-adaptive heterogeneous data federated collaborative training system, characterized in that: Includes the following modules: The initialization module is used to initialize a global model on the server, deploy the global model to each client as a local model, create and train a heterogeneous data encoding model to deploy to each client, and create a knowledge graph-guided contrastive learning model. The privacy weight calculation module is used by each client to dynamically calculate the privacy weight of each data in the local dataset based on data sensitivity, trust scores of participants, and differences in data distribution. The local dataset includes structured data, semi-structured data, or unstructured data; The de-identification encoding module is used by each client to perform differential privacy operations on the local dataset based on the privacy weight to obtain a de-identified dataset. Through the deployed heterogeneous data encoding model, the module performs encoding operations on the de-identified dataset to obtain an initial structured encoding vector, an initial semi-structured encoding vector, or an initial unstructured encoding vector. The semantic alignment module is used by each client to perform a collaborative semantic alignment operation on the initial structured encoding vector, initial semi-structured encoding vector, or initial unstructured encoding vector through the contrastive learning model created by the server, to obtain aligned structured encoding vectors, aligned semi-structured encoding vectors, or aligned unstructured encoding vectors, and to construct an aligned vector set; The local training module is used by each client to train the local model through the alignment vector set, record local training logs, generate local gradients of the local model, extract local model parameters of the trained local model, obtain local difference parameters by comparing the local model parameters extracted in the two training rounds with the local training logs, compress and encrypt the privacy weights, local gradients and local difference parameters into a local encrypted data packet, and upload the local encrypted data packet to the server based on a preset elastic participation strategy. The online training module is used by the server to decrypt and decompress the received local encrypted data packets to obtain privacy weights, local gradients, and local difference parameters. Based on each privacy weight, the local gradients are globally aggregated to obtain the global gradient. Based on the received local difference parameters and global gradient, the global model is collaboratively trained. During the training process, optimization is performed through a preset composite loss function. The global model parameters of the trained global model are extracted, and the global model parameters are compressed and encrypted into an online encrypted data packet, which is then sent to each client. The federated collaborative training module is used by each client to decrypt and decompress the received online encrypted data packets to obtain global model parameters, and to train the local model based on the global model parameters until the preset training rounds or early stopping conditions are met.

7. The privacy-weighted adaptive heterogeneous data federated collaborative training system as described in claim 6, characterized in that: The initialization module is specifically used for: The server initializes a global model based on a neural network, and deploys the global model to each client as a local model through a pre-configured federated gateway. It also creates and trains a heterogeneous data encoding model, deploys the heterogeneous data encoding model to each client through the federated gateway, and creates a knowledge graph-guided contrastive learning model. The heterogeneous data encoding model is constructed based on a structured data encoder, a semi-structured data encoder, and an unstructured data encoder. The structured data encoder is used to encode structured data through a first graph attention network to obtain a structured encoding vector. The semi-structured data encoder is used to encode semi-structured data through a temporal convolutional network and a Transformer to obtain a semi-structured encoding vector. The unstructured data encoder is used to encode unstructured data through a three-dimensional residual network to obtain an unstructured encoding vector. The contrastive learning model is constructed based on a knowledge graph guidance module, a contrastive learning module, and a collaborative optimization module. The knowledge graph guidance module is constructed based on an entity alignment unit, a relation reasoning unit, and a second graph attention network. The entity alignment unit is used to perform similarity matching between entity embedding vectors in the knowledge graph and structured encoding vectors, semi-structured encoding vectors, and unstructured encoding vectors, respectively, to generate entity association weights. The relation reasoning unit is used to reason about the potential semantic relationships between structured encoding vectors, semi-structured encoding vectors, and unstructured encoding vectors based on relation paths in the knowledge graph, to obtain semantic relation vectors. The second graph attention network is used to aggregate neighborhood node information in the knowledge graph based on the entity association weights and semantic relation vectors through a multi-head attention mechanism to generate graph embedding vectors. The contrastive learning module is constructed based on a cross-modal projection unit, a dynamic negative sampling unit, and a contrastive loss calculation unit. The cross-modal projection unit maps structured encoding vectors, semi-structured encoding vectors, unstructured encoding vectors, and corresponding graph embedding vectors to a unified contrastive space, performs cross-modal contrastive learning within the contrastive space, and constructs positive and negative sample pairs to perform semantic alignment operations. The dynamic negative sampling unit calculates the semantic similarity of each positive and negative sample pair, and dynamically selects difficult negative sample pairs from each negative sample pair based on the semantic similarity. The contrastive loss calculation unit calculates the contrastive loss between each positive and negative sample pair using the InfoNCE loss function, and adjusts the discriminative power between the positive and difficult negative sample pairs using the temperature coefficient in the InfoNCE loss function. The collaborative optimization module is built based on a gradient coordinator and an adaptive weight allocator. The gradient coordinator is used to balance the optimization direction of the guidance loss of the knowledge graph guidance module and the contrastive loss of the contrastive learning module through a gradient reversal mechanism. The adaptive weight allocator is used to dynamically adjust the loss weights of the guidance loss and the contrastive loss according to the alignment difficulty of structured encoding vectors, semi-structured encoding vectors and unstructured encoding vectors.

8. The privacy-weighted adaptive heterogeneous data federated collaborative training system as described in claim 6, characterized in that: In the privacy weight calculation module, the formula for calculating the privacy weight is: ; in, This represents the privacy weight of the i-th data item on the client side. Indicates data sensitivity; Indicates the trust score of the participants; Indicates differences in data distribution; This represents the i-th data; This represents the smoothing coefficient used to prevent division by zero errors; and All represent adjustment coefficients; This indicates the number of personally identifiable information entries contained in the i-th data item; This represents the total amount of data for the i-th data item; This indicates the percentage of sensitive information. This represents the entropy value of the data distribution; In the de-identification encoding module, the differential privacy operation specifically includes: The Gaussian noise added to the local dataset is dynamically adjusted based on the privacy weights. The formula for the Gaussian noise is: ; in, represents the Gaussian noise function; 0 represents the expectation of the Gaussian noise, which takes the value of zero; The variance of Gaussian noise is represented; This represents the fundamental intensity of Gaussian noise; This represents the privacy weight of the i-th data point.

9. A privacy-weight-adaptive heterogeneous data federated collaborative training system as described in claim 6, characterized in that: The semantic alignment module is specifically used for: Each client compresses the initial vector set, including the initial structured encoding vector, the initial semi-structured encoding vector, or the initial unstructured encoding vector, into an initial vector compressed package using preset compression rules. The initial vector compressed package is then encrypted into an initial vector encrypted package using preset encryption rules. The initial vector encrypted package is then uploaded to the server through the federated gateway. The server decrypts and decompresses the initial vector encryption packets sent by each client using the preset encryption and compression rules to obtain an initial vector set. The server then calls the preset knowledge graph through the contrastive learning model to perform a collaborative semantic alignment operation on the initial vector set to obtain aligned structured encoded vectors, aligned semi-structured encoded vectors, or aligned unstructured encoded vectors, thereby constructing corresponding aligned vector sets for each client. The server compresses the alignment vector set into an alignment vector compressed package using the compression rules, encrypts the alignment vector compressed package into an alignment vector encrypted package using the encryption rules, and sends the alignment vector encrypted package to each client through the federated gateway. Each client decrypts and decompresses the aligned vector encryption packet sent by the server using the encryption and compression rules to obtain the aligned vector set; The compression rules are specifically as follows: CRC calculation is performed on the data to be compressed to obtain a CRC checksum. The data to be compressed is normalized using a preset normalization parameter to obtain normalized data. The normalized data is quantized using a preset quantization step size to obtain quantized data. Spatiotemporal differential encoding is performed on the quantized data to obtain differential coded data. A sparse matrix is ​​constructed based on the position and value of non-zero values ​​in the differential coded data. The sparse matrix is ​​divided into blocks using a preset block size to obtain sparse sub-matrices. Using a preset global dictionary, repeated sparse sub-matrices in the sparse matrix are replaced with corresponding short codes to obtain a first-level compressed sparse matrix. The short code is obtained by performing Huffman coding on the dictionary index, which is used to match sparse submatrices from the global dictionary; Run-length encoding is performed on the positions in the first-level compressed sparse matrix, and adaptive arithmetic encoding is performed on the values ​​to obtain the second-level compressed sparse matrix. The secondary compressed sparse matrix, CRC check code, normalization parameter, quantization step size, and global dictionary are encapsulated into a compressed data packet. The encryption rules are as follows: Generate a master key and a random nonce value of 12 bytes, obtain the data to be encrypted, calculate a 128-bit authentication tag for the data to be encrypted, the master key and the random nonce value using the HMAC algorithm, and perform SHA3-512 calculation on the random nonce value as the encryption seed; The master key is encrypted using the encryption seed via the Argon2id algorithm to obtain a first-level encryption key, and then the first-level encryption key is encrypted using the NTRUEncrypt quantum algorithm to obtain a second-level encryption key. The master key is called using the AES-256-GCM algorithm to encrypt the data to be encrypted into first-level encrypted data. The first-level encrypted data is Base85 encoded to obtain Base85 encoded data. The characters in the Base85 encoded data are replaced by a dynamic character replacement table to obtain second-level encrypted data. The secondary encrypted data, the secondary encryption key, the random Nonce value, and the 128-bit authentication tag are concatenated to obtain concatenated data. The concatenated data is then subjected to packet permutation encryption using a Feistel network structure to obtain an encrypted data packet. The dynamic character replacement table is stored on the blockchain platform in the form of a smart contract and is updated every 5 minutes based on preset update rules. The dynamic character replacement table is stored on the blockchain platform in the form of a smart contract. The update function of the smart contract is called periodically by an external script to update the dynamic character replacement table, and the hash value of the updated dynamic character replacement table is stored on the blockchain.

10. A privacy-weighted adaptive heterogeneous data federated collaborative training system as described in claim 6, characterized in that: The local training module is specifically used for: Each client uses hardware acceleration technology to train its local model by calling the alignment vector set. During training, it records local training logs in real time, including at least model information, training data information, model structure, model parameters, training process, hyperparameter adjustment process, and anomaly information. It generates local gradients of the local model through backpropagation algorithm, extracts local model parameters of the trained local model, and obtains local difference parameters by comparing the local model parameters extracted from the local training logs with those from the previous and next training rounds. It compresses the privacy weights, local gradients, and local difference parameters into a local compressed data packet using preset compression rules, encrypts the local compressed data packet into a local encrypted data packet using preset encryption rules, and uploads the local encrypted data packet to the server through the federated gateway based on a preset elastic participation strategy. The elastic participation strategy specifically involves: scoring the hardware performance of the client to obtain a performance score, and converting the performance score into the training frequency for participating in federated collaborative training; The online training module is specifically used for: The server decrypts and decompresses the local encrypted data packets sent by each client using preset encryption and compression rules to obtain privacy weights, local gradients, and local difference parameters. Based on the privacy weights, it globally aggregates the local gradients to obtain the global gradient and aggregates the local difference parameters uploaded by each client to obtain aggregated difference parameters. The server then uses the Adam optimizer to update the global model with the global gradients and aggregated difference parameters for collaborative training of the global model. During training, the server optimizes the model using a preset composite loss function and extracts the global model parameters after training. The server then compresses the global model parameters into an online compressed data packet using the compression rules and encrypts the online compressed data packet into an online encrypted data packet using the encryption rules. Finally, the server distributes the encrypted data packet to each client through the federated gateway. The formula for the global gradient is: ; in, This represents the global gradient; n represents the total number of participants. This represents the global privacy weight of the j-th client, achieved by... The average value is obtained. This represents the privacy weight of the i-th data item on the client side. This represents the gradient clipping operation; This represents the local gradient of the j-th client; Indicates the gradient clipping threshold; represents the Gaussian noise function; 0 represents the expectation of the Gaussian noise, which takes the value of zero; The variance of Gaussian noise is represented; This represents the fundamental intensity of Gaussian noise; The formula for the composite loss function is: ; in, This represents the loss value of the composite loss function; The task loss is represented by the cross-entropy loss function; Indicates the smoothing coefficient; Indicates privacy constraints; The federated collaborative training module is specifically used for: Each client decrypts and decompresses the received online encrypted data packets using preset encryption and compression rules to obtain global model parameters. Each client then updates the global model parameters to its local model using the Adam optimizer and trains the local model using hardware acceleration technology until the preset training rounds or early stopping conditions are met.

Citation Information

Patent Citations

  • Federal learning model training method and device, electronic equipment and storage medium

    CN118153105A

  • Trans-department data collaborative modeling method and system based on federal learning

    CN120408727A