Multi-omics fusion and communication-restricted federated learning method based on diffusion model, system and application
Patent Information
- Application Number
- CN202611006824.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-08
AI Technical Summary
在联邦学习上行阶段,空中计算利用无线叠加特性可在一次同步上行中完成多终端求和聚合,从而显著降低通信时延与调度复杂度;然而其接收结果天然包含信道失配、同步误差与噪声叠加等引起的聚合失真
1. 通过跨模态注意力实现内容自适应的深度对齐;通过“条件扩散重构”对统一潜表示施加生成先验,在特征层直接实现缺失模态的补全与噪声抑制,显著提升融合表示的完整性、判别性与鲁棒性。
Smart Images

Figure CN122511339B_ABST
Abstract
Description
Technical Field
[0001] This invention application relates to the field of artificial intelligence, and more particularly to a multi-omics data fusion method and its application in communication-restricted federated learning for cancer subtype identification. Background Technology
[0002] Lung adenocarcinoma is an important type of non-small cell lung cancer, and different molecular subtypes show significant differences in prognosis, targeted therapy, and immunotherapy response. Clinically, relying solely on a single omics (such as transcriptomics or methylation) for subtyping often leads to blurred subtype boundaries due to incomplete information, noise interference, and heterogeneity. Multi-omics tasks are valuable in lung adenocarcinoma subtype identification, but real-world scenarios suffer from modality loss, noise, and heterogeneity. Multi-omics fusion can simultaneously utilize complementary information such as transcriptomics (RNA expression), epigenetics (DNA methylation), genomic structural variation information (such as CNV), and clinical characteristics to improve the stability and robustness of subtyping; however, existing fusion strategies still have shortcomings. Most methods fail to adequately characterize the differences in contributions of different omics during the fusion stage, lack an interpretable and trainable "importance / confidence" modeling mechanism, and struggle to adaptively assign weights to "strong signal omics" and "weak signal omics," thus easily introducing fusion bias under conditions of heterogeneous samples and noisy omics. Many methods rely on similarity matrix fusion or shallow feature splicing, lacking end-to-end cross-omics semantic alignment and generative reconstruction constraints. They are difficult to form a stable unified representation under real data forms such as "missing modalities, partial missing features, and noise pollution", which in turn limits the performance of downstream subtype discrimination. In real-world applications, multi-omics data is often scattered across multiple medical institutions, making centralized training difficult due to privacy and compliance requirements. Federated learning can achieve cross-institutional collaborative modeling without data leaving the domain, but it still faces challenges in wireless network environments, such as high communication overhead, uplink aggregation distortion caused by link noise / fading, and unstable training convergence. These problems are particularly pronounced when the fusion model has a large number of parameters and many participating institutions.
[0003] Task-oriented communication emphasizes a transmission philosophy centered on task performance, meaning it no longer prioritizes bit-by-bit reliable recovery as the sole objective, but rather focuses on optimizing the final learning / inference task metrics (such as subtype classification accuracy, convergence speed, and stability). In the federated learning uplink phase, over-the-air computation leverages wireless overlay capabilities to perform multi-terminal summation and aggregation in a single synchronous uplink, significantly reducing communication latency and scheduling complexity. However, the received results naturally contain aggregation distortion caused by channel mismatch, synchronization errors, and noise superposition. Directly applying this aggregated distortion to global model updates can easily lead to oscillations in the optimization trajectory, slower convergence, or even divergence. Furthermore, traditional schemes relying on strict power control or channel coding to achieve "bit-level reliability" often further increase latency and system complexity, making it difficult to balance communication efficiency and learning performance. Existing technologies typically handle "feature fusion robustness" and "aggregate update resistance to distortion" separately, lacking a unified collaborative mechanism oriented towards the same task objective.
[0004] The diffusion model learns the prior knowledge of data distribution generation through a Markov chain process of "stepwise noise addition and stepwise noise reduction," demonstrating good feasibility and robustness in signal reconstruction, missing data completion, and noise correction. Based on this, existing technologies urgently need a fusion and transmission collaboration mechanism capable of addressing lung adenocarcinoma subtype discrimination tasks and adapting to communication-constrained federated training scenarios. On the one hand, on the client side, conditional diffusion reconstruction constraints are introduced into the unified latent representation obtained from cross-omics fusion, enabling latent representation completion and feature purification even under missing or noisy omics conditions, thus enhancing the fidelity of information useful for subtype discrimination. On the other hand, on the server side, aerial aggregation is treated as "observations with structural perturbations," introducing an aggregation distortion reconstruction module centered on diffusion denoising. Communication statistics such as signal-to-noise ratio, quantization bit width, aggregation variance, and the number of participating clients can be used as conditional inputs to achieve adaptive correction for different link qualities and compression intensities, thereby maintaining the convergence quality of federated training and improving lung adenocarcinoma subtype identification performance under communication-constrained conditions. Summary of the Invention
[0005] To address the aforementioned shortcomings in existing technologies and to resolve the core contradiction of "difficult data fusion" and "difficult privacy communication" in multi-center medical collaboration, this invention aims to provide a multi-omics fusion and communication-constrained federated learning method, system, and application based on a diffusion model. This achieves a federated multi-omics learning technical solution of "diffusion prior + task-oriented compressed transmission + over-the-air aggregation + diffusion denoising and aggregation correction," thereby simultaneously improving the quality of fusion representation under missing / noisy modal conditions and the stability of federated training under disturbed aggregation update conditions.
[0006] This invention discloses a multi-omics fusion and communication-constrained federated learning method based on a diffusion model, the method specifically comprising the following steps: S1: The local client obtains data samples from multiple omics systems; S2: Input the multi-omics data samples into a multimodal fusion network to obtain a unified latent representation; the multimodal fusion network includes a modality-specific encoder, a cross-modal fusion module, and a conditional diffusion reconstruction module; S3: Based on the unified latent representation, the task prediction result is obtained through a classifier; S4: Calculate the local model update amount on the client, perform task-oriented compression on the local model update amount, obtain the compressed signal, and upload it via the wireless channel; S5: Through over-the-air computation, synchronous uplink signals from multiple clients are received and superimposed in the same time slot to obtain the equivalent aggregated update amount; S6: Input the equivalent aggregation update amount into the diffusion denoising aggregation correction module on the server side to reconstruct the corrected aggregation update amount; S7: Update the global model using the corrected aggregate update amount, and broadcast the updated global model to each client.
[0007] Preferably, step S2 further includes: S201: Input the data of each omics in the multi-omics data sample into the corresponding modality-specific encoder to obtain the feature representation of each modality; S202: Project and concatenate the feature representations of each modality, and add modality type embedding and position encoding to form a unified token; S203: Input the unified token into the cross-modal attention backbone of the cross-modal fusion module, perform feature fusion and alignment through the cross-modal attention mechanism, and output the unified latent representation; S204: Apply conditional diffusion reconstruction priors to the unified latent representation to achieve missing mode completion and noise suppression.
[0008] Preferably, step S204 further includes: S2041: Receive the unified latent representation output by the cross-modal fusion module; S2042: Perform a forward diffusion noise addition process on the unified latent representation; S2043: Based on conditional information, a reverse denoising and reconstruction process is performed on the noisy latent representation through a denoising network to obtain the reconstructed latent representation; wherein, the conditional information includes at least a modal availability mask for indicating the availability status of each omics data. S2044: When the modality availability mask indicates the presence of missing omics data, the corresponding missing part of the unified latent representation is completed during the reverse denoising reconstruction process; S2045: Output the reconstructed latent representation for the classifier to make predictions.
[0009] Preferably, step S4 further includes: S401: Define the local model update amount as the local increment after each client trains its local model after the server broadcasts the global model in each round. S402: Perform importance-aware group sparsification on the local model update amount to obtain the sparsified update amount; S403: Perform residual encoding and quantization on the sparsification update amount to generate a digitized compressed signal; S404: The compressed signal is mapped to an analog transmission symbol and synchronously transmitted uplink in the agreed communication time slot so that the server can receive and superimpose it through over-the-air computation.
[0010] Preferably, the diffusion denoising aggregation correction module on the server side in step S6 is configured as follows: Using the equivalent aggregate update amount as the noisy input and the ideal aggregate update amount as the reconstruction target, train the noise prediction network. During inference, the currently received equivalent aggregate update and the current communication statistics are used as conditions to perform few-step denoising through the noise prediction network, and the corrected aggregate update is output.
[0011] Preferably, the total training loss function of the multimodal fusion network further includes a cross-modal contrastive loss term; the cross-modal contrastive loss term is calculated based on the InfoNCE loss function, and its construction method includes: The coding features of the first group were used as anchor features; The reconstructed features output by the conditional diffusion reconstruction module of the second group of learning are used as positive sample features; The second omics reconstruction features of other samples in the same training batch are used as negative sample features; The cross-modal contrast loss term is obtained by calculating the contrast loss between the anchor feature, the positive sample feature, and multiple negative sample features.
[0012] To implement the above method, this application provides a multi-omics fusion and communication-constrained federated learning system based on a diffusion model. The system includes multiple clients and a central server. The client includes: A multimodal fusion network is used to extract and fuse features from local multi-omics data samples and output task prediction results. The multimodal fusion network integrates a conditional diffusion reconstruction module. The local training module is used to calculate the model loss and the amount of local model updates based on the task prediction results. A communication compression module is used to perform task-oriented compression of the local model update data and upload it. The server includes: The over-the-air aggregation receiving module is used to receive and superimpose the signals uploaded by each client through over-the-air computing to obtain the equivalent aggregated update amount; The aggregation correction module includes a diffusion denoiser, used to reconstruct and correct the equivalent aggregation update amount to obtain the corrected aggregation update amount; The global update module is used to update the global model using the corrected aggregate update amount and distribute it to each client.
[0013] This application also provides an application of a diffusion-based multi-omics fusion and communication-restricted federated learning method, which is applied to the auxiliary identification of lung adenocarcinoma subtypes. The task prediction result is lung adenocarcinoma subtype discrimination or classification. The multimodal fusion network includes a modality-specific encoder, specifically comprising: Transcriptome (RNA) encoder E rna A multilayer perceptron (MLP) block encoder is used, with gene expression vectors as input; DNA methylation encoder E meth The encoder end of the sparse high-dimensional MLP encoder or autoencoder is used, and the input is the methylation intensity vector of the site or region. CNV encoder E cnv An MLP encoder is used, with the input being gene-level or fragment-level CNV vectors; Clinical Variable Encoder E cli It adopts an embedding + MLP structure, and the input includes structured variables containing individual patient information and medical records.
[0014] The beneficial effects of this invention are as follows: 1. Achieve content-adaptive deep alignment through cross-modal attention; apply generative priors to the unified latent representation through "conditional diffusion reconstruction" to directly achieve missing modal completion and noise suppression at the feature layer, significantly improving the integrity, discriminability and robustness of the fused representation.
[0015] 2. At the communication and aggregation level: The client compresses the model update by "importance-aware group sparsity and residual quantization" and uploads it synchronously through "over-the-air computing" technology to achieve efficient aggregation; the server treats the received distorted aggregated update as a noisy observation and sends it to another diffusion denoiser conditioned on communication statistics for correction in order to restore the near-ideal model update and ensure training stability under poor channel conditions.
[0016] 3. Under the premise of strictly adhering to the principle that data does not leave the local machine (federated learning), this approach significantly improves the accuracy of multi-center lung adenocarcinoma molecular subtype identification, robustness to missing and noisy data, and overall communication and training efficiency. This solution can be extended to other multimodal medical diagnostic and distributed learning scenarios. Attached Figure Description
[0017] Figure 1 This is the overall system architecture diagram of the present invention.
[0018] Figure 2 This is a schematic diagram of the diffusion fusion network structure and diffusion process.
[0019] Figure 3 This is a schematic diagram of the communication and aggregation correction process for a single round of federated training. Detailed Implementation
[0020] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0021] This invention discloses a multi-omics fusion and communication-constrained federated learning method based on a diffusion model, the method specifically comprising the following steps: S1: The local client obtains data samples from multiple omics systems; S2: Input the multi-omics data samples into a multimodal fusion network to obtain a unified latent representation; the multimodal fusion network includes a modality-specific encoder, a cross-modal fusion module, and a conditional diffusion reconstruction module; S3: Based on the unified latent representation, the task prediction result is obtained through a classifier; S4: Calculate the local model update amount on the client, perform task-oriented compression on the local model update amount, obtain the compressed signal, and upload it via the wireless channel; S5: Through over-the-air computation, synchronous uplink signals from multiple clients are received and superimposed in the same time slot to obtain the equivalent aggregated update amount; S6: Input the equivalent aggregation update amount into the diffusion denoising aggregation correction module on the server side to reconstruct the corrected aggregation update amount; S7: Update the global model using the corrected aggregate update amount, and broadcast the updated global model to each client.
[0022] Preferably, step S2 further includes: S201: Input the data of each omics in the multi-omics data sample into the corresponding modality-specific encoder to obtain the feature representation of each modality; S202: Project and concatenate the feature representations of each modality, and add modality type embedding and position encoding to form a unified token; the token is the smallest basic feature unit for the model to process multimodal information and is used for subsequent feature fusion and semantic modeling.
[0023] S203: Input the unified token into the cross-modal attention backbone of the cross-modal fusion module, perform feature fusion and alignment through the cross-modal attention mechanism, and output the unified latent representation; S204: Applying conditional diffusion to reconstruct priors on unified latent representations, missing mode completion, and noise suppression.
[0024] This application also provides an application of a multi-omics fusion and communication-restricted federated learning method based on a diffusion model. The above method is applied to the auxiliary identification of lung adenocarcinoma subtypes, and the task prediction result is the lung adenocarcinoma subtype discrimination or classification. This embodiment takes lung adenocarcinoma subtype identification as an example to further illustrate the technical solution of this application.
[0025] In step 201 of this embodiment, the multimodal fusion network includes a modality-specific encoder, specifically comprising: Transcriptome (RNA) encoder E rna A multilayer perceptron (MLP) group encoder is used, with gene expression vectors as input; the DNA methylation encoder E... meth The encoder employs a sparse high-dimensional MLP encoder or an autoencoder, with the input being a methylation intensity vector of a site or region; CNV encoder E cnv An MLP encoder is used, with input being gene-level or fragment-level CNV vectors; the clinical variable encoder E... cli It adopts an embedding + MLP structure, and the input includes structured variables containing individual patient information and medical records.
[0026] Specifically, in this embodiment, the modality-specific encoder adopts the following method: Transcriptome (RNA) encoder E rna RNA modal input is gene expression vector G rna The number of genes is represented. Preferably, the expression level is transformed by log(1+x) and then normalized within or across samples before being input into the encoder.
[0027] In this embodiment, a multilayer perceptron (MLP) group encoder is preferably used as the E... rna Specifically, the process involves first grouping and aggregating genes according to a preset gene set / pathway to obtain L. rna Each input group is then mapped to dimension d using either a shared or non-shared MLP, thus obtaining...
[0028] Among them, L rna优选 32–256; the preferred number of MLP layers is 2–4; the preferred width of the hidden layer is 512–4096; the preferred activation function is GELU / ReLU; the preferred dropout value is 0.1–0.5; the preferred normalization method is LayerNorm / BatchNorm.
[0029] DNA methylation encoder E meth In an embodiment of the present invention, the methylation mode input is a methylation intensity vector of a site or region. Since methylation dimensionality is typically high, this invention preferably performs site / region screening and grouping aggregation (e.g., variance-based screening, differential screening, or aggregation by gene region) on the client side, compressing the high-dimensional input to L... meth Each group-level feature is then input into the encoder.
[0030] The present invention preferably uses a sparse high-dimensional MLP encoder or an autoencoder as the encoding end of the E. meth Mapping group-level inputs to a unified dimension d yields...
[0031] Among them, L meth优选 32–256; MLP layer count preferably 2–5; hidden layer width preferably 1024–8192; Dropout preferably 0.1–0.5. The encoder adopts a stacked structure of "linear layer + normalization + activation" to enhance robustness to noise and batch effects.
[0032] CNV encoder E cnv In this embodiment of the invention, the CNV modal input is a gene-level or fragment-level CNV vector. Its elements can be copy values, logarithmic ratios, or discrete states.
[0033] The present invention preferably uses an MLP encoder as the E cnv It can optionally segment CNV features by chromosome or gene set, and output token sequence representation:
[0034] Among them, L cnv优选 16–128; the preferred number of MLP layers is 2–4; the preferred width of the hidden layer is 256–2048; the preferred activation function is GELU / ReLU.
[0035] Clinical Variable Encoder E cli In one embodiment of the present invention, the clinical modality input includes structured variables (including age, gender, stage, smoking history, past medical history, etc.). Continuous variables are preferably standardized, and categorical variables are preferably embedded. The present invention preferably uses an embedding + MLP structure as the E... cli Output clinical token sequence representation:
[0036] Among them, L cli可取值为 1–32, when using a single token representation, L cli =1; When clinical variables are grouped by field to form multiple tokens, L cli The preferred values are 8–32. The preferred number of MLP layers is 1–3; the preferred width of the hidden layer is 128–1024.
[0037] In a preferred embodiment of the present invention, step 202 involves performing uniform dimension alignment on the outputs of each modal encoder to ensure consistent input to the cross-modal fusion backbone network. Since each modal encoder in this invention directly outputs a token sequence of dimension d, these sequences can be directly concatenated to form the total token sequence.
[0038] Further research on H k Modality type embedding and position encoding are superimposed to distinguish different modalities and token order, thus serving as input to the subsequent cross-modality attention fusion backbone.
[0039] Among them, the preferred values for the unified token dimension d are 256, 512, or 768; on the resource-constrained client side, 256 or 512 are preferred; on the server side, which pursues expressive power or under conditions of sufficient computing power, 768 is preferred.
[0040] In a preferred embodiment of the present invention, step 203 specifically involves using a cross-modal attention backbone F(·) to output a unified latent representation z0, and achieving content relevance and geometric adaptive alignment through a sparse sampling set.
[0041]
[0042]
[0043]
[0044] Based on a unified token representation U, a cross-modal attention fusion backbone F(·) is used to output a unified latent representation z0. The fusion backbone preferably includes a deformable cross-modal attention mechanism: for the i-th query vector q...i Instead of performing intensive attention computation on the entire set of key-value pairs, it determines a sparse sampling set S(i) in the key-value pairs based on content relevance and (optionally) learnable reference points / offsets, and computes the attention weight α only within this set. ij It also completes information aggregation, thereby obtaining a more cross-modal consistent fusion representation.
[0045] Semantic cross-modal alignment: By establishing the association between different omics / modal tokens in a unified token space through cross-modal attention, the consistency and discriminativeness of the fused representation are improved.
[0046] Sparse, content-related adaptive aggregation: Attention is calculated only within S(i) for each query, enabling the model to dynamically focus on the cross-modal information most relevant to the current token and reduce interference from irrelevant / noisy tokens.
[0047] Geometric / structural adaptation (deformable properties): Through learnable sampling positions / offsets (or equivalent dynamic index selection), the attention receptive field is adaptively adjusted according to changes in sample content and modal structure, thereby improving alignment robustness under modal differences, token arrangement differences, and missing / noise conditions.
[0048] Computational complexity is controllable: O(L) time complexity of full attention. 2 The computational cost is approximately reduced to O(L⋅∣S(i)∣), which reduces computational overhead and improves deployability while maintaining fusion capabilities.
[0049] Specifically, in step S204, a conditional diffusion model is introduced into the latent representation layer. In this embodiment of the invention, the objective is to discriminate / classify lung adenocarcinoma subtypes, with the output being a subtype probability distribution. Training primarily uses classification cross-entropy loss, with diffusion reconstruction loss and contrast loss used as regularization constraints. This invention does not use unsupervised clustering results as the final output.
[0050] Forward noise addition process:
[0051]
[0052] Inverse denoising employs noise prediction parameterization, where condition c can include modality availability mask, clinical prior embedding, noise intensity, etc.
[0053]
[0054]
[0055] The training objective is expressed in the form of mean squared error:
[0056] Further includes: S2041: Receive the unified latent representation output by the cross-modal fusion module; Specifically, features are extracted from the multi-omics / multimodal inputs of the same sample by corresponding encoders, and a unified latent representation z0 is obtained through a cross-modal attention fusion backbone network. Unlike the shared / specific latent variable inference based on variational autoencoders, this invention preferably uses a cross-modal attention fusion backbone to interactively align the tokens of each modality to form a unified latent representation z0. Then, a conditional diffusion prior is applied to this unified latent representation. The unified latent representation can be in the form of a token sequence or a vector, and is used to characterize the comprehensive information of the sample in the multimodal semantic space.
[0057] S2042: Perform a forward diffusion noise addition process on the unified latent representation; This step involves performing a T-step forward diffusion process on the unified latent representation z0, gradually injecting Gaussian noise to obtain the noise latent variable z. t The noise intensity is determined by the diffusion time step t and a preset noise schedule parameter. The forward process is used to construct a continuous perturbation trajectory "from clean to noisy" so as to train the inverse denoising network to learn the prior data.
[0058] Specifically, let the number of diffusion steps be T, and the noise sequence be {β}. t}, and define α t With cumulative product ᾱ t .
[0059]
[0060] Forward noise addition process:
[0061]
[0062] S2043: Based on conditional information, a reverse denoising and reconstruction process is performed on the noisy latent representation through a denoising network to obtain the reconstructed latent representation; wherein, the conditional information includes at least a modal availability mask for indicating the availability status of each omics data. Configure the inverse denoising network ε on the server or client side (depending on the deployment method). θ (⋅), where the input is the noise latent variable z. t Given time step t and condition information c, the output is the predicted value of the noise term (or equivalent denoising direction information), and the reconstruction result z^0 of the clean latent representation is obtained accordingly.
[0063]
[0064] Wherein, the condition information c includes at least: Modal availability mask: used to indicate whether each omics has missing data, the location of the missing data, and the corresponding token range, and may include: 1) Prior clinical information: such as structured variables like age, gender, smoking history, stage, and medical history; 2) Noise statistics or communication statistics: such as signal-to-noise ratio, quantization bit width, aggregation variance, number of participating clients, etc., are used to enhance reconstruction stability under different noise environments. (Note: Those skilled in the art should understand that the noise level of the diffusion process is characterized by time step t, and the conditional information c is used to introduce external prior or observational environment information, and is not limited to "noise intensity" itself.) 3) In the context of constrained communication in federated learning, condition c can further include statistics related to transmission / aggregation distortion (such as SNR, quantization bit width, aggregation variance, and number of participating clients) to improve the robustness of reconstruction under non-ideal communication conditions.
[0065] 4) In the implementation that includes server-side aggregation correction, the conditional diffusion reconstruction module and the server-side aggregation denoising diffusion module share or collaboratively optimize noise priors and conditional representations, so that feature layer reconstruction and update layer aggregation correction form an end-to-end consistent robust training closed loop.
[0066] S2044: When the modality availability mask indicates the presence of missing omics data, the corresponding missing part of the unified latent representation is completed during the reverse denoising reconstruction process; When a missing omics / modality is detected, the latent representation substructure corresponding to the missing modality (e.g., missing modality token subsequence or missing feature block) is taken as the part to be reconstructed. Under the constraints of the latent representations of the other observed modalities and the conditional information c, the latent representation of the missing part is generated through the diffusion reconstruction process in steps S2042–S2043, thereby realizing the completion of the missing modality in the feature layer / latent space.
[0067] Preferably, the completion result does not require the recovery of the original modality data (such as the original image pixels or the original sequencing reads), but rather the recovery of an effective feature representation for downstream tasks, in order to ensure privacy and computational efficiency.
[0068] S2045: Output the reconstructed latent representation for the classifier to make predictions.
[0069] The downstream discrimination task preferably uses a Softmax classification head to output the probabilities of each subtype, and uses cross-entropy as the main optimization objective to reconstruct the probability. As input for downstream discrimination tasks, it is used for lung adenocarcinoma subtype classification. The diffusion denoising process has a statistically significant inhibitory effect on random perturbations and system noise, which can improve the signal-to-noise ratio and robustness of the fused representation. Furthermore, the diffusion reconstruction loss can be introduced as a generative regularization term into the training objective to constrain the distribution of the fused representation to closely approximate the data prior, thereby improving the representation fidelity and generalization ability while improving classification performance.
[0070] Pooling is performed on the reconstructed clean latent representation ẑ0 (Pool (ẑ0)), the input is the classification head C (・), and the output is the probability of each subtype of lung adenocarcinoma through Softmax; pooling can be [CLS] token, mean pooling, or attention pooling; the classification head C (·) is usually a linear layer or a two-layer MLP; the joint loss is: L total = L cls (Cross-entropy classification loss) + λ diff •L diff (Diffusion reconstruction loss) +λ con •L con (Cross-modal contrast alignment loss).
[0071] In this embodiment, the core task is directly implemented: the classifier converts the fused features into "subtype probabilities", which directly meets the clinical needs of "lung adenocarcinoma molecular subtype identification" - for example, outputting a clear result of "subtype A probability 85% and subtype B probability 12%".
[0072] Preferably, step S4 involves uploading the local update volume from the client to the server via task-oriented compression and over-the-air aggregation.
[0073] In this embodiment, step S4 further includes: S401: Define the local model update amount as the local increment after each client trains its local model after the server broadcasts the global model in each round. S402: Perform importance-aware group sparsification on the local model update amount to obtain the sparsified update amount; S403: Perform residual encoding and quantization on the sparsification update amount to generate a digitized compressed signal; S404: The compressed signal is mapped to an analog transmission symbol and synchronously transmitted uplink in the agreed communication time slot so that the server can receive and superimpose it through over-the-air computation.
[0074] Regarding step S401, in this embodiment, after the server broadcasts the global model w_t in the t-th round, the client trains locally to obtain the local model w_(k,t), defines the local increment, and gives the ideal aggregation form.
[0075]
[0076] This allows the client to transmit only relative updates without needing to transmit complete model parameters, thus providing a unified operation object for subsequent sparsification, quantization, and task-oriented compression, and reducing the uplink communication burden.
[0077]
[0078] Where p k Aggregate weights (preferably proportional to the local sample size or taking...) The aforementioned It serves as a reference target for subsequent task-oriented compressed transmission, over-the-air aggregation, and server-side noise correction, used to measure the deviation caused by compression and channel distortion and guide the design of correction strategies.
[0079] Quantify the value of local training: Transform the client's "local training results" into "model increments" that can be transferred and aggregated - no longer need to transfer the complete local model, only the "updated part" is transferred, reducing the amount of data transmission from the source.
[0080] For step S404, the quantization vector is mapped to an analog transmission symbol and synchronized uplink. The server receives the superimposed signal and calibrates and scales it to obtain an equivalent aggregation update.
[0081]
[0082]
[0083]
[0084] The diffusion denoising aggregation correction module on the server side in step S6 is configured as follows: Using the equivalent aggregate update amount as the noisy input and the ideal aggregate update amount as the reconstruction target, train the noise prediction network. During inference, the currently received equivalent aggregate update and the current communication statistics are used as conditions to perform few-step denoising through the noise prediction network, and the corrected aggregate update is output.
[0085] The aggregated update observed by the server can be regarded as the sum of the ideal aggregate and the equivalent perturbation, and a diffusion denoiser is introduced to output the corrected update.
[0086]
[0087]
[0088] The ideal update is treated as a "clean sample". A diffusion chain is constructed and a noise prediction network is trained. The training loss is in the form of mean square error. Inference uses a few-step backward denoising to control the latency.
[0089] Preferably, the total training loss function of the multimodal fusion network further includes a cross-modal contrastive loss term; the cross-modal contrastive loss term is calculated based on the InfoNCE loss function, and its construction method includes: The coding features of the first group were used as anchor features; The reconstructed features output by the conditional diffusion reconstruction module of the second group of learning are used as positive sample features; The second omics reconstruction features of other samples in the same training batch are used as negative sample features; The cross-modal contrast loss term is obtained by calculating the contrast loss between the anchor feature, the positive sample feature, and multiple negative sample features.
[0090] This step introduces cross-modal InfoNCE contrast constraints to improve the discriminativeness and semantic consistency of the fused representation. Using the representation of modality A as the anchor point, the diffused reconstructed representation of modality B is considered a positive sample, and other samples within the batch are considered negative samples. To implement the above method, this application provides a multi-omics fusion and communication-constrained federated learning system based on a diffusion model. The system includes multiple clients and a central server. The client includes: A multimodal fusion network is used to extract and fuse features from local multi-omics data samples and output task prediction results. The multimodal fusion network integrates a conditional diffusion reconstruction module. The local training module is used to calculate the model loss and the amount of local model updates based on the task prediction results. A communication compression module is used to perform task-oriented compression of the local model update data and upload it. The server includes: The over-the-air aggregation receiving module is used to receive and superimpose the signals uploaded by each client through over-the-air computing to obtain the equivalent aggregated update amount; The aggregation correction module includes a diffusion denoiser, used to reconstruct and correct the equivalent aggregation update amount to obtain the corrected aggregation update amount; The global update module is used to update the global model using the corrected aggregate update amount and distribute it to each client.
[0091] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A multi-omics fusion and communication-constrained federated learning method based on a diffusion model, characterized in that, The method specifically includes the following steps: S1: The local client obtains data samples from multiple omics systems; S2: Input the multi-omics data samples into a multimodal fusion network to obtain a unified latent representation; the multimodal fusion network includes a modality-specific encoder, a cross-modal fusion module, and a conditional diffusion reconstruction module; S3: Based on the unified latent representation, the task prediction result is obtained through a classifier; S4: Calculate the local model update amount on the client, perform task-oriented compression on the local model update amount, obtain the compressed signal, and upload it via the wireless channel; S5: Through over-the-air computation, synchronous uplink signals from multiple clients are received and superimposed in the same time slot to obtain the equivalent aggregated update amount; S6: Input the equivalent aggregation update amount into the diffusion denoising aggregation correction module on the server side to reconstruct the corrected aggregation update amount; S7: Update the global model using the corrected aggregate update amount, and broadcast the updated global model to each client; Step S2 further includes: S201: Input the data of each omics in the multi-omics data sample into the corresponding modality-specific encoder to obtain the feature representation of each modality; S202: Project and concatenate the feature representations of each modality, and add modality type embedding and position encoding to form a unified token; S203: Input the unified token into the cross-modal attention backbone of the cross-modal fusion module, perform feature fusion and alignment through the cross-modal attention mechanism, and output the unified latent representation; S204: Apply conditional diffusion reconstruction priors to the unified latent representation to achieve missing mode completion and noise suppression; Step S204 further includes: S2041: Receive the unified latent representation output by the cross-modal fusion module; S2042: Perform a forward diffusion noise addition process on the unified latent representation; S2043: Based on conditional information, a reverse denoising and reconstruction process is performed on the noisy latent representation through a denoising network to obtain the reconstructed latent representation; wherein, the conditional information includes at least a modal availability mask for indicating the availability status of each omics data. S2044: When the modality availability mask indicates the presence of missing omics data, the corresponding missing part of the unified latent representation is completed during the reverse denoising reconstruction process; S2045: Output the reconstructed latent representation and use it for prediction by the classifier; The diffusion denoising aggregation correction module on the server side in step S6 is configured as follows: Using the equivalent aggregate update amount as the noisy input and the ideal aggregate update amount as the reconstruction target, train the noise prediction network. During inference, the currently received equivalent aggregate update and the current communication statistics are used as conditions to perform few-step denoising through the noise prediction network, and the corrected aggregate update is output.
2. The method according to claim 1, characterized in that, Step S4 further includes: S401: Define the local model update amount as the local increment after each client trains its local model after the server broadcasts the global model in each round. S402: Perform importance-aware group sparsification on the local model update amount to obtain the sparsified update amount; S403: Perform residual encoding and quantization on the sparsification update amount to generate a digitized compressed signal; S404: The compressed signal is mapped to an analog transmission symbol and synchronously transmitted uplink in the agreed communication time slot so that the server can receive and superimpose it through over-the-air computation.
3. The method according to claim 1, characterized in that, The total training loss function of the multimodal fusion network also includes a cross-modal contrastive loss term; the cross-modal contrastive loss term is calculated based on the InfoNCE loss function, and its construction method includes: The coding features of the first group were used as anchor features; The reconstructed features output by the conditional diffusion reconstruction module of the second group of learning are used as positive sample features; The second omics reconstruction features of other samples in the same training batch are used as negative sample features; The cross-modal contrast loss term is obtained by calculating the contrast loss between the anchor feature, the positive sample feature, and multiple negative sample features.
4. A multi-omics fusion and communication-restricted federated learning system based on a diffusion model, the system being used to execute the multi-omics fusion and communication-restricted federated learning method based on a diffusion model as described in claim 1, the system comprising multiple clients and a central server. The client includes: A multimodal fusion network is used to extract and fuse features from local multi-omics data samples and output task prediction results. The multimodal fusion network integrates a conditional diffusion reconstruction module. The local training module is used to calculate the model loss and the amount of local model updates based on the task prediction results. A communication compression module is used to perform task-oriented compression of the local model update data and upload it. The server includes: The over-the-air aggregation receiving module is used to receive and superimpose the signals uploaded by each client through over-the-air computing to obtain the equivalent aggregated update amount; The aggregation correction module includes a diffusion denoiser, used to reconstruct and correct the equivalent aggregation update amount to obtain the corrected aggregation update amount; The global update module is used to update the global model using the corrected aggregate update amount and distribute it to each client.
5. An application of a multi-omics fusion and communication-constrained federated learning method based on a diffusion model, characterized in that, When any one of the methods of claims 1-3 is applied to the auxiliary identification of lung adenocarcinoma subtypes, the task prediction result is lung adenocarcinoma subtype discrimination or classification, wherein the multimodal fusion network includes a modality-specific encoder, specifically including: Transcriptome (RNA) encoder E rna A multilayer perceptron (MLP) block encoder is used, with gene expression vectors as input; DNA methylation encoder E meth The encoder end of the sparse high-dimensional MLP encoder or autoencoder is used, and the input is the methylation intensity vector of the site or region. CNV encoder E cnv An MLP encoder is used, with the input being gene-level or fragment-level CNV vectors; Clinical Variable Encoder E cli It adopts an embedding + MLP structure, and the input contains structured variables containing individual patient information and medical records.
Citation Information
Patent Citations
Federal learning algorithm based on diffusion model and weight adaptive knowledge distillation
CN116665000A
Distributed federated learning method based on homomorphic encryption and air computing
CN119886286A