A cross-domain federated medical image classification method and system based on fedssac
Patent Information
- Application Number
- CN202610913358.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-24
AI Technical Summary
[0005]针对现有技术的不足,本发明提供一种基于FedSSAC的跨域联邦医学影像分类方法及系统,解决跨域联邦MRI影像分类中多源域偏移、标签空间异构、自然-医疗域适配差及临床硬件受限的问题,在保护医疗数据隐私的前提下,提升多中心脑肿瘤MRI分类准确率与模型鲁棒性,降低部署门槛,为多中心医疗 AI协同诊断提供高效可行的方案
1、本发明通过空间-光谱SSP适配器实现MRI影像空间-光谱双域特征校准,有效缓解多源域偏移,弥补现有方法忽略光谱域特性的缺陷,提升病理特征提取精度与域不变性。
Smart Images

Figure CN122455265B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a cross-domain federated medical image classification method and system based on FedSSAC, belonging to the field of medical image processing and federated learning technology. Background Technology
[0002] Magnetic resonance imaging (MRI) has become a core imaging technology for the clinical diagnosis and classification assessment of brain tumors due to its advantages of no ionizing radiation and high soft tissue resolution. The accuracy of its classification diagnosis directly affects clinical treatment decisions. Federated learning, as a privacy-preserving AI technology, enables collaborative training of models from multiple medical institutions, addressing the challenges of medical data privacy leaks and data silos. It is a mainstream direction for multi-center medical image analysis. However, cross-domain federated MRI image classification still faces many key technical challenges: different medical institutions' scanning protocols and field strength differences lead to non-independent and identically distributed data, causing severe multi-source domain shifts, and spatial convolution struggles to distinguish pathological textures from high-frequency scanner noise; when visual language models such as CLIP, pre-trained on natural images, are directly transferred to medical scenarios, there is a natural-medical domain adaptation gap, making it impossible to accurately capture high-frequency diagnostic features such as tumor boundaries; different medical institutions have different diagnostic label systems, and traditional fixed-dimensional classifiers cannot adapt to the cross-site label spatial heterogeneity problem, limiting the effectiveness of collaborative training; full fine-tuning of visual language models places stringent demands on GPU memory, exceeding the hardware capabilities of clinical workstations, and existing lightweight adaptation methods only focus on spatial domain features, ignoring MRI spectral domain characteristics, making it difficult to balance performance and deployment efficiency.
[0003] In existing technologies, classic federated learning methods such as FedAvg and FedProx are not designed for the spatial-spectral features of medical images, exhibit poor robustness to non-independent and identically distributed data, and show significant attenuation in classification accuracy after cross-center aggregation. Lightweight adaptation methods such as FedCLIP only make simple fine-tuning to CLIP, without spectral domain modeling, failing to address scanner-induced domain shifts and neglecting the need for label spatial heterogeneity. Other adapter methods only focus on spatial domain feature alignment, lacking customized feature normalization and noise suppression designs for medical scenarios, making it difficult to accurately extract pathological features. Furthermore, existing federated aggregation strategies often employ equal weighting or data volume weighting methods, making them susceptible to interference from low-precision client parameters, resulting in poor global model convergence and failing to meet the practical needs of high-precision classification of multi-center brain tumor MRI images in decentralized clinical scenarios.
[0004] To address the aforementioned issues, a classification method specifically designed for cross-domain federated medical image analysis is needed. This method should strictly protect the privacy of medical data while simultaneously resolving problems such as multi-source domain offset, label spatial heterogeneity, poor domain adaptability, and limitations of clinical hardware. It should achieve high-precision classification of MRI images across institutions, lower the threshold for model deployment, and provide a feasible technical solution for multi-center medical AI collaborative diagnosis. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a cross-domain federated medical image classification method and system based on FedSSAC. It solves the problems of multi-source domain offset, label spatial heterogeneity, poor natural-medical domain adaptation, and clinical hardware limitations in cross-domain federated MRI image classification. While protecting medical data privacy, it improves the accuracy and robustness of multi-center brain tumor MRI classification, lowers the deployment threshold, and provides an efficient and feasible solution for multi-center medical AI collaborative diagnosis.
[0006] The technical solution of the present invention is as follows: A cross-domain federal medical image classification method based on FedSSAC, the steps of which are as follows: (1) Each medical client retrieves locally stored, privacy-compliant labeled brain tumor MRI slice datasets; (2) Construct the FedSSAC spatial-spectral adaptive collaborative framework model, which includes a frozen CLIP visual backbone, a spatial-spectral SSP adapter, a dynamic dimension adaptive DDAC classifier, and a performance-aware global aggregation module, respectively realizing basic visual feature extraction, spatial-spectral dual-domain feature calibration, heterogeneous label spatial adaptation, and dynamic weighted aggregation of client parameters; (3) Each medical client completes the training and verification of the FedSSAC spatial-spectral adaptive collaborative framework model based on local data, selects qualified parameters and uploads them to the server. The server updates the global model parameters and distributes them through a performance-aware global aggregation strategy, and iterates the training until the model converges. (4) Input the preprocessed MRI images to be classified into the trained FedSSAC spatial-spectral adaptive collaborative framework model and output the brain tumor classification results. The FedSSAC spatial-spectral adaptive collaborative framework model also supports zero-sample cross-site classification reasoning for clinical sites with unknown labels.
[0007] According to a preferred embodiment of the present invention, in step (1), the slice dataset is multicenter brain tumor T1-weighted MRI slice data, with a single medical client having no less than 2000 data entries. The slice dataset covers 1.5T / 3.0T field strength scan results and includes four categories of labels: glioma, meningioma, pituitary adenoma, and no tumor. Using a fixed grouping and allocation standard, all medical clients are divided into three groups, corresponding to three label overlap levels of 70%, 50%, and 30%, respectively. The data is evenly distributed according to the number of medical clients to simulate the high, medium, and low label sharing levels and construct a non-independent and identically distributed scenario with heterogeneous label space. All data in the slice dataset are de-identified, comply with HIPAA and GDPR privacy specifications, and Z-score standardization preprocessing is completed.
[0008] The data from different medical clients exhibit non-independent and identically distributed characteristics. Specifically, there are differences in MRI field strength, image contrast, noise level, and scanning protocol among different medical clients. At the same time, the proportions of the four types of tumor labels are uneven, with label overlap of 70%, 50%, and 30%, respectively. This results in the data feature distribution and label distribution being different from each other, forming a typical multi-center Non-IID data scenario.
[0009] According to a preferred embodiment of the present invention, in step (2), the frozen CLIP visual backbone uses a pre-trained ResNet-50 or ViT-B / 32 as the base network, with the lower layer weights completely frozen and the last 2-3 upper layers lightly fine-tuned, with the fine-tuning learning rate set to 10. −5 It interacts with the spatial-spectral SSP adapter through feature stitching, reducing GPU memory usage by 40% and adapting to clinical standard workstation hardware.
[0010] The basic visual feature extraction steps for achieving a frozen CLIP visual backbone are as follows: 1. Directly reuse pre-trained representations (Forward Feature Extraction). Standardized preprocessed local brain tumor MRI slices are input into a pre-trained CLIP visual network (such as ResNet-50 or ViT-B / 32).
[0011] Images are forward-propagated through the network, directly reusing CLIP's powerful general visual representation capabilities learned on massive amounts of natural images (such as capturing boundaries, structures, and basic anatomical spatial context).
[0012] 2. Lower-level weights are completely frozen (Frozen Mechanism); The algorithm completely freezes the lower network parameters of the CLIP visual backbone (by setting `requires_grad=False`), and does not calculate the gradients of these layers or update their weights during local federated training.
[0013] Function: This mechanism transforms the most computationally intensive basic feature extraction process into pure forward computation, significantly reducing GPU memory overhead by 40% (from 9.0GB to 5.4GB on an RTX 4090), enabling it to be deployed directly on standard clinical workstations.
[0014] 3. Lightweight Fine-tuning at the Top Level To bridge the domain gap between natural images and medical MRI images, the framework only unfroze the last 2-3 layers (top layer) of the CLIP visual backbone.
[0015] With a very small learning rate ( These layers are fine-tuned so that the features output by the model can be adaptively projected from the "natural general representation" and transformed into more accurate medical pathology-based spatial features (such as preliminary anchoring of tumor boundaries and space-occupying information), and then passed to the subsequent spatial-spectral SSP adapter.
[0016] According to a preferred embodiment of the present invention, in step (2), the spatial-spectral SSP adapter includes a medical instance center normalization submodule and a frequency-gated feedforward (FFN) submodule. The medical instance center normalization submodule alleviates cross-client domain offset. In the frequency-gated feedforward (FFN) submodule, frequency domain feature modulation is achieved through Fourier transform and inverse transform, which suppresses high-frequency noise of the scanner, refines pathological textures, and outputs spatial-spectral dual-domain calibration fusion features.
[0017] According to a preferred embodiment of the present invention, in step (2), the Dynamic Dimension Adaptive DDAC classifier is a 2-layer MLP structure. Local label space dynamic mapping is achieved by multiplying the learnable dimension mapping coefficients with the spatial-spectral dual-domain calibration fusion features element-wise. The mean squared error (MSE) loss is used to calculate the difference between the local dimension mapping coefficients and the global reference coefficients. (The local dimension mapping coefficients are learned autonomously through backpropagation by each medical client (local site) during training. Initial state: In the first round of federated learning or when a client is first added, the local dimension mapping coefficients are initialized to a vector of all 1s or a random feature vector in the interval (0,1], whose dimension is consistent with the number of channels of the spatial-spectral fusion features.) Local training optimization: In each round of local training, this coefficient is used as a learnable parameter matrix / vector (defined as nn.Parameter along with the network in the code). Gradient update: When local data is input into the model, the gradient propagates forward by calculating the total loss function of the classification task (including cross-entropy loss and alignment loss). Since the learning rate of the Dynamic Dimension Adaptive Classifier (DDAC) is set relatively high (e.g., ... This coefficient is updated synchronously with the network weights through backpropagation gradient updates. Physical meaning: Through local training, it automatically and dynamically adjusts the weight scaling of each channel in the feature layer, enabling the unified spatial-spectral features to be accurately mapped to the hospital's unique, heterogeneous label space. This achieves global alignment of label spaces across sites, and the dynamic dimension-adaptive DDAC classifier has a higher learning rate than the adapter and CLIP fine-tuning layers, enabling rapid adaptation to the local label space.
[0018] According to a preferred embodiment of the present invention, in step (2), the performance-aware global aggregation module includes dynamic weight allocation and parameter weighted aggregation, achieving efficient parameter fusion based on the local verification accuracy of the medical client: Dynamic weight allocation: Aggregated weights are assigned based on the local verification accuracy (ACCI) uploaded by each medical client, while upper and lower limits are set to prevent a single medical client from dominating the weights. The formula is as follows:
[0019] In the formula, n is the number of medical clients that uploaded qualified parameters, and w min The lower bound for the weight is 0.1, w. max This represents the upper limit of the weight, with a value of 0.4; w i Here, represents the aggregate weight of the i-th medical client; `clip()` is the numerical truncation function, restricting the result to a specified range; `acc`... j Let acc be the local verification accuracy of the j-th client; i The local validation accuracy obtained by the i-th medical client on the local validation set is used to evaluate the classification performance of the local model in the current round of the client. Parameter weighted aggregation: The parameters of the spatial-spectral SSP adapter and the dynamically dimensional adaptive DDAC classifier uploaded by each medical client are weighted and summed separately, while the global dimension mapping reference coefficient is updated. The formula is as follows:
[0020]
[0021]
[0022] In the formula, The parameters for the global spatial-spectral SSP adapter in round t+1; The parameters are the global dynamic dimension adaptive DDAC classifier parameters for the (t+1)th round; w is the reference coefficient for the global dimension mapping in round t+1; i Let be the aggregate weight of the i-th medical client; The local parameters of the spatial-spectral SSP adapter uploaded for the i-th medical client; The local parameters of the dynamically dimension-adaptive DDAC classifier uploaded for the i-th medical client; is the local dimension mapping coefficient for the i-th medical client; n is the number of medical clients that have uploaded qualified parameters; t is the current federated training round; Gradient clipping is applied to the aggregated global dimension mapping reference coefficients to ensure model convergence stability.
[0023] According to a preferred embodiment of the present invention, in step (3), the total local training loss of each medical client is a weighted fusion of the labeled smooth cross-entropy classification loss and the cross-site alignment loss, calculated as follows:
[0024] In the formula, The true label for the MRI image; λ=0.0 is the alignment loss weighting coefficient; This represents the total loss during local training. The classification loss is the labeled smooth cross-entropy; y k Predict the probability of the k-th pathological classification; For cross-site tag space alignment loss; For the k-th medical client local dimension mapping coefficient; The reference coefficient for the global dimension mapping in round t; The AdamW optimizer and cosine annealing learning rate strategy are used. After local training, the model is validated. Only when the local validation accuracy exceeds a set threshold for global accuracy are the parameters of the local spatial-spectral (SSP) adapter and the dynamically dimensional adaptive DDAC classifier uploaded to the server. This avoids low-precision parameters contaminating the global model.
[0025] In the formula, To perform the action (upload parameters); It is a logical condition (if and only if); > This is a condition for judgment.
[0026] According to a preferred embodiment of the present invention, in step (3), the convergence criterion of the FedSSAC spatial-spectral adaptive collaborative framework model is that the absolute value of the difference in global verification accuracy over 5 consecutive rounds is less than 0.5%, and the maximum number of federated iteration rounds is 50. The FedSSAC spatial-spectral adaptive collaborative framework model supports random sampling training with 1-5 medical clients. When a single medical client participates, the accuracy is ≥91.2%, and when 3 or more medical clients participate, the optimal average accuracy reaches 93.7%, with a performance variance <0.2%.
[0027] According to a preferred embodiment of the present invention, in step (4), the MRI image to be classified is a clinical standard T1-weighted brain tumor MRI slice. Before input, Z-score standardization preprocessing is completed. The FedSSAC spatial-spectral adaptive collaborative framework model outputs four classification labels and corresponding confidence scores for glioma, meningioma, pituitary adenoma, and no tumor. The confidence score threshold is 0.5. If the confidence score is lower than the threshold, it is marked as needing to be reviewed. If the confidence score is higher than the threshold, the classification result is directly output. For clinical sites with unknown labels, zero-sample classification is achieved through global dimension mapping coefficients, and the inference accuracy is ≥85%.
[0028] A cross-domain federal medical image classification system based on FedSSAC includes: The client-side local data module is used to retrieve the labeled brain tumor MRI slice dataset from the local machine of each medical client. The model building module is used to build the FedSSAC spatial-spectral adaptive collaborative framework model. It includes a frozen CLIP visual backbone, a spatial-spectral SSP adapter, a dynamic dimension adaptive DDAC classifier, and a performance-aware global aggregation module, which respectively realize basic visual feature extraction, spatial-spectral dual-domain feature calibration, heterogeneous label space adaptation, and dynamic weighted aggregation of client parameters. The training module enables each medical client to complete the training and validation of the FedSSAC spatial-spectral adaptive collaborative framework model based on local data, select qualified parameters and upload them to the server. The server updates the global model parameters through a performance-aware global aggregation strategy and distributes them, iterating the training until the model converges. The inference module inputs the preprocessed MRI images to be classified into the trained FedSSAC spatial-spectral adaptive collaborative framework model and outputs brain tumor classification results. The FedSSAC spatial-spectral adaptive collaborative framework model also supports zero-sample cross-site classification inference for clinical sites with unknown labels.
[0029] The present invention also provides a cross-domain federated medical image classification device based on FedSSAC, including a memory and a processor, wherein: The memory is used to store computer programs that can run on the processor, as well as MRI image data and model weight files. The memory can be a solid-state drive, a hard disk drive, server memory, or other storage media. The processor is used to execute the steps of the FedSSAC-based cross-domain federated medical image classification method described above when running the computer program. The processor may be a central processing unit (CPU), a graphics processing unit (GPU), or an application-specific integrated circuit (ASIC). In particular, it is recommended to use an NVIDIA RTX 4090 or higher GPU that supports CUDA acceleration to achieve high efficiency in model training and real-time inference.
[0030] Based on the same inventive concept, the present invention also provides a storage medium storing a computer program, which, when executed by at least one processor, implements the steps of the aforementioned FedSSAC-based cross-domain federated medical image classification method. The storage medium may be a non-volatile computer-readable storage medium, including but not limited to disk storage, CD-ROM, optical storage, USB flash drive, portable hard drive, etc., which can realize long-term storage and cross-device operation of the computer program.
[0031] The beneficial effects of this invention are as follows: 1. This invention achieves spatial-spectral dual-domain feature calibration of MRI images through a spatial-spectral SSP adapter, effectively alleviating multi-source domain shift, compensating for the shortcomings of existing methods that ignore spectral domain characteristics, and improving the accuracy of pathological feature extraction and domain invariance.
[0032] 2. This invention achieves dynamic adaptation and global alignment of heterogeneous label spaces through a dynamic dimension adaptive DDAC classifier, solving the cross-site adaptation problem of traditional fixed-dimensional classifiers and improving the model's versatility.
[0033] 3. This invention adopts a frozen CLIP + lightweight component fine-tuning architecture, which reduces GPU memory overhead by 40%, is compatible with clinical standard workstations, and significantly lowers the deployment threshold.
[0034] 4. The performance-aware aggregation strategy of this invention improves the robustness of the model to non-independent and identically distributed data, supports random sampling training from a large number of clients, and has excellent deployment flexibility and stability.
[0035] 5. The model of this invention supports zero-sample cross-site inference for unknown labeled spatial sites with an accuracy of ≥85%, improves cross-institutional adaptability, and helps to facilitate the large-scale implementation of multi-center collaborative diagnosis.
[0036] 6. The architecture of this invention strictly adheres to privacy protection standards, with data stored locally and the server only aggregating parameters, fundamentally solving the problems of medical data silos and privacy leaks. Moreover, the model achieves a classification accuracy of 93.7%, demonstrating extremely high clinical translational value. Attached Figure Description
[0037] Figure 1 This is a flowchart of the present invention; Figure 2 This is a schematic diagram of the overall structure of the FedSSAC space-spectrum adaptive collaborative framework model of the present invention; Figure 3 This is a comparison of the convergence curves of the training accuracy of various methods on a multicenter brain tumor MRI dataset. Figure 3 (a) is a comparison of the convergence curves of training accuracy for each method on client 1; Figure 3 (b) is a comparison of the convergence curves of the training accuracy of each method on client 2; Figure 3 (c) is a comparison of the convergence curves of training accuracy for each method on client 3; Figure 3 (d) is a comparison of the convergence curves of the global average training accuracy of each method; Figure 4 A comparison chart showing the classification accuracy of various OfficeHome methods under different numbers of participating medical clients. Figure 4 (a) is a comparison chart of the classification accuracy of OfficeHome using different methods with the participation of two medical clients; Figure 4 (b) is a comparison chart of the classification accuracy of OfficeHome using different methods with the participation of three medical clients; Figure 4 (c) is a comparison chart of the classification accuracy of OfficeHome using different methods with the participation of four medical clients; Figure 4 The middle (d) chart is a comparison of the classification accuracy of OfficeHome using different methods with the participation of 5 medical clients; Figure 5 A comparison chart of MRI classification accuracy rates for different methods under varying numbers of medical client participants. Figure 5 (a) is a comparison of the classification accuracy of various MRI methods with the participation of two medical clients; Figure 5 (b) is a comparison chart of the classification accuracy of different MRI methods with the participation of three medical clients; Figure 5 (c) is a comparison chart of the classification accuracy of different MRI methods with the participation of four medical clients; Figure 5 The middle (d) chart is a comparison of the classification accuracy of different MRI methods with the participation of 5 medical clients. Detailed Implementation
[0038] Example 1: like Figure 1 As shown, this embodiment provides a cross-domain federal medical image classification method based on FedSSAC, with the following steps: (1) Each medical client retrieves the locally stored multicenter brain tumor T1-weighted MRI slice dataset. Specifically, it retrieves locally labeled brain tumor MRI slice data from each medical center to construct a multicenter brain tumor MRI clinical benchmark dataset containing 8640 slices. At the same time, the OfficeHome general cross-domain benchmark dataset is selected as an auxiliary verification. Among them, the brain tumor MRI dataset has a single client data volume of ≥2000 slices, covering data acquired by 1.5T and 3.0T field strength scanners, and includes four categories of labels: glioma, meningioma, pituitary adenoma, and no tumor. In order to simulate the real clinical non-independent identically distributed scenario, each medical client dataset is set with a label overlap of 70%, 50%, or 30%, respectively. All MRI image data are de-identified and privacy-desensitized after approval, which complies with the HIPA and GDPR medical data privacy compliance requirements. At the same time, all images are uniformly preprocessed with Z-score standardization to eliminate intensity deviation between different devices and ensure the consistency of data input for each medical client. The standardization calculation formula is:
[0039] In the formula, X represents the pixel value of the original MRI slice, and μ represents the pixel mean of the single-client dataset. Z represents the standard deviation of pixels in a single-client dataset, and Z represents the standardized pixel values. (2) Construct the FedSSAC space-spectral adaptive collaborative framework model, such as Figure 2 As shown, it includes a frozen CLIP visual backbone, a spatial-spectral SSP adapter, a dynamic dimension-adaptive DDAC classifier, and a performance-aware global aggregation module, which respectively realize basic visual feature extraction, spatial-spectral dual-domain feature calibration, heterogeneous label space adaptation, and dynamic weighted aggregation of client parameters; The frozen CLIP visual backbone uses a pre-trained ResNet-50 or ViT-B / 32 as the base network, with the lower-layer weights completely frozen and the last 2-3 upper layers lightly fine-tuned. The fine-tuning learning rate is set to 10. −5 Feature stitching interaction is achieved through a feature fusion interface and a spatial-spectral SSP adapter. Freezing the backbone significantly reduces GPU memory overhead. The memory optimization rate is calculated using the following formula:
[0040] Achieve a 40% reduction in memory overhead on the NVIDIA RTX 4090 GPU, down from 9.0GB to 5.4GB, perfectly suited for clinical standard workstation hardware environments.
[0041] The spatial-spectral SSP adapter includes a medical instance center normalization submodule and a frequency-gated feedforward (FFN) submodule, enabling spatial-spectral dual-domain calibration of features. The core calculations are as follows: The Medical Instance Center Normalization Submodule: performs instance-level mean and variance normalization on the input basic visual feature batch z, anchoring the pathological feature distribution to mitigate cross-client domain bias. The formula is:
[0042] In the formula, , For instance-level mean and variance of the feature batch, To prevent the division by zero error from reaching a minimum, , These are learnable parameters; Frequency-gated feedforward (FFN) submodule: The features are mapped to the frequency domain through Fourier transform, modulated by the gating coefficient (value 0.1), and then reconstructed back to the spatial domain through inverse Fourier transform. The dropout value is set to 0.2 to achieve spectral domain pathological texture refinement and scanner-induced high-frequency noise suppression. By randomly deactivating 20% of neurons, the model is forced to learn robust pathological texture features, while reducing dependence on device noise. Finally, the spatial-spectral dual-domain calibrated fusion features are output.
[0043] The Dynamic Dimension Adaptive DDAC classifier is a 2-layer MLP structure, including dimension mapping and cross-site alignment loss; Dimension mapping: through learnable dimension mapping coefficients The ∈(0,1] is multiplied element-wise with the fused features to achieve dynamic mapping of the local label space, as shown in the formula:
[0044] In the formula, f(x) represents the spatial-spectral dual-domain calibrated fusion feature obtained after spectral dimension calibration and pathological texture refinement of the CLIP backbone output feature using spatial-spectral fusion features, which serves as the input to the dimension mapper; , , , For learnable parameters of MLP, It is the ReLU activation function. For element-wise multiplication; y k The classification prediction probability for the kth type of brain tumor; Cross-site alignment loss: The local dimension mapping coefficients are calculated using the mean squared error (MSE) loss. With global reference coefficient The difference is used to achieve global alignment of the label space across clients; the learning rate of the classifier module is set to... It surpasses the SSP adapter module and CLIP backbone fine-tuning layer, enabling rapid adaptation of local tag space.
[0045] The performance-aware global aggregation module includes dynamic weight allocation and parameter weighted aggregation, achieving efficient parameter fusion based on the accuracy of local verification on the medical client. Dynamic weight allocation: Aggregated weights are assigned based on the local verification accuracy (ACCI) uploaded by each medical client, while upper and lower limits are set to prevent a single medical client from dominating the weights. The formula is as follows:
[0046] In the formula, n is the number of medical clients that uploaded qualified parameters, and w min The lower bound for the weight is 0.1, w. max This represents the upper limit of the weight, with a value of 0.4; w i Here, represents the aggregate weight of the i-th medical client; `clip()` is the numerical truncation function, restricting the result to a specified range; `acc`... j Let be the local verification accuracy of the j-th client; Parameter weighted aggregation: The parameters of the spatial-spectral SSP adapter and the dynamically dimensional adaptive DDAC classifier uploaded by each medical client are weighted and summed separately, while the global dimension mapping reference coefficient is updated. The formula is as follows:
[0047]
[0048]
[0049] In the formula, The parameters for the global spatial-spectral SSP adapter in round t+1; The parameters are the global dynamic dimension adaptive DDAC classifier parameters for the (t+1)th round; w is the reference coefficient for the global dimension mapping in round t+1; i Let be the aggregate weight of the i-th medical client; The local parameters of the spatial-spectral SSP adapter uploaded for the i-th medical client; The local parameters of the dynamically dimension-adaptive DDAC classifier uploaded for the i-th medical client; is the local dimension mapping coefficient for the i-th medical client; n is the number of medical clients that have uploaded qualified parameters; t is the current federated training round; Gradient clipping is applied to the aggregated global dimension mapping reference coefficients to ensure model convergence stability.
[0050] (3) Each medical client completes the training and verification of the FedSSAC spatial-spectral adaptive collaborative framework model based on local data, selects qualified parameters and uploads them to the server. The server updates the global model parameters and distributes them through a performance-aware global aggregation strategy, and iterates the training until the model converges. Each medical client initializes its local model with global parameters, employs the AdamW optimizer (weight decay of 0.01) and a cosine annealing learning rate scheduling strategy, and sets the maximum number of local training epochs to 50. The total loss function for local training is a weighted fusion of the labeled smooth cross-entropy classification loss and the cross-site alignment loss, calculated as follows:
[0051] In the formula, The true label for the MRI image; λ=0.0 is the alignment loss weighting coefficient; This represents the total loss during local training. The classification loss is the labeled smooth cross-entropy; y k Predict the probability of the k-th pathological classification; For cross-site tag space alignment loss; For the k-th medical client local dimension mapping coefficient; The reference coefficient for the global dimension mapping in round t; The AdamW optimizer and cosine annealing learning rate strategy are used. After local training, the model is validated. Only when the local validation accuracy exceeds a set threshold for global accuracy are the parameters of the local spatial-spectral (SSP) adapter and the dynamically dimensional adaptive DDAC classifier uploaded to the server. This avoids low-precision parameters contaminating the global model.
[0052] In the formula, To perform the action (upload parameters); It is a logical condition (if and only if); > This is a condition for judgment.
[0053] The model convergence criterion is that the absolute value of the difference in accuracy between five consecutive rounds of global model validation is less than 0.5%, i.e. ; The accuracy of the global model is validated in round t; t is the current global round of federated training. The maximum number of federated iteration rounds is set to 50; during federated training, 1-5 clients are randomly sampled to participate in the training, adapting to the network and hardware constraints in actual clinical deployments.
[0054] (4) Input the standardized preprocessed MRI images to be classified into the FedSSAC spatial-spectral adaptive collaborative framework model after training convergence, and output accurate brain tumor classification results. The specific reasoning process, performance verification and experimental results are as follows: The clinical standard T1 weighted brain tumor MRI slices to be classified need to be preprocessed by Z-score standardization before inputting into the model; the model outputs four classification labels and corresponding confidence scores for glioma, meningioma, pituitary adenoma and no tumor. The confidence score threshold is set to 0.5. If the confidence score is higher than the threshold, the classification result is directly output. If the confidence score is lower than the threshold, it is marked as needing to be verified; for clinical sites with unknown label space, the model automatically infers the target label space through the dimension mapping coefficient α of the global DDAC classifier, and zero-sample cross-site classification reasoning can be achieved without additional fine-tuning.
[0055] To objectively evaluate the effectiveness of the FedSSAC method, classification accuracy (mean ± standard deviation) was selected as the core evaluation metric. The method was compared and validated against mainstream federated learning methods such as Local, FedCLIP, FedAvg, FedProx, and MOON on a multicenter brain tumor MRI clinical dataset and the OfficeHome general dataset. Deployment robustness experiments were also designed with different client sampling numbers. All experiments were implemented on an NVIDIA RTX 4090 GPU using PyTorch 1.13.1 / CUDA12.6. The results are the average of three independent experiments. Specific experimental results are as follows: Table 1: Performance Comparison of Various Federated Learning Methods on the OfficeHome and Brain Tumor MRI Datasets
[0056] As shown in Table 1, FedSSAC significantly outperforms the comparison methods across all clients and average metrics on both datasets. It achieves the best average classification accuracy of 93.7 ± 0.2% on the brain tumor MRI dataset, which is 3.2% higher than the second-best method, FedProx, thus verifying the model's superiority in cross-domain federated medical image classification tasks.
[0057] Table 2: Performance Comparison of Methods under Different Client Sampling Quantities
[0058] As shown in Table 2, FedSSAC maintains excellent classification performance and robustness across different client sampling numbers. Even with a single client participating in training, it maintains a classification accuracy of over 91.2%, and achieves an optimal average accuracy of 93.7% with three or more clients participating, with a performance variance of less than 0.2%, far superior to other comparative methods. On the OfficeHome general dataset, FedSSAC also achieves optimal classification performance across various sampling numbers, with an accuracy of 92.7% when n=5, verifying the model's high adaptability in cross-dataset generalization and decentralized deployment in clinical practice.
[0059] Meanwhile, the effectiveness of each core module was verified through ablation experiments. Removing the frequency-gated feedforward (FFN) submodule or replacing it with a fixed-dimensional classifier significantly reduced the model's classification accuracy, demonstrating that spatial-spectral dual-domain calibration and dynamic dimension adaptive classification are the core designs for solving cross-domain offset and label spatial heterogeneity. The modules work together to achieve optimal model performance. Furthermore, Figure 3 shows a comparison of the training accuracy convergence curves of this embodiment with other existing technologies on a multi-center brain tumor MRI dataset. Figure 4 shows an accuracy comparison under different client sampling numbers. Figure 4 As shown, the performance advantages of FedSSAC compared to existing methods are clearly demonstrated, including faster convergence speed, higher classification accuracy, and stronger deployment robustness. A comparison of the classification accuracy of this embodiment with other existing technologies on the OfficeHome general cross-domain dataset with different client sampling numbers is provided. Figure 4 As shown, the classification accuracy on the brain tumor MRI clinical dataset varies depending on the number of samples taken from different clients. Figure 5 As shown, FedSSAC clearly demonstrates its performance advantages of having the best accuracy, the strongest robustness, and the best deployment adaptability in both general cross-domain scenarios and medical imaging scenarios.
[0060] Example 2: This embodiment provides a cross-domain federal medical image classification system based on FedSSAC, including: The client-side local data module is used to retrieve the labeled brain tumor MRI slice dataset from the local machine of each medical client. The model building module is used to build the FedSSAC spatial-spectral adaptive collaborative framework model. It includes a frozen CLIP visual backbone, a spatial-spectral SSP adapter, a dynamic dimension adaptive DDAC classifier, and a performance-aware global aggregation module, which respectively realize basic visual feature extraction, spatial-spectral dual-domain feature calibration, heterogeneous label space adaptation, and dynamic weighted aggregation of client parameters. The training module enables each medical client to complete the training and validation of the FedSSAC spatial-spectral adaptive collaborative framework model based on local data, select qualified parameters and upload them to the server. The server updates the global model parameters through a performance-aware global aggregation strategy and distributes them, iterating the training until the model converges. The inference module inputs the preprocessed MRI images to be classified into the trained FedSSAC spatial-spectral adaptive collaborative framework model and outputs brain tumor classification results. The FedSSAC spatial-spectral adaptive collaborative framework model also supports zero-sample cross-site classification inference for clinical sites with unknown labels.
[0061] Embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0062] Contents not described in detail in this specification are prior art known to those skilled in the art. It is hereby indicated that the above description is intended to help those skilled in the art understand this invention, but does not limit the scope of protection of this invention. Any equivalent substitutions, modifications, improvements, or simplifications of the above descriptions that do not depart from the essential content of this invention fall within the scope of protection of this invention.
[0063] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0064] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0065] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0066] Contents not described in detail in this specification are prior art known to those skilled in the art. It is hereby indicated that the above description is intended to help those skilled in the art understand this invention, but does not limit the scope of protection of this invention. Any equivalent substitutions, modifications, improvements, or simplifications of the above descriptions that do not depart from the essential content of this invention fall within the scope of protection of this invention.
Claims
1. A cross-domain federal medical image classification method based on FedSSAC, characterized in that, The steps are as follows: (1) Each medical client retrieves the local labeled brain tumor MRI slice dataset; (2) Construct the FedSSAC spatial-spectral adaptive collaborative framework model, which includes a frozen CLIP visual backbone, a spatial-spectral SSP adapter, a dynamic dimension adaptive DDAC classifier, and a performance-aware global aggregation module, respectively realizing basic visual feature extraction, spatial-spectral dual-domain feature calibration, heterogeneous label spatial adaptation, and dynamic weighted aggregation of client parameters; The Spatial-Spectral SSP adapter includes a medical instance center normalization submodule and a frequency-gated feedforward FFN submodule. The dynamic dimension adaptive DDAC classifier is a 2-layer MLP structure. It achieves local label spatial dynamic mapping by multiplying the learnable dimension mapping coefficients with the spatial-spectral dual-domain calibration fusion features element-wise. (3) Each medical client completes the training and verification of the FedSSAC spatial-spectral adaptive collaborative framework model based on local data, selects qualified parameters and uploads them to the server. The server updates the global model parameters and distributes them through a performance-aware global aggregation strategy, and iterates the training until the model converges. (4) Input the preprocessed MRI images to be classified into the trained FedSSAC spatial-spectral adaptive collaborative framework model and output the brain tumor classification results. The FedSSAC spatial-spectral adaptive collaborative framework model also supports zero-sample cross-site classification reasoning for clinical sites with unknown labels.
2. The cross-domain federal medical image classification method based on FedSSAC as described in claim 1, characterized in that, In step (1), the slice dataset is multicenter brain tumor T1-weighted MRI slice data, with no less than 2000 data samples per medical client. The slice dataset covers 1.5T / 3.0T field strength scan results and includes four categories of labels: glioma, meningioma, pituitary adenoma, and no tumor. Using a fixed grouping and allocation standard, all medical clients are divided into three groups, corresponding to three label overlap levels of 70%, 50%, and 30%, respectively. The data is evenly distributed according to the number of medical clients to simulate the high, medium, and low label sharing levels and construct a non-independent and identically distributed scenario with heterogeneous label space. All data in the slice dataset are de-identified and Z-score standardized preprocessing is completed.
3. The cross-domain federal medical image classification method based on FedSSAC as described in claim 2, characterized in that, In step (2), the frozen CLIP visual backbone uses a pre-trained ResNet-50 or ViT-B / 32 as the base network, with the lower layer weights completely frozen and the last 2-3 upper layers lightly fine-tuned. The fine-tuning learning rate is set to 10. −5 It interacts with the spatial-spectral SSP adapter through feature splicing.
4. The cross-domain federal medical image classification method based on FedSSAC as described in claim 3, characterized in that, In step (2), the cross-client domain offset is mitigated by the medical instance center normalization submodule. In the frequency-gated feedforward FFN submodule, frequency domain feature modulation is achieved through Fourier transform and inverse transform to suppress high-frequency noise of the scanner, refine pathological texture, and output spatial-spectral dual-domain calibration fusion features.
5. The cross-domain federal medical image classification method based on FedSSAC as described in claim 4, characterized in that, In step (2), the difference between the local dimension mapping coefficient and the global reference coefficient is calculated using mean squared error (MSE) loss to complete the global alignment of the cross-site label space. The learning rate of the dynamic dimension adaptive DDAC classifier is higher than that of the adapter and CLIP fine-tuning layer, thus achieving rapid adaptation of the local label space.
6. The cross-domain federal medical image classification method based on FedSSAC as described in claim 5, characterized in that, In step (2), the performance-aware global aggregation module includes dynamic weight allocation and parameter weighted aggregation, achieving efficient parameter fusion based on the local verification accuracy of the medical client: Dynamic weight allocation: Aggregated weights are assigned based on the local verification accuracy (ACCI) uploaded by each medical client, while upper and lower limits are set to prevent a single medical client from dominating the weights. The formula is as follows: ; In the formula, n is the number of medical clients that uploaded qualified parameters, and w min The lower bound for the weight is 0.1, w. max This represents the upper limit of the weight, with a value of 0.4; w i Let be the aggregate weight of the i-th medical client; The `clip()` function is a numeric truncation function that restricts the result to a specified range; acc j Let acc be the local verification accuracy of the j-th client; i The local validation accuracy obtained by the i-th medical client on the local validation set is used to evaluate the classification performance of the local model in the current round of the client. Parameter weighted aggregation: The parameters of the spatial-spectral SSP adapter and the dynamically dimensional adaptive DDAC classifier uploaded by each medical client are weighted and summed separately, while the global dimension mapping reference coefficient is updated. The formula is as follows: ; ; ; In the formula, The parameters for the global spatial-spectral SSP adapter in round t+1; The parameters are the global dynamic dimension adaptive DDAC classifier parameters for the (t+1)th round; The reference coefficients for global dimension mapping in round t+1; The local parameters of the spatial-spectral SSP adapter uploaded for the i-th medical client; The local parameters of the dynamically dimension-adaptive DDAC classifier uploaded for the i-th medical client; Let be the local dimension mapping coefficient for the i-th medical client; t is the current federated training round. Gradient clipping is applied to the aggregated global dimension mapping reference coefficients to ensure model convergence stability.
7. The cross-domain federal medical image classification method based on FedSSAC as described in claim 6, characterized in that, In step (3), the total local training loss for each medical client is a weighted fusion of the labeled smooth cross-entropy classification loss and the cross-site alignment loss, calculated using the following formula: ; In the formula, True labeling for MRI images; λ=0.0 is the alignment loss weight coefficient; This represents the total loss during local training. The classification loss is the labeled smooth cross-entropy; y k Predict the probability of the k-th pathological classification; For cross-site tag space alignment loss; For the k-th medical client local dimension mapping coefficient; The reference coefficient for the global dimension mapping in round t; The AdamW optimizer and cosine annealing learning rate strategy are used. After local training, the model is validated. Only when the local validation accuracy exceeds a set threshold for global accuracy are the parameters of the local spatial-spectral (SSP) adapter and the dynamically dimensional adaptive DDAC classifier uploaded to the server. This avoids low-precision parameters contaminating the global model. ; In the formula, To perform the action; For logical conditions; > This is a condition for judgment.
8. The cross-domain federal medical image classification method based on FedSSAC as described in claim 7, characterized in that, In step (3), the convergence criterion of the FedSSAC spatial-spectral adaptive collaborative framework model is that the absolute value of the difference in global validation accuracy over 5 consecutive rounds is less than 0.5%, and the maximum number of federated iteration rounds is 50. The FedSSAC spatial-spectral adaptive collaborative framework model supports random sampling training with 1-5 medical clients. When a single medical client participates, the accuracy is ≥91.2%, and when 3 or more medical clients participate, the optimal average accuracy reaches 93.7%, with a performance variance of <0.2%.
9. The cross-domain federal medical image classification method based on FedSSAC as described in claim 8, characterized in that, In step (4), the MRI images to be classified are clinical standard T1-weighted brain tumor MRI slices. Before input, Z-score standardization preprocessing is completed. The FedSSAC spatial-spectral adaptive collaborative framework model outputs four classification labels and corresponding confidence scores for glioma, meningioma, pituitary adenoma, and no tumor. The confidence score threshold is 0.
5. If it is lower than the threshold, it is marked as needing to be reviewed. If it is higher than the threshold, the classification result is directly output. For clinical sites with unknown labels, zero-sample classification is achieved through global dimension mapping coefficients. The inference accuracy is ≥85%.
10. A cross-domain federal medical image classification system based on FedSSAC, characterized in that, The cross-domain federal medical image classification method based on FedSSAC as described in claim 1 includes: The client-side local data module is used to retrieve the labeled brain tumor MRI slice dataset from the local machine of each medical client. The model building module is used to build the FedSSAC spatial-spectral adaptive collaborative framework model. It includes a frozen CLIP visual backbone, a spatial-spectral SSP adapter, a dynamic dimension adaptive DDAC classifier, and a performance-aware global aggregation module, which respectively realize basic visual feature extraction, spatial-spectral dual-domain feature calibration, heterogeneous label space adaptation, and dynamic weighted aggregation of client parameters. The training module enables each medical client to complete the training and validation of the FedSSAC spatial-spectral adaptive collaborative framework model based on local data, select qualified parameters and upload them to the server. The server updates the global model parameters through a performance-aware global aggregation strategy and distributes them, iterating the training until the model converges. The inference module inputs the preprocessed MRI images to be classified into the trained FedSSAC spatial-spectral adaptive collaborative framework model and outputs brain tumor classification results. The FedSSAC spatial-spectral adaptive collaborative framework model also supports zero-shot cross-site classification inference for clinical sites with unknown labels.
Citation Information
Patent Citations
Multivariable time series prediction method fusing federated learning and pre-training large model fine tuning
CN118690900A
Medical image classification method based on federal semi-supervised learning
CN120495750A