Medical 3D image feature fusion method, device and system, and storage medium
Through multimodal federated learning and self-supervised distillation technology, combined with deep feature extraction and knowledge graph enhancement, the problems of multimodal data fusion and data heterogeneity in medical image processing systems are solved, a high-precision, privacy-protected imaging diagnostic assistance system is realized, and diagnostic efficiency and report quality are improved.
Patent Information
- Application Number
- CN202510881281.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-23
AI Technical Summary
Existing medical image processing systems have problems such as insufficient multimodal data fusion, poor model generalization due to data heterogeneity, data silos, high dependence on annotation resources, and difficulty in generating high-level semantic representations.
It adopts multimodal federated learning and self-supervised distillation technology, combined with deep feature extraction, multi-view feature encoding, cross-modal attention module and knowledge graph enhancement, to achieve automatic extraction and semantic alignment of image features, optimize the model generalization ability through adaptive aggregation algorithm, and generate high-precision diagnostic reports.
It enables cross-institutional collaborative modeling, protects data privacy, improves the accuracy and efficiency of imaging diagnosis, meets HIPAA/GDPR compliance, and assists radiologists in quickly writing high-quality diagnostic reports.
Smart Images

Figure CN120689711A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical artificial intelligence technology, and in particular relates to a medical 3D image feature fusion method and device, system, and storage medium. Background Art
[0002] Medical image processing and diagnosis primarily involves radiologists analyzing medical image files (such as X-rays, CT scans, and MRI scans) based on their specialized radiology expertise to produce a description and diagnosis of the image files. The results of this system can be used to assist clinicians in assessing a patient's health status.
[0003] However, traditional medical image processing systems have the following limitations: insufficient multimodal data fusion. Medical images are often accompanied by multimodal information such as text reports and labels, but existing methods mostly focus on single-modal classification tasks and do not fully explore cross-modal associations. Data heterogeneity conflicts with model generalization. The data distribution of medical institutions varies significantly. Traditional federated aggregation algorithms use homogenized weight distribution and are easily interfered with by a small number of low-quality data nodes. Data heterogeneity causes global models to perform poorly in personalized tasks, and adaptive aggregation strategies are needed to balance global sharing and local characteristics. Traditional centralized learning relies on medical institutions to share raw data. The high sensitivity of medical data leads to a low willingness to share data between institutions, forming data silos, and centralized storage is vulnerable to data leakage attacks, limiting the possibility of cross-institutional collaborative modeling. Annotation resources are dependent and costly. Medical image annotation requires expert participation, which is time-consuming and labor-intensive. In addition, small-scale annotated data can easily lead to insufficient model generalization capabilities. Although existing self-supervised methods can utilize unlabeled data, they rely on pixel-level reconstruction or local area enhancement, making it difficult to generate high-semantic-level representations, which limits the adaptability of pre-trained models to downstream tasks (such as segmentation and report generation). Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a medical 3D image feature fusion method and device, system, and storage medium.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A medical 3D image feature fusion method, comprising:
[0007] Step S1, preprocessing medical 3D image data;
[0008] Step S2: extracting image features from the pre-processed medical 3D image data;
[0009] Step S3, performing multimodal feature fusion on the extracted image features;
[0010] Step S4: Based on the multimodal fusion feature as the conditional input, a multimodal joint feature encoder / decoder is used to gradually generate structured results based on the autoregressive generation strategy.
[0011] Preferably, the preprocessing in step S1 includes: DICOM format check, data enhancement, normalization and three-dimensional reconstruction.
[0012] Preferably, in step S2, image features are extracted from the preprocessed medical 3D image data through a deep feature extraction network, multi-view feature encoding and a cross-modal attention module.
[0013] Preferably, in step S3, multimodal feature fusion is performed on the extracted image features through knowledge graph enhancement and feature alignment and fusion processing.
[0014] The present invention also provides a medical 3D image feature fusion device, comprising:
[0015] A first processing module, for preprocessing medical 3D image data;
[0016] The second processing module is used to extract image features from the pre-processed medical 3D image data;
[0017] The third processing module is used to perform multimodal feature fusion on the extracted image features;
[0018] The fourth processing module is used to gradually generate structured results based on the autoregressive generation strategy using the multimodal joint feature encoder and decoder according to the multimodal fusion feature as the conditional input.
[0019] Preferably, the preprocessing in step S1 includes: DICOM format check, data enhancement, normalization and three-dimensional reconstruction.
[0020] Preferably, the second processing module extracts image features from the preprocessed medical 3D image data through a deep feature extraction network, multi-view feature encoding and a cross-modal attention module.
[0021] Preferably, the third processing module performs multimodal feature fusion on the extracted image features through knowledge graph enhancement and feature alignment and fusion processing.
[0022] The present invention also provides a medical 3D image feature fusion system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes the medical 3D image feature fusion method when executed by the processor.
[0023] The present invention also provides a storage medium having a computer program stored thereon, and the computer program executes the medical 3D image feature fusion method when running.
[0024] Based on multimodal federated learning and self-supervised distillation technology, this invention constructs a 3D medical image processing system. Through cross-institutional collaborative modeling, it realizes automatic extraction and semantic alignment of privacy-protected image features, combines the adaptive attention aggregation algorithm to optimize the generalization ability of the global model, and uses the masked self-distillation mechanism to learn high-order semantic representations from unlabeled images. Ultimately, it realizes the automatic generation of processing results of multimodal (CT / MRI images, text reports and labels) fusion, assisting radiologists to quickly complete the writing of high-precision diagnostic reports while ensuring the data privacy compliance of medical institutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0026] Figure 1 This is a flow chart of a medical 3D image feature fusion method according to an embodiment of the present invention;
[0027] Figure 2 This is a feature extraction flow chart of an embodiment of the present invention;
[0028] Figure 3 This is a feature fusion flow chart of an embodiment of the present invention;
[0029] Figure 4 This is a flowchart of intelligent diagnosis of medical 3D image sequences according to an embodiment of the present invention;
[0030] Figure 5 This is a flowchart of 3D image multi-dimensional feature extraction and self-supervised distillation according to an embodiment of the present invention;
[0031] Figure 6 This is a flowchart of multimodal feature fusion and adaptive federated aggregation in an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the drawings and specific embodiments.
[0033] Example 1:
[0034] like Figure 1 As shown, an embodiment of the present invention provides a medical 3D image feature fusion method, comprising:
[0035] Step S1, preprocessing medical 3D image data;
[0036] Step S2: extracting image features from the pre-processed medical 3D image data;
[0037] Step S3, performing multimodal feature fusion on the extracted image features;
[0038] Step S4: Based on the multimodal fusion feature as the conditional input, a multimodal joint feature encoder / decoder is used to gradually generate structured results based on the autoregressive generation strategy.
[0039] As an implementation of an embodiment of the present invention, in step S1, preprocessing includes: DICOM format check, data enhancement, normalization and three-dimensional reconstruction.
[0040] DICOM format check: Using a specialized format verification algorithm, the system performs a comprehensive integrity check on the file header tags (covering key information such as Modality and PatientID) of uploaded medical 3D imaging data, accurately filtering out files that do not comply with the DICOM standard, and ensuring that the data format for subsequent processing is standardized and unified.
[0041] Data augmentation: We perform data augmentation on image data based on random image transformations such as rotation, scaling, and translation. Specifically, we set reasonable rotation angle ranges (e.g., ±30 degrees), scaling ratio ranges (e.g., 0.8x-1.2x), and translation pixel distances (e.g., ±10 pixels in both the x and y directions) to expand the dataset and improve the model's generalization capabilities for different types of image changes.
[0042] Normalization: Use the linear normalization method to uniformly map the pixel values of the image data to the [0,1] interval. At the same time, standardize the image parameters such as size and resolution to ensure that they strictly meet the data specification requirements of subsequent processing links.
[0043] 3D reconstruction: For applicable types of images, such as abdominal CT images, multi-planar reconstruction technology is used to generate cross-sectional, sagittal, and coronal plane sequences based on the original image data, presenting image information from different anatomical perspectives to facilitate subsequent multi-angle analysis.
[0044] As an implementation method of the embodiment of the present invention, in step S2, as Figure 2 As shown, specifically including:
[0045] Deep Feature Extraction Network: The ResNet-101 deep learning network architecture is used to perform deep feature extraction from images. Taking ResNet-101 as an example, it automatically learns key image features such as edges, textures, and shapes through multi-layer convolution operations (with kernel sizes ranging from 3×3 to 1×1). Each convolution operation is followed by a nonlinear activation function (such as ReLU) to enhance feature representation. Pooling operations (such as max pooling and average pooling) are used to reduce the dimensionality of the feature space, reducing computational effort while preserving core feature information.
[0046] Multi-view feature encoding: The image sequence data of the transverse, sagittal, and coronal planes are input into the above-mentioned deep feature extraction network respectively. After the network forward propagation calculation, the deep feature vector corresponding to each view is obtained. In this way, independent encoding of multi-view features is achieved, laying a solid foundation for subsequent fusion operations.
[0047] Cross-modal attention module: Constructs a feature fusion architecture based on the attention mechanism. Its core idea is to calculate the similarity weights between features from different perspectives. The specific formula is: Similarity weight matrix S = softmax(W1·tanh(W2·f i +W3·f j )), where f i and f j Representing the feature vectors of different perspectives, through the fully connected layer (weight matrix W1, W2, W3)
[0048] The weight S is calculated by combining the activation functions (softmax and tanh), and the features are weighted and fused accordingly, so that the model focuses on important feature information with more diagnostic value and outputs multi-dimensional image feature coding.
[0049] As an implementation method of the embodiment of the present invention, in step S3, Figure 3 As shown, specifically including:
[0050] Knowledge graph enhancement: Based on the diagnostic label, the graph neural network (GNN) algorithm is used to retrieve the pathological characteristics, treatment plans and other related nodes in the knowledge graph to generate a knowledge embedding vector. For example, for the liver cancer diagnostic label, the typical pathological characteristics of liver cancer at different stages, common treatment methods and other node information are retrieved. The knowledge embedding vector k is obtained through GNN aggregation calculation and is spliced and fused with the image feature code f to form a fusion feature f. concat =[f;k], realizing the organic combination of image information and domain knowledge.
[0051] Feature alignment and fusion: With the help of feature alignment technology, image features and knowledge encoding are mapped to the same semantic space. Linear transformation (such as full connection layer) is used to transform image features f img and knowledge encoding ktext Transform it so that it is in the same feature space and obtain the aligned feature f' img =W img ·f img +b img , k' text =W text ·k text +b text , and then dynamically adjust the weights (the weight calculation formula can be determined based on the statistical results of the importance of each feature in the training data, such as Where score is the pre-set feature importance score), through the weighted summation formula f fuse =w1·f' img +w2·k' text Fusion is performed to finally obtain the multimodal fusion feature f fuse .
[0052] As an implementation method of an embodiment of the present invention, in step S4, using multimodal fusion features as conditional input, a multimodal joint feature encoder / decoder is used to gradually generate structured results, such as the text description of the diagnostic report, based on an autoregressive generation strategy. At the same time, a text correction mechanism based on a rule engine is introduced to correct the generated text according to pre-established clinical standard rules (such as standard expressions of diagnostic terms, report format specifications, etc.), ensuring that the output report accurately meets actual clinical needs, and ultimately providing radiologists with high-quality auxiliary diagnosis basis.
[0053] This invention is a close-knit process from data preprocessing, feature extraction, feature fusion to result output, fully utilizing multimodal federated learning and self-supervised distillation technology, which not only ensures data privacy but also improves the accuracy and efficiency of medical 3D imaging diagnosis, and is expected to play an important role in actual medical scenarios.
[0054] By combining a multimodal federated learning framework with self-supervised distillation technology, an efficient and secure medical 3D image feature fusion processing system has been constructed. Doctors can log in to the platform through a unified authentication system (SSO), create patient profiles, and upload image sequences such as abdominal CT scans. The system automatically performs encrypted preprocessing and privacy-protected transmission of DICOM-formatted images. Masked self-distillation is used to extract high-level semantic representations from unlabeled data, and a teacher-student model architecture is used to optimize segmentation accuracy. The uploaded images undergo multi-view feature encoding and cross-modal attention fusion to generate joint features. Knowledge graph retrieval (e.g., liver cancer staging criteria) is also used to enhance semantic alignment and form a multimodal feature encoding. The federated learning framework employs an adaptive aggregation algorithm to dynamically allocate model weights to address data heterogeneity, and differential privacy noise and homomorphic encryption techniques are used to mitigate privacy risks. Finally, the system generates structured processing results using a multimodal joint feature encoder-decoder. Based on a knowledge distillation compression model, the system achieves lightweight deployment, assisting radiologists in rapidly producing high-precision diagnostic reports while ensuring data privacy compliance within medical institutions. While protecting data privacy, the system greatly improves the work efficiency of radiologists and provides high-precision and high-efficiency automation solutions for cross-institutional collaboration. It also meets HIPAA / GDPR compliance requirements and promotes the clinical implementation and standardized application of medical AI technology.
[0055] The medical 3D image feature fusion method according to the embodiment of the present invention is used to perform intelligent diagnosis of medical 3D image sequences, such as Figure 4 As shown, including:
[0056] 1. Upload CT image sequence
[0057] Patient file management supports doctors to create / modify patient files (including basic information and medical history data). The file management system automatically detects whether the uploaded files comply with the DICOM standard and encrypts the transmission of image files.
[0058] 2. Liver region extraction
[0059] The image files are automatically identified and processed, and then segmented and masked using an adaptive segmentation algorithm to obtain 3D image sequence data of the liver area with an adaptive number of layers, which are then submitted to the image coding processing module.
[0060] 3. Image coding processing
[0061] The image features of the cross-sectional, sagittal, and coronal perspectives are extracted respectively, and combined with the cross-modal feature alignment module to perform feature dimensionality reduction and fusion to obtain multi-dimensional image fusion feature encoding.
[0062] 4. Classification of imaging diagnosis
[0063] A multi-label classifier based on multi-dimensional image fusion feature encoding training, combined with a three-layer fully connected network and a weighted cross entropy loss function, performs diagnostic classification on the input 3D image sequence.
[0064] 5. Knowledge Graph Enhancement
[0065] According to the image diagnosis classification results, semantic coding is obtained by combining node matching.
[0066] 6. Multimodal feature fusion
[0067] Calculate the similarity matrix between image features and knowledge encoding to generate alignment weights. Weighted fusion of image features and knowledge encoding is performed to output a multimodal vector, and the fused features are normalized.
[0068] 7. Image processing result generation
[0069] Based on multimodal fusion features and using a large medical language model, 3D image processing results are generated, and post-processing such as terminology correction and structured output is performed.
[0070] like Figure 5 As shown, the image encoding process in step 3 includes:
[0071] Step 31: Initialize the multimodal federated learning framework. Each medical institution registers with the central server via a RESTful API and is assigned a unique node ID. Next, integrate a compliance detection module to automatically scan sensitive fields in patient records and dynamically desensitize data fields that do not comply with privacy regulations.
[0072] Step 32: Pre-train the self-supervised distillation model to process the complete image, optimizing feature predictions with a smoothed L1 loss. Input the transverse, sagittal, and coronal image sequences into the model, outputting a multidimensional feature vector that is then aligned using cross-modal attention.
[0073] Step 33: The server calculates the attention weight based on the node model performance, the edge device loads the encrypted model parameters, encodes the image features under the three perspectives, and inputs them into the multidimensional feature fuser to generate multidimensional image fusion features.
[0074] like Figure 6 As shown, the multimodal feature fusion in step 6 includes:
[0075] Step 61: Input the multi-dimensional image fusion features into the multi-label classifier. The input layer receives the multimodal image feature encoding (from the cross-modal attention fusion module). The hidden layer uses a 3-layer fully connected network to perform data dimensionality reduction. The output layer activates the output diagnostic label to obtain the image processing classification result.
[0076] Step 62: Use the classification result output by the image diagnosis classification module as a knowledge node retrieval input parameter, and obtain the knowledge feature code through the knowledge search engine and semantic coding conversion.
[0077] Step 63: Fuse the spliced image features with the attention weight output and obtain multimodal fusion features through normalization.
[0078] Example 2:
[0079] An embodiment of the present invention further provides a medical 3D image feature fusion device, comprising:
[0080] A first processing module, for preprocessing medical 3D image data;
[0081] The second processing module is used to extract image features from the pre-processed medical 3D image data;
[0082] The third processing module is used to perform multimodal feature fusion on the extracted image features;
[0083] The fourth processing module is used to gradually generate structured results based on the autoregressive generation strategy using the multimodal joint feature encoder and decoder according to the multimodal fusion feature as the conditional input.
[0084] As an implementation of the embodiment of the present invention, the preprocessing in step S1 includes: DICOM format check, data enhancement, normalization and three-dimensional reconstruction.
[0085] As an implementation method of an embodiment of the present invention, the second processing module extracts image features from the preprocessed medical 3D image data through a deep feature extraction network, multi-view feature encoding and a cross-modal attention module.
[0086] As an implementation method of an embodiment of the present invention, the third processing module performs multimodal feature fusion on the extracted image features through knowledge graph enhancement and feature alignment and fusion processing.
[0087] Example 3:
[0088] An embodiment of the present invention also provides a medical 3D image feature fusion system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a medical 3D image feature fusion method when executed by the processor.
[0089] Example 4:
[0090] An embodiment of the present invention further provides a storage medium having a computer program stored thereon, and the computer program executes the medical 3D image feature fusion method when running.
[0091] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A medical 3D image feature fusion method, characterized in that: include: Step S1, preprocessing medical 3D image data; Step S2: extracting image features from the pre-processed medical 3D image data; Step S3, performing multimodal feature fusion on the extracted image features; Step S4: Based on the multimodal fusion feature as the conditional input, a multimodal joint feature encoder / decoder is used to gradually generate structured results based on the autoregressive generation strategy.
2. The medical 3D image feature fusion method according to claim 1, wherein: The preprocessing of step S1 includes: DICOM format check, data enhancement, normalization and 3D reconstruction.
3. The medical 3D image feature fusion method according to claim 1, wherein: In step S2, image features are extracted from the preprocessed medical 3D image data through a deep feature extraction network, multi-view feature encoding, and cross-modal attention module.
4. The medical 3D image feature fusion method according to claim 3, wherein: In step S3, multimodal feature fusion is performed on the extracted image features through knowledge graph enhancement and feature alignment and fusion processing.
5. A medical 3D image feature fusion device, characterized in that: include: A first processing module, for preprocessing medical 3D image data; The second processing module is used to extract image features from the pre-processed medical 3D image data; The third processing module is used to perform multimodal feature fusion on the extracted image features; The fourth processing module is used to gradually generate structured results based on the autoregressive generation strategy using the multimodal joint feature encoder and decoder according to the multimodal fusion feature as the conditional input.
6. The medical 3D image feature fusion device according to claim 5, wherein: The preprocessing of step S1 includes: DICOM format check, data enhancement, normalization and 3D reconstruction.
7. The medical 3D image feature fusion device according to claim 6, wherein: The second processing module extracts image features from the preprocessed medical 3D imaging data through a deep feature extraction network, multi-view feature encoding, and cross-modal attention module.
8. The medical 3D image feature fusion device according to claim 7, wherein: The third processing module performs multimodal feature fusion on the extracted image features through knowledge graph enhancement and feature alignment and fusion processing.
9. A medical 3D image feature fusion system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the medical 3D image feature fusion method according to any one of claims 1 to 4 is executed.
10. A storage medium, characterized in that: The storage medium stores a computer program, which, when run, executes the medical 3D image feature fusion method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Image modal conversion method and system based on multi-scale cross-modal alignment network
CN118982735A
Multi-modal protein sequence generation method based on deep learning
CN119943134A