Multi-modal medical image data analysis method and device, equipment and medium

Through the split task processing of the edge-cloud collaborative architecture and hybrid model, combined with the Transformer and 3D convolutional neural network modules, the modal weights are dynamically adjusted, which solves the problems of insufficient complex lesion recognition and poor adaptability of multimodal fusion in the existing system, and realizes efficient and accurate multimodal medical image analysis.

CN120726019APending Publication Date: 2025-09-30KANG JIAN INFORMATION TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511053388.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing medical image analysis systems lack sensitivity in identifying complex lesions, and their algorithms have poor adaptability and high misjudgment rates during multimodal image fusion analysis.

Method used

An edge-cloud collaborative architecture is adopted to split image analysis tasks into preprocessing and real-time preliminary screening tasks and in-depth analysis tasks. A simplified hybrid model is deployed on edge devices for preprocessing and preliminary screening, and a complete hybrid model is deployed on the cloud for in-depth analysis. The hybrid model includes a Transformer module and a 3D convolutional neural network module, and a dynamic weight allocation mechanism is introduced to adjust the modal weights.

Benefits of technology

It significantly improves the sensitivity of identifying complex lesions, reduces the misjudgment rate, and realizes efficient and accurate collaborative analysis of multimodal images to meet clinical real-time needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726019A_ABST
    Figure CN120726019A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal medical image data analysis method and device, equipment and a medium, relates to the technical field of medical treatment, can be applied to a medical health business scene, and comprises the following steps: obtaining an image analysis task corresponding to multi-modal medical image data; the image analysis task is divided into a first computing power task and a second computing power task, the first computing power task is a preprocessing and real-time preliminary screening task, the second computing power task is a deep analysis task, and the output of the first computing power task serves as the input of the second computing power task; scheduling a first hybrid model deployed at the edge device side to execute a first computing power task, and scheduling a second hybrid model deployed at the cloud side to execute a second computing power task; and outputting a collaborative image analysis result of the first mixed model and the second mixed model for the multi-modal medical image data. According to the invention, the complex focus recognition sensitivity can be improved, and the misjudgment rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical technology, and in particular to a multimodal medical imaging data analysis method, apparatus, equipment, and medium. Background Art

[0002] With the rapid development of medical imaging, multimodal imaging technologies such as computed tomography (CT), magnetic resonance imaging (MRI), and positron emission tomography-computed tomography (PET-CT) have become the core basis for disease diagnosis, playing a key role in early tumor screening, cardiovascular and cerebrovascular disease assessment, and other fields. At the same time, the popularity of telemedicine has promoted the demand for cross-institutional collaboration, and the dependence of grassroots hospitals on expert resources and intelligent auxiliary diagnosis has increased significantly. In this context, medical image analysis systems based on artificial intelligence (AI) have gradually become the core tools connecting imaging data, clinical diagnosis, and remote collaboration. Their performance directly affects the accuracy and efficiency of diagnosis and treatment.

[0003] Existing medical image analysis systems often use traditional image processing algorithms or AI models trained on a single modality, while remote assisted diagnosis functions primarily rely on image transmission and basic annotation. Although some systems attempt to integrate multimodal data or enhance remote interaction, overall technical limitations persist: traditional models lack sensitivity for identifying complex lesions (such as early-stage tumors and small nodules), and algorithms for multimodal image fusion analysis suffer from poor adaptability and high misjudgment rates. Summary of the Invention

[0004] In view of this, the present application provides a multimodal medical imaging data analysis method, apparatus, equipment and medium, which can improve the sensitivity of complex lesion identification and reduce the misjudgment rate.

[0005] According to a first aspect of the present application, a multimodal medical imaging data analysis method is provided, comprising:

[0006] Obtain image analysis tasks corresponding to multimodal medical imaging data;

[0007] Split the image analysis task into a first computing task and a second computing task, wherein the first computing task is a preprocessing and real-time initial screening task, and the second computing task is a deep analysis task, and the output of the first computing task serves as the input of the second computing task;

[0008] Scheduling a first hybrid model deployed on the edge device to perform the first computing task, and scheduling a second hybrid model deployed on the cloud to perform the second computing task. The first hybrid model is a simplified version of the hybrid model, and the second hybrid model is a complete version of the hybrid model. The hybrid model includes a Transformer module for extracting global correlation features and a 3D convolutional neural network module for extracting local spatial features. A dynamic weight allocation mechanism is introduced to adjust the weight coefficients of different modal images.

[0009] Output collaborative image analysis results of the first hybrid model and the second hybrid model on the multimodal medical image data.

[0010] According to a second aspect of the present application, a multimodal medical image data analysis device is provided, comprising:

[0011] An acquisition module is used to obtain image analysis tasks corresponding to multimodal medical image data;

[0012] a splitting module, configured to split the image analysis task into a first computing task and a second computing task, wherein the first computing task is a preprocessing and real-time initial screening task, and the second computing task is a deep analysis task, and the output of the first computing task serves as the input of the second computing task;

[0013] An execution module is configured to schedule a first hybrid model deployed on an edge device to execute the first computing task, and schedule a second hybrid model deployed on the cloud to execute the second computing task. The first hybrid model is a simplified version of the hybrid model, and the second hybrid model is a complete version of the hybrid model. The hybrid model includes a Transformer module for extracting global correlation features and a 3D convolutional neural network module for extracting local spatial features, and introduces a dynamic weight allocation mechanism to adjust the weight coefficients of different modal images.

[0014] An output module is used to output collaborative image analysis results of the first hybrid model and the second hybrid model on the multimodal medical image data.

[0015] According to a third aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned multimodal medical imaging data analysis method is implemented.

[0016] According to a fourth aspect of the present application, an electronic device is provided, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the above-mentioned multimodal medical imaging data analysis method when executing the program.

[0017] With the help of the above-mentioned technical solution, the multimodal medical imaging data analysis method, device, equipment and medium provided by this application, by adopting a hybrid model including a Transformer module and a 3D convolutional neural network module, can not only capture the global correlation features of multimodal images through the Transformer, but also extract the local spatial features of lesions with the help of a 3D convolutional neural network, which can significantly improve the recognition sensitivity of complex lesions; the introduction of a dynamic weight allocation mechanism to dynamically adjust the weight coefficients of different modal images can enhance the adaptability of the algorithm during multimodal image fusion analysis and reduce the misjudgment rate; at the same time, through the edge-cloud collaborative architecture, the simplified version of the edge device quickly completes preprocessing and real-time initial screening, and the full version of the cloud model performs in-depth analysis, which can improve diagnostic accuracy while ensuring real-time performance, overcome the technical shortcomings of traditional systems in complex lesion identification and multimodal fusion, and realize efficient and accurate collaborative analysis of multimodal medical images.

[0018] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 A schematic diagram of a process for analyzing multimodal medical imaging data provided in an embodiment of the present application is shown;

[0021] Figure 2 A schematic diagram of a process for analyzing multimodal medical image data according to another embodiment of the present application is shown;

[0022] Figure 3 A schematic diagram of the structure of a multimodal medical image data analysis device provided in an embodiment of the present application is shown;

[0023] Figure 4 A schematic structural diagram of a multimodal medical image data analysis device provided in another embodiment of the present application is shown. DETAILED DESCRIPTION

[0024] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0025] Existing medical image analysis systems often use traditional image processing algorithms or AI models trained on a single modality, while remote assisted diagnosis functions primarily rely on image transmission and basic annotation. Although some systems attempt to integrate multimodal data or enhance remote interaction, overall technical limitations persist: traditional models lack sensitivity for identifying complex lesions (such as early-stage tumors and small nodules), and algorithms for multimodal image fusion analysis suffer from poor adaptability and high misjudgment rates.

[0026] In order to solve the above technical problems, the embodiment of the present invention provides a multimodal medical image data analysis method, such as Figure 1 As shown, the method includes:

[0027] Step 110: Obtain an image analysis task corresponding to the multimodal medical image data.

[0028] Among them, multimodal medical imaging data refers to a collection of images of the same subject obtained through two or more imaging technologies (such as brain images scanned jointly by CT and MRI, and tumor images fused by PET and CT). Different modal data complement each other in reflecting the characteristics of lesions from the dimensions of structure, function, metabolism, etc.; image analysis tasks refer to specific processing goals for medical imaging data, including but not limited to lesion detection (locating lesion location), qualitative analysis (determining whether it is benign or malignant), quantitative measurement (calculating size / volume), efficacy evaluation (comparing changes before and after treatment), etc., which need to be dynamically determined according to the clinical scenario.

[0029] Step 120: Split the image analysis task into a first computing task and a second computing task. The first computing task is a preprocessing and real-time initial screening task, and the second computing task is a deep analysis task. The output of the first computing task serves as the input of the second computing task.

[0030] Among them, the first computing task refers to the lightweight processing task performed by the edge device, including preprocessing (noise reduction, standardization, etc.) and real-time initial screening of multimodal images. The core goal is to quickly extract key information of suspected lesion areas and provide a basis for subsequent in-depth analysis. It is characterized by fast response speed and low computing power requirements; the second computing task refers to the in-depth analysis task performed by the cloud. Based on the initial screening results output by the first computing task, it can perform complex calculations such as three-dimensional reconstruction, lesion nature determination, and growth trend simulation. It relies on high-performance computing power support on the cloud and aims to provide accurate and comprehensive diagnostic basis; preprocessing is the preliminary processing step after the multimodal image enters the analysis process, including noise reduction (removing the image acquisition process) The real-time initial screening task is to use a simplified hybrid model on the edge device to quickly scan the pre-processed multimodal images, locate the suspected lesions and extract voxel data. It focuses on real-time performance and can provide rapid response support for emergency scenarios. The deep analysis task is an advanced analysis task performed on the cloud based on the initial screening results. The Transformer module and 3D convolutional neural network of the complete hybrid model can be used to calculate the global correlation features and local fine features of the lesions, and the dynamic weight distribution mechanism can be used to determine the nature of the lesions, simulate the development trend, and output comprehensive diagnostic results.

[0031] For the disclosed embodiment, the type of image analysis task (such as emergency initial screening, deep diagnosis of tumors) and the corresponding multimodal image data characteristics (such as data volume, modality combination, and urgency) can be first identified, and then the task can be split based on the hierarchical processing logic of the edge-cloud collaborative architecture. The first computing task focuses on preprocessing and real-time initial screening, including basic processing such as noise reduction and standardization of multimodal images, and the rapid positioning of suspected lesions through a simplified hybrid model deployed on edge devices, and outputs initial screening results containing voxel data of the lesion area; the second computing task is in-depth analysis, which uses the initial screening results as input and is executed by the complete hybrid model on the cloud, including complex calculations such as three-dimensional reconstruction, lesion property determination, and growth trend simulation. During the splitting process, the scenario requirements will be dynamically adapted. For example, emergency cases will prioritize compressing the initial screening time, and large-volume image data will automatically increase the cloud computing power allocation weight to ensure smooth task connection.

[0032] This task-splitting mechanism significantly improves the efficiency and accuracy of multimodal image analysis through a collaborative model of "lightweight processing at the edge + deep computing in the cloud." Edge devices quickly complete pre-processing and initial screening, reducing the amount of invalid data transmitted to the cloud and lowering network load. The cloud focuses on high-computing deep analysis, leveraging GPU clusters to shorten the time required for complex tasks and avoid bottlenecks caused by insufficient computing power on edge devices.

[0033] Step 130: Schedule the first hybrid model deployed on the edge device to perform the first computing task, and schedule the second hybrid model deployed on the cloud to perform the second computing task. The first hybrid model is a simplified version of the hybrid model, and the second hybrid model is a complete version of the hybrid model. The hybrid model includes a Transformer module for extracting global correlation features and a 3D convolutional neural network module for extracting local spatial features, and introduces a dynamic weight allocation mechanism to adjust the weight coefficients of different modal images.

[0034] Among them, the first hybrid model is a simplified hybrid model deployed on edge devices. It is small in size, retains the core functions of the Transformer and 3D convolutional neural network modules but simplifies the calculation process, and focuses on processing preprocessing and real-time initial screening tasks. It is characterized by fast response and low computing power requirements; the second hybrid model is a complete hybrid model deployed on the cloud. It integrates the complete Transformer and 3D convolutional neural network modules through containerization technology, supports complex global correlation feature calculations and local feature extraction, and is suitable for deep analysis tasks with high computing power requirements; the Transformer module is the component in the hybrid model responsible for extracting global correlation features, and is good at It can capture the correlation between multimodal images (such as CT and MRI) (such as the corresponding positions of lesions in different modalities) and strengthen the effective correlation through the multi-head self-attention mechanism; the 3D convolutional neural network module is the component in the hybrid model responsible for extracting local spatial features, focusing on capturing the three-dimensional local details of the lesions (such as nodule shape and boundary clarity) to improve the recognition ability of complex lesions such as small nodules and early tumors; the dynamic weight allocation mechanism is a mechanism in the hybrid model that automatically adjusts the weights of different modal images according to the lesion type. For example, when analyzing lung cancer, the CT weight is increased, and when analyzing brain tumors, the MRI weight is increased. This can maximize the advantages of each modality and reduce the misjudgment rate.

[0035] When pre-training the hybrid model, the embodiment steps may include: obtaining model parameters obtained after each medical node trains a single-modal sub-model based on local single-modal medical imaging data, the single-modal sub-model includes a Transformer module and a 3D convolutional neural network module, the Transformer module is used to convert the local single-modal medical image into a feature vector sequence, and calculate the global correlation features through a multi-head self-attention mechanism, and the 3D convolutional neural network module is used to extract local features based on the global correlation features to obtain local spatial features; calling the dynamic weight allocation mechanism, adjusting the weight coefficient of each modality based on the feature contribution of each single-modal sub-model in local training, fusing the model parameters of each single-modal sub-model according to the adjusted weight coefficient, and realizing multi-modal feature association through a cross-modal attention mechanism to obtain a trained hybrid model.

[0036] This training process, through the collaborative model of "local single-modality refinement + cross-institutional parameter fusion", can effectively solve the pain points of the "island problem" of medical data and the poor adaptability of multimodal fusion: the federated learning framework allows each hospital to contribute training experience while protecting data privacy, so that the model can learn richer multi-center data features; the dynamic weight allocation mechanism adjusts the weight according to the actual contribution of each modality, which can avoid invalid modal interference, and combined with the cross-modal attention mechanism to enhance the use of complementary information between modalities, it can significantly improve the hybrid model's ability to recognize complex lesions.

[0037] Accordingly, when calling the dynamic weight allocation mechanism and adjusting the weight coefficient of each modality based on the feature contribution of each single-modal sub-model in local training, first, the initial weight coefficient of each modality can be set according to medical prior knowledge (such as the display advantage of CT for lung structure in lung cancer diagnosis, and the high-resolution characteristics of MRI for soft tissue in brain tumor analysis); secondly, the feature value is quantified through the two dimensions of local feature discriminability and global correlation strength. For local feature discriminability, the local spatial features (such as nodule size and density) extracted by the 3D convolutional neural network in each unimodal sub-model can be matched with the local pathological annotation results (such as biopsy reports), and the matching accuracy can be used to measure the ability of the modality to independently identify lesion details. For global correlation strength, the average correlation weight of the global correlation features (such as the spatial correlation between the lesion and surrounding tissue) output by the Transformer in each unimodal sub-model and the features of other modalities is calculated. The higher the weight, the more significant the role of the modality in cross-modal collaboration. Finally, the local feature discriminability (reflecting the reliability of unimodal diagnosis) and the global correlation strength (reflecting the value of cross-modal collaboration) are combined. The feature contribution is calculated through a weighted algorithm (such as 60% for local feature discriminability and 40% for global correlation strength). The weight coefficient is dynamically adjusted according to the contribution (for example, if a modality has high discriminability and strong correlation, its weight is increased, and vice versa), so that the weight distribution is consistent with clinical laws and adapted to actual data characteristics.

[0038] Accordingly, when calling the dynamic weight allocation mechanism and adjusting the weight coefficient of each modality based on the feature contribution of each single-modal sub-model in local training, the following implementation steps may be included: setting the initial weight coefficient of each modality based on medical prior knowledge; determining the local feature discrimination of each single-modal sub-model based on the matching accuracy by matching the local spatial features output by the 3D convolutional neural network module in each modality sub-model with the locally annotated pathological results; calculating the average correlation weight of the global correlation features output by the Transformer module in each single-modal sub-model and the features of other modalities, and determining the global correlation strength corresponding to the average correlation weight; determining the feature contribution of each modality based on the local feature discrimination and the global correlation strength, and adjusting the weight coefficient of each modality based on the feature contribution.

[0039] In specific application scenarios, after the hybrid model is trained, model versions adapted to different scenarios can be generated and deployed through targeted processing. For the first hybrid model on the edge device side, model optimization technology (such as TensorRT compression) can be used to lightweight the hybrid model, streamline the computing link, retain the core feature extraction functions (such as basic Transformer global correlation calculation and 3D convolution local feature extraction), and make the compressed model volume adapt to the architecture of edge devices such as CT machines and ultrasound machines to ensure that it can run quickly locally; for the second hybrid model on the cloud, containerization technology (similar to packaging program) can be used to encapsulate the complete version of the hybrid model, integrate all functional modules (such as complete cross-modal attention mechanism, dynamic weight distribution logic), rely on Kubernetes to achieve elastic scheduling of GPU resources, and deploy on high-performance server clusters to support deep analysis tasks with high computing power requirements. After deployment is completed, the edge device and the cloud are connected through the hospital intranet or 5G network to form a collaborative closed loop of "local fast response + remote deep computing". Accordingly, the steps of the embodiment may include: using model optimization technology to lightweight the hybrid model, and deploying the obtained first hybrid model on the edge device end; packaging the complete version of the hybrid model through containerization technology, and deploying the obtained second hybrid model on the cloud.

[0040] For the disclosed embodiment, the edge device end (such as a CT machine or an ultrasound machine with its own terminal) can deploy a lightweight and optimized first hybrid model (such as a volume <50MB). For the first computing task (preprocessing and real-time initial screening), a simplified version of the Transformer module is used to quickly capture the basic global correlation of multimodal images (such as the position correspondence of lesions in different modalities), and a simplified 3D convolutional neural network module is used to extract local key features (such as nodule size and density) to complete noise reduction, standardization, and initial screening of suspected lesions. The cloud deploys a complete version of the second hybrid model through containerization technology. After receiving the initial screening results output by the edge device, the complete Transformer module is started to deeply calculate the global correlation features of the multimodal images (such as the three-dimensional correlation between the lesion and the surrounding tissue), and with the help of the 3D convolutional neural network module, fine local feature extraction (such as the internal texture of the tumor) is performed. At the same time, the dynamic weight allocation mechanism adjusts the modality contribution according to the lesion type (such as lung cancer prioritizes CT weighting, and brain tumors prioritize MRI weighting), and finally completes the deep analysis task. The entire process synchronizes data in real time through the network, forming a closed loop of fast edge processing and deep cloud computing.

[0041] This scheduling mechanism significantly improves the efficiency and accuracy of multimodal image analysis through "model hierarchical deployment + functional collaboration." Specifically, the simplified model on edge devices quickly responds to real-time needs, reducing invalid data transmission and avoiding network congestion. The complete cloud-based model, relying on GPU clusters, shortens the time required for in-depth analysis and overcomes the computing power limitations of edge devices. At the same time, the hybrid model's Transformer and 3D convolutional neural network modules work together, combined with a dynamic weight allocation mechanism to specifically enhance the feature contributions of key modalities. Experiments have shown that the sensitivity for identifying early-stage lung cancer reaches 98.2%. This effectively addresses the low sensitivity of traditional single models for identifying complex lesions and poor adaptability to multimodal fusion, while simultaneously addressing both clinical real-time requirements and diagnostic accuracy.

[0042] Step 140: Output collaborative image analysis results of the first hybrid model and the second hybrid model on the multimodal medical image data.

[0043] Among them, the collaborative image analysis result is the final output generated by the first hybrid model on the edge device and the second hybrid model on the cloud through division of labor and cooperation to jointly analyze the multimodal medical imaging data.

[0044] Specifically, the first hybrid model (simplified hybrid model) of the edge device can first pre-process and perform real-time initial screening of multimodal images, extract voxel data of suspected lesion areas, and generate initial screening results including lesion location and preliminary morphology; the second hybrid model on the cloud uses the initial screening results as input, calculates global correlation features (such as the three-dimensional correlation between lesions and surrounding tissues) through a complete Transformer module, and extracts local fine features (such as tumor texture) through a 3D convolutional neural network module, and combines a dynamic weight allocation mechanism to generate in-depth analysis results including lesion size, nature and diagnostic judgment; finally, the edge device integrates the initial screening results and the in-depth analysis results to form a collaborative image analysis result that is both real-time and accurate.

[0045] In summary, the multimodal medical imaging data analysis method provided by the present invention adopts a hybrid model including a Transformer module and a 3D convolutional neural network module, which can not only capture the global correlation features of multimodal images through the Transformer, but also extract the local spatial features of lesions with the help of a 3D convolutional neural network, which can significantly improve the recognition sensitivity of complex lesions; the introduction of a dynamic weight allocation mechanism to dynamically adjust the weight coefficients of different modal images can enhance the adaptability of the algorithm during multimodal image fusion analysis and reduce the misjudgment rate; at the same time, through the edge-cloud collaborative architecture, the simplified version of the edge device quickly completes preprocessing and real-time initial screening, and the full version of the cloud model performs in-depth analysis, which can improve diagnostic accuracy while ensuring real-time performance, overcome the technical shortcomings of traditional systems in complex lesion recognition and multimodal fusion, and realize efficient and accurate collaborative analysis of multimodal medical images.

[0046] Furthermore, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the implementation of this embodiment, this embodiment also provides another multimodal medical imaging data analysis method, such as Figure 2 As shown, the method includes:

[0047] Step 210: Obtain an image analysis task corresponding to the multimodal medical image data.

[0048] Step 220: Split the image analysis task into a first computing task and a second computing task. The first computing task is a preprocessing and real-time initial screening task, and the second computing task is a deep analysis task. The output of the first computing task serves as the input of the second computing task.

[0049] For the embodiment of the present disclosure, the specific implementation process can be found in the relevant description of step 120 of the embodiment, which will not be repeated here.

[0050] Step 230: Schedule the first hybrid model deployed on the edge device to preprocess and perform real-time preliminary screening on the multimodal medical imaging data, extract voxel data of the suspected lesion area, and encrypt and transmit the preliminary screening results containing the voxel data to the cloud. The preprocessing includes at least noise reduction and standardization.

[0051] Among them, voxel data is the basic unit data that constitutes the lesion area in three-dimensional medical images, which contains spatial coordinates and grayscale value information, and is the core data that describes the three-dimensional morphology of the lesion; the initial screening result is the preliminary analysis result of the image by the edge device, which includes the suspected lesion location, voxel data and basic morphological information, etc., which serves as the input for cloud-based deep analysis.

[0052] For the embodiment of the present disclosure, the first hybrid model deployed on the edge device can start the preprocessing process after receiving the multimodal medical imaging data, and complete noise reduction (removing interference signals from image acquisition) and standardization (unifying the formats and grayscale parameters of different devices and modalities) through the built-in algorithm to ensure that the data format is suitable for subsequent analysis; then, the first hybrid model can call the simplified Transformer module and the 3D convolutional neural network module to quickly capture the basic global correlation (such as the cross-modal lesion position correspondence) and local features (such as nodule size and density) of suspected lesions in the image, locate and extract voxel data (pixel information in three-dimensional space) of the suspected lesion area, and generate a preliminary screening result including the lesion location and preliminary morphology; finally, the encrypted preliminary screening results (including voxel data) can be transmitted to the cloud via the hospital intranet or 5G network, which can provide input for subsequent in-depth analysis. The response time of the entire process is short and can meet real-time requirements.

[0053] This process is processed locally on edge devices, enabling rapid response and data reduction for multimodal image analysis. Specifically, the preprocessing step eliminates data noise and format differences, providing consistent input for subsequent model analysis and avoiding interference from invalid features. Real-time initial screening uses simplified models to quickly locate suspected lesions, reducing the amount of data transmitted to the cloud (only voxel data in the lesion area is transmitted), reducing network load and transmission delays. Encrypted transmission protects patient privacy and data security, meeting medical data compliance requirements.

[0054] In a specific application scenario, before the first hybrid model deployed on the scheduling edge device performs preprocessing and real-time initial screening on the multimodal medical imaging data, as an optimal implementation method, an adaptive desensitization process can be started. Specifically, the metadata of the multimodal medical imaging data (such as DICOM files) can be first identified, and the information containing the patient's identity (such as name, ID number, medical record number, and other DICOM tag fields) can be located through the preset sensitive field library; then, field masking technology (such as replacing with anonymous identifiers, deleting sensitive fields) is used to desensitize this information, while retaining the image feature parameters necessary for lesion analysis (such as pixel grayscale value, layer thickness, modality type, etc.); after desensitization is completed, the processing results are automatically verified to ensure that there is no identity information remaining and the image feature parameters are complete, and then the processed data is passed to the first hybrid model for subsequent preprocessing and initial screening. The entire desensitization process is embedded in the link of the data access edge device, without adding significant delay. Through adaptive desensitization processing, the leakage path of patient identity information can be blocked before the data enters the analysis process, which can significantly improve the security and compliance of medical data.

[0055] Accordingly, the embodiment steps may also include: performing adaptive desensitization processing on multimodal medical imaging data, shielding the DICOM tag field containing patient identity information, and retaining only the image feature parameters required for lesion analysis. Among them, the DICOM tag field is a field used to store image metadata in the DICOM (Digital Imaging and Communications in Medicine) standard, which contains patient identity information (such as name, ID), examination information (such as equipment model, layer thickness), etc., and is an important part of medical imaging data; image feature parameters refer to key image data information used for lesion analysis, such as pixel grayscale value, three-dimensional coordinates of the lesion area, image layer thickness, modality type (such as CT, MRI), etc., and are the core input for the first hybrid model to perform preprocessing and initial screening.

[0056] Step 240: Schedule the second hybrid model deployed in the cloud to perform a deep analysis of the suspected lesion area based on the voxel data to obtain a deep analysis result including the lesion size, lesion nature and diagnostic judgment, and encrypt the deep analysis result and transmit it back to the edge device end, so that the edge device end can integrate the initial screening result and the deep analysis result to obtain the final image analysis result.

[0057] For the embodiment of the present disclosure, after receiving the voxel data of the suspected lesion area encrypted and transmitted by the edge device, the second hybrid model (the complete version of the hybrid model) deployed in the cloud can start the complete Transformer module and the 3D convolutional neural network module for in-depth analysis. Specifically: the Transformer module can calculate the global correlation features of the voxel data in the multimodal image (such as the three-dimensional spatial relationship between the lesion and the surrounding tissue), and the 3D convolutional neural network module can extract the local fine spatial features of the lesion (such as the internal texture of the tumor and the clarity of the boundary); combined with the dynamic weight allocation mechanism (such as increasing the CT modality weight in lung cancer analysis and focusing on the MRI modality in brain tumor analysis), the global and local features are comprehensively used to determine the size, nature (benign / malignant tendency) and diagnostic judgment (such as "the possibility of early lung cancer") of the lesion to generate in-depth analysis results; then, the results are transmitted back to the edge device through an encrypted transmission protocol (such as an encrypted channel in the hospital intranet), and the edge device integrates it with the local initial screening results to form a final image analysis result containing complete lesion information for the doctor's reference.

[0058] This in-depth analysis and result feedback mechanism can significantly improve the accuracy and process integrity of multimodal imaging diagnosis through the "cloud computing power advantage + edge integration closed loop", which is specifically reflected in: the cloud-based full version of the model relies on the GPU cluster to shorten the time required for complex tasks such as three-dimensional reconstruction and lesion nature determination, and can break through the computing power limitations of edge devices; the dynamic weight allocation mechanism allows the advantages of different modalities to be accurately exerted, improves the sensitivity of complex lesion identification, and can solve the problem of high misjudgment rate of complex lesions in traditional models; encrypted feedback and edge integration can ensure data security and result integrity, forming a full-process closed loop of "initial screening → in-depth analysis → integration", meeting the collaborative needs of rapid initial judgment and accurate diagnosis in clinical practice, and at the same time providing efficient and secure technical support for remote assisted diagnosis.

[0059] As a preferred method, the implementation process of step 240 can introduce homomorphic encryption technology in cloud-based deep analysis: after the voxel data (including information on suspected lesion areas) encrypted and transmitted by the edge device reaches the cloud, the second hybrid model (full version) does not need to be decrypted and the calculation is started directly in the ciphertext state, that is, the global correlation feature calculation of the ciphertext voxel data is performed through the complete Transformer module (capturing the cross-modal correlation between the lesion and the surrounding tissue), and at the same time, the 3D convolutional neural network module is called to extract local features in the ciphertext state (such as lesion boundaries and internal textures); the dynamic weight distribution mechanism adjusts the modal contribution in the ciphertext environment according to the lesion type, and finally generates a deep analysis result including the lesion size, nature and diagnostic judgment based on the ciphertext calculation result; after the result is generated, the system encrypts it and transmits it back to the edge device through the hospital intranet or 5G. During the entire process, the original data always exists in ciphertext form, and only the plaintext result is output. Feature calculation in the encrypted state can avoid the risk of privacy leakage in the data decryption process, and solve the trust barriers in cross-institutional data sharing; the second hybrid model can still fully call the functions of Transformer and 3D convolutional neural network in the encrypted environment, combined with the dynamic weight distribution mechanism. Experiments show that the sensitivity of identifying early lung cancer remains at 98.2%, and the analysis accuracy is not reduced due to encryption processing.

[0060] Correspondingly, step 240 of the embodiment may also include: using homomorphic encryption technology to schedule the second hybrid model deployed in the cloud to perform global correlation feature calculation and local feature extraction on the suspected lesion area based on voxel data in an encrypted state, to obtain in-depth analysis results including lesion size, lesion nature and diagnostic judgment.

[0061] As another preferred method, after receiving the initial screening results transmitted by the edge device, the cloud can also judge its complexity through the built-in evaluation module (such as based on indicators such as the number of suspected lesions, volume, modality combination diversity, edge model confidence, etc.), and compare it with the preset threshold (such as the volume of a single lesion exceeds 500 voxels, and the number of multimodal image superposition layers is greater than 20 layers); when the complexity exceeds the standard, the system automatically triggers the parallel computing mode of the cloud GPU cluster, and decomposes the second computing power task into multiple subtasks based on the Kubernetes containerized scheduling mechanism (such as independent computing units such as lesion three-dimensional reconstruction, cross-modal feature matching, and property determination), and executes subtasks in parallel through multi-process calls to copies of the second hybrid model deployed on different GPU nodes (such as simultaneously calculating the local features of the CT modality and the global correlation of the MRI modality); after all subtasks are completed, the cloud integrates the results to generate a unified deep analysis result, which is encrypted and transmitted back to the edge device.

[0062] Correspondingly, step 240 of the embodiment may also include: judging the complexity of the initial screening results, and when the complexity is judged to be greater than a preset threshold, controlling the cloud to automatically enable the parallel computing mode of the GPU cluster, decomposing the second computing task into multiple subtasks, and calling the second hybrid model through multiple processes to execute multiple subtasks in parallel.

[0063] This preferred approach significantly improves the efficiency of in-depth analysis of complex cases through a dynamic scheduling mechanism of "complexity triggering + parallel computing." Specifically, for highly complex tasks (such as multiple lesions and large-volume images), GPU cluster parallel computing can shorten processing time and resolve the efficiency bottleneck of traditional cloud-based serial processing. Task decomposition and multi-process calling mechanisms can fully utilize the cloud's elastic computing resources, avoid idle resources, and further improve resource utilization.

[0064] Step 250: Output collaborative image analysis results of the first hybrid model and the second hybrid model on the multimodal medical image data.

[0065] In summary, the technical solution in this application achieves collaborative processing between the edge and the cloud through task splitting. The edge quickly completes preprocessing and initial screening and encrypts the transmission results. The cloud performs in-depth analysis of voxel data and then encrypts the data before transmitting it back, which can improve the efficiency and accuracy of multimodal medical imaging data analysis. At the same time, adaptive desensitization before preprocessing shields patient identity information, and cloud-based analysis uses homomorphic encryption technology to perform feature calculation and extraction in a ciphertext state, which can enhance privacy protection while ensuring data availability, thereby meeting medical data security and compliance requirements and effectively balancing the efficiency, accuracy, and security of multimodal medical imaging analysis.

[0066] Further, as Figure 1 and Figure 2The present embodiment provides a multimodal medical image data analysis device, such as Figure 3 As shown, the device includes: an acquisition module 31, a splitting module 32, an execution module 33, and an output module 34.

[0067] An acquisition module 31 may be used to acquire image analysis tasks corresponding to multimodal medical image data;

[0068] A splitting module 32 is configured to split the image analysis task into a first computing task and a second computing task. The first computing task is a pre-processing and real-time screening task, and the second computing task is a deep analysis task. The output of the first computing task serves as the input of the second computing task.

[0069] An execution module 33 may be used to schedule a first hybrid model deployed on an edge device to execute a first computing task and a second hybrid model deployed on the cloud to execute a second computing task. The first hybrid model is a simplified version of the hybrid model, and the second hybrid model is a complete version of the hybrid model. The hybrid model includes a Transformer module for extracting global correlation features and a 3D convolutional neural network module for extracting local spatial features. A dynamic weight allocation mechanism is introduced to adjust the weight coefficients of different modal images.

[0070] The output module 34 may be configured to output collaborative image analysis results of the first hybrid model and the second hybrid model on the multimodal medical image data.

[0071] In some embodiments of the present application, Figure 4 As shown, the device further includes: a training module 35;

[0072] The training module 35 can be used to obtain the model parameters obtained after each medical node trains a single-modal sub-model based on local single-modal medical imaging data. The single-modal sub-model includes a Transformer module and a 3D convolutional neural network module. The Transformer module is used to convert the local single-modal medical image into a feature vector sequence and calculate the global correlation feature through a multi-head self-attention mechanism. The 3D convolutional neural network module is used to extract local features based on the global correlation feature to obtain local spatial features; call the dynamic weight allocation mechanism to adjust the weight coefficient of each modality based on the feature contribution of each single-modal sub-model in local training, and fuse the model parameters of each single-modal sub-model according to the adjusted weight coefficient, and realize multi-modal feature association through a cross-modal attention mechanism to obtain a trained hybrid model.

[0073] In some embodiments of the present application, when calling the dynamic weight allocation mechanism and adjusting the weight coefficient of each modality based on the feature contribution of each single-modal sub-model in local training, the training module 35 can be specifically used to set the initial weight coefficient of each modality based on medical prior knowledge; by matching the local spatial features output by the 3D convolutional neural network module in each modality sub-model with the locally annotated pathological results, the local feature discrimination of each single-modal sub-model is determined according to the matching accuracy; the average correlation weight of the global correlation features output by the Transformer module in each single-modal sub-model and the features of other modalities is calculated, and the global correlation strength corresponding to the average correlation weight is determined; the feature contribution of each modality is determined based on the local feature discrimination and the global correlation strength, and the weight coefficient of each modality is adjusted based on the feature contribution.

[0074] In some embodiments of the present application, Figure 4 As shown, the apparatus further includes: a deployment module 36;

[0075] The deployment module 36 can be specifically used to: use model optimization technology to lightweight the hybrid model, and deploy the obtained first hybrid model on the edge device end; encapsulate the complete version of the hybrid model through containerization technology, and deploy the obtained second hybrid model on the cloud.

[0076] In some embodiments of the present application, the execution module 33 can be specifically used to schedule the first hybrid model deployed on the edge device to preprocess and perform real-time preliminary screening on the multimodal medical imaging data, extract voxel data of the suspected lesion area, and encrypt and transmit the preliminary screening results containing the voxel data to the cloud. The preprocessing includes at least noise reduction and standardization. The second hybrid model deployed on the cloud is scheduled to perform in-depth analysis on the suspected lesion area based on the voxel data to obtain a deep analysis result including the lesion size, lesion nature and diagnostic judgment, and the deep analysis result is encrypted and transmitted back to the edge device, so that the edge device can integrate the preliminary screening result and the deep analysis result to obtain the final image analysis result.

[0077] In some embodiments of the present application, Figure 4 As shown, the device may further include: a processing module 37;

[0078] The processing module 37 can be specifically used to perform adaptive desensitization processing on the multimodal medical image data, shielding the DICOM tag field containing the patient's identity information and retaining only the image feature parameters required for lesion analysis.

[0079] In some embodiments of the present application, when the second hybrid model deployed on the cloud is scheduled to perform an in-depth analysis of the suspected lesion area based on voxel data to obtain an in-depth analysis result including the lesion size, lesion nature and diagnostic judgment, the execution module 33 can be specifically used to adopt homomorphic encryption technology to schedule the second hybrid model deployed on the cloud to perform global correlation feature calculation and local feature extraction on the suspected lesion area based on voxel data in an encrypted state to obtain an in-depth analysis result including the lesion size, lesion nature and diagnostic judgment.

[0080] It should be noted that for other corresponding descriptions of the functional units involved in the multimodal medical image data analysis device provided in this embodiment, please refer to Figure 1 and Figure 2 The corresponding description in will not be repeated here.

[0081] Based on the above Figure 1 and Figure 2 The method shown in FIG. 1 is a method for performing the above-mentioned operation. Accordingly, this embodiment further provides a storage medium on which a computer program is stored. When the program is executed by a processor, the above-mentioned Figure 1 and Figure 2 The multimodal medical imaging data analysis method shown.

[0082] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), and includes a number of instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of the present application.

[0083] Based on the above Figure 1 and Figure 2 The method shown, and Figure 3 、 4 In order to achieve the above-mentioned purpose, the embodiment of the present application further provides an electronic device, which can be a personal computer, a tablet computer, a server, or other network equipment, etc. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-mentioned Figure 1 and Figure 2 The multimodal medical imaging data analysis method shown.

[0084] Optionally, the physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a Wi-Fi module, and the like. The user interface may include a display, an input unit such as a keyboard, and the like. The optional user interface may also include a USB interface, a card reader interface, and the like. The network interface may optionally include a standard wired interface, a wireless interface (such as a Wi-Fi interface), and the like.

[0085] Those skilled in the art will understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or a combination of certain components, or different component arrangements.

[0086] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the physical device, supporting the execution of information processing programs and other software and / or programs. The network communication module is used to enable communication between components within the storage medium, as well as with other hardware and software within the physical information processing device.

[0087] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or by hardware.

[0088] The embodiment of the present invention adopts a hybrid model including a Transformer module and a 3D convolutional neural network module, which can not only capture the global correlation features of multimodal images through the Transformer, but also extract the local spatial features of lesions with the help of a 3D convolutional neural network, which can significantly improve the recognition sensitivity of complex lesions; the introduction of a dynamic weight allocation mechanism to dynamically adjust the weight coefficients of images of different modalities can enhance the adaptability of the algorithm during multimodal image fusion analysis and reduce the misjudgment rate; at the same time, through the edge-cloud collaborative architecture, the simplified version of the edge device quickly completes preprocessing and real-time initial screening, and the complete version of the cloud model performs in-depth analysis, which can improve diagnostic accuracy while ensuring real-time performance, overcome the technical shortcomings of traditional systems in complex lesion recognition and multimodal fusion, and realize efficient and accurate collaborative analysis of multimodal medical images.

[0089] Those skilled in the art will understand that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily required to implement the present application. Those skilled in the art will understand that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the implementation scenario description, or can be changed accordingly and located in one or more devices different from the implementation scenario. The modules of the above-mentioned implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0090] The serial numbers of the above application are for descriptive purposes only and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosure only discloses several specific implementation scenarios of the present application, but the present application is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present application.

Claims

1. A multimodal medical image data analysis method, characterized in that: include: Obtain image analysis tasks corresponding to multimodal medical imaging data; Split the image analysis task into a first computing task and a second computing task, wherein the first computing task is a preprocessing and real-time initial screening task, and the second computing task is a deep analysis task, and the output of the first computing task serves as the input of the second computing task; Scheduling a first hybrid model deployed on the edge device to perform the first computing task, and scheduling a second hybrid model deployed on the cloud to perform the second computing task. The first hybrid model is a simplified version of the hybrid model, and the second hybrid model is a complete version of the hybrid model. The hybrid model includes a Transformer module for extracting global correlation features and a 3D convolutional neural network module for extracting local spatial features. A dynamic weight allocation mechanism is introduced to adjust the weight coefficients of different modal images. Output collaborative image analysis results of the first hybrid model and the second hybrid model on the multimodal medical image data.

2. The method according to claim 1, characterized in that The method also includes a training process of the hybrid model: Obtaining model parameters obtained by training a unimodal sub-model based on local unimodal medical imaging data at each medical node, wherein the unimodal sub-model includes a Transformer module and a 3D convolutional neural network module. The Transformer module is used to convert the local unimodal medical image into a feature vector sequence and calculate global correlation features through a multi-head self-attention mechanism. The 3D convolutional neural network module is used to extract local features based on the global correlation features to obtain local spatial features. The dynamic weight allocation mechanism is called to adjust the weight coefficient of each modality based on the feature contribution of each unimodal sub-model in local training. The model parameters of each unimodal sub-model are fused according to the adjusted weight coefficient, and multimodal feature association is realized through the cross-modal attention mechanism to obtain a trained hybrid model.

3. The method according to claim 2, characterized in that The calling of the dynamic weight allocation mechanism to adjust the weight coefficient of each modality based on the feature contribution of each single-modality sub-model in local training includes: Set the initial weight coefficient of each modality based on medical prior knowledge; By matching the local spatial features output by the 3D convolutional neural network module in each modality sub-model with the locally annotated pathological results, the local feature discrimination power of each of the single-modality sub-models is determined according to the matching accuracy; Calculating the average correlation weights of the global correlation features output by the Transformer modules in each of the single-modal sub-models and the features of other modalities, and determining the global correlation strength corresponding to the average correlation weights; The feature contribution of each modality is determined based on the local feature discrimination power and the global correlation strength, and the weight coefficient of each modality is adjusted based on the feature contribution.

4. The method according to claim 1, wherein The method further comprises: Lightweighting the hybrid model using model optimization technology, and deploying the obtained first hybrid model on an edge device; The complete version of the hybrid model is encapsulated by containerization technology, and the obtained second hybrid model is deployed in the cloud.

5. The method according to claim 1, wherein The scheduling of the first hybrid model deployed on the edge device to execute the first computing task and the scheduling of the second hybrid model deployed on the cloud to execute the second computing task includes: Scheduling the first hybrid model deployed on the edge device to perform preprocessing and real-time preliminary screening on the multimodal medical imaging data, extracting voxel data of suspected lesion areas, and encrypting and transmitting preliminary screening results containing the voxel data to the cloud, wherein the preprocessing includes at least noise reduction and normalization; The second hybrid model deployed in the cloud is scheduled to perform an in-depth analysis of the suspected lesion area based on the voxel data to obtain an in-depth analysis result including lesion size, lesion nature and diagnostic judgment, and the in-depth analysis result is encrypted and transmitted back to the edge device end, so that the edge device end integrates the initial screening result and the in-depth analysis result to obtain the final image analysis result.

6. The method according to claim 5, characterized in that Before scheduling the first hybrid model deployed on the edge device to perform preprocessing and real-time preliminary screening on the multimodal medical image data, the method further includes: Adaptive desensitization processing is performed on the multimodal medical image data to shield the DICOM tag field containing the patient's identity information and retain only the image feature parameters required for lesion analysis.

7. The method according to claim 5, characterized in that The scheduling of the second hybrid model deployed in the cloud performs an in-depth analysis on the suspected lesion area based on the voxel data to obtain an in-depth analysis result including lesion size, lesion nature, and diagnosis judgment, including: Using homomorphic encryption technology, the second hybrid model deployed in the cloud is scheduled to perform global correlation feature calculation and local feature extraction on the suspected lesion area based on the voxel data in an encrypted state, and obtain in-depth analysis results including lesion size, lesion nature and diagnostic judgment.

8. A multimodal medical image data analysis device, characterized in that: include: An acquisition module is used to obtain image analysis tasks corresponding to multimodal medical image data; a splitting module, configured to split the image analysis task into a first computing task and a second computing task, wherein the first computing task is a preprocessing and real-time initial screening task, and the second computing task is a deep analysis task, and the output of the first computing task serves as the input of the second computing task; An execution module is configured to schedule a first hybrid model deployed on an edge device to execute the first computing task, and schedule a second hybrid model deployed on the cloud to execute the second computing task. The first hybrid model is a simplified version of the hybrid model, and the second hybrid model is a complete version of the hybrid model. The hybrid model includes a Transformer module for extracting global correlation features and a 3D convolutional neural network module for extracting local spatial features, and introduces a dynamic weight allocation mechanism to adjust the weight coefficients of different modal images. An output module is used to output collaborative image analysis results of the first hybrid model and the second hybrid model on the multimodal medical image data.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Slice analysis method and device and storage medium

    CN122048948A

  • Slice analysis method, device, and storage medium

    CN122048948B