Multi-modal image processing method and system based on deep learning

Through multimodal image processing methods and systems based on deep learning, the challenges of feature alignment and fusion problems, efficiency reduction and causal understanding in traditional methods are solved, and efficient and accurate multimodal image analysis and processing are achieved, which is suitable for resource-constrained environments.

CN120126697APending Publication Date: 2025-06-10JIANGSU PROVINCE HOSPITAL (THE FIRST AFFILIATED HOSPITAL OF NANJING MEDICAL UNIVERSITY)
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510259071.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional multimodal image processing methods face the difficulties of feature alignment and fusion, the decline in efficiency when data scale and complexity increases, the challenges of privacy protection and distributed data sharing, and the lack of understanding and dynamic adaptation of causal relationships between modals.

Method used

Multimodal image processing methods and systems based on deep learning are adopted, including adaptive perception and generation module, intelligent interactive processing module, inter-modal intelligent inference module and adaptive learning and optimization module. Through reverse modal generation, cross-modal characteristic unification, dynamic modal reconstruction and data simulation experiments, combined with causal reasoning, knowledge transfer and adversarial learning, data diversity expansion, characteristic analysis and inter-modal consistency optimization are achieved.

Benefits of technology

It significantly improves the efficiency and accuracy of multimodal image analysis, broadens the scope of application in resource-constrained environments, realizes a highly flexible and interactive image processing process, and reduces the complexity of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126697A_ABST
    Figure CN120126697A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a multi-modal image processing method and system based on deep learning, and the system comprises a self-adaptive perception and generation module which is used for exploring the new characteristics of image data, enhancing the perception capability through a generation and simulation technology, and discovering that details are difficult to capture in an existing processing mode. According to the invention, through reverse modal generation, cross-modal characteristic unification, dynamic modal reconstruction and data simulation experiments, data diversity expansion and accurate characteristic analysis are realized; through causal reasoning, knowledge migration and adversarial learning, inter-modal characteristic consistency and analysis performance are optimized; according to the multi-modal image analysis method and system, the multi-modal image analysis efficiency and accuracy are remarkably improved by combining the task-driven model switching and federated learning technology and supporting distributed model training of dynamic adaptation and privacy protection, and meanwhile, the application range of medical image analysis in a resource-constrained environment is widened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a multi-modal image processing method and system based on deep learning. Background Art

[0002] Multi-modal image processing technology is widely used in medical image analysis, industrial inspection, and intelligent monitoring fields, aiming to provide more comprehensive solutions for complex problems by integrating data from different modalities (such as CT, MRI, PET images). In the medical field, multi-modal image analysis is particularly important, as it can combine the advantages of different image modalities for more accurate diagnosis and treatment planning. For example, CT images provide tissue density information, MRI images have high soft tissue resolution, and PET images can provide tissue metabolic activity information. By combining these three modalities, validating and complementing each other, the accuracy of disease detection can be significantly improved, and the ability to qualitatively and locate lesions can be enhanced. However, traditional multi-modal image processing methods face many challenges.

[0003] Existing technologies usually use feature engineering or shallow learning methods for multi-modal data integration, which have the following limitations: 1. The characteristics between different modalities vary greatly, making it difficult to achieve effective feature alignment and fusion; 2. When the data scale and complexity increase, the model training and inference efficiency significantly decreases; 3. The requirements for privacy protection and distributed data sharing are difficult to meet in resource-constrained environments. In addition, traditional methods lack in-depth understanding of the causal relationship between modalities and dynamic adaptation capabilities, and are difficult to handle multi-task requirements and changes in complex scenarios.

[0004] To solve the above problems, we therefore provide a multi-modal image processing method and system based on deep learning. Summary of the Invention

[0005] Aiming at the above-mentioned shortcomings of the prior art, the first object of the present invention is to provide a multi-modal image processing method and system based on deep learning to solve the problems in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A multi-modal image processing method and system based on deep learning, including an adaptive perception and generation module for exploring new features of image data; an intelligent interactive processing module for combining user and system collaborative operations to dynamically optimize the image processing process through intelligent interaction; an inter-modal intelligent reasoning module for discovering potential patterns and knowledge through cross-modal association and causal reasoning; and an adaptive learning and optimization module for completing dynamic optimization through system-level learning. The adaptive perception and generation module includes: an inverse modality generation module for generating other modality extended data through diffusion GAN and according to a single modality image to diversify data and verify modality consistency; an image feature encoding and decoding module for creating a general image feature description language and using a specific model to convert image features into codes to achieve cross-modal semantic unity; a dynamic modality reconstruction module for dynamically selecting specific modality features for reconstruction according to real-time task requirements; and a data simulation experiment module for constructing a virtual environment and combining physical simulation and a generation model to generate synthetic data.

[0008] The intelligent interactive processing module includes a human-machine collaborative feature extraction module for fine-tuning the preliminary extraction of image features through voice, gesture or text, a programmable processing chain module for providing a modular processing chain configuration tool to allow users to freely design the image processing process, a personalized visual analysis module for dynamically adjusting the model's attention area and processing method based on specific goals input by the user, and a semantic-level decision-making dialogue module for using a large language model to have a semantic dialogue with the user about the image processing results to explain the results and provide selection suggestions.

[0009] The present invention is further configured as: the inter-modal intelligent reasoning module includes a causal graph generation and reasoning module for constructing an inter-modal causal relationship graph to reveal the hidden causal mechanism of data and provide assistance for image analysis and decision-making, an inter-modal dynamic collaboration module for dynamically selecting the most effective modal features for joint reasoning according to task requirements, a cross-modal knowledge transfer module for transferring the knowledge of one modality to another modality through Transformer cross-modal, and an inter-modal adversarial learning module for using GAN to optimize inter-modal consistency and eliminate noise in multi-modal data.

[0010] The adaptive learning and optimization module includes a task-driven model switching module for automatically switching or combining models according to specific tasks, an energy consumption adaptive optimization module for automatically adjusting the calculation amount according to the hardware resource situation and selecting a lightweight or fine-grained model to adapt to different processing platforms, a continuous multi-modal learning module for continuously optimizing the model according to new modal data to form a long-term accumulated multi-modal knowledge base, and a federated image learning module for supporting cross-institutional data privacy protection learning and jointly training models through federated learning with multiple data sources.

[0011] The present invention is further configured such that: the programmable processing chain module includes a data preprocessing and enhancement module for performing preprocessing and enhancement of image data, a feature extraction and representation module for extracting key features of the image and converting them into a representation form suitable for subsequent processing, an inter-modal conversion and fusion module for performing cross-modal data conversion and integrating information from different modalities to enhance the analysis ability, an advanced model training and optimization module for optimizing the entire processing flow for model training and tuning, a model verification and evaluation module for evaluating the performance of the model to ensure the output quality and overall effect of each step, and a final output report generation module for generating the final image processing result and automatically generating a report.

[0012] The present invention is further configured such that: on the basis that the data preprocessing and enhancement module of the feature extraction and representation module is responsible for cleaning, standardizing, and enhancing the original image data, the feature extraction and representation module receives the image data after denoising, alignment, and artifact removal, extracts spatial and semantic features through the Transformer deep learning algorithm, and converts them into a high-dimensional feature representation suitable for subsequent analysis; on the basis that the inter-modal conversion and fusion module extracts the spatial texture, structural features, and semantic information of the image, the inter-modal conversion and fusion module integrates multi-modal features (spectral images) through a cross-modal attention mechanism to achieve cross-modal correlation analysis and feature unified representation; on the basis that the inter-modal conversion and fusion module integrates different modal data and generates a unified representation, the advanced model training and optimization module uses the generated multi-modal feature representation for training of the task model, and optimizes the model structure and parameters through neural architecture search and adaptive learning methods.

[0013] The present invention is further configured such that: on the basis that the advanced model training and optimization module completes model training and generates preliminary model parameters, the model verification and evaluation module receives the trained model, and comprehensively evaluates it through multi-dimensional performance indicators such as precision, recall, and F1 score and provides optimization suggestions; on the basis that the model verification and evaluation module comprehensively evaluates the model performance and generates index results, the final output report generation module receives the evaluation report and the processed image result, and automatically generates an image processing report including analysis conclusions, visualization charts, and multi-modal feature descriptions, providing intuitive decision-making support for users.

[0014] The present invention is further configured such that: the image feature encoding and decoding module performs unified encoding and semantic mapping on the generated multi-modal data features on the basis that the reverse mode generation module generates other modal data through Diffusion GAN to ensure semantic consistency and interoperability between modalities; the dynamic modal reconstruction module dynamically selects and reconstructs specific modal features according to real-time task requirements on the basis that the image feature encoding and decoding module completes feature encoding and cross-modal unified representation; the data simulation experiment module constructs a virtual environment through physical simulation and a generation model on the basis of the specific modal data features generated by the dynamic modal reconstruction module;

[0015] The programmable processing chain module provides users with a flexible processing chain design on the basis that the human-computer collaborative feature extraction module fine-tunes the preliminary image features through voice, gesture or text, and users can create an exclusive image processing process according to the optimized features; the personalized visual analysis module embeds a target-oriented dynamic adjustment mechanism in the process configured by the programmable processing chain module and adjusts the attention area and processing method of the model by using the specific target input by the user; the semantic-level decision-making dialogue module uses a large language model to conduct a semantic dialogue with the user on the basis of the dynamic adjustment of the personalized visual analysis module, interprets the analysis results and provides diverse selection suggestions.

[0016] The present invention is further configured such that: the inter-modal dynamic collaboration module selects the most contributive modal features according to task requirements on the basis of the inter-modal causal relationship graph constructed by the causal graph generation and reasoning module to achieve causal relationship-driven joint reasoning; the cross-modal knowledge transfer module uses Transformer cross-modal to transfer the knowledge of one modality to another modality on the basis of the specific modal features selected by the inter-modal dynamic collaboration module; the inter-modal adversarial learning module optimizes the consistency of features between modalities through a generative adversarial network and eliminates noise in multi-modal data after the cross-modal knowledge transfer module realizes knowledge transfer;

[0017] The energy consumption adaptive optimization module dynamically adjusts the amount of computation according to the hardware resource situation on the basis that the task-driven model switching module automatically switches or merges models, and selects a lightweight or fine-grained model to adapt to different processing platforms to ensure the balance between task execution efficiency and resource consumption; the continuous multi-modal learning module continuously optimizes the model by using newly added modal data on the basis of the model adjusted by the energy consumption adaptive optimization module and forms a multi-modal knowledge base; the federated image learning module supports privacy-preserving learning across institutions on the basis of the model optimized by the continuous multi-modal learning module, and conducts distributed training by combining multiple data sources through federated learning technology.

[0018] The present invention is further configured such that: the method of the processing system includes the following steps:

[0019] S1. Multi-modal feature extension and data generation optimization technology;

[0020] S2. Semantic interaction-driven multi-modal feature reconstruction and optimization;

[0021] S3. Cross-modal reasoning and knowledge transfer based on causal models;

[0022] S4. Intelligent resource adaptation and dynamic model optimization and allocation;

[0023] S5. Multi-modal feature fusion and deep learning adaptive training.

[0024] The present invention is further configured as follows: in the step S1, through high-dimensional embedding and diffusion-based generative model (StableDiffusion), multi-modal expansion and feature optimization of image data are realized. The generative adversarial network is used to create virtual modal data to make up for the shortage of existing data, and the authenticity and distribution rationality of the generated data are ensured through modal consistency checking. In the experiment, based on the BraTS2021 brain tumor dataset, ISIC skin lesion dataset, and HyperSpectral Imaging data, a multi-modal medical image generation framework is constructed, and distributed training is carried out in a multi-GPU environment (NVIDIA A100);

[0025] In the step S2, through semantic parsing technology and user active interaction, image feature reconstruction and dynamic guidance are realized. The user can interact with the system through voice, text, or gesture to deeply optimize the extraction and analysis of image features. The HuggingFace Transformers semantic model converts the user input into a feature adjustment target and real-time feedbacks the processing result. The experimental design adopts the ChestXray14 medical image database and RSNA image dataset, combined with a touch interactive screen and a microphone array device, to test the multi-modal interaction effect in a simulated real medical scenario;

[0026] In the step S3, by constructing a causal graph and a cross-modal reasoning network (graph neural network GNN and DoWhy framework), the deep causal relationship between different image modalities is revealed. Combining the cross-modal attention mechanism and knowledge transfer technology, feature transfer and consistency optimization from one modality to another are realized. The experiment generates a causal relationship graph based on brain CT and MRI scan data (BraTS dataset) and uses the PyTorch Geometric and DoWhy libraries for causal inference analysis, and uses a high-performance computing cluster for parallelized modeling.

[0027] The present invention is further configured such that in step S4, through resource adaptation and task allocation techniques, dynamic adjustment and optimization of the model among different computing environments are achieved. A lightweight model (MobileNet) is used for real-time inference on edge devices, while a high-performance model with a deep Transformer architecture runs on the cloud server. The experiment is based on the AWS cloud platform and NVIDIA Jetson Nano edge devices to test the dynamic adaptation ability from cloud-edge collaboration to edge-independent operation, and the TensorFlow Lite and NVIDIA TensorRT frameworks are used for optimization;

[0028] In step S5, through cross-modal feature fusion techniques and deep learning optimization frameworks, spatial, texture, and semantic features of different modalities are integrated and uniformly characterized. An attention mechanism and a Transformer architecture are used to train a multi-modal learning model, and a hyperparameter search algorithm (Bayesian optimization) is adopted to improve the model performance. In the experiment, based on the Multi-modal Brain Tumor Segmentation Challenge (BraTS) dataset and the MS-COCO multi-modal dataset, efficient model training and optimization are completed on the NVIDIA DGX Server using the Horovod distributed training environment.

[0029] Beneficial effects

[0030] Adopting the technical solution provided by the present invention, compared with the known public technology, it has the following beneficial effects:

[0031] 1. Through inverse modality generation, cross-modal feature unification, dynamic modality reconstruction, and data simulation experiments, the present invention realizes data diversity expansion and accurate feature analysis; through causal reasoning, knowledge transfer, and adversarial learning, it optimizes the consistency of features among modalities and analysis performance; and by combining task-driven model switching and federated learning techniques, it supports distributed model training for dynamic adaptation and privacy protection. These functions significantly improve the efficiency and accuracy of multi-modal image analysis, and at the same time broaden the application scope of medical image analysis in resource-constrained environments.

[0032] 2. The present invention relies on the Transformer deep learning algorithm, cross-modal attention mechanism, and generative adversarial network (GAN) to dynamically extract and fuse multi-modal features. At the same time, the model performance is evaluated through precision, recall, and F1-score metrics, and an intuitive analysis report is generated for users. This highly flexible and interactive design makes the image processing process more personalized and efficient, and significantly reduces the complexity of data processing.

[0033] 3. The present invention optimizes data quality by using high-dimensional embedding and diffusion-based generative models, adjusts the feature extraction targets in real time through semantic interaction technology, conducts causal inference and feature transfer by combining graph neural networks (GNNs) and the DoWhy framework, and optimizes model resource allocation through cloud-edge collaboration. Finally, it fuses multimodal features through the Transformer architecture and optimizes model performance by combining distributed training and hyperparameter search. These methods demonstrate excellent flexibility and accuracy in medical image analysis, can adapt to multi-scenario and multi-task requirements, and significantly improve the quality and efficiency of image analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the system modules of a multimodal image processing system based on deep learning according to the present invention;

[0035] Figure 2 It is a schematic diagram of the programmable processing chain module of a multimodal image processing system based on deep learning according to the present invention;

[0036] Figure 3 It is a schematic flowchart of a multimodal image processing method based on deep learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0038] It should be further noted that the drawings and embodiments of the present invention mainly describe and explain the concept of the present invention. On the basis of this concept, the specific forms and settings of some connection relationships, positional relationships, power mechanisms, power supply systems, hydraulic systems and control systems may not be fully described. However, on the premise that those skilled in the art understand the concept of the present invention, those skilled in the art can implement the above specific forms and settings in a well-known manner.

[0039] When an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.

[0040] The orientation terms "inside" and "outside" refer to the inside and outside relative to the outline of each component itself. The terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", and "outside" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.

[0041] For the convenience of description, spatial relative terms such as "above", "over", "on the upper surface", "upper" can be used here to describe the spatial positional relationship between a device or feature shown in the drawings and other devices or features. It should be understood that the spatial relative terms are intended to include different orientations in use or operation in addition to the orientation described in the drawings for the device. For example, if the device in the drawing is inverted, the device described as "above or over other devices or structures" will then be positioned "below or under other devices or structures". Thus, the exemplary term "above" can include both the orientations of "above" and "below". The device can also be positioned in other different ways, and corresponding interpretations are made for the spatial relative descriptions used here.

[0042] The terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a plurality" is two or more, and the meaning of "several" is one or more, unless otherwise specifically defined.

[0043] Now, a multi-modal image processing method and system based on deep learning provided by the present invention will be described.

[0044] Embodiment 1

[0045] As Figure 1 shown, the present invention provides a technical solution: a multi-modal image processing system based on deep learning, including an adaptive perception and generation module for exploring new features of image data, enhancing the perception ability through generation and simulation techniques, and discovering details that are difficult to capture by existing processing methods, an intelligent interactive processing module for combining user and system collaborative operations to dynamically optimize the image processing process through intelligent interaction, an inter-modal intelligent reasoning module for discovering potential patterns and knowledge through cross-modal association and causal reasoning, and an adaptive learning and optimization module for completing dynamic optimization through system-level learning.

[0046] The adaptive perception and generation module includes an inverse modality generation module for generating other modality extended data diversity through a diffusion GAN and based on single-modal images, and verifying modality consistency, an image feature encoding and decoding module for creating a general image feature description language and converting image features into codes using a specific model to achieve cross-modal semantic unity, a dynamic modality reconstruction module for dynamically selecting specific modality features for reconstruction according to real-time task requirements, and a data simulation experiment module for constructing a virtual environment and generating synthetic data by combining physical simulation and a generation model;

[0047] The intelligent interactive processing module includes a human-machine collaborative feature extraction module for preliminarily extracting image features through voice, gesture or text fine-tuning, a programmable processing chain module for providing a modular processing chain configuration tool to allow users to freely design the image processing flow, a personalized visual analysis module for dynamically adjusting the model's attention area and processing method based on specific targets input by users, and a semantic-level decision-making dialogue module for using a large language model to have a semantic dialogue with users about the image processing results, interpret the results, and provide selection suggestions;

[0048] The inter-modal intelligent reasoning module includes a causal graph generation and reasoning module for constructing an inter-modal causal relationship graph to reveal the hidden causal mechanism of data and provide assistance for image analysis and decision-making, an inter-modal dynamic collaboration module for dynamically selecting the most effective modal features for joint reasoning according to task requirements, a cross-modal knowledge transfer module for transferring the knowledge of one modality to another through Transformer cross-modal, and an inter-modal adversarial learning module for using GAN to optimize inter-modal consistency and eliminate noise in multi-modal data;

[0049] The adaptive learning and optimization module includes a task-driven model switching module for automatically switching or merging models according to specific tasks, an energy consumption adaptive optimization module for automatically adjusting the amount of calculation according to the hardware resource situation and selecting a lightweight or fine-grained model to adapt to different processing platforms, a continuous multi-modal learning module for continuously optimizing the model according to new modal data to form a long-term accumulated multi-modal knowledge base, and a federated image learning module for supporting cross-institutional data privacy protection learning and jointly training models by federated learning with multiple data sources;

[0050] The image feature encoding and decoding module performs unified encoding and semantic mapping on the characteristics of the generated multi-modal data based on the reverse mode generation module generating other modal data through Diffusion GAN to ensure semantic consistency and interoperability between modalities; the dynamic modal reconstruction module dynamically selects and reconstructs specific modal characteristics according to real-time task requirements based on the image feature encoding and decoding module completing feature encoding and cross-modal unified representation; the data simulation experiment module constructs a virtual environment through physical simulation and generation models based on the specific modal data characteristics generated by the dynamic modal reconstruction module;

[0051] The programmable processing chain module provides users with a flexible processing chain design based on the human-computer collaborative feature extraction module fine-tuning the preliminary image features through voice, gestures or text, and users can create a dedicated image processing process according to the optimized features; the personalized visual analysis module embeds a target-oriented dynamic adjustment mechanism in the process configured by the programmable processing chain module, and uses the specific target input by the user to adjust the attention area and processing method of the model; the semantic-level decision-making dialogue module uses a large language model to conduct semantic dialogue with the user based on the dynamic adjustment of the personalized visual analysis module, explains the analysis results and provides diverse selection suggestions;

[0052] The inter-modal dynamic collaboration module selects the most contributing modal features according to task requirements based on the inter-modal causal relationship graph constructed by the causal graph generation and reasoning module to achieve causal relationship-driven joint reasoning; the cross-modal knowledge transfer module uses Transformer cross-modal to transfer the knowledge of one modality to another modality based on the specific modal features selected by the inter-modal dynamic collaboration module; the inter-modal adversarial learning module optimizes the consistency of features between modalities through a generative adversarial network after the cross-modal knowledge transfer module realizes knowledge transfer, and eliminates the noise in the multi-modal data;

[0053] The energy consumption adaptive optimization module dynamically adjusts the computational amount according to the hardware resource situation based on the task-driven model switching module automatically switching or merging models, and selects lightweight or fine-grained models to adapt to different processing platforms to ensure the balance between task execution efficiency and resource consumption; the continuous multi-modal learning module continuously optimizes the model using the newly added modal data based on the model adjusted by the energy consumption adaptive optimization module and forms a multi-modal knowledge base; the federated image learning module supports privacy-preserving learning across institutions based on the model optimized by the continuous multi-modal learning module, and conducts distributed training by combining multiple data sources through federated learning technology.

[0054] In the above embodiments, the adaptive perception and generation module uses reverse modality generation technology to expand data diversity, realizes cross-modal semantic unity through image feature encoding and decoding, and meets the real-time task requirements through dynamic modality reconstruction. Synthetic data is generated by combining data simulation experiments. In terms of intelligent interaction, users can fine-tune image features through voice, gestures or text, and design personalized processing flows with the help of the programmable processing chain module; the personalized visual analysis module dynamically adjusts the attention area and processing method of the model, and finally the semantic-level decision-making dialogue module uses the large language model to interpret the results and give optimization suggestions. The inter-modal intelligent reasoning module reveals the potential mechanism of the data through the causal graph, realizes joint reasoning of multi-modal features through inter-modal dynamic collaboration and cross-modal knowledge transfer, and optimizes feature consistency through inter-modal adversarial learning. The system's adaptive learning and optimization module flexibly adjusts the model by task-driven model switching, dynamically adapts to different platforms in combination with energy consumption optimization, and constructs a long-term multi-modal knowledge base and privacy-protected distributed training ability through continuous multi-modal learning and federated image learning. The overall system integrates perception, generation, reasoning and interaction in dynamic optimization, comprehensively improving the analysis ability and processing efficiency of multi-modal images.

[0055] Embodiment 2

[0056] As Figure 1-2 shown, the present invention provides a technical solution: a multi-modal image processing system based on deep learning. The programmable processing chain module includes a data preprocessing and enhancement module for preprocessing and enhancing image data, a feature extraction and representation module for extracting key features of the image and converting them into a representation form suitable for subsequent processing, an inter-modal conversion and fusion module for converting and fusing cross-modal data to integrate information from different modalities to enhance the analysis ability, a high-level model training and optimization module for optimizing the entire processing flow for model training and tuning, a model verification and evaluation module for evaluating the performance of the model to ensure the output quality of each step and the overall effect, and a final output report generation module for generating the final image processing result and automatically generating a report.

[0057] Based on the data preprocessing and enhancement module responsible for cleaning, standardizing, and enhancing the original image data, the feature extraction and representation module receives the image data after denoising, alignment, and artifact removal, extracts spatial and semantic features through the Transformer deep learning algorithm, and converts them into high-dimensional feature representations suitable for subsequent analysis; based on the feature extraction and representation module extracting the spatial texture, structural features, and semantic information of the image, the inter-modal conversion and fusion module integrates multi-modal features (spectral images) through a cross-modal attention mechanism to achieve cross-modal correlation analysis and feature unified representation; based on the inter-modal conversion and fusion module integrating different modal data and generating a unified representation, the advanced model training and optimization module uses the generated multi-modal feature representation to train the task model, and optimizes the model structure and parameters through neural architecture search and adaptive learning methods;

[0058] Based on the advanced model training and optimization module completing the model training and generating preliminary model parameters, the model verification and evaluation module receives the trained model, conducts a comprehensive evaluation through multi-dimensional performance indicators such as precision, recall, and F1-score, and provides optimization suggestions; based on the model verification and evaluation module comprehensively evaluating the model performance and generating index results, the final output report generation module receives the evaluation report and the processed image results, and automatically generates an image processing report containing analysis conclusions, visualization charts, and multi-modal feature descriptions, providing intuitive decision-making support for users.

[0059] In the above embodiment, the data preprocessing and enhancement module is responsible for cleaning, standardizing, and enhancing the original image data, denoising and removing artifacts, laying a foundation for subsequent processing. The feature extraction and representation module uses the Transformer algorithm to extract the spatial, texture, structural, and semantic features of the image, and converts them into high-dimensional feature representations suitable for analysis. The inter-modal conversion and fusion module integrates multi-modal information through a cross-modal attention mechanism to achieve unified representation and correlation analysis of cross-modal features. Based on these multi-modal features, the advanced model training and optimization module trains the task model, and optimizes the structure and parameters of the model through neural architecture search and adaptive learning methods. The model verification and evaluation module evaluates the performance of the trained model through precision, recall, and F1-score indicators and provides optimization suggestions. The final output report generation module automatically generates a report containing analysis conclusions, visualization charts, and multi-modal feature descriptions based on the evaluation results, providing intuitive decision-making support for users. The entire process realizes the automation, intelligence, and high efficiency of image processing through seamless cooperation between modules.

[0060] Among them: The advanced model training and optimization module calls the useful callable data in the data preprocessing module through the inter-modal conversion and fusion module.

[0061] Wherein: the data preprocessing and enhancement module extracts features from the data called by the advanced model training and optimization module through the feature extraction and representation module, and transfers the data that can be called to the advanced model training and optimization module after conversion and fusion through the inter-modal conversion and fusion module.

[0062] Embodiment 3

[0063] As Figure 3 shown, the present invention provides a technical solution: a multi-modal image processing method based on deep learning. The method of this processing system includes the following steps:

[0064] S1. Multi-modal feature extension and data generation optimization technology;

[0065] S2. Semantic interaction-driven multi-modal feature reconstruction optimization;

[0066] S3. Cross-modal reasoning and knowledge transfer based on causal models;

[0067] S4. Intelligent resource adaptation and dynamic model optimization allocation;

[0068] S5. Multi-modal feature fusion and deep learning adaptive training;

[0069] In the step S1, through high-dimensional embedding and diffusion-based generative model (Stable Diffusion), the multi-modal extension and feature optimization of image data are realized. The generative adversarial network is used to create virtual modal data to make up for the lack of existing data, and the authenticity and reasonable distribution of the generated data are ensured through modal consistency checking. In the experiment, based on the BraTS2021 brain tumor dataset, ISIC skin lesion dataset and HyperSpectral Imaging data, a multi-modal medical image generation framework is constructed and distributed training is carried out in a multi-GPU environment (NVIDIA A100);

[0070] In the step S2, through semantic parsing technology and user active interaction, the image feature reconstruction and dynamic guidance are realized. The user can interact with the system through voice, text or gesture, and deeply optimize the extraction and analysis of image features. The HuggingFace Transformers semantic model converts the user input into a feature adjustment target and real-time feedbacks the processing result. The experimental design adopts the ChestXray14 medical image database and RSNA image dataset, combined with a touch interaction screen and a microphone array device, to test the multi-modal interaction effect in a simulated real medical scenario;

[0071] In step S3, by constructing a causal graph and a cross-modal inference network (graph neural network GNN and DoWhy framework), the deep causal associations between different imaging modalities are revealed. Combining the cross-modal attention mechanism and knowledge transfer technology, the feature transfer and consistency optimization from one modality to another are realized. The experiment generates a causal relationship graph based on brain CT and MRI scan data (BraTS dataset) and uses the PyTorch Geometric and DoWhy libraries for causal inference analysis, and parallelized modeling is carried out using a high-performance computing cluster;

[0072] In step S4, through resource adaptation and task allocation technology, the dynamic adjustment and optimization of the model among different computing environments are realized. A lightweight model (MobileNet) is used for real-time inference on edge devices, while a high-performance model with a deep Transformer architecture runs on the cloud server. The experiment is based on the AWS cloud platform and NVIDIA Jetson Nano edge devices, tests the dynamic adaptation ability from cloud-edge collaboration to edge-independent operation, and is optimized using the TensorFlow Lite and NVIDIA TensorRT frameworks;

[0073] In step S5, through cross-modal feature fusion technology and a deep learning optimization framework, the spatial, texture, and semantic features of different modalities are integrated and uniformly represented. The attention mechanism and Transformer architecture are used to train a multi-modal learning model, and at the same time, a hyperparameter search algorithm (Bayesian optimization) is adopted to improve the model performance. In the experiment, based on the Multi-modal Brain Tumor Segmentation Challenge (BraTS) dataset and the MS-COCO multi-modal dataset, efficient model training and optimization are completed on the NVIDIA DGX Server using the Horovod distributed training environment.

[0074] In the above embodiments, high-dimensional embedding, diffusion-based generative models (Stable Diffusion), and generative adversarial networks (GANs) are used to generate virtual modality data, optimizing the diversity and consistency of the data. On the basis of data preprocessing, semantic parsing technology (HuggingFace Transformers) and active user interaction (voice, text, gesture) are used to achieve the reconstruction and optimization of image characteristics, and the processing results are fed back in real time. In the inference stage, a causal graph is constructed by combining graph neural networks (GNNs) and the DoWhy framework. Through cross-modal attention mechanisms and knowledge transfer technologies, the consistency optimization of characteristics between modalities and knowledge transfer are completed. At the same time, the system dynamically adapts the model according to the computing environment, runs lightweight models (such as MobileNet) on edge devices, runs high-performance Transformer architectures in the cloud, and uses TensorFlow Lite and TensorRT to optimize performance. The spatial, texture, and semantic characteristics are fused through the combination of Transformer and attention mechanisms, and Bayesian optimization is used for hyperparameter search. The multi-modal learning model is trained in a Horovod distributed environment to improve the model performance. The verification and feedback links evaluate the model performance through multi-dimensional scoring of precision, recall, and F1 score, interpret the model decisions with the help of Grad-CAM, and dynamically optimize the process through reinforcement learning. Finally, the processing results are displayed through augmented reality (AR) and dynamic graphics technologies (Matplotlib and Plotly), and an intelligent analysis report is generated in combination with artificial intelligence, providing efficient decision-making support for users. The overall system demonstrates excellent flexibility and accuracy throughout the entire process from data generation to result display, and is applicable to complex multi-modal image processing scenarios.

[0075] Working principle

[0076] As Figure 1-3 shown, in the actual use process of the present invention, firstly, through multi-modal feature extension and data generation optimization technology, high-dimensional embedding, diffusion-based generative models (Stable Diffusion), and generative adversarial networks (GANs) are used to generate virtual modality data on a multi-GPU environment (NVIDIA A100), effectively making up for the deficiencies of existing data. At the same time, the authenticity and reasonable distribution of the generated data are ensured through modal consistency checks, improving the data quality and diversity. For the medical image application scenario, the system constructs a multi-modal generation framework based on BraTS and ISIC medical image datasets, providing high-quality basic data for subsequent analysis.

[0077] The system combines semantic parsing technology (HuggingFace Transformers) with active user interaction (voice, text, or gesture input) to achieve the reconstruction and optimization of image features with the support of a touch interaction screen and a microphone array device. Users can dynamically adjust the goals and strategies of image processing through semantic interaction, and the system provides real-time feedback on the processing results to further optimize the image feature extraction and analysis effects. In practical applications, using the ChestXray14 and RSNA datasets, the system can significantly improve the user experience and the accuracy of image analysis in real medical scenarios.

[0078] In the cross-modal reasoning stage, the system constructs a causal graph through a graph neural network (GNN) and the DoWhy framework to reveal the deep causal relationships between different image modalities. Combining cross-modal attention mechanisms and knowledge transfer technologies, it optimizes the consistency of features between modalities and completes modality feature transfer. In addition, the system dynamically adapts the model according to the computing environment, running a lightweight model (MobileNet) on edge devices (NVIDIA Jetson Nano) to support real-time reasoning, while running a deep Transformer architecture on the cloud (AWS cloud platform) and combining TensorFlow Lite and TensorRT to achieve performance optimization, so as to operate efficiently in different environments.

[0079] Finally, through multi-modal feature fusion and deep learning adaptive training technology, the system uses the Transformer architecture and attention mechanism to uniformly represent spatial, texture, and semantic features, and combines Bayesian optimization for hyperparameter search. Based on the Horovod distributed training environment (NVIDIA DGX Server), the system completes efficient model training and optimization. The model is verified through multi-dimensional metrics such as precision, recall, and F1 score, combined with Grad-CAM to visually explain the model's decisions, and uses reinforcement learning to dynamically optimize the process. The final results are presented through augmented reality (AR) and dynamic graphics technologies (Matplotlib and Plotly), combined with an analysis report generated by artificial intelligence, providing users with intuitive decision support and insights. This process significantly improves the processing efficiency, analysis ability, and user experience of multi-modal images.

[0080] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0081] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of the stated features, steps, operations, devices, components, and / or combinations thereof.

[0082] Unless otherwise specifically stated, the relative arrangements of the components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present application. At the same time, it should be understood that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods, and devices should be regarded as part of the authorized specification. In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

Claims

1. A multimodal image processing system based on deep learning, characterized in that: include: An adaptive perception and generation module, wherein the adaptive perception and generation module is used to explore new characteristics of image data; An intelligent interactive processing module, which is used to dynamically optimize the image processing process through intelligent interaction in combination with user and system collaborative operations; An inter-modal intelligent reasoning module, which is used to discover potential patterns and knowledge through cross-modal association and causal reasoning; An adaptive learning and optimization module, wherein the adaptive learning and optimization module is used to complete dynamic optimization through system-level learning; The adaptive perception and generation module comprises: The reverse modality generation module is used to generate other modalities based on single modality images through diffusion GAN, expand data diversity and verify modality consistency; Image feature encoding and decoding module, which is used to create a universal image feature description language and use a specific model to convert image features into codes to achieve cross-modal semantic unification; A dynamic modal reconstruction module, which is used to dynamically select specific modal characteristics for reconstruction according to real-time task requirements; and a data simulation experiment module, which is used to build a virtual environment combining physical simulation and generative modeling to generate synthetic data; The intelligent interactive processing module includes a human-machine collaborative feature extraction module for preliminarily extracting image features through voice, gesture or text fine-tuning, a programmable processing chain module for providing a modular processing chain configuration tool to allow users to freely design image processing processes, a personalized visual analysis module for dynamically adjusting the model's focus area and processing method based on specific goals input by the user, and a semantic-level decision-making dialogue module for using a large language model to conduct semantic dialogue with the user on the image processing results to explain the results and provide selection suggestions.

2. The multimodal image processing system based on deep learning according to claim 1, characterized in that: The inter-modal intelligent reasoning module includes a causal graph generation reasoning module for constructing an inter-modal causal relationship graph, an inter-modal dynamic collaboration module for dynamically selecting the most effective modal features for joint reasoning according to task requirements, a cross-modal knowledge transfer module for transferring knowledge from one modality to another modality through Transformer cross-modality, and an inter-modal adversarial learning module for eliminating noise in multi-modal data; The adaptive learning optimization module includes a task-driven model switching module for automatically switching or merging models, an energy consumption adaptive optimization module for adapting to different processing platforms according to hardware resource conditions, a continuous multimodal learning module for continuously optimizing models according to new modal data to form a long-term accumulated multimodal knowledge base, and a federated image learning module for supporting cross-institutional model training by combining multiple data sources through federated learning.

3. The multimodal image processing system based on deep learning according to claim 1, characterized in that: The programmable processing chain module includes a data preprocessing enhancement module for performing image data preprocessing and enhancement, a feature extraction and representation module for extracting key features of an image and converting them into a representation form suitable for subsequent processing, an inter-modal conversion and fusion module for conversion and fusion integration, an advanced model training optimization module for optimizing the entire processing flow for model training and tuning, a model validation and evaluation module for evaluating model performance, and a final output report generation module for generating final image processing results and automatically generating a report.

4. The multimodal image processing system based on deep learning according to claim 3, characterized in that: The feature extraction and representation module receives the image data that has been denoised, aligned and artifacts removed, based on the data preprocessing and enhancement module being responsible for cleaning, standardizing and enhancing the original image data, and extracts spatial and semantic features through the Transformer deep learning algorithm, and converts them into high-dimensional feature representations suitable for subsequent analysis.

5. A multimodal image processing system based on deep learning according to claim 3 or 4, characterized in that: The inter-modal conversion fusion module integrates multimodal features through a cross-modal attention mechanism to achieve cross-modal correlation analysis and unified feature representation; the advanced model training optimization module uses the generated multimodal feature representation to train the task model.

6. The multimodal image processing system based on deep learning according to claim 1, characterized in that: The image characteristic encoding and decoding module uniformly encodes and semantically maps the generated multimodal data characteristics based on other modal data generated by diffusion GAN to ensure semantic consistency and interoperability between modalities; the dynamic modal reconstruction module dynamically selects and reconstructs specific modal characteristics according to real-time task requirements; the data simulation experiment module constructs a virtual environment through physical simulation and generative modeling.

7. The multimodal image processing system based on deep learning according to claim 1, characterized in that: The programmable processing chain module provides users with flexible processing chain design tools; the personalized visual analysis module uses the specific goals input by the user to adjust the model's focus area and processing method; the semantic-level decision-making dialogue module uses a large language model to conduct semantic dialogue with the user on the basis of dynamic adjustment of the personalized visual analysis module and explain the analysis results and provide a variety of selection suggestions.

8. A multimodal image processing system based on deep learning according to any one of claims 1 to 7, characterized in that: The method of processing the system comprises the following steps: S1. Multimodal feature expansion and data generation optimization technology, through high-dimensional embedding and diffusion generation model, to build a multimodal medical image generation framework; S2, semantic interaction driven multimodal feature reconstruction and optimization, through semantic parsing technology and user active interaction, combined with touch interactive screen and microphone array equipment to test the multimodal interaction effect in a simulated real medical scenario; S3, cross-modal reasoning and knowledge transfer based on causal models, by constructing causal graphs and cross-modal reasoning networks, revealing the deep causal relationship between different imaging modalities; S4, intelligent resource adaptation and dynamic model optimization allocation, through resource adaptation and task allocation technology, to achieve dynamic adjustment and optimization of models in different computing environments; S5. Multimodal feature fusion and deep learning adaptive training. Through cross-modal feature fusion technology and deep learning optimization framework, the spatial, texture and semantic characteristics of different modalities are integrated and represented in a unified manner.

9. The multimodal image processing method based on deep learning according to claim 8, characterized in that: In step S1, a high-dimensional embedding and diffusion generation model is used to create virtual modality data using a generative adversarial network to make up for the lack of existing data, and the authenticity and distribution rationality of the generated data are ensured through a modality consistency check to build a multimodal medical image generation framework; In step S2, semantic parsing technology and active user interaction are used to achieve image feature reconstruction and dynamic guidance, and real-time feedback of processing results. The multimodal interaction effect is tested in a simulated real medical scenario in combination with a touch interactive screen and a microphone array device.

10. The multimodal image processing method based on deep learning according to claim 8, characterized in that: In step S3, by constructing a causal graph and a cross-modal reasoning network, the deep causal relationship between different image modalities is revealed, combining the cross-modal attention mechanism with the knowledge transfer technology.

Citation Information

Cited By

  • Multi-modal medical image imaging information processing method and system

    CN121545689A

  • Multimodal medical image imaging information processing method and system

    CN121545689B