Multi-modal data fusion diagnosis method and system based on knowledge graph and large model

By constructing a multimodal data fusion diagnostic method that combines knowledge graphs and large models, the problem that single-modal data cannot fully reflect a patient's condition is solved. This enables in-depth collaborative analysis of multimodal data, improves the accuracy and efficiency of diagnosis, and enhances the interpretability of clinical decisions.

CN120954671APending Publication Date: 2025-11-14NANTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510891342.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Current medical diagnostic methods mainly rely on single-modal data, which cannot fully reflect the patient's condition. Furthermore, large models struggle to understand the decision-making process in multimodal data fusion diagnosis, leading to distorted or misleading generated content.

Method used

By combining knowledge graphs and large models, multimodal data fusion diagnosis is achieved. Graph neural network detection is used to detect contradictory data. The interaction between dynamic instance graphs and static knowledge graphs is constructed to capture implicit relationships and realize in-depth collaborative analysis of multimodal data.

Benefits of technology

It improves the accuracy and efficiency of medical diagnosis and enhances the interpretability of clinical decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954671A_ABST
    Figure CN120954671A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal data fusion diagnosis method and system based on a knowledge graph and a large model. The method comprises the following steps: collecting multi-modal data; constructing a medical knowledge graph; pre-training a large model and enhancing knowledge; and performing multi-modal data fusion diagnosis and decision generation. According to the method, a large model can integrate multi-dimensional information through fusion of medical multi-modal data, so that the comprehensiveness and accuracy of diagnosis are improved, and clinical decisions have interpretability in combination with logical reasoning of a knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical diagnostic technology, and in particular to a multimodal data fusion diagnostic method and system based on knowledge graphs and large models. Background Technology

[0002] Traditional medical diagnostic methods primarily rely on single-modal data, such as electronic medical records, medical imaging, or physiological monitoring data. This single-modal data analysis often has limitations and cannot comprehensively reflect a patient's condition. For example, while textual information in electronic medical records contains important information such as the patient's medical history and symptoms, it lacks a direct visual representation of the patient's internal structures and functions; medical imaging data can provide high-resolution anatomical information, but may struggle with early disease diagnosis and detection of subtle lesions; physiological monitoring data focuses on reflecting the patient's physiological state in real time, but it is difficult to provide deeper information such as the etiology and pathological mechanisms of diseases. The fusion of multimodal medical data (such as text, imaging, physiological signals, and genomic data) can provide multi-dimensional information complementarity, and multimodal models outperform single-modal models in disease classification, staging, and prediction.

[0003] In recent years, with the rapid development of artificial intelligence technology, large-scale models have achieved remarkable results in fields such as natural language processing and image recognition, providing new ideas and methods for the fusion diagnosis of multimodal medical data. However, the internal mechanisms of large-scale models are complex, and their decision-making processes are often difficult to understand, which may lead to distorted or misleading content generated, thereby affecting clinical decision-making. Knowledge graph technology can provide a structured representation of various knowledge and data in the medical field, offering rich background knowledge and logical reasoning support for medical diagnosis. However, how to effectively integrate the advantages of knowledge graphs and large-scale models to achieve deep collaborative analysis of multimodal data in the fusion diagnosis of medical data remains an urgent problem to be solved. Summary of the Invention

[0004] This invention aims to enhance the comprehensiveness and accuracy of diagnosis by integrating multidimensional information into large models through the fusion of medical multimodal data, and to make clinical decisions more interpretable by combining logical reasoning with knowledge graphs.

[0005] According to one aspect of the present invention, a multimodal data fusion diagnostic method based on knowledge graphs and large models is provided, comprising the following steps:

[0006] Multimodal data collection;

[0007] Building a medical knowledge graph:

[0008] Large model pre-training and knowledge augmentation;

[0009] Multimodal data fusion diagnosis and decision generation.

[0010] Furthermore, the specific steps for multimodal data collection are as follows:

[0011] Collect multimodal medical data from patients, including structured and unstructured data;

[0012] The collected multimodal medical data is preprocessed. For text data, medical entity recognition technology based on the BioBERT-CRF model is used to automatically correct spelling errors and non-standardized terms in electronic medical records. For image data, adaptive filtering algorithms are used to eliminate noise and motion artifacts in CT / MRI images, and GAN (Generative Adversarial Network) is used to repair low-quality images. For time-series data, dynamic time warping is used to align physiological monitoring data with different sampling rates, and Kalman filtering is used to eliminate baseline drift of electrocardiogram signals.

[0013] Construct association constraint rules between different modalities of data, and detect contradictory data through graph neural networks.

[0014] Furthermore, the specific steps for constructing a medical knowledge graph are as follows:

[0015] Adapt to multi-source heterogeneous data sources, integrate structured data, semi-structured data and unstructured data, and use a federated learning framework to achieve cross-institutional data collaboration;

[0016] Multimodal knowledge extraction and semantic alignment are implemented. A clinical BERT model is used for nested entity recognition on text data, and dependency parsing is used to extract relationships between entities. Object detection algorithms are employed to segment key regions in images and extract morphological parameters. Deep semantic features are extracted using 3D ResNet50. Temporal signals are extracted using a Transformer encoder and aligned with features described in the text. UMLS standard terminology is used to map entities from different modalities to a unified semantic space. Attribute weight calculation and constraint settings are introduced to enhance the clinical interpretability of the knowledge graph.

[0017] Construct graph convolutional interactions between dynamic instance graphs and static knowledge graphs to capture implicit relationships.

[0018] Furthermore, the specific steps for large-scale model pre-training and knowledge augmentation are as follows:

[0019] We chose a Transformer-based model and adopted a knowledge graph-guided pre-training strategy. During the training of the large model, we injected knowledge graph entity embeddings as prior knowledge into the attention layer and optimized the semantic alignment between modalities through contrastive learning.

[0020] A multi-task fine-tuning mechanism is adopted, which combines tasks and utilizes clinical annotation data to achieve professional model adaptation;

[0021] When new evidence emerges, the knowledge graph is dynamically expanded using a graph neural network embedding propagation algorithm.

[0022] Furthermore, the steps involved in multimodal data fusion diagnosis and decision generation are as follows:

[0023] A weighted attention mechanism is used to achieve deep association between image features and text and gene data; preprocessed multimodal data is input into a trained large model, and the feature extraction and fusion capabilities of the large model are used to generate a comprehensive feature representation;

[0024] By combining raw data stitching, graph neural network interaction, and decision-making layer integration, a hybrid fusion strategy is achieved.

[0025] Based on the diagnostic results, a detailed diagnostic report is generated.

[0026] According to another aspect of the present invention, a multimodal data fusion diagnostic system based on knowledge graphs and large models is provided, comprising:

[0027] The preprocessing and standardization module is used to collect patients' multimodal medical data, preprocess the collected multimodal medical data, construct association constraint rules between different modalities of data, and detect contradictory data through graph neural networks;

[0028] The dynamic construction module for medical knowledge graphs is used for adapting to multi-source heterogeneous data sources, extracting multimodal knowledge and semantic alignment, using the clinical BERT model to perform nested entity recognition on text data, extracting relationships between entities through dependency parsing, constructing graph convolution interaction between dynamic instance graphs and static knowledge graphs, and capturing implicit associations.

[0029] The large model pre-training and knowledge enhancement module is used to inject knowledge graph entity embeddings as prior knowledge into the attention layer and optimize the semantic alignment between modalities through contrastive learning; a multi-task fine-tuning mechanism is adopted to combine tasks and use clinical annotation data to achieve professional model adaptation; when new evidence appears, the knowledge graph is dynamically expanded through graph neural network embedding propagation algorithm.

[0030] The multimodal data fusion diagnosis and decision generation module is used to achieve deep correlation between image features and text and gene data using a weighted attention mechanism; it inputs preprocessed multimodal data into a trained large model, and uses the feature extraction and fusion capabilities of the large model to generate a comprehensive feature representation; it combines raw data stitching, graph neural network interaction and decision layer integration to realize a hybrid fusion strategy; and it generates a detailed diagnostic report based on the diagnostic results.

[0031] According to one aspect of the present invention, a storage medium is provided, wherein the storage medium stores instructions that, when read by a computer, cause the computer to execute the multimodal data fusion diagnostic method based on knowledge graphs and large models as described above.

[0032] According to another aspect of the present invention, an electronic device is provided, comprising a processor and the aforementioned storage medium, wherein the processor executes instructions in the storage medium.

[0033] Compared with existing technologies, the beneficial effects of the present invention are as follows: The present invention can effectively integrate the advantages of knowledge graphs and large models to achieve in-depth collaborative analysis of medical multimodal data, thereby improving the accuracy and efficiency of medical diagnosis. Attached Figure Description

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention.

[0035] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0036] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0037] To better understand the technical content of this invention, specific embodiments are described below in conjunction with the accompanying drawings. Various aspects of this invention are described with reference to the accompanying drawings, which illustrate numerous illustrative embodiments. The embodiments of this invention are not limited to those shown in the drawings. It should be understood that this invention is implemented through any of the various concepts and embodiments described above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed in this invention are not limited to any particular implementation. Furthermore, some aspects of this invention can be used alone or in any suitable combination with other aspects disclosed in this invention.

[0038] Example 1: As Figure 1 As shown, this embodiment takes a hospital's medical diagnostic system as an example to explain in detail the specific implementation process of a medical multimodal data fusion diagnostic method and system based on knowledge graphs and large models.

[0039] This embodiment provides a multimodal data fusion diagnostic method based on knowledge graphs and large models, including the following steps:

[0040] S1. Data Collection and Preprocessing:

[0041] Collect patients' electronic medical records, medical images (such as CT, MRI, etc.), and physiological monitoring data (such as electrocardiogram, blood pressure, etc.).

[0042] Text data in electronic medical records is cleaned to remove irrelevant information and erroneous data; medical image data is denoised to improve image quality; and physiological monitoring data is normalized to make it comparable.

[0043] S2. Knowledge Graph Construction:

[0044] By integrating professional knowledge, clinical guidelines, and research findings in the medical field, a medical knowledge graph can be constructed. For example, information such as disease symptoms, causes, and treatments can be used as nodes in the knowledge graph, and the relationships between them can be used as edges.

[0045] Using natural language processing technology, medical knowledge is automatically extracted from medical literature and clinical reports and added to a knowledge graph.

[0046] S3, Large Model Training:

[0047] We chose a Transformer-based model as the large model architecture because it has powerful multimodal data processing capabilities.

[0048] The large model is trained using a massive medical multimodal dataset, including electronic medical records, medical images, and physiological monitoring data for various diseases. Through training, the model learns the correlations and feature representations between different modalities of data.

[0049] S4. Multimodal data fusion diagnosis:

[0050] The preprocessed multimodal data is input into a trained large model, which then extracts and fuses features from the data to generate a comprehensive feature representation.

[0051] By combining medical knowledge from the knowledge graph, reasoning and analysis are performed on the comprehensive feature representation. For example, based on the relationship between diseases and symptoms in the knowledge graph, the type of disease a patient may have can be determined.

[0052] Based on the diagnosis results, a detailed diagnostic report is generated, including information such as the type of disease, severity, and possible causes, for doctors to refer to and make decisions.

[0053] Through the above methods and systems, this invention can effectively integrate the advantages of knowledge graphs and large models to achieve in-depth collaborative analysis of medical multimodal data, thereby improving the accuracy and efficiency of medical diagnosis.

[0054] A multimodal data fusion diagnostic system based on knowledge graphs and large models includes:

[0055] The preprocessing and standardization module is used to collect patients' multimodal medical data, preprocess the collected multimodal medical data, construct association constraint rules between different modalities of data, and detect contradictory data through graph neural networks;

[0056] The dynamic construction module for medical knowledge graphs is used for adapting to multi-source heterogeneous data sources, extracting multimodal knowledge and semantic alignment, using the clinical BERT model to perform nested entity recognition on text data, extracting relationships between entities through dependency parsing, constructing graph convolution interaction between dynamic instance graphs and static knowledge graphs, and capturing implicit associations.

[0057] The large model pre-training and knowledge enhancement module is used to inject knowledge graph entity embeddings as prior knowledge into the attention layer and optimize the semantic alignment between modalities through contrastive learning; a multi-task fine-tuning mechanism is adopted to combine tasks and use clinical annotation data to achieve professional model adaptation; when new evidence appears, the knowledge graph is dynamically expanded through graph neural network embedding propagation algorithm.

[0058] The multimodal data fusion diagnosis and decision generation module is used to achieve deep correlation between image features and text and gene data using a weighted attention mechanism; it inputs preprocessed multimodal data into a trained large model, and uses the feature extraction and fusion capabilities of the large model to generate a comprehensive feature representation; it combines raw data stitching, graph neural network interaction and decision layer integration to realize a hybrid fusion strategy; and it generates a detailed diagnostic report based on the diagnostic results.

[0059] Example 2:

[0060] The computer-readable storage medium of this embodiment stores a computer program that, when executed by a processor, implements the steps of the multimodal data fusion diagnostic method based on knowledge graphs and large models in Embodiment 1.

[0061] The computer-readable storage medium in this embodiment can be an internal storage unit of the terminal, such as the terminal's hard disk or memory; the computer-readable storage medium in this embodiment can also be an external storage device of the terminal, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc. equipped on the terminal; furthermore, the computer-readable storage medium can include both the terminal's internal storage unit and external storage devices.

[0062] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0063] Example 3:

[0064] The computer device of this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the multimodal data fusion diagnostic method based on knowledge graphs and large models in Embodiment 1.

[0065] In this embodiment, the processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The memory can include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory can also include non-volatile random access memory. For example, the memory can also store device type information.

[0066] Those skilled in the art will understand that the content disclosed in the embodiments can be provided as a method, system, or computer program product. Therefore, this solution can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this solution can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage) containing computer-usable program code.

[0067] This solution is described with reference to flowchart illustrations and / or block diagrams of methods and computer program products according to embodiments of this solution. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0068] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0069] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0070] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0071] The examples described herein are merely preferred embodiments of the invention and are not intended to limit the concept and scope of the invention. Any modifications and improvements made by those skilled in the art to the technical solutions of the invention without departing from the design concept of the invention should fall within the protection scope of the invention.

Claims

1. A multimodal data fusion diagnostic method based on knowledge graphs and large models, characterized in that, Includes the following steps: Multimodal data collection; Building a medical knowledge graph: Large model pre-training and knowledge augmentation; Multimodal data fusion diagnosis and decision generation.

2. The method according to claim 1, characterized in that, The specific steps for multimodal data collection are as follows: Collect multimodal medical data from patients, including structured and unstructured data; The collected multimodal medical data is preprocessed. For text data, medical entity recognition technology based on the BioBERT-CRF model is used to automatically correct spelling errors and non-standardized terms in electronic medical records. For image data, adaptive filtering algorithms are used to eliminate noise and motion artifacts in CT / MRI images, and GAN (Generative Adversarial Network) is used to repair low-quality images. For time-series data, dynamic time warping is used to align physiological monitoring data with different sampling rates, and Kalman filtering is used to eliminate baseline drift of electrocardiogram signals. Construct association constraint rules between different modalities of data, and detect contradictory data through graph neural networks.

3. The method according to claim 1, characterized in that, The specific steps for constructing a medical knowledge graph are as follows: Adapt to multi-source heterogeneous data sources, integrate structured data, semi-structured data and unstructured data, and use a federated learning framework to achieve cross-institutional data collaboration; Multimodal knowledge extraction and semantic alignment: The clinical BERT model is used to perform nested entity recognition on text data, and the relationships between entities are extracted through dependency parsing; the key regions in the image are segmented using object detection algorithms to extract morphological parameters; and deep semantic features are extracted using 3D ResNet50. The Transformer encoder is used to extract temporal signals and align them with the features of the text description; UMLS standard terminology is used to map entities of different modalities to a unified semantic space; attribute weight calculation and constraint setting are introduced to enhance the clinical interpretability of the knowledge graph. Construct graph convolutional interactions between dynamic instance graphs and static knowledge graphs to capture implicit relationships.

4. The method according to claim 3, characterized in that, The specific steps of large model pre-training and knowledge augmentation are as follows: We chose a Transformer-based model and adopted a knowledge graph-guided pre-training strategy. During the training of the large model, we injected knowledge graph entity embeddings as prior knowledge into the attention layer and optimized the semantic alignment between modalities through contrastive learning. A multi-task fine-tuning mechanism is adopted, which combines tasks and utilizes clinical annotation data to achieve professional model adaptation; When new evidence emerges, the knowledge graph is dynamically expanded using a graph neural network embedding propagation algorithm.

5. The method according to claim 4, characterized in that, The steps involved in multimodal data fusion diagnosis and decision generation are as follows: A weighted attention mechanism is used to achieve deep association between image features and text and gene data; preprocessed multimodal data is input into a trained large model, and the feature extraction and fusion capabilities of the large model are used to generate a comprehensive feature representation; By combining raw data stitching, graph neural network interaction, and decision-making layer integration, a hybrid fusion strategy is achieved. Based on the diagnostic results, a detailed diagnostic report is generated.

6. A multimodal data fusion diagnostic system based on knowledge graphs and large models, characterized in that, include: The preprocessing and standardization module is used to collect patients' multimodal medical data, preprocess the collected multimodal medical data, construct association constraint rules between different modalities of data, and detect contradictory data through graph neural networks; The dynamic construction module for medical knowledge graphs is used for adapting to multi-source heterogeneous data sources, extracting multimodal knowledge and semantic alignment, using the clinical BERT model to perform nested entity recognition on text data, extracting relationships between entities through dependency parsing, constructing graph convolution interaction between dynamic instance graphs and static knowledge graphs, and capturing implicit associations. The large model pre-training and knowledge augmentation module is used to inject knowledge graph entity embeddings as prior knowledge into the attention layer and optimize the semantic alignment between modalities through contrastive learning. A multi-task fine-tuning mechanism is adopted, which combines tasks and utilizes clinically labeled data to achieve professional model adaptation; when new evidence emerges, the knowledge graph is dynamically expanded through graph neural network embedding propagation algorithm. The multimodal data fusion diagnosis and decision generation module is used to achieve deep correlation between image features and text and gene data using a weighted attention mechanism; it inputs preprocessed multimodal data into a trained large model, and uses the feature extraction and fusion capabilities of the large model to generate a comprehensive feature representation; it combines raw data stitching, graph neural network interaction and decision layer integration to realize a hybrid fusion strategy; and it generates a detailed diagnostic report based on the diagnostic results.

7. A storage medium, characterized in that, The storage medium stores instructions that, when read by a computer, cause the computer to execute the multimodal data fusion diagnostic method based on knowledge graphs and large models as described in claims 1-5.

8. An electronic device, characterized in that, It includes a processor and the storage medium of claim 7, wherein the processor executes instructions in the storage medium.

Citation Information

Cited By

  • Virtual reality data acquisition and restoration method

    CN121707869A

  • Medical informatization data intelligent analysis system based on large model

    CN121905543A