Multimodal intelligent medical scene dynamic reasoning method and device based on medical knowledge

By introducing medical knowledge graphs and dynamic inference trees into intelligent medical systems, the shortcomings of existing systems in multimodal data processing and inference path selection are solved, efficient feature fusion and personalized inference path selection are achieved, and the intelligence level and inference accuracy of the system are improved.

CN119294534BActive Publication Date: 2025-05-02ZHEJIANG YISHAN SMART MEDICAL RES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411824509.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-02
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

The existing intelligent medical system has shortcomings in multimodal data processing and inference path selection, and cannot fully adapt to the personalized needs of medical scenarios, lacks effective utilization of medical knowledge, and the inference path optimization methods are rigid, so it is unable to cope with complex scenarios and individual differences.

Method used

A multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge is adopted, and feature enhancement is achieved through the introduction of medical knowledge graphs, adaptive feature fusion is achieved by combining dynamic importance weights and attention weights, and graph neural networks and deep reinforcement learning are used to construct a dynamic reasoning tree to optimize inference path selection.

Benefits of technology

It realizes efficient feature extraction and fusion of multimodal data, enhances the scientificity and accuracy of medical reasoning, meets the needs of personalized, high concurrency and real-time feedback in intelligent medical scenarios, and improves the overall reasoning efficiency and intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119294534B_ABST
    Figure CN119294534B_ABST
Patent Text Reader

Abstract

The present application proposes a multimodal intelligent medical scene dynamic reasoning method and device based on medical knowledge, which obtains the user's multimodal data, extracts the multimodal data into multimodal features, inputs the text modal features into a graph neural network preset with a medical knowledge graph for feature enhancement to obtain medical knowledge enhanced text features, inputs the text modal features into a memory network stored in the user's personalized health record for feature enhancement to obtain archive enhanced text features, splices the medical knowledge enhanced text features and the archive enhanced text features to obtain context enhanced text features; inputs the context enhanced text features and the multimodal features into a pre-trained multimodal fusion model to output fusion features; inputs the fusion features into a dynamic reasoning tree constructed by multiple reasoning nodes in a graph structure to determine the optimal reasoning node, and uses the optimal reasoning node to reason on the fusion features to achieve dynamic reasoning adapted to medical scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart medical care, and in particular to a multimodal smart medical scenario dynamic reasoning method and device based on medical knowledge. Background Art

[0002] In the process of the continuous development of intelligent medical systems, multimodal data processing and reasoning path selection have become key technical fields. In terms of multimodal feature extraction and fusion: existing methods usually use convolutional neural networks (CNN) to process image data to extract features, use natural language processing models (such as BERT) to encode text, and use simple connections or fixed weight linear weighting methods for multimodal feature fusion. Each modality is simply merged after independent processing. This method does not fully consider the particularity and personalized needs of medical scenarios, and cannot adaptively adjust the weights of multimodal features according to the dynamic changes of medical scenarios. In terms of reasoning path selection and optimization: most existing methods are based on predefined decision trees, fixed edge / cloud node switching strategies, and Markov decision processes to determine the reasoning path. However, these methods have obvious shortcomings when facing complex scenarios and personalized needs in the medical field. There is a lack of dynamic evaluation of the importance of reasoning nodes in different scenarios, and the path optimization means are rigid. It does not fully consider the personalized needs of medical reasoning tasks and changes in system load.

[0003] In other words, although there are already solutions for multimodal medical reasoning in the field of intelligent medicine, most of them still have many defects in actual medical application scenarios:

[0004] First, there is a lack of personalized adaptability to medical scenarios: multimodal feature extraction and fusion methods do not pay attention to the characteristics of medical scenarios. It is difficult to dynamically adjust the importance of modal features according to the patient's condition, scenario characteristics (such as emergency or routine physical examinations) and personalized historical health data, resulting in poor feature fusion effects and inability to efficiently support medical reasoning tasks;

[0005] Second, medical knowledge is not fully utilized to assist reasoning: Existing technologies mostly process multimodal data independently, and fail to systematically integrate medical knowledge graphs into the feature extraction and reasoning process, resulting in insufficient understanding of medical concepts and their associations. Given that concepts such as diseases, symptoms, and drugs are closely related in the medical field, this lack will reduce the accuracy of reasoning;

[0006] Third, the reasoning path selection method is too simple: it is mostly based on the Markov decision process (MDP), whose decision-making strategy relies on historical experience and a fixed reward mechanism, and cannot be adaptively optimized. It has significant limitations in dealing with individual differences and real-time changing scenarios. In addition, MDP does not combine the medical knowledge graph to evaluate the priority of the reasoning node, and cannot reflect the relevance of medical knowledge, making the reasoning path selection unscientific.

[0007] Fourth, there is a lack of consideration for complex relationships between reasoning nodes: the existing MDP and reasoning path selection mechanism is difficult to cope with the complex relationships between reasoning nodes and the dynamic changes of resources. In medical reasoning, the interdependence of reasoning nodes is complex and cannot meet the requirements of the optimal reasoning path under high concurrency and high-dimensional state space.

[0008] In summary, existing technical solutions are difficult to fully meet the requirements of personalization, high concurrency and real-time feedback in intelligent medical scenarios. There is an urgent need for innovations in feature extraction, knowledge enhancement, reasoning path optimization, etc., especially solutions that enhance reasoning capabilities by combining dynamic adjustment of medical scenarios and deep embedding of knowledge graphs. Summary of the invention

[0009] The present application scheme provides a multimodal intelligent medical scenario dynamic reasoning method and device based on medical knowledge, introduces a medical knowledge graph to perform feature enhancement on multimodal data, and combines dynamic importance weights and attention weights to achieve adaptive feature fusion to obtain high-quality fused feature representation; introduces a dynamic reasoning tree optimization module that combines graph neural networks and deep reinforcement learning to achieve dynamic reasoning of medical tasks and improve the accuracy of the overall reasoning results.

[0010] To achieve the above objectives, in the first aspect, this solution provides a multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge, comprising the following steps:

[0011] Acquire multimodal data of the user, wherein the multimodal data includes text modal data, voice modal data, and image modal data;

[0012] Inputting the multimodal data into a trained feature extraction module to extract multimodal features, wherein the multimodal features include text modality features, speech modality features, and image modality features;

[0013] Input the text modality features into a graph neural network with a preset medical knowledge graph for feature enhancement to obtain medical knowledge enhanced text features, input the text modality features into a memory network stored in a user's personalized health record for feature enhancement to obtain record enhanced text features, and concatenate the medical knowledge enhanced text features and the record enhanced text features to obtain context enhanced text features;

[0014] Input context-enhanced text features and multimodal features into the pre-trained multimodal fusion model to output fusion features;

[0015] The fused features are input into a dynamic inference tree constructed by a graph structure of multiple inference nodes to determine the optimal inference node, wherein the approximation between the medical knowledge enhanced text features and the fused features of each inference node is calculated, and the priority of the current inference node is determined based on the approximation of the current inference node, and the inference node with the highest score is selected in combination with the priority of the current inference node and the historical quality value;

[0016] The fused features are inferred using the inference node with the highest score.

[0017] Secondly, this solution provides a multi-modal intelligent medical scene dynamic reasoning device based on medical knowledge, including:

[0018] A data acquisition unit, used to acquire multimodal data of a user, wherein the multimodal data includes text modal data, voice modal data, and image modal data;

[0019] A modal feature conversion unit, used for inputting multimodal data into a trained feature extraction module to extract multimodal features, wherein the multimodal features include text modal features, speech modal features, and image modal features;

[0020] A context-enhanced text feature acquisition unit is used to input text modal features into a graph neural network preset with a medical knowledge graph to perform feature enhancement to obtain medical knowledge enhanced text features, input text modal features into a memory network stored in a user's personalized health record to perform feature enhancement to obtain record enhanced text features, and concatenate the medical knowledge enhanced text features and the record enhanced text features to obtain context-enhanced text features;

[0021] A fusion feature acquisition unit, used for inputting context-enhanced text features and multimodal features into a pre-trained multimodal fusion model to output fusion features;

[0022] An inference node selection unit, used for inputting the fusion feature into a dynamic inference tree constructed by a plurality of inference nodes in a graph structure to determine the optimal inference node, wherein the approximation between the medical knowledge enhanced text feature and the fusion feature of each inference node is calculated, and the priority of the current inference node is determined based on the approximation of the current inference node, and the inference node with the highest score is selected in combination with the priority of the current inference node and the historical quality value;

[0023] The reasoning unit uses the highest-scoring reasoning node to reason about the fused features.

[0024] In a third aspect, the present solution provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute a multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge.

[0025] Different from the prior art, the main contributions and innovations of the present invention are as follows:

[0026] This application proposes a dynamic reasoning method for multimodal intelligent medical scenarios based on medical knowledge, which provides an efficient and accurate feature extraction and fusion method for multimodal data in intelligent medical scenarios. It combines personalized health records with medical knowledge graphs for multimodal data input such as voice, text, and images for feature enhancement, and realizes adaptive feature fusion through dynamic weights and attention weights, thereby providing high-quality fusion feature representation for subsequent reasoning modules, and provides a deep optimization method that combines adaptive dynamic adjustment of medical scenarios and uses medical knowledge graphs and personalized archives to assist in fusion. In addition, this solution also proposes a dynamic reasoning tree that combines graph neural networks and deep reinforcement learning, which aims to intelligently select the optimal reasoning path based on multimodal fusion features to meet the needs of personalization, high concurrency, and real-time feedback in intelligent medical scenarios. The introduction of medical knowledge graphs can effectively evaluate the priority of reasoning nodes, and combine historical quality value comprehensive scores for path selection. At the same time, the reinforcement learning method is used to reward feedback and update the quality of the reasoning path, so that the reasoning process has the characteristics of self-adaptation and knowledge-driven, ensuring the optimal match between the reasoning node and the current medical task, thereby improving the efficiency and accuracy of the overall reasoning of the system, thereby greatly improving the system's understanding and reasoning ability of multimodal data in intelligent medical scenarios, effectively responding to personalized, diversified, and real-time medical tasks, and ultimately improving the accuracy of medical reasoning and the intelligence level of the overall system. It has the following many beneficial effects:

[0027] Highly adaptable to different medical scenarios: Dynamic weights and attention weights are combined with medical knowledge to adjust the importance of multimodal data in different medical scenarios. The dynamic reasoning tree driven by medical knowledge ensures that the reasoning path can be dynamically adjusted in different medical scenarios such as emergency and routine follow-up. The system can adaptively select the optimal reasoning path according to changes in the patient's condition, thereby effectively supporting the personalized needs of different medical scenarios and effectively solving the problem of insufficient adaptability of existing technologies.

[0028] Accurate personalized feature fusion: By introducing personalized health records and medical knowledge graphs, and using dynamic importance weight calculation and adaptive attention mechanism, we can accurately capture the key characteristics of patients' specific conditions to obtain representative fusion features, improve the defects of existing methods in personalized feature fusion, and improve fusion accuracy.

[0029] Enhanced utilization of medical reasoning knowledge: Introducing medical knowledge graphs to enhance fusion features and combining them with dynamic selection of dynamic reasoning trees, the system can make full use of disease, symptom, and drug-related information to solve the problem of knowledge gaps, improve the scientificity and accuracy of reasoning, and meet complex reasoning needs.

[0030] Optimize reasoning path selection: Combine graph neural networks and deep reinforcement learning to build a dynamic reasoning tree, optimize the reasoning path based on comprehensive evaluation of medical knowledge graphs and node historical quality values, enhance dynamic adaptability, overcome the lack of dynamics in existing path selection, and improve reasoning efficiency and accuracy.

[0031] Effectively handle inference node associations: Using graph neural networks to learn the global structure of dynamic inference trees and combining deep reinforcement learning to optimize node associations can cope with high concurrency and high complexity inference requirements, solve the problem that existing technologies do not handle complex inference node associations and resource changes well, and achieve efficient and stable inference task execution.

[0032] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0034] Figure 1 It is a logical diagram of the dynamic reasoning method of multimodal intelligent medical scenarios based on medical knowledge provided by this solution.

[0035] Figure 2 It is a schematic diagram of the structure of a multimodal intelligent medical scenario dynamic reasoning device based on medical knowledge.

[0036] Figure 3 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0037] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0038] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0039] Embodiment 1

[0040] like Figure 1 As shown, the embodiment of the present application provides a multimodal intelligent medical scene dynamic reasoning method based on medical knowledge, comprising the following steps:

[0041] Acquire multimodal data of the user, wherein the multimodal data includes text modal data, voice modal data, and image modal data;

[0042] Inputting the multimodal data into a trained feature extraction module to extract multimodal features, wherein the multimodal features include text modality features, speech modality features, and image modality features;

[0043] Input the text modality features into a graph neural network with a preset medical knowledge graph for feature enhancement to obtain medical knowledge enhanced text features, input the text modality features into a memory network stored in a user's personalized health record for feature enhancement to obtain record enhanced text features, and concatenate the medical knowledge enhanced text features and the record enhanced text features to obtain context enhanced text features;

[0044] Input context-enhanced text features and multimodal features into the pre-trained multimodal fusion model to output fusion features;

[0045] The fused features are input into a dynamic inference tree constructed by a graph structure of multiple inference nodes to determine the optimal inference node, wherein the approximation between the medical knowledge enhanced text features and the fused features of each inference node is calculated, and the priority of the current inference node is determined based on the approximation of the current inference node, and the inference node with the highest score is selected in combination with the priority of the current inference node and the historical quality value;

[0046] The fused features are inferred using the inference node with the highest score.

[0047] The dynamic reasoning method for multimodal intelligent medical scenarios based on medical knowledge provided by this solution is different from the traditional multimodal medical reasoning solution. It uses a graph attention network combined with a medical knowledge graph to enhance the text modal features, and uses a memory network combined with personalized health records to perform additional feature enhancement on the text modal features. The multimodal fusion model is used to achieve adaptive feature fusion of multimodal features and context-enhanced features, thereby providing high-quality fusion feature representation for subsequent reasoning modules; at the same time, a dynamic reasoning tree is used in combination with medical knowledge to enhance text features to intelligently select the optimal reasoning path, meeting the needs of personalization, high concurrency and real-time feedback in intelligent medical scenarios. Through the introduction of the medical knowledge graph, the dynamic reasoning tree can effectively evaluate the priority of the reasoning path, and select the optimal reasoning path in combination with the comprehensive score of the historical quality value, ensuring the optimal selectivity of the reasoning path.

[0048] This solution is aimed at users of the medical system to perform reasonable reasoning on medical tasks. Users of this solution can input multimodal data into the multimodal intelligent medical scenario dynamic reasoning system provided by this solution through mobile devices, computers and other terminals. The text modal data in the multimodal data refers to text modal data, such as health questionnaires and text medical records; voice modal data refers to voice modal data, such as voice data describing symptoms; and image modal data refers to image modal data, such as physical examination images.

[0049] In order to process multimodal data of different modalities, this solution needs to input the multimodal data into a trained feature extraction module to extract multimodal features of different modalities respectively.

[0050] For the speech modal data in the multimodal data, the speech modal data is input into the trained feature extraction module to extract the Mel frequency cepstrum coefficient as the speech modal feature. Specifically, the speech modal data is converted into a frequency domain representation through short-time Fourier transform, the power spectrum is calculated based on the frequency domain representation, the power spectrum is filtered by the Mel filter bank to extract the filtered frequency feature, the filtered frequency feature Mm is first logarithmically operated and then the Mel frequency cepstrum coefficient feature is obtained by discrete cosine transform as the speech modal feature.

[0051] The specific process of extracting speech modal features is as follows:

[0052]

[0053] Where w(n) is the window function, represents frequency domain representation, f is frequency, t is time, x(n) represents speech modal data, n represents discrete time index, represents a complex exponential function;

[0054] ;

[0055] in is the power spectrum;

[0056] ;

[0057] in represents the coefficient of the mth mel filter, N represents the total number of filters, Indicates frequency representation, Indicates the input signal frequency The power spectrum at ;

[0058] ;

[0059] in It represents the Mel frequency cepstral coefficient feature, which corresponds to the speech modal feature. log() represents logarithmic operation, and DCT() represents discrete cosine transform.

[0060] For the text modality data in the multimodal data, the text modality data is input into the trained feature extraction module to convert each word into a context-dependent embedding vector and form a text modality feature. In some embodiments, the trained feature extraction module for extracting text modality features is ClinicalBERT specialized for medical scenarios. ClinicalBERT is a pre-trained language model based on the BERT architecture, which is optimized and trained specifically for medical scenarios. ClinicalBERT captures the dependency between words through a multi-layer Transformer structure, which can enhance the feature extraction module's ability to understand medical terms in clinical texts, thereby generating context-dependent text modality features.

[0061] The specific text modality feature extraction is as follows: ;

[0062] where d is the dimension of the embedding vector, Represents the embedding vector of the nth word.

[0063] For the image modality data in the multimodal data, the image modality data is input into a trained feature extraction module for feature extraction to obtain image modality features. In some embodiments, the trained feature extraction module for extracting image modality features is a pre-trained convolutional neural network that has been fine-tuned for medical fields.

[0064] The specific image modality feature extraction is expressed as follows:

[0065] ;

[0066] Where h, w, c represent the height, width and number of channels of the convolutional feature map of the image, respectively, and V represents the image modality feature.

[0067] In order to enhance the semantic understanding and medical relevance of multimodal features, this scheme inputs text modal features into a graph attention network with a preset medical knowledge graph for feature enhancement to obtain medical knowledge enhanced text features, where medical knowledge is taken as the nodes of the medical knowledge graph, and the relationships between medical knowledge are taken as the edges between nodes. Medical knowledge includes diseases, symptoms, drugs, and treatment methods, etc.

[0068] The medical knowledge graph of this solution is input into the graph neural network and converted into a knowledge graph vector. The text modality features are input into the graph neural network and the graph attention operation is performed on the knowledge graph vector to obtain the medical knowledge enhanced text features. By introducing the medical knowledge graph, this solution enables the system to more effectively understand the medical terms in the text input and their associations, and adapt to the unique logical relationships in medicine.

[0069] Specifically, the medical knowledge graph obtained by this scheme can be represented in the graph neural network as a knowledge graph vector composed of embedded representations of different knowledge graph nodes, expressed as:

[0070] ;

[0071] KG is the medical knowledge graph, and GNN () is the graph neural network, which can obtain the structural features of the entire medical knowledge graph by iteratively aggregating the neighbor information of the nodes of the medical knowledge graph, and finally generate an embedded representation of each node (such as disease, symptom, etc.). Represented as a knowledge graph vector.

[0072] The text modality features are input into the graph neural network with a preset medical knowledge graph to extract the text modality subgraph. The text modality subgraph is embedded with the knowledge graph vector through the graph attention operation to obtain the medical knowledge enhanced text features, which are expressed as:

[0073] ;

[0074] in Represented as text modality features, GAT() represents graph attention operation, Representing medical knowledge enhanced text features.

[0075] Furthermore, this solution also needs to pay attention to the personalized characteristics of users. Correspondingly, this solution loads the personalized health record of the current user into the memory network, where the personalized health record includes but is not limited to text content such as historical medical history, drug records, and genetic information. This method can dynamically extract user historical information related to the current medical scenario, so that the model has the ability to adapt to individual health data.

[0076] Specifically, the text modality features are input into the memory network that stores the user's personalized health records for feature enhancement to obtain the record-enhanced text features represented as:

[0077] ;

[0078] MemNet() represents the memory network, A represents the user's personalized health record, Represents archival enhanced text features.

[0079] This solution combines the medical knowledge enhanced text features and the archive enhanced text features to obtain the context enhanced text features, so as to obtain the feature representation in a specific medical scenario that combines the user's individual historical information and medical information, which is expressed as: ;

[0080] in Represents context-enhanced text features.

[0081] In the intelligent medical scenario, the importance of multimodal data of different modalities changes dynamically with the changes of the patient's clinical context, medical scenario and condition. In order to maximize the use of key information in the multimodal feature fusion process, this solution introduces a dynamic importance weight calculation module based on the medical scenario in the pre-trained multimodal fusion model. The dynamic importance weight calculation module can dynamically adjust the corresponding importance of different multimodal features (such as text, image, voice, etc.) according to the patient's current condition and medical scenario (such as emergency, routine physical examination, etc.), ensuring that key medical information is given priority attention, thereby improving the accuracy and adaptability of medical decision-making. For example, in an actual emergency scenario, assuming that the patient complains of acute chest pain and uploads relevant imaging data and medical history information, through medical knowledge graph reasoning, the system finds that chest pain is highly correlated with cardiovascular disease, so the correlation measurement of imaging data is assigned a higher value. In this scenario, the weight of the image feature will be significantly higher than that of the text feature or voice feature, thereby giving priority to the contribution of the imaging information to the final diagnosis result. At the same time, the pre-trained multimodal fusion model also needs to consider the attention weights to achieve the fusion of multimodal features through the attention mechanism. This process ensures that the fused features after fusion not only reflect the dynamic importance of each modality in the current scenario, but also consider the relative contribution of each feature.

[0082] Specifically, the multimodal fusion model of the present scheme includes a dynamic weight allocation module for generating dynamic weights, an attention weight allocation module for generating attention weights, and a fusion module, which are connected in sequence. The context-enhanced text features and multimodal features are input into the dynamic weight allocation module to generate their respective dynamic weights. The context-enhanced text features and multimodal features are input into the attention weight allocation module to generate their respective attention weights. The context-enhanced text features and multimodal features are fused based on the attention weights and dynamic weights to obtain fused features.

[0083] Regarding the dynamic weight assignment module, the cosine similarity of each multimodal feature is calculated based on the context-enhanced text feature, and the normalized value of the cosine similarity of each multimodal feature is used as the dynamic weight of each multimodal feature. This solution quantifies the importance of each multimodal feature in the current medical scenario by calculating the cosine similarity between each multimodal feature and the context-enhanced text feature. Generally speaking, the higher the cosine similarity, the greater the importance of the current multimodal feature in the current medical scenario. For example, in acute conditions, the relevance of imaging features may be higher than that of text features, while in chronic disease follow-up, text features may have a higher weight.

[0084] Specifically, the calculation formula of dynamic weight is as follows;

[0085] ;

[0086] in represents the nth multimodal feature, represents context-enhanced text features, Represents the cosine correlation of the nth multimodal feature;

[0087] ;

[0088] in represents the dynamic weight of the nth multimodal feature, N represents the total number of multimodal features, and this formula ensures that the importance of each multimodal feature is quantified as a normalized dynamic weight.

[0089] Regarding the attention weight allocation module, the context-enhanced text features and multimodal features are mapped to the same spatial dimension and then merged to obtain the multimodal feature matrix, and the attention weight of each multimodal feature is calculated through the attention mechanism.

[0090] Specifically, mapping the context-enhanced text features and multimodal features to the same spatial dimension can ensure that each feature has the same representation when fused, so as to facilitate subsequent weighted calculation, which is specifically expressed as:

[0091] ;

[0092] ;

[0093] ;

[0094] ;

[0095] ;

[0096] in , , , as well as They are respectively represented as the projection matrices of medical knowledge enhanced text features, archive enhanced text features, text modality features, speech modality features, and image modality features. , , , , They are respectively represented as bias items of medical knowledge enhanced text features, archive enhanced text features, text modality features, speech modality features, and image modality features. Each feature is mapped to a unified dimensional space to ensure that each modality feature can be fused and the final features have a unique dimension. .

[0097] The features mapped to the same spatial dimension are merged to obtain a multimodal feature matrix, which integrates the feature information of all modalities and facilitates the subsequent calculation of the attention weights between different modalities. It is expressed as:

[0098] ;

[0099] Where M is a multimodal feature matrix, and each column corresponds to the feature representation of a mode, which are dimensional vector.

[0100] The attention weight of each multimodal feature is calculated through the attention mechanism to indicate the importance of the feature representation of the current modality, which is expressed as:

[0101] ;

[0102] in is the learnable weight matrix, represents the i-th multimodal feature or context-enhanced text feature, N is the total number of multimodal features and context-enhanced text features, is the attention weight of the i-th multimodal feature or context-enhanced text feature, indicating the contribution of the multimodal feature or context-enhanced text feature in the current fusion task.

[0103] It should be noted that the dynamic weight of the context-enhanced text feature is 1. This scheme takes the product of the dynamic weight of each context-enhanced text feature or multimodal feature and the attention weight as the weighting coefficient, and performs weighted fusion with the weighting coefficients of the context-enhanced text feature and the multimodal feature to obtain the fused feature. This ensures that the fusion process takes into account both the importance of the dynamic features of the current medical scenario and the adaptive contribution of each modal feature.

[0104] Specifically, the acquisition of fusion features is expressed as:

[0105]

[0106] ;

[0107] in represents the attention weight of the mth multimodal feature or context-enhanced text feature, represents the dynamic weight of the mth multimodal feature or context-enhanced text feature, represents the weight coefficient of the mth multimodal feature or context-enhanced text feature, M represents the total number of multimodal features and context-enhanced text features, Represents fusion features.

[0108] This solution adopts a dynamic reasoning tree constructed with a graph structure composed of multiple reasoning nodes, combined with fusion features to select the optimal reasoning path, thereby realizing the requirements of personalization, high concurrency and real-time feedback in smart medical scenarios.

[0109] Regarding the dynamic reasoning tree of this solution,

[0110] The dynamic inference tree designed in this solution is an adaptive inference structure based on hierarchical design, which can efficiently select the appropriate inference path according to the dynamic characteristics of the fusion features. Specifically, the dynamic inference tree of this solution is modeled using a graph structure, where each inference node is a node of the graph and the inference path is an edge in the graph. The graph neural network is used to learn the characteristics of the entire graph structure to further help optimize the selection of the inference path.

[0111] Specifically, the dynamic reasoning tree includes multiple reasoning nodes, which are organized into a multi-level structure. According to the type of reasoning nodes, the reasoning nodes can be divided into: edge nodes, cloud nodes and adaptive nodes. The edge nodes are responsible for simple computing tasks with low latency, such as basic health monitoring and daily vital sign monitoring; the cloud nodes are responsible for complex and large-scale reasoning tasks, such as image analysis and disease prediction; the adaptive nodes have adaptive adjustment capabilities to flexibly switch between edge nodes and cloud nodes according to the real-time requirements of the task and the resource status of the system.

[0112] The node status of each inference node includes node processing capability and resource availability. Node processing capability is used to measure the efficiency of the node in processing inference tasks in the current scenario, which is related to the node's hardware configuration, computing power and current load. Resource availability is used to measure the current system resource status of the node (such as remaining CPU, memory, etc.), thereby reflecting its spare capacity for processing new tasks. Correspondingly, the selection of the optimal inference path is to combine the status of the inference node and the current task requirements to find the best inference path that meets the task requirements, making the inference process more efficient.

[0113] When this solution uses the dynamic reasoning tree to select the optimal reasoning node, the dynamic reasoning tree can be combined with the medical knowledge graph to evaluate the priority of the reasoning node to achieve intelligent reasoning driven by the medical knowledge graph.

[0114] Specifically, in the selection process of the reasoning node, each reasoning node needs to calculate its relevance to the current reasoning task based on the medical knowledge enhanced text features to guide the setting of the priority of the reasoning node. Correspondingly, for each reasoning node, the approximation between the medical knowledge enhanced text features of each reasoning node and the fusion features is calculated, and the weighted value based on the approximation of the current reasoning node, the node processing capacity and the resource availability is used as the priority of the current reasoning node, and the weighted sum of the priority of the current reasoning node and the historical quality value is used as the comprehensive score of the current reasoning node, and the reasoning node with the highest comprehensive score is selected.

[0115] The specific calculation formula is as follows:

[0116]

[0117]

[0118] in represents the medical knowledge enhanced text features, F represents the fusion features of the current reasoning task, is a similarity measurement function used to measure the matching degree between the medical knowledge enhanced text features and the fusion features. This similarity measurement is used to evaluate the correlation between the inference node and the inference task, so as to decide whether to select the inference node for inference. Represents the similarity of the i-th inference node.

[0119] ;

[0120] in represents the computing power of the i-th inference node, represents the resource availability of the i-th inference node, represents the priority of the i-th inference node, λ1, λ2, λ3 are weight coefficients that respectively represent the influence of computing power, resource availability and similarity on the priority. By combining computing power, resource availability and similarity to comprehensively evaluate each inference node, reliable guidance is provided for the selection of inference nodes.

[0121] ;

[0122] in is the historical quality value of the i-th inference node, ω is the weight coefficient, which controls the relative importance of the priority and historical quality value in the comprehensive score, Indicates the overall rating.

[0123] It should be noted that each inference node is assigned a historical quality value, which is used to reflect the performance of the inference node in historical tasks. When the dynamic inference tree is first deployed or for a newly added inference node, the quality value Qi of the inference node is initialized to 0 by default:

[0124] .

[0125] The historical quality value Qi of the inference node is one of the important factors used to guide the selection of the inference path. It reflects the performance of the inference node in historical tasks. During the optimization process of the dynamic inference tree, the system will combine the historical quality value Qi to evaluate the performance of each inference node, and give priority to those inference nodes with excellent performance when selecting paths. As the system runs, the historical quality value Qi of the inference node will be continuously updated through the reinforcement learning strategy to ensure that the inference path selection can adapt to the changing scenarios and user needs, and gradually optimize the inference efficiency and accuracy of the overall system.

[0126] After obtaining the comprehensive score of each inference node on the dynamic inference tree, this scheme can select the inference node with the highest comprehensive score, and pass the fusion features to the inference model of the inference node for reasoning. The specific inference node can be health assessment or disease prediction. The reasoning task of the specific inference model is adjusted according to the actual situation.

[0127] In order to realize the feedback of the reasoning node, the multimodal intelligent medical scene dynamic reasoning method based on medical knowledge provided by this solution further includes the following steps:

[0128] A reward function is set based on the inference accuracy, response time, and user satisfaction of the highest-scoring inference node, and the historical quality value of the highest-scoring inference node is updated through feedback based on the reward function.

[0129] Specifically, the reward function is set as:

[0130] ;

[0131] in For reasoning accuracy, it means rewarding the accuracy of the reasoning results; is the response time, the shorter the response time, the higher the reward; For user satisfaction, the inference results are rewarded or punished according to user feedback, where , as well as They represent the weight coefficients of inference accuracy, response time, and user satisfaction on the reward value, respectively. is the reward value.

[0132] The way to feedback and update the historical quality value of the highest-scoring inference node based on the reward function is as follows:

[0133] ;

[0134] Where α is the learning rate, which controls the impact of new rewards on historical quality values.

[0135] In addition, this scheme updates the computing power and resource availability of the inference node according to the degree of completion of the inference task to reflect the resource changes of the inference node after executing the inference task. If the load of the inference node increases, the computing power is reduced, and the current available resources are updated according to the resource consumption of the inference node in the task.

[0136] In order to better adapt to the ever-changing medical tasks, the system will also dynamically optimize the inference tree structure to ensure that the selection of the inference path can match the current medical scenario and task type. Correspondingly, the multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge provided by this solution further includes the following steps:

[0137] Use graph neural networks to analyze the overall structure of the dynamic reasoning tree and the connection status of each reasoning node. If the historical quality value of an reasoning node is low for a long time, it will be removed. If the frequency of a certain type of reasoning task increases, the corresponding reasoning node will be added to the dynamic reasoning tree.

[0138] This solution can also regularly enrich and update the medical knowledge graph, using the rich relationships among diseases, symptoms, examinations and treatments in the medical knowledge graph, combined with the graph neural network (GNN) to optimize the dynamic reasoning tree, ensuring that the reasoning tree structure is consistent with the medical field knowledge. This knowledge graph-based optimization enables the dynamic reasoning tree to better adapt to the knowledge relevance in medical scenarios, thereby improving the rationality and effectiveness of the reasoning path.

[0139] Embodiment 2

[0140] like Figure 2As shown, this embodiment provides a multimodal intelligent medical scene dynamic reasoning device based on medical knowledge, including the following structure:

[0141] A data acquisition unit, used to acquire multimodal data of a user, wherein the multimodal data includes text modal data, voice modal data, and image modal data;

[0142] A modal feature conversion unit, used for inputting multimodal data into a trained feature extraction module to extract multimodal features, wherein the multimodal features include text modal features, speech modal features, and image modal features;

[0143] A context-enhanced text feature acquisition unit is used to input text modal features into a graph neural network preset with a medical knowledge graph to perform feature enhancement to obtain medical knowledge enhanced text features, input text modal features into a memory network stored in a user's personalized health record to perform feature enhancement to obtain record enhanced text features, and concatenate the medical knowledge enhanced text features and the record enhanced text features to obtain context-enhanced text features;

[0144] A fusion feature acquisition unit, used for inputting context-enhanced text features and multimodal features into a pre-trained multimodal fusion model to output fusion features;

[0145] An inference node selection unit, used for inputting the fusion feature into a dynamic inference tree constructed by a plurality of inference nodes in a graph structure to determine the optimal inference node, wherein the approximation between the medical knowledge enhanced text feature and the fusion feature of each inference node is calculated, and the priority of the current inference node is determined based on the approximation of the current inference node, and the inference node with the highest score is selected in combination with the priority of the current inference node and the historical quality value;

[0146] The reasoning unit uses the highest-scoring reasoning node to reason about the fused features.

[0147] In some embodiments, the multimodal intelligent medical scene dynamic reasoning device based on medical knowledge further includes:

[0148] The feedback unit is used to set a reward function based on the reasoning accuracy, response time and user satisfaction of the reasoning node with the highest score, and to feedback and update the historical quality value of the reasoning node with the highest score based on the reward of the reward function.

[0149] In some embodiments, the multimodal intelligent medical scene dynamic reasoning device based on medical knowledge further includes:

[0150] The optimization unit is used to use graph neural networks to analyze the overall structure of the dynamic reasoning tree and the connection status of each reasoning node. If the historical quality value of a certain reasoning node is low for a long time, it will be removed. If the frequency of a certain type of reasoning task increases, the corresponding reasoning node will be added to the dynamic reasoning tree.

[0151] The same contents as those in the first embodiment will not be described in detail here.

[0152] Embodiment 3

[0153] This embodiment also provides an electronic device, referring to Figure 3 , including a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge.

[0154] Specifically, the processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0155] The memory 404 may include a large-capacity memory 404 for data or instructions. The memory 404 may be used to store or cache various data files that need to be processed and / or used for communication, as well as possible computer program instructions executed by the processor 402.

[0156] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any one of the multimodal intelligent medical scenario dynamic reasoning methods based on medical knowledge in the above embodiments.

[0157] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .

[0158] The transmission device 406 may be used to receive or send data via a network. Specific examples of the above network may include a wired or wireless network provided by a communication provider of the electronic device.

[0159] The input / output device 408 is used to input or output information. In this embodiment, the input information may be multimodal data, etc., and the output information may be inference results, etc.

[0160] Optionally, in this embodiment, the processor 402 may be configured to perform the following steps through a computer program:

[0161] Acquire multimodal data of the user, wherein the multimodal data includes text modal data, voice modal data, and image modal data;

[0162] Inputting the multimodal data into a trained feature extraction module to extract multimodal features, wherein the multimodal features include text modality features, speech modality features, and image modality features;

[0163] Input the text modality features into a graph neural network with a preset medical knowledge graph for feature enhancement to obtain medical knowledge enhanced text features, input the text modality features into a memory network stored in a user's personalized health record for feature enhancement to obtain record enhanced text features, and concatenate the medical knowledge enhanced text features and the record enhanced text features to obtain context enhanced text features;

[0164] Input context-enhanced text features and multimodal features into the pre-trained multimodal fusion model to output fusion features;

[0165] The fused features are input into a dynamic inference tree constructed by a graph structure of multiple inference nodes to determine the optimal inference node, wherein the approximation between the medical knowledge enhanced text features and the fused features of each inference node is calculated, and the priority of the current inference node is determined based on the approximation of the current inference node, and the inference node with the highest score is selected in combination with the priority of the current inference node and the historical quality value;

[0166] The fused features are inferred using the inference node with the highest score.

[0167] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0168] In general, various embodiments may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the boxes, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0169] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware.

[0170] Those skilled in the art should understand that the technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge, characterized in that: The following steps are involved: Acquire multimodal data of the user, wherein the multimodal data includes text modal data, voice modal data, and image modal data; Inputting the multimodal data into a trained feature extraction module to extract multimodal features, wherein the multimodal features include text modality features, speech modality features, and image modality features; Input the text modality features into a graph neural network with a preset medical knowledge graph for feature enhancement to obtain medical knowledge enhanced text features, input the text modality features into a memory network stored in a user's personalized health record for feature enhancement to obtain record enhanced text features, and concatenate the medical knowledge enhanced text features and the record enhanced text features to obtain context enhanced text features; Inputting the context-enhanced text features and the multimodal features into a pre-trained multimodal fusion model to output fusion features, wherein the multimodal fusion model includes a dynamic weight allocation module for generating dynamic weights, an attention weight allocation module for generating attention weights, and a fusion module that are sequentially connected, inputting the context-enhanced text features and the multimodal features into the dynamic weight allocation module to generate respective dynamic weights, inputting the context-enhanced text features and the multimodal features into the attention weight allocation module to generate respective attention weights, and fusing the context-enhanced text features and the multimodal features based on the attention weights and the dynamic weights to obtain fusion features; The fused features are input into a dynamic inference tree constructed by a plurality of inference nodes in a graph structure to determine the inference node with the highest score, wherein the approximation between the medical knowledge enhanced text features of each inference node and the fused features is calculated, and the priority of the current inference node is determined based on the approximation of the current inference node, and the inference node with the highest score is selected in combination with the priority of the current inference node and the historical quality value; The fused features are inferred using the inference node with the highest score.

2. The multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge according to claim 1 is characterized in that: The medical knowledge graph is input into the graph neural network and converted into a knowledge graph vector. The text modal features are input into the graph neural network and the graph attention operation is performed with the knowledge graph vector to obtain medical knowledge enhanced text features, where medical knowledge is used as the nodes of the medical knowledge graph and the relationships between medical knowledge are used as the edges between the nodes.

3. The multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge according to claim 1 is characterized in that: Calculate the cosine similarity of each multimodal feature based on the context-enhanced text feature, and use the normalized value of the cosine similarity of each multimodal feature as the dynamic weight of each multimodal feature; The context-enhanced text features and multimodal features are mapped to the same spatial dimension and then merged to obtain the multimodal feature matrix. The attention weight of each multimodal feature is calculated through the attention mechanism.

4. The multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge according to claim 1 is characterized in that: The dynamic weight of the context-enhanced text feature is 1. The product of the dynamic weight of each context-enhanced text feature or multimodal feature and the attention weight is taken as the weighting coefficient. The context-enhanced text feature and the multimodal feature are weightedly fused with their respective weighting coefficients to obtain the fused feature.

5. The multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge according to claim 1 is characterized in that: The dynamic inference tree includes multiple inference nodes, which are organized into a multi-level structure. According to the type of inference node, the inference nodes can be divided into: edge nodes, cloud nodes and adaptive nodes. Edge nodes are responsible for simple computing tasks with low latency, cloud nodes are responsible for complex and large data volume inference tasks, and adaptive nodes switch between edge nodes and cloud nodes.

6. The multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge according to claim 1 is characterized in that: A reward function is set based on the inference accuracy, response time, and user satisfaction of the highest-scoring inference node, and the historical quality value of the highest-scoring inference node is updated through feedback based on the reward function.

7. The multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge according to claim 1 is characterized in that: Use graph neural networks to analyze the overall structure of the dynamic reasoning tree and the connection status of each reasoning node. If the historical quality value of an reasoning node is low for a long time, it will be removed. If the frequency of a certain type of reasoning task increases, the corresponding reasoning node will be added to the dynamic reasoning tree.

8. A multimodal intelligent medical scene dynamic reasoning device based on medical knowledge, characterized in that: Includes the following: A data acquisition unit, used to acquire multimodal data of a user, wherein the multimodal data includes text modal data, voice modal data, and image modal data; A modal feature conversion unit, used for inputting multimodal data into a trained feature extraction module to extract multimodal features, wherein the multimodal features include text modal features, speech modal features, and image modal features; A context-enhanced text feature acquisition unit is used to input text modal features into a graph neural network preset with a medical knowledge graph to perform feature enhancement to obtain medical knowledge enhanced text features, input text modal features into a memory network stored in a user's personalized health record to perform feature enhancement to obtain record enhanced text features, and concatenate the medical knowledge enhanced text features and the record enhanced text features to obtain context-enhanced text features; A fusion feature acquisition unit, used for inputting context-enhanced text features and multimodal features into a pre-trained multimodal fusion model to output fusion features, wherein the multimodal fusion model includes a dynamic weight allocation module for generating dynamic weights, an attention weight allocation module for generating attention weights, and a fusion module connected in sequence, inputting context-enhanced text features and multimodal features into the dynamic weight allocation module to generate respective dynamic weights, inputting context-enhanced text features and multimodal features into the attention weight allocation module to generate respective attention weights, and fusing context-enhanced text features and multimodal features based on the attention weights and dynamic weights to obtain fusion features; An inference node selection unit, used for inputting the fusion feature into a dynamic inference tree constructed by a plurality of inference nodes in a graph structure to determine the inference node with the highest score, wherein the approximation between the medical knowledge enhanced text feature of each inference node and the fusion feature is calculated, and the priority of the current inference node is determined based on the approximation of the current inference node, and the inference node with the highest score is selected in combination with the priority of the current inference node and the historical quality value; The reasoning unit uses the highest-scoring reasoning node to reason about the fused features.

9. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to run the computer program to execute the multimodal intelligent medical scenario dynamic reasoning method based on medical knowledge as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Knowledge graph reasoning method and system, electronic equipment and medium

    CN118780361A

  • Multi-source knowledge graph construction method and system for chronic disease diagnosis and treatment

    CN118820486A