A knowledge guide-based medical large model fine-tuning optimization method and related device
By constructing a multimodal-multi-level connection graph and using an expert consensus-driven approach, the diagnostic limitations and uninterpretability of large medical models under single-modal data were addressed, enabling efficient, accurate, and interpretable diagnostic results from large medical models in the diagnosis of neuropsychiatric diseases.
Patent Information
- Application Number
- CN202411951055.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing large medical models rely on single-modal data when diagnosing complex diseases, failing to fully integrate multimodal data. This results in incomplete and inaccurate diagnostic results, and the lack of clear logical reasoning paths makes the models uninterpretable and the diagnostic results unreliable.
A knowledge-guided approach is adopted, which constructs a multimodal-multi-level connection graph through a multimodal-multi-level connection graph fusion module, an expert consensus-driven module, an analogy deduction module, and an explicit inference rule calculation module. Combined with expert knowledge and logical chain hints, explicit inference rules are generated to enhance the interpretability and diagnostic accuracy of the model.
It enables the complementary use of multimodal data, improves the accuracy and interpretability of medical large models in the diagnosis of neuropsychiatric diseases, ensures the logic and credibility of model output, and improves the accuracy of early diagnosis.
Smart Images

Figure CN119830219B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to a medical large model fine-tuning optimization method, and particularly relates to a knowledge-guided medical large model fine-tuning optimization method and related device. BACKGROUND
[0002] With the acceleration of global population aging, the number of patients with neuropsychiatric diseases has increased significantly, bringing great pressure to society and the medical system. Diseases such as Alzheimer's disease, depression, and anxiety are representative, and more and more elderly and neuropsychiatric patients have an urgent need for efficient and precise medical services.
[0003] Under this background, artificial intelligence technology, especially medical large models, has emerged as a powerful tool and solution for the medical field. Medical large models can integrate multi-modal data such as MRI, PET, genomic data, and electronic medical records to support early diagnosis and precise treatment of complex diseases. However, most existing medical large models mainly have three problems: (1) only using single-modal data, failing to fully integrate the advantages of multi-modal data, resulting in incomplete diagnosis results or insufficient diagnostic accuracy. (2) Since medical large models rely on complex deep learning architectures, they cannot provide clear logical reasoning paths, making it difficult for doctors and patients to trust the conclusions. (3) Due to the large differences between individual patients, the accuracy of existing medical large models in early diagnosis is still not ideal. SUMMARY
[0004] The present application provides a knowledge-guided medical large model fine-tuning optimization method and related device to address the technical problems of incomplete diagnosis results, low diagnostic accuracy, and unclear logical reasoning paths in current medical large models.
[0005] To achieve the above purpose, the following technical solutions are adopted in the present application:
[0006] In a first aspect, the present application provides a knowledge-guided medical large model fine-tuning optimization method, comprising:
[0007] Obtaining multi-modal data to be diagnosed;
[0008] Inputting the multi-modal data to be diagnosed into a knowledge-guided model to obtain a thinking guidance information part in the logical chain prompt, which is used to input into a medical large model to obtain a normalized output after chain prompt;
[0009] The knowledge-guided model includes a multi-modal-multi-level connection graph fusion module, an expert consensus driven module, an analogical deduction module, an explicit derivation rule calculation module, and a progressive thinking guidance module.
[0010] The multi-modal-multi-level connection graph fusion module is configured to obtain overall embedding features as states according to the modality data graphs of the modality data of the disease to be diagnosed.
[0011] The expert consensus driving module is configured to extract expert consensus and clinical experience from an expert knowledge base of the disease to be diagnosed, combine the state semantic labels corresponding to the overall embedding features, determine state relationship labels between the overall embedding features, extract relationship embedding feature representations from the state relationship labels, and combine the overall embedding features, the state semantic labels, the state relationship labels, and the relationship embedding feature representations to construct a state level graph, and then combine the modality data graphs of the modality data to obtain a global multi-modal-multi-level connection graph.
[0012] The analogy deduction module is configured to construct an individual multi-modal-multi-level connection graph based on analogy deduction of a plurality of triplets according to the global multi-modal-multi-level connection graph, wherein the triplets include two overall embedding features and corresponding embedding feature representations.
[0013] The explicit derivation rule calculation module is configured to obtain complete reasoning rules of all nodes according to the nodes of the individual multi-modal-multi-level connection graph.
[0014] The progressive thinking guidance module is configured to combine the modality categories in the global multi-modal-multi-level connection graph and the complete reasoning rules of all nodes to obtain a thinking guidance information part in a logic chain prompt.
[0015] Further, obtaining overall embedding features as states according to the modality data graphs of the modality data of the disease to be diagnosed includes:
[0016] The feature aggregation layer based on the graph neural network aggregates the neighborhood information of each node in the modality data graph of the modality data of the disease to be diagnosed by using an aggregation function Aggregate The neighborhood information of each node in the modality data graph of the modality data of the disease to be diagnosed is aggregated to form a global representation of the node, and the overall embedding features of the modality data graph of the modality data of the disease to be diagnosed are extracted.
[0017] Further, the analogy deduction based on a plurality of triplets to construct an individual multi-modal-multi-level connection graph includes:
[0018] According to the global multi-modal-multi-level connection graph, an example triplet is obtained.
[0019] The to-be-predicted triplet and the example triplet are embedded and represented by using an analogy feature embedding layer to learn feature mapping between modalities by using a cross-modal adaptive interaction layer, and the embedding representations corresponding to each modality are obtained.
[0020] The embedding representations corresponding to each modality are integrated by a multi-modal information fusion layer, and then the predicted triplets after prediction and the feature representations of the predicted triplets after prediction are obtained through a reasoning consistency normalization layer and a prediction output layer.
[0021] According to the predicted triplets after prediction and the feature representations of the predicted triplets after prediction, an individual multi-modal-multi-layer connection graph is constructed.
[0022] Further, the complete reasoning rules of all nodes are obtained according to each node of the individual multi-modal-multi-layer connection graph, including:
[0023] The nodes of the individual multi-modal-multi-layer connection graph are sequentially input into a strategy function of reasoning rule calculation;
[0024] The strategy function includes a state coding layer, a plurality of feature extraction layers and a strategy selection layer connected in sequence, and an activation layer is arranged before the state coding layer, the plurality of feature extraction layers and the strategy selection layer, and the output of the strategy selection layer further includes a normalization processing.
[0025] Further, the thinking guidance information part in the logic chain prompt is obtained by combining the modality types in the complete reasoning rules of all nodes and the global multi-modal-multi-layer connection graph, including:
[0026] Representative data of different modalities is extracted from the state level graph of the global multi-modal-multi-layer connection graph, and a question and answer example library with progressive reasoning is constructed for different modal data combinations; then, according to the modality types in the complete reasoning rules of all nodes, corresponding question and answer examples are extracted from the question and answer example library as the thinking guidance information part in the logic chain prompt.
[0027] Further, the loss function used by the analogy and deduction module during training is:
[0028]
[0029] Wherein, is a relaxation loss function, is the total number of triplets in the training set, is a sine similarity function, is the hidden feature representation of the example triplet, is the hidden feature representation of the predicted triplet, is a maximum calculation formula, and are the hidden feature representations of the head and tail entities in the predicted pair, respectively.
[0030] Further, the reward function used by the explicit derivation rule calculation module during training is:
[0031]
[0032]
[0033]
[0034]
[0035] wherein, is a total reward function, is a global goal reward, is a path efficiency reward, is a path diversity reward, is the last node of the individual multi-modal multi-hierarchical connection graph, is a path length, is a set of explored paths, is a cosine similarity function, is a coefficient of is a coefficient of is a coefficient of is a coefficient of is a coefficient of is a coefficient of
[0036] The explicit derivation rule calculation module adopts a Monte Carlo policy gradient method to update the policy function parameters during training.
[0037] Further, the loss function used by the knowledge guidance model during training is:
[0038]
[0039] wherein, is an overall loss function during training of the knowledge guidance model, is a loss function involved in updating the policy function parameters using the Monte Carlo policy gradient method, denoted as a path optimization loss, is a reasoning process of the thinking guidance information part in the logical chain prompt and a label calculation semantic loss function:
[0040]
[0041] wherein, is a reasoning process label, is a semantic embedding representation, is a semantic embedding representation, is a semantic embedding representation, is a semantic embedding representation, is a reasoning process of the thinking guidance information part in the logical chain prompt.
[0042] In a second aspect, the application provides a medical large model fine-tuning optimization system based on knowledge guidance, comprising:
[0043] A data acquisition module is configured to acquire multi-modal data to be diagnosed.
[0044] A prompt output module is configured to input the multi-modal data to be diagnosed into a knowledge guidance model to obtain a thinking guidance information part in a logic chain prompt, and input the thinking guidance information part into a medical large model to obtain an output normalized by the chain prompt.
[0045] The knowledge guidance model comprises a multi-modal-multi-layer connection graph fusion module, an expert consensus driving module, an analogical deduction module, an explicit derivation rule calculation module, and a progressive thinking guidance module.
[0046] The multi-modal-multi-layer connection graph fusion module is configured to obtain overall embedding features as states according to a multi-modal data graph of each modality data of a disease to be diagnosed.
[0047] The expert consensus driving module is configured to extract expert consensus and clinical experience from an expert knowledge base of the disease to be diagnosed, determine state relationship labels between the overall embedding features in combination with state semantic labels corresponding to the overall embedding features, extract relationship embedding feature representations from the state relationship labels, and construct a state level graph in combination with the overall embedding features, the state semantic labels, the state relationship labels, and the relationship embedding feature representations, and obtain a global multi-modal-multi-layer connection graph in combination with a multi-modal data graph of the modality data.
[0048] The analogical deduction module is configured to construct an individual multi-modal-multi-layer connection graph based on analogical deduction of a plurality of triplets according to the global multi-modal-multi-layer connection graph; the triplets comprise two overall embedding features and corresponding embedding feature representations.
[0049] The explicit derivation rule calculation module is configured to obtain complete reasoning rules of all nodes according to each node of the individual multi-modal-multi-layer connection graph.
[0050] The progressive thinking guidance module is configured to obtain the thinking guidance information part in the logic chain prompt in combination with the global multi-modal-multi-layer connection graph and modality categories in the complete reasoning rules of all nodes.
[0051] In a third aspect, the application provides a computer program product comprising a computer program, which, when executed by a processor, implements the medical large model fine-tuning optimization system based on knowledge guidance.
[0052] Compared with the prior art, the application has the following beneficial effects:
[0053] The application provides a medical large model fine-tuning optimization method based on knowledge guidance. The thinking guidance information part in the logical chain prompt is obtained by means of a knowledge guidance model, is input into a medical large model, and is used to obtain a normalized output after chain prompting, so that the medical large model is fine-tuned and optimized. The knowledge guidance module comprises a multi-modal-multi-level connection graph fusion module, an expert consensus driving module, an analogical deduction module, an explicit derivation rule calculation module and a progressive thinking guidance module. The multi-modal-multi-level connection graph fusion module converts medical data of different modalities into graph structure data, extracts node features of different modalities of data, and realizes the complementation of multi-modal information. The expert consensus driving module extracts medical consensus from an expert knowledge base, generates semantic labels of different modalities and relationship labels therebetween. The analogical deduction module realizes feature mapping and fusion between different modalities, supports analogical deduction of multi-modal data, improves reasoning ability and individualized diagnosis accuracy. The explicit derivation rule calculation module extracts the explicit derivation rule of a disease, provides the logical reasoning rule part of the logical chain prompt, provides explicit logical support for the reasoning of the medical large model, and enhances the explainability of the medical large model. The progressive thinking guidance module can realize sparse sample learning of the medical large model by using the chain prompting technology, guide the medical large model to reason step by step, and finally improve the diagnosis accuracy of the medical large model. BRIEF DESCRIPTION OF DRAWINGS
[0054] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.
[0055] Figure 1 It is a flowchart of the medical large model fine-tuning optimization method based on knowledge guidance of the application;
[0056] Figure 2 It is a second flowchart of the medical large model fine-tuning optimization method based on knowledge guidance of the application
[0057] Figure 3 It is a principle diagram corresponding to the second flowchart of the medical large model fine-tuning optimization method based on knowledge guidance of the application;
[0058] Figure 4 It is a principle diagram of the analogical deduction module based on multi-modal-multi-level connection graph in the embodiments of the application;
[0059] Figure 5 It is a principle diagram of the explicit derivation rule calculation module based on multi-modal-multi-level connection graph in the embodiments of the application;
[0060] Figure 6 An illustrative diagram of a knowledge-guided medical large model fine-tuning optimization system. DETAILED DESCRIPTION
[0061] To make the purposes, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.
[0062] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by a person of ordinary skill in the art based on the embodiments in the present application without creative labor are within the scope of protection of the present application.
[0063] It should be noted that: similar reference numbers and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0064] With the acceleration of global population aging, the number of patients with neuropsychiatric diseases has increased significantly, bringing great pressure to society and the medical system. Diseases such as Alzheimer's disease, depression, and anxiety are representative, and more and more elderly and neuropsychiatric patients have an urgent need for efficient and precise medical services. The increase in the number of these patients not only affects people's health, but also exacerbates the tension of social medical resources.
[0065] In this context, the rise of artificial intelligence technology, especially medical large models, provides powerful tools and solutions for the medical field. Medical large models can integrate multi-modal data (such as MRI, PET, genomic data, electronic medical records, etc.) to provide support for early diagnosis and precise treatment of complex diseases. Compared to traditional diagnosis methods based on a single data modality, the fusion of multi-modal data can more comprehensively reflect the complex mechanisms of diseases and improve the accuracy of diagnosis.
[0066] However, medical large models face a major challenge in application, which is the "black box" problem. The decision-making process of medical large models usually relies on complex deep learning architectures and large parameter quantities, based on data-driven implicit learning mode, rather than explicit rules. Although these models perform well in handling complex data and prediction tasks, their decision-making process lacks credibility and cannot provide clear explanations for doctors and patients. This makes the reasoning process of the model difficult to understand, limiting its application in the clinic.
[0067] In the medical field, explainability is crucial for the application of models. Medical large models must be explainable so that doctors can trust the model's diagnostic results and make clinical decisions based on the information provided. Explainability not only helps ensure the safety and reliability of the model's diagnosis, but also helps medical institutions improve patient treatment outcomes through the model's output. Especially in the diagnosis and treatment of neuropsychiatric diseases, due to the complexity of the disease itself and the differences between patients, models lacking explainability may lead to inaccurate diagnoses and even pose medical risks. Existing explainability methods (such as LIME and SHAP) can provide some explanations, but these methods are usually suitable for simpler models. When faced with the complex nonlinear relationships and high-dimensional data of medical large models, they can only provide local or fragmented explanations, making it difficult to deal with the complex logic and large-scale data processing involved in medical large models. Therefore, existing explainability techniques cannot meet the needs of medical large models.
[0068] Therefore, the current medical large models mainly have the following problems:
[0069] Current medical large models rely on single modality data when dealing with complex disease diagnosis. However, single modality data often cannot fully reflect the multi-dimensional characteristics of diseases, resulting in limited diagnostic accuracy. Multi-modal data (such as brain imaging, electroencephalogram, genome, and electronic medical records) can provide complementary information sources, but current medical large models are mostly difficult to effectively integrate these different data sources, resulting in incomplete diagnosis. In addition, the current medical large model structure is highly complex, with a large number of parameters, relying on the processing and transmission of high-dimensional data through nonlinear transformation and multi-layer network. This implicit learning mechanism makes the reasoning and decision-making process of the model difficult to be understood by the outside world, resulting in the "black box" problem, which seriously affects the explainability of medical large models and limits their application in medical clinics. Existing explainability techniques (such as LIME and SHAP) can provide local explanations for relatively simple models, but when dealing with highly complex nonlinear relationships and multi-dimensional data, these techniques can only provide fragmented explanations and are difficult to cover the overall decision logic of the model, resulting in insufficient explanation support when dealing with complex medical large models. Furthermore, although existing medical large models have certain diagnostic capabilities when dealing with complex diseases, due to the lack of deep understanding of the disease domain and effective reasoning mechanisms, their diagnostic accuracy is still not ideal.
[0070] The specific research related mainly includes:
[0071] The Chinese invention patent with publication number CN118802369A proposes a collaborative enhancement method for APT knowledge graph and large language model. Specifically, it discloses a collaborative enhancement method for APT knowledge graph and large language model, including the following steps: S1, constructing an APT knowledge graph focusing on the network security field; S2, designing a chain prompt according to the user's input question and passing it to the large language model for step-by-step answering. The large language model uses the knowledge provided by the APT knowledge graph to query and locate the problem; S3, combining the located APT knowledge nodes with the user's input question content, and using the strongly related nodes of subgraph retrieval to enrich the context awareness prompt; S4, the large language model generates a comprehensive answer for the user based on the context awareness prompt. This patent application enhances the semantic understanding, reasoning, and prediction ability of the large language model in the APT scenario. The large language model can better analyze and infer attacker behavior and attack paths, improve accuracy in complex network environments, and enhance the defense and response capabilities against APT attacks. Furthermore, the Chinese patent application with publication number CN118674056A proposes an intelligent solution method, system, device, and medium for multi-step reasoning problems, specifically designing a multi-step reasoning problem solving framework. It generates explanation steps through a step-by-step approach to inspire model reasoning as a multi-step thinking chain, enhancing multi-step reasoning ability and having strong practical value in realizing strong artificial intelligence. Moreover, the application establishes a knowledge fusion mechanism and a relevance evaluation mechanism in the multi-step reasoning problem solving framework, which helps to improve the logical reasonableness and factual accuracy of the generated reasoning process, thereby improving the accuracy of problem solving. At the same time, the external knowledge introduced by the knowledge fusion mechanism can provide additional information for many question and answer related machine learning tasks, and the relevance evaluation mechanism can assist the education platform in making reasonable evaluations and feedback for the solution process, providing more personalized online learning services based on this.
[0072] However, the current research has the following problems:
[0073] Most existing medical models only use single-modal data (such as brain images, electroencephalogram data, genetic data, or electronic medical records), and fail to fully integrate the advantages of multi-modal data. Single-modal data is difficult to fully capture the multi-dimensional characteristics of complex diseases, especially diseases such as Alzheimer's disease, which often require multi-modal data to fully reveal the biomarkers and development mechanisms of the disease. The use of single-modal data results in incomplete diagnostic results and insufficient diagnostic accuracy. Although existing medical large models perform well in handling complex data, they rely on complex deep learning architectures and cannot provide explicit logical reasoning paths, making their conclusions difficult for doctors and patients to trust. In particular, in the diagnosis of diseases, doctors cannot clearly understand the reasoning basis through the decision-making process of the model, affecting the credibility and application of the diagnostic results. Medical large models fail to provide sufficient explicit reasoning basis when faced with non-linear, multi-dimensional complex data, and the "black box" nature of the model remains a major problem. Due to the complexity of the pathogenesis of neuropsychiatric diseases (such as Alzheimer's disease and depression), and the large differences between individual patients, the accuracy of existing models in early diagnosis is still not ideal. Current methods often rely on limited pathological data and pathological markers, and cannot provide high-accuracy diagnosis in the early stages of the disease. At the same time, existing models fail to fully utilize expert knowledge and medical consensus to enhance diagnostic accuracy, and lack a systematic explicit reasoning mechanism, making the diagnostic results lack logical-based explanations and support.
[0074] To overcome the shortcomings of the prior art, the present application proposes a knowledge-guided medical large model fine-tuning optimization method and related device, which constructs an explainable enhancement method for medical large models based on "multi-modal data + expert knowledge", focusing on improving the diagnostic accuracy and explainability of the model in the disease field. The present application will be described in detail below in conjunction with the embodiments and drawings.
[0075] As shown in Figure 1 , it is a flowchart of the knowledge-guided medical large model fine-tuning optimization method of the present application, which can include:
[0076] S101, acquiring multi-modal data to be diagnosed.
[0077] S102, inputting the multi-modal data to be diagnosed into a knowledge-guided model to obtain a thinking guidance information part in the logical chain prompt, which is used to input into a medical large model to obtain a normalized output after chain prompt.
[0078] The knowledge-guided model includes a multi-modal-multi-level connection graph fusion module, an expert consensus driven module, an analogical deduction module, an explicit derivation rule calculation module, and a progressive thinking guidance module.
[0079] The multi-modal-multi-layer joint graph fusion module is used to obtain the overall embedding features as states according to the modality data graphs of the modality data of the disease to be diagnosed. The multi-modal-multi-layer joint graph fusion module is the basis of the entire optimization method, responsible for integrating multi-modal data from different medical examination methods (such as CT, MRI images, pathological reports, physiological monitoring data, etc.). These data are each represented in the form of a graph, containing rich disease information. The fused graph not only retains the unique information of each modality data, but also captures the correlation and complementarity between them, providing a comprehensive description of the disease state for subsequent analysis. The multi-modal-multi-layer joint graph fusion module integrates multi-modal data sources, thereby providing comprehensive disease characteristics, overcoming the information loss and insufficient diagnostic accuracy problems caused by single modality data, and effectively revealing the complementary information between different modality data through multi-layer joint graph construction, fully capturing the multi-dimensional characteristics of the disease, and solving the limitations of single modality data.
[0080] The expert consensus driven module is used to extract expert consensus and clinical experience from the expert knowledge base of the disease to be diagnosed, combine the state semantic labels corresponding to the overall embedding features, determine the state relationship labels between the overall embedding features, extract the embedding feature representation of the relationship from the state relationship labels, and combine the overall embedding features, state semantic labels, state relationship labels and embedding feature representation of the relationship to construct a state level graph. Then, the modality data graphs of the modality data are combined to obtain a global multi-modal-multi-layer joint graph. The expert consensus driven module can rely on a large expert knowledge base, which contains the consensus, clinical experience and research results of medical experts on various diseases. First, the overall embedding features are extracted from the multi-modal data of the disease to be diagnosed, and then the expert consensus and clinical experience related to the overall embedding features are searched in the knowledge base. Through semantic matching and similarity calculation, accurate state semantic labels and state relationship labels are assigned to these features. These labels not only describe the state of the disease, but also reveal the logical relationship between the states. Finally, the module uses these labels and relationships to construct a state level graph, which reflects the possible paths and stages of disease development. The expert consensus driven module introduces expert knowledge to ensure that the reasoning rules extracted during model reasoning are explicit.
[0081] The analogical deduction module is used to construct an individual multi-modal-multi-layer joint graph based on the analogical deduction of multiple groups of triples according to the global multi-modal-multi-layer joint graph. The triples include two overall embedding features and corresponding embedding feature representations. Analogical deduction is a logical reasoning method that infers new triples based on known triples (i.e., two overall embedding features and their relationships). Through multiple iterations and reasoning, a graph reflecting the individual patient's disease state is gradually constructed. This graph not only contains the patient's specific condition information, but also embodies the individualized characteristics of disease development.
[0082] The explicit derivation rule calculation module is configured to obtain complete reasoning rules of all nodes according to the nodes of the individualized multi-modal-multi-layer connection graph. The explicit derivation rule calculation module generates complete reasoning rules according to the individualized multi-modal-multi-layer connection graph. These rules are expressed in the form of logical expressions and describe the causal relationship and evolution law between disease states. The explicit derivation rule calculation module can extract all logical relationships by traversing the nodes and edges in the graph and convert them into executable reasoning rules. These rules provide clear guidance for subsequent decision-making. The design of the explicit derivation rule calculation module based on the multi-modal-multi-layer connection graph extracts explicit reasoning rules from the connection graph and provides them to the medical large model, thereby enhancing the interpretability of the model and solving the problem of low reliability of diagnostic results caused by the "black box" problem of the medical large model. The analogical deduction module and the explicit derivation rule calculation module based on the multi-modal-multi-layer connection graph extract key explicit reasoning rules in the decision-making process by mining the paths between different data modalities and diagnostic conclusions from the multi-modal-multi-layer connection graph, thereby making up for the shortcomings of existing technologies in explaining complex models.
[0083] The progressive thinking guidance module is configured to obtain the thinking guidance information part in the logic chain prompt by combining the modality types in the global multi-modal-multi-layer connection graph and the complete reasoning rules of all nodes. The progressive thinking guidance module is the final output link of the optimization method. It combines the global multi-modal-multi-layer connection graph and the complete reasoning rules of all nodes to generate the thinking guidance information in the logic chain prompt. This information is presented to doctors or users in a structured manner to guide them to think and make decisions according to the logic chain. The design of the progressive thinking guidance module guides the medical large model to gradually generate diagnostic results, enhances the diagnostic accuracy of the model in the context of neuropsychiatric diseases, and ensures that the output of the model is logical and interpretable. This significantly improves the accuracy of early diagnosis and helps the model better utilize the knowledge it has acquired about diseases, thereby enhancing the precision and reliability of diagnosis.
[0084] The application proposes a logical chain prompt construction method based on "multi-modal data + expert knowledge". By constructing a connection graph and calculating the path of explicit reasoning rules, the prompt of the medical large model is further constructed, and the diagnosis accuracy of the medical large model and the explainability of the diagnosis result are improved. The multi-modal-multi-level connection graph fusion module integrates the graph structure data of different modal medical data (such as brain imaging, electroencephalogram, genome, electronic medical record, etc.), and aggregates and extracts the node features of different modal neuropsychiatric disease data through a graph neural network. The complementary of multi-modal information is realized. The expert consensus driven module extracts medical consensus from the expert knowledge base, generates semantic labels of different modal states and their relationship labels, and uses a natural language processing model to extract semantic embedding features of the relationship, providing semantic support for the connection graph of multi-modal data. The mechanism of combining expert knowledge and multi-modal connection graph, the mechanism of combining modal level graph and state level graph, solves the problem of how the semantic embedding model effectively processes expert knowledge and organically combines it with the feature embedding of multi-modal data, ensuring that the multi-dimensional features of different modalities can be integrated into a unified multi-modal-multi-level connection graph. The analogical reasoning module based on multi-modal-multi-level connection graph realizes the feature mapping and fusion between different modalities through the design of a cross-modal adaptive interaction layer, supports the analogical reasoning of multi-modal data, and can predict the relationship between states and further construct individualized multi-modal-multi-level connection graphs, improving the reasoning ability and individualized diagnosis accuracy. The explicit derivation rule calculation module based on multi-modal-multi-level connection graph extracts the explicit derivation rules of neuropsychiatric diseases from the state level graph based on deep reinforcement learning technology, provides the logical chain prompt logical reasoning rule part, and provides explicit logical support for the reasoning of the medical large model, enhancing the explainability of the large model.
[0085] Existing medical large models often rely on single-modal data, leading to incomplete understanding of diseases and insufficient diagnostic accuracy. The present application proposes a multi-modal-multi-layer connection graph fusion module to integrate multiple modal data and extract and aggregate features through graph neural networks. Compared with existing single-modal or simple multi-modal fusion techniques, the present application can more comprehensively reflect the complex mechanisms of neuropsychiatric diseases and improve the accuracy and comprehensiveness of diagnosis. Existing medical large models have a "black box" problem, which cannot clearly explain the decision-making process, making it difficult for the model to gain the trust of doctors in clinical practice. Existing explainability methods have certain applications in simple models, but they can only provide fragmented explanations when dealing with complex nonlinear relationships and high-dimensional data. The present application introduces an expert consensus driven module, a multi-modal-multi-layer connection graph based analogical deduction module, and an explicit derivation rule calculation module, combined with medical consensus in the expert knowledge base, to generate explicit reasoning rules between states and build explicit reasoning paths, significantly improving the explainability of the model, so that doctors can understand and trust the diagnostic results of the model. Most existing medical models use standard training methods and lack step-by-step reasoning ability for complex tasks, especially in the field of neuropsychiatric diseases, where data is scarce and models are difficult to learn effectively through traditional methods. The present application proposes a progressive thinking guidance module that combines the professional knowledge of clinical doctors to build a question and answer example library with progressive reasoning, and guides the medical large model to learn from sparse samples through chain prompts. This technology enables the model to reason step by step with a small amount of training data, improving its learning efficiency and diagnostic ability in neuropsychiatric disease diagnosis.
[0086] As shown in Figure 2 , it is a second process schematic diagram of the knowledge-guided medical large model fine-tuning optimization method of the present application. As shown in Figure 3 , it is a principle schematic diagram corresponding to the second process schematic diagram of the knowledge-guided medical large model fine-tuning optimization method of the present application. Specifically, it can include:
[0087] S201, model construction
[0088] The model constructed in the present application mainly includes five modules: multi-modal-multi-layer connection graph fusion module, expert consensus driven module, multi-modal-multi-layer connection graph based analogical deduction module, multi-modal-multi-layer connection graph based explicit derivation rule calculation module, and progressive thinking guidance module. The specific construction method of each model is as follows:
[0089] (1) Constructing a multi-modal-multi-layer connection graph fusion module
[0090] Step a1: input the modality data graph of a certain modality as input into the input layer of the multi-modal-multi-layer connection graph fusion module.
[0091] It should be noted that the modality in this application specifically refers to different medical data, such as brain images, electroencephalogram data, electronic medical records, etc., which are all modal data. The modal data can be represented as a graph structure, and the graph structure formed by nodes and edges forms a modal data graph.
[0092] Step a2: The feature aggregation layer based on the graph neural network aggregates the neighborhood information of the nodes through an aggregation function to form a global representation of the node, and then extracts the overall embedding features of the modal data graph (hereinafter also referred to as "state"):
[0093] The graph neural network is a deep learning method specially designed for graph data, which can perform end-to-end learning and inference on graph structure data. The core components of the graph neural network include graph convolution layer, message passing mechanism and multi-layer perceptron. The graph convolution layer is responsible for updating the node features in the graph, which updates the features of the center node by aggregating the features of the neighboring nodes, and applies a linear transformation and a nonlinear activation function to update the features of the center node. The message passing mechanism allows nodes to pass and aggregate information between nodes, thereby capturing global graph structure information. The multi-layer perceptron is used to further process and transform the node features to improve the model's expression ability.
[0094] Step a3: The modal data graph in step a1 is the modal level graph, and the overall embedding features obtained by step a2 are the nodes of the state level graph. And as the output of this modality, together with the output of the expert consensus driven module, constitute the multi-modal-multi-level connection graph.
[0095] (2) Constructing an expert consensus driven module
[0096] Step b1: Taking Alzheimer's disease as an example, extract expert consensus and clinical experience from the expert knowledge base in the field of Alzheimer's disease, and according to the state semantic label corresponding to the overall embedding feature obtained in step a2 (for example, in the structural image modality, the overall embedding feature may be "hippocampal atrophy", and the corresponding state semantic label is "hippocampal atrophy"), find the relationship label between states , the relationship label can be "pathological changes caused by", "symptoms associated with pathology", etc. The form of expert knowledge extraction is a textual description.
[0097] Step b2: using a semantic extraction model trained based on a large-scale medical text dataset , with the semantic features of the relationship label as the embedding features of the relationship, denoted as:
[0098]
[0099] wherein, represents the embedding features of the relationship, i.e., the state relationship.
[0100] Step b3: combining the overall embedding features in step a3 , the state semantic label , the state relationship label , and the state relationship obtained in step b2 to construct a state hierarchical atlas . Finally, a global multi-modal-multilevel connection atlas is obtained.
[0101] (3) Constructing an analogy and deduction module based on the multi-modal-multilevel connection atlas.
[0102] As shown in Figure 4 , it is a schematic diagram of the analogy and deduction module based on the multi-modal-multilevel connection atlas.
[0103] Step c1: when training the analogy and deduction module based on the multi-modal-multilevel connection atlas, the input is the training triple , and the relationship between them, as well as the example triple composed of state pairs , and their relationship . When actually applied to prediction, the input is the example triple , and the prediction triple composed of state pairs , and their to-be-predicted relationship (masked as [MASK]). .
[0104] Step c2: taking the prediction process as an example, this step is explained. The input training triple and example triple are embedded and represented by analogy feature embedding layer, and the feature mapping between modalities is learned by cross-modal adaptive interaction layer. The embedding representation , , , is obtained.
[0105] Step c3: After generating the node embedding representation, the integration of multi-modal data is performed through a multi-modal information fusion layer. Then, through the reasoning consistency normalization layer and the prediction output layer, the predicted triplets (predicted prediction triplets) and the feature representation of the predicted triplets are obtained 、 .
[0106] Step c4: According to the obtained predicted triplets , combined with the feature representation 、 , the individual multi-modal-multilayer connection graph is constructed.
[0107] (4) Constructing an explicit derivation rule calculation module based on the multi-modal-multilayer connection graph.
[0108] As shown in Figure 5 , it is a principle diagram of the explicit derivation rule calculation module based on the multi-modal-multilayer connection graph.
[0109] Step d1: input the nodes of the individual multi-modal-multilayer connection graph into the model.
[0110] Step d2: the strategy function of the reasoning rule calculation is , which represents the next step relationship selection under the current state . The in the function is a learnable parameter. The strategy function is composed of a state coding layer, several feature extraction layers, and a strategy selection layer. An activation layer is added after each layer, and the output of the strategy selection layer is normalized. Step d3: the model selects the next node of the reasoning rule path starting from . Repeat steps d1 and d2 until the output node of step d2 is , that is, the complete reasoning rule from to is obtained.
[0111] (5) Constructing a progressive thinking guidance module
[0112] Step e1: extract representative data of different modalities from the state level graph of the global multi-modal-multilayer connection graph, and construct a question and answer example library with progressive reasoning under the guidance of clinical doctors for different modal data groups. Each example should contain the following contents: simulation input part and simulation output part. The simulation input part includes: simulation instruction requirement, simulation patient information description, simulation strong logical evidence chain, and simulation weak logical evidence chain. The simulation output part includes: simulation detailed diagnosis report and individualized intervention strategy, simulation reasoning process, and simulation thinking logic tree.
[0113] Step e2: According to the modal category in the reasoning rule obtained from step d3, extract the corresponding question and answer examples from the question and answer example library with progressive reasoning constructed in step e1 as the thinking guidance information part in the logical chain prompt.
[0114] S202, model training
[0115] (1) Construct a global multi-modal-multi-level connection graph
[0116] Step f1: Collect multi-modal data of a number of Alzheimer's disease patients, and divide them into training set, validation set and test set according to the proportion. And make basic training parameter settings, including training times, etc.
[0117] Step f2: Preprocess the multi-modal data in the training set in step f1 to obtain the graph structure data corresponding to each modality data of each individual in the training set , which can be performed according to the content in the multi-modal-multi-level connection graph fusion module in step S201 to obtain corresponding state .
[0118] Step f3: Perform the steps in the expert consensus driven module in step S201 to obtain a global multi-modal-multi-level connection graph composed of training set data .
[0119] (2) Construct individual multi-modal-multi-level connection graph
[0120] Step g1: The graph structure data of each modality of the individual obtained in step f2 in the training set corresponding state , constitute a number of triples as prediction triples. Extract triples from the global multi-modal-multi-level connection graph obtained in step f3 as training triples.
[0121] Step g2: input and into the analogy deduction module based on multi-modal-multi-level connection graph, execute the specific process in the analogy deduction module based on multi-modal-multi-level connection graph in step S201, obtain prediction triples . Combine all triples obtained after reasoning the state of each modality data of the individual in step f2 , to obtain the individual multi-modal-multi-level connection graph .
[0122] Step g3: construct a relaxation loss, which is minimized by example Predicting pairs of head and tail entities with hidden feature representations while maximizing the hidden feature representations of head and tail entities in the predicted pairs Implementation:
[0123]
[0124] (3) Mining explicit inference rules
[0125] Step h1: The individual multi-modal multi-level joint graph obtained in step g2 The features of the disease detection standard modality in the graph, such as "scale shows cognitive impairment", are used as terminal nodes , and other nodes are used as starting nodes .
[0126] Step h2: input the explicit deduction rule calculation module based on the multi-modal multi-level joint graph During the training of the strategy function, three reward functions are defined as follows:
[0127]
[0128]
[0129]
[0130] Among them, the path efficiency reward improves the inference efficiency by limiting the path length of the environment interaction of reinforcement learning, represents the path length. The path diversity reward uses negative cosine similarity to improve the diversity of the paths searched by the model.
[0131] The total reward function is . The Monte Carlo policy gradient is used to update the strategy function parameters, as follows:
[0132]
[0133] The loss function involved in this step is represented by , which can be represented as:
[0134] .
[0135] Step h3: Repeat step d3 to obtain complete inference rules from to
[0136] Step h4: Repeat step h3 to obtain all starting nodes to terminal nodes The reasoning path, the strong logical evidence chain and the weak logical evidence chain as part of the explicit reasoning rule.
[0137] (4) Construction of logical chain prompt
[0138] Step i1: execute the steps of the progressive thinking guidance module in step S201 to obtain progressive thinking guidance information.
[0139] Step i2: construct a patient information description according to the medical record of the patient (individual) in step g2. The instruction requirements should include detailed requirements for diagnostic reports and personalized intervention strategies in clinical applications.
[0140] Step i3: the instruction requirements in step i2, the patient information description, the explicit reasoning rule obtained in step h4, and the progressive thinking guidance information obtained in step i1 are collectively constructed into a logical chain prompt.
[0141] (5) Model optimization
[0142] Step j1: select a medical large model, execute steps (1) to (4) in step S202 to obtain a logical chain prompt. In the above steps, the dimensions of all latent space variables are aligned with the semantic space of the large model.
[0143] Step j2: input the logical chain prompt into the medical large model to obtain the output normalized by the chain prompt.
[0144] Step j3: the state of the reasoning rule part of the logical chain prompt and relationship are replaced by their corresponding text labels and , which are used as the overall reasoning process label . The reasoning process of the output part in step j2 is calculated with the semantic loss score as:
[0145]
[0146] Wherein, respectively represent each embedding representation produced after the semantic encoding model.
[0147] Step j4: the loss function of the model as a whole is:
[0148]
[0149] Update all modules and latent space parameters according to the overall loss function.
[0150] Step j5: Train the model with the training set according to the number of training times in step f1, and validate it with the validation set in each iteration process. After the iteration is completed, select the best model according to the validation result.
[0151] S203, model testing
[0152] Step k1: According to the test set individual data in step f1, execute the steps in the multi-modal-multi-layer joint graph fusion module in step S201 to obtain the state corresponding to each modality data of the test individual .
[0153] Step k2: Perform steps (2)-(4) in step S202 to obtain the logical chain prompt.
[0154] Step k3: Input the logical chain prompt into the medical large model selected in step j1 to obtain the explainable diagnosis report and personalized intervention strategy of the patient (test individual).
[0155] It should be noted that the foregoing is only an example of Alzheimer's disease for explaining the present application, and the present application has wide applicability and can be extended to other complex disease fields such as oncology and cardiovascular disease. Specifically, in the application process of the model, only the multi-modal data of the corresponding disease (such as tumor images, heart ultrasound, blood biomarkers, etc.) needs to be replaced as input in the construction stage of the multi-modal-multi-layer joint graph fusion module, so that it can be applied to diagnosis and treatment decision-making in other medical fields. By inputting multi-modal data of different diseases into this method, the present application can construct personalized multi-modal joint graph for specific pathological mechanisms of various complex diseases and generate corresponding explicit reasoning rules and diagnostic prompts, thereby providing effective support for precise diagnosis and treatment of different diseases.
[0156] As shown in Figure 6 , it is a kind of schematic diagram of the knowledge-guided medical large model fine-tuning optimization system of the present application, which can include:
[0157] A data acquisition module for acquiring multi-modal data to be diagnosed;
[0158] A prompt output module for inputting the multi-modal data to be diagnosed into the knowledge-guided model to obtain the thinking guidance information part in the logical chain prompt, which is used as input into the medical large model to obtain the output standardized by the chain prompt;
[0159] Among them, the knowledge-guided model includes a multi-modal-multi-layer joint graph fusion module, an expert consensus driven module, an analog reasoning module, an explicit derivation rule calculation module and a progressive thinking guidance module.
[0160] The multi-modal-multi-level connection graph fusion module is configured to obtain the overall embedding features as states according to the modality data graphs of the modal data of the disease to be diagnosed.
[0161] The expert consensus driving module is configured to extract expert consensus and clinical experience from an expert knowledge base of the disease to be diagnosed, determine state relationship labels between the overall embedding features in combination with state semantic labels corresponding to the overall embedding features, extract relationship embedding feature representations from the state relationship labels, and construct a state level graph in combination with the overall embedding features, the state semantic labels, the state relationship labels, and the relationship embedding feature representations, and obtain a global multi-modal-multi-level connection graph in combination with the modality data graphs of the modal data.
[0162] The analogy deduction module is configured to construct an individual multi-modal-multi-level connection graph based on analogy deduction of a plurality of triples according to the global multi-modal-multi-level connection graph, wherein the triples include two overall embedding features and corresponding embedding feature representations.
[0163] The explicit derivation rule calculation module is configured to obtain complete reasoning rules of all nodes according to the nodes of the individual multi-modal-multi-level connection graph.
[0164] The progressive thinking guidance module is configured to obtain a thinking guidance information part in the logic chain prompt in combination with the modality categories in the global multi-modal-multi-level connection graph and the complete reasoning rules of all nodes.
[0165] It should be noted that in several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the system embodiments described above are only illustrative, for example, the division of each module is only a logical function division, and actual implementation can have another division mode, for example, a plurality of modules can be combined or integrated into another device, or some features can be ignored or not executed. The modules described as separate components can be or can not be physically separated, and the components displayed as modules can be one physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs.
[0166] In addition, each module in each embodiment of the present application can be integrated in one processing unit, or each module can exist physically, or two or more modules can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0167] The application further provides a computer program product comprising instructions for implementing the steps of the above-mentioned knowledge-guided medical large model fine-tuning optimization method.
[0168] The computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application.
[0169] The carrier for implementing the above-mentioned computer program product can be a computer device. The computer program product can also be stored in a computer storage medium.
[0170] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server and the like. The computer device can include, but is not limited to, a processor and a memory.
[0171] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0172] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the computer device by running or executing the computer program and / or modules stored in the memory, and calling the data stored in the memory.
[0173] The modules / units integrated in the computer device, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0174] The above only is the preferred embodiment of the present application and is not used to limit the present application. The present application can have various changes and variations for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A knowledge guide-based medical large model fine-tuning optimization method, characterized in that, The method comprises the following steps: Obtaining multi-modal data to be diagnosed; The multi-modal data to be diagnosed is image or text type multi-modal data from different medical examination means; Inputting the multi-modal data to be diagnosed into a knowledge guiding model to obtain a thinking guiding information part in a logic chain prompt, which is used for inputting into a medical large model to obtain an output normalized by chain prompting; The knowledge guiding model comprises a multi-modal-multi-level connection graph fusion module, an expert consensus driven module, an analogy deduction module, an explicit derivation rule calculation module and a progressive thinking guidance module; The multi-modal-multi-level connection graph fusion module is used for obtaining overall embedding features as states according to a modal data graph of each modal data of a disease to be diagnosed; The expert consensus driven module is used for extracting expert consensus and clinical experience from an expert knowledge base of the disease to be diagnosed, combining state semantic labels corresponding to the overall embedding features, determining state relationship labels between the overall embedding features, extracting embedding feature representations of relationships from the state relationship labels, and combining the overall embedding features, the state semantic labels, the state relationship labels and the embedding feature representations of the relationships to construct a state level graph, and then combining a modal data graph of the modal data to obtain a global multi-modal-multi-level connection graph; The analogy deduction module is used for constructing an individual multi-modal-multi-level connection graph based on analogy deduction of a plurality of groups of triples according to the global multi-modal-multi-level connection graph; the triples comprise two overall embedding features and corresponding embedding feature representations; The explicit derivation rule calculation module is used for obtaining complete reasoning rules of all nodes according to each node of the individual multi-modal-multi-level connection graph; The progressive thinking guidance module is used for combining modal categories in the global multi-modal-multi-level connection graph and the complete reasoning rules of all nodes to obtain a thinking guiding information part in a logic chain prompt: Representative data of different modalities is extracted from a state level graph of the global multi-modal-multi-level connection graph, and a question and answer example library with progressive reasoning is constructed for different modal data combinations; then, corresponding question and answer examples are extracted from the question and answer example library according to the modal categories in the complete reasoning rules of all nodes as the thinking guiding information part in the logic chain prompt.
2. The method of claim 1, wherein the method further comprises: According to the modal data graph of each modal data of the disease to be diagnosed, the overall embedding features as states are obtained, comprising: a feature aggregation layer based on a graph neural network, by an aggregation function Aggregate The neighborhood information of each node in the modality data graph of each modality data of the disease to be diagnosed is aggregated, a global representation of the node is formed by aggregation, and the overall embedding features of the modality data graph of each modality data of the disease to be diagnosed are extracted.
3. The method of claim 1, wherein the method further comprises: The analogy deduction module is used for constructing an individual multi-modal-multi-level connection graph based on analogy deduction of a plurality of groups of triples according to the global multi-modal-multi-level connection graph; the triples comprise two overall embedding features and corresponding embedding feature representations; According to the global multi-modal-multi-level connection graph, example triples are obtained; The to-be-predicted triples and the example triples are embedded and represented by a analogy feature embedding layer, the feature mapping between modalities is learned by a cross-modal adaptive interaction layer, and the embedding representations corresponding to each modality are obtained; The embedding representations corresponding to each modality are integrated by a multi-modal information fusion layer, and then the predicted triples and the feature representations of the predicted triples after prediction are obtained by a reasoning consistency normalization layer and a prediction output layer. According to the predicted prediction triplets and the feature representations of the predicted prediction triplets, an individual multi-modal-multi-layer joint graph is constructed.
4. The method of claim 1, wherein the method further comprises: According to the nodes of the individual multi-modal-multi-layer joint graph, complete reasoning rules of all nodes are obtained, including: The nodes of the individual multi-modal-multi-layer joint graph are sequentially input into a strategy function of the reasoning rule calculation; The strategy function includes a state coding layer, a plurality of feature extraction layers, and a strategy selection layer connected in sequence, and an activation layer is arranged before each of the state coding layer, the plurality of feature extraction layers, and the strategy selection layer, and the output of the strategy selection layer further includes normalization processing.
5. The method of claim 1, wherein the method further comprises: The loss function used by the analogical deduction module during training is: wherein, is a relaxed loss function, is the total number of triplets in the training set, is a sinusoidal similarity function, is the hidden feature representation of the example triplet, is the hidden feature representation of the predicted triplet, is a maximization computation formula, and are the hidden feature representations of the head and tail entities in the predicted pair, respectively.
6. The method of claim 1, wherein the method further comprises: The reward function used by the explicit derivation rule calculation module during training is: wherein, is a total reward function, is a global goal reward, is a path efficiency reward, is a path diversity reward, is the last node of the individual multimodal-multilevel connection graph, is a path length, is a set of explored paths, is a cosine similarity function, is a coefficient of is a coefficient of is a coefficient of is a coefficient of is a coefficient of is a coefficient of The update of the strategy function parameters of the explicit derivation rule calculation module during training adopts a Monte Carlo policy gradient method.
7. The method of claim 1, wherein the method further comprises: The loss function used by the knowledge guidance model during training is: wherein, is the overall loss function for the knowledge guiding model training, is the loss function involved in updating the parameters of the policy function using the Monte Carlo policy gradient method, denoted as the path optimization loss, is the reasoning process of the thinking guiding information part in the logical chain prompt and the semantic loss function of the label calculation, is the relaxation loss function: wherein, is a reasoning process label, is is each embedding representation produced after a semantic encoding model, is is each embedding representation produced after a semantic encoding model, is a reasoning process for a thought leading information portion in a logical chain hint, is a sine similarity function.
8. A knowledge guide-based medical large model fine-tuning optimization system, characterized in that, including: A data acquisition module is configured to acquire multi-modal data to be diagnosed. The multi-modal data to be diagnosed is image or text type multi-modal data from different medical examination methods. A prompt output module is configured to input the multi-modal data to be diagnosed into a knowledge guidance model to obtain a thinking guidance information part in a logic chain prompt, and input the thinking guidance information part into a medical large model to obtain an output normalized by the chain prompt. The knowledge guidance model includes a multi-modal-multi-layer joint graph fusion module, an expert consensus driving module, an analogical deduction module, an explicit derivation rule calculation module, and a progressive thinking guidance module. The multi-modal-multi-layer joint graph fusion module is configured to obtain overall embedding features as states according to a modality data graph of each modality data of a disease to be diagnosed. The expert consensus driving module is configured to extract expert consensus and clinical experience from an expert knowledge base of the disease to be diagnosed, determine state relationship labels between the overall embedding features in combination with state semantic labels corresponding to the overall embedding features, extract relationship embedding feature representations from the state relationship labels, and construct a state level graph in combination with the overall embedding features, the state semantic labels, the state relationship labels, and the relationship embedding feature representations, and obtain a global multi-modal-multi-layer joint graph in combination with a modality data graph of the modality data. The analogical deduction module is configured to construct an individual multi-modal-multi-layer joint graph based on analogical deduction of a plurality of triplets according to the global multi-modal-multi-layer joint graph, the triplets including two overall embedding features and corresponding embedding feature representations. The explicit derivation rule calculation module is configured to obtain complete reasoning rules of all nodes according to nodes of the individual multi-modal-multi-layer joint graph. The progressive thinking guidance module is configured to obtain a thinking guidance information part in a logic chain prompt in combination with modality categories in the global multi-modal-multi-layer joint graph and the complete reasoning rules of all nodes. Representative data of different modalities is extracted from the state level atlas of the global multi-modal multi-level connection atlas, and a question and answer example library with progressive reasoning is constructed for different modal data combinations; and according to the modal categories in the complete reasoning rules of all nodes, corresponding question and answer examples are extracted from the question and answer example library as the thinking guide information part in the logic chain prompt.
9. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the knowledge-guided medical large model fine-tuning optimization according to any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent answering method, system and device for multi-step reasoning problem and medium
CN118674056A
Cooperative enhancement method oriented to APT knowledge graph and large language model
CN118802369A
Fine adjustment method, system and equipment based on large language model and medium
CN117290480A
Tumor disease auxiliary diagnosis and treatment system based on multi-modal big data model
CN118262893A