Medical consultation reply generation method, device and equipment based on multi-modal fusion

By employing a multimodal fusion-based medical consultation response generation method, which utilizes visual and textual intelligent agents for feature extraction and joint encoding, and combines diagnostic and medical knowledge intelligent agents for reasoning analysis, the problem of low generation quality in large medical models is solved, and computational efficiency and response accuracy are improved.

CN120878255APending Publication Date: 2025-10-31SUN YAT SEN MEMORIAL HOSPITAL SUN YAT SEN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510936369.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing large-scale medical models suffer from high computational error rates, weak ability to interpret image information, and insufficient knowledge generalization ability when processing medical consultation response generation, especially in mathematical reasoning and medical entity recognition tasks. Furthermore, they fail to effectively integrate multimodal information such as text and images, resulting in low generation quality.

Method used

By receiving multimodal data, visual and text-based intelligent agents are selected for feature extraction, and a latent attention caching mechanism is used for joint encoding. Inference analysis is then performed by diagnostic and medical knowledge intelligent agents to verify the consistency of the output and finally generate a medical consultation response.

Benefits of technology

It improves the computational efficiency and resource utilization of large medical models in complex scenarios, and ensures the quality of generated medical consultation responses, especially in terms of accuracy and detail when processing multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120878255A_ABST
    Figure CN120878255A_ABST
Patent Text Reader

Abstract

The invention discloses a medical consultation reply generation method, device and equipment based on multi-modal fusion, and belongs to the field of text generation, and the method comprises the steps: selecting a plurality of first agent experts and a plurality of second agent experts according to multi-modal data input by a user; extracting medical data features and numeric symbol features from the multi-modal data through a first agent expert, and performing joint coding on the medical data features and the numeric symbol features to obtain a multi-modal feature vector; inputting the output of the first agent expert and the multi-modal feature vector to a corresponding second agent expert, and calling the second agent expert according to a preset calling sequence; the second agent experts are connected according to a preset calling sequence; and if the consistency verification of the medical knowledge output by each agent expert and the mathematical logic passes, aggregating the output of each agent expert, and generating a medical consultation reply. The problem of low quality of medical consultation reply generated by a large medical model can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text generation, and in particular to a method, apparatus and device for generating medical consultation responses based on multimodal fusion. Background Technology

[0002] Since current large medical models (such as HuatuoGPT and ChatMed) have demonstrated preliminary capabilities in tasks such as disease diagnosis, image analysis, and drug dosage inference, they are widely used in generating medical consultation responses to improve the efficiency of such responses.

[0003] However, due to the complexity of medical scenarios, large-scale medical models still face challenges in handling mathematical reasoning and medical entity recognition tasks. For example, when dealing with drug dosage calculations or symbolic logic problems, the models exhibit high computational error rates, weak ability to interpret image information, and insufficient knowledge generalization capabilities. Furthermore, existing models mostly employ a single-modal processing flow, failing to effectively integrate multimodal information such as text and images. These issues result in lower quality final medical consultation responses.

[0004] Therefore, improving the quality of medical consultation response generation by large-scale medical models is a technical problem that needs to be solved. Summary of the Invention

[0005] This application provides a method, apparatus, and device for generating medical consultation responses based on multimodal fusion, which can solve the problem of low quality in the generation of medical consultation responses from large medical models in the prior art.

[0006] One embodiment of this application provides a method for generating medical consultation responses based on multimodal fusion, including:

[0007] Receive multimodal data input by the user, and select multiple first intelligent agent experts for feature extraction and multiple second intelligent agent experts for reasoning analysis from a preset pool of intelligent agent experts based on the multimodal data.

[0008] The first intelligent agent expert extracts medical data features and digital symbol features from the multimodal data, and uses a latent attention caching mechanism to jointly encode the medical data features and digital symbol features to obtain a multimodal feature vector.

[0009] The output of the first intelligent agent expert and the multimodal feature vector are input to the corresponding second intelligent agent expert, and each second intelligent agent expert is called in a preset calling order to obtain the output of each second intelligent agent expert; wherein, the second intelligent agent experts are connected in the preset calling order.

[0010] The consistency between the medical knowledge and mathematical logic outputs of each first intelligent agent expert and each second intelligent agent expert is verified separately. When all verifications pass, the outputs of each first intelligent agent expert and each second intelligent agent expert are aggregated to generate a medical consultation response.

[0011] Compared with existing technologies, the above embodiments have the following beneficial effects: By using multimodal data, the intelligent agent experts used in this medical consultation response generation task are rationally selected, thereby avoiding invalid intelligent agent expert invocation operations and improving the computational efficiency and resource utilization of the medical big model; furthermore, medical data features and digital symbol features are extracted from multimodal data respectively, and the medical data features and digital symbol features are jointly encoded through a latent attention caching mechanism to achieve deep integration of medical data and mathematical reasoning, improve the comprehensive performance of the medical big model in complex scenarios, and verify the consistency between the medical knowledge and mathematical logic output by the second intelligent agent expert, ensuring the quality of the final generated medical consultation response; by calling the second intelligent agent experts in sequence and connecting them according to the calling order, the second intelligent agent experts called later can receive the output of the second intelligent agent experts called earlier, improving the degree of information sharing and the rationality of information transmission routes between different intelligent agent experts during information transmission, giving full play to the capabilities of each second intelligent agent expert, and improving the quality of the final generated medical consultation response.

[0012] Further, the step of selecting multiple first agent experts for feature extraction and multiple second agent experts for inference analysis from a preset pool of agent experts based on the multimodal data includes:

[0013] The multimodal data includes: text data and image data;

[0014] The first intelligent agent expert includes at least: a visual intelligent agent expert and a text intelligent agent expert; the second intelligent agent expert includes at least: the text intelligent agent expert.

[0015] When the image data contains a DICOM image, the DICOM agent expert is used as the first agent expert;

[0016] Extract user demand keywords from the text data, and select the corresponding reasoning agent expert as the second agent expert based on the user demand keywords.

[0017] Compared with existing technologies, the above embodiments have the following beneficial effects: by including visual intelligent agent experts and text intelligent agent experts in the first intelligent agent expert, the medical big model's ability to process multimodal data such as text and images is improved; by adding text intelligent agent experts to the second intelligent agent expert, the medical big model's text generation capability is guaranteed; in addition, the DICOM intelligent agent expert is activated only when DICOM images are present in the multimodal data, and the user's needs are confirmed from the text data and the inference intelligent agent expert used in this medical consultation response generation task is reasonably selected, thereby avoiding invalid intelligent agent expert call operations and improving the computational efficiency and resource utilization of the medical big model.

[0018] Furthermore, the step of extracting medical data features and digital symbol features from the multimodal data through the first intelligent agent expert includes:

[0019] The visual intelligent agent expert extracts medical image semantic features from the image data;

[0020] The text-based intelligent agent expert extracts query semantic features from the text data;

[0021] When a DICOM image is present in the image data, the DICOM intelligent agent expert extracts medical metadata features from the DICOM image.

[0022] By concatenating the medical image semantic features, the query semantic features, and the medical metadata features, the medical data features are obtained.

[0023] Quantitative medical parameters are extracted from the medical data features to obtain the digital symbol features.

[0024] Compared with existing technologies, the above embodiments have the following beneficial effects: by receiving text data and image data, and extracting relevant medical data features and mathematical symbol features from the text data and image data through different intelligent agent experts, it is convenient to perform subsequent joint feature encoding, realize the deep fusion of multimodal information of text and image data, and improve the comprehensive performance of the medical big model in complex scenarios; furthermore, in the feature extraction process, the corresponding intelligent agent expert is activated according to the data type included in the multimodal data, which can effectively improve the computational efficiency of the medical big model.

[0025] Further, the step of inputting the output of the first intelligent agent expert and the multimodal feature vector to the corresponding second intelligent agent expert, and calling each second intelligent agent expert according to a preset calling order, includes:

[0026] The reasoning agent experts include: diagnostic agent experts and medical knowledge agent experts;

[0027] The multimodal feature vector is input to each of the inference agent experts among the plurality of second agent experts;

[0028] If the diagnostic intelligent agent expert is among the plurality of second intelligent agent experts, then the semantic features of the medical image are input into the diagnostic intelligent agent expert.

[0029] If the medical knowledge intelligent agent expert is among the plurality of second intelligent agent experts and the image data contains a DICOM image, then the medical metadata features are input to the medical knowledge intelligent agent expert.

[0030] The outputs of each of the reasoning agent experts among the plurality of second agent experts are input into the text agent expert to obtain a comprehensive processing result.

[0031] Compared with existing technologies, the above embodiments have the following beneficial effects: by rationally dividing the task of generating medical consultation responses among intelligent agent experts with different capabilities, the computational efficiency and resource utilization of the large medical model are improved; at the same time, by transmitting information among intelligent agent experts, information sharing among different intelligent agent experts is realized, giving full play to the capabilities of each intelligent agent expert and improving the quality of the final generated medical consultation responses; in addition, during reasoning analysis, the corresponding reasoning intelligent agent expert is activated as needed according to user requirements, saving the computational resources of the large medical model.

[0032] Furthermore, the step of verifying the consistency between the medical knowledge and mathematical logic output by each of the first and second intelligent agent experts includes:

[0033] The consistency of medical knowledge output by each first intelligent agent expert and each second intelligent agent expert is verified by using a pre-set medical knowledge graph.

[0034] The mathematical logic consistency of the outputs of each first intelligent agent expert and each second intelligent agent expert is verified by a preset symbolic logic engine.

[0035] Calculate the confidence level of the outputs of each first agent expert and each second agent expert, and verify whether each confidence level meets a preset threshold.

[0036] Compared with existing technologies, the above embodiments have the following beneficial effects: by introducing a dual logic verification mechanism, it ensures that the outputs of each second intelligent expert conform to medical standards and avoids calculation errors, thus solving the accuracy problem of medical models in mathematical reasoning and medical entity recognition in existing technologies, reducing the error rate, and improving the generation quality of subsequent medical consultation responses; furthermore, by verifying the confidence level of the outputs of each second intelligent expert, low-confidence results are filtered out, preventing uncertain or low-quality intermediate results from affecting the credibility of the final medical consultation response.

[0037] Furthermore, after verifying the consistency between the medical knowledge and mathematical logic output by each of the first and second intelligent agent experts, the process further includes:

[0038] If any verification fails, a backup agent expert is added to the preset agent experts, and multiple first agent experts and multiple second agent experts are reselected from the preset agent experts to regenerate the medical consultation response.

[0039] Compared with the prior art, the above embodiments have the following beneficial effects: when verification fails, by automatically adding backup intelligent agents and reorganizing expert processes, the system's ability to cope with complex or abnormal data is enhanced, ensuring the quality of the final report generation.

[0040] Furthermore, after aggregating the outputs of each of the first and second intelligent agent experts to generate a medical consultation response, the process also includes:

[0041] If the medical consultation response also needs to generate medical images, the medical images are generated through an improved convolutional neural network; wherein, the improved convolutional neural network is obtained by parsing the textual conditional constraints from the medical consultation response through a conditional diffusion model, and then embedding the textual conditional constraints into a preset U-Net convolutional neural network through an attention mechanism.

[0042] Compared with existing technologies, the above embodiments have the following beneficial effects: when it is necessary to generate medical images, the textual conditional constraints are parsed from the medical consultation response through the conditional diffusion model, and the textual conditional constraints are embedded into the U-Net convolutional neural network based on the attention mechanism, so that it has the ability to understand text, thereby enabling the model to have the ability to generate text and images in a multimodal manner, while improving the accuracy and detail of the generated images, especially when dealing with complex lesion descriptions.

[0043] Another embodiment of this application also provides a medical consultation response generation device based on multimodal fusion, including: an agent expert selection module, a joint encoding module, an agent expert invocation module, and a verification module;

[0044] The intelligent agent expert selection module is used to receive multimodal data input by the user, and select multiple first intelligent agent experts for feature extraction and multiple second intelligent agent experts for reasoning analysis from a preset pool of intelligent agent experts based on the multimodal data.

[0045] The joint encoding module is used to extract medical data features and digital symbol features from the multimodal data through the first intelligent agent expert, and to jointly encode the medical data features and digital symbol features through a latent attention caching mechanism to obtain a multimodal feature vector.

[0046] The intelligent agent expert invocation module is used to input the output of the first intelligent agent expert and the multimodal feature vector to the corresponding second intelligent agent expert, and invoke each second intelligent agent expert according to a preset invocation order to obtain the output of each second intelligent agent expert; wherein, the second intelligent agent experts are connected to each other according to the preset invocation order;

[0047] The verification module is used to verify the consistency of the medical knowledge and mathematical logic output by each of the first intelligent agent experts and each of the second intelligent agent experts. When all verifications pass, the outputs of each of the first intelligent agent experts and each of the second intelligent agent experts are aggregated to generate a medical consultation response.

[0048] Furthermore, the agent expert selection module is used to select, based on the multimodal data, multiple first agent experts for feature extraction and multiple second agent experts for inference analysis from a preset pool of agent experts, including:

[0049] The multimodal data includes: text data and image data;

[0050] The first intelligent agent expert includes at least: a visual intelligent agent expert and a text intelligent agent expert; the second intelligent agent expert includes at least: the text intelligent agent expert.

[0051] When the image data contains a DICOM image, the DICOM agent expert is used as the first agent expert;

[0052] Extract user demand keywords from the text data, and select the corresponding reasoning agent expert as the second agent expert based on the user demand keywords.

[0053] Another embodiment of this application also provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the steps of the medical consultation response generation method based on multimodal fusion as described in this application. Attached Figure Description

[0054] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0055] Figure 1 This is a flowchart illustrating a method for generating medical consultation responses based on multimodal fusion, provided in some embodiments of this application.

[0056] Figure 2 This is another flowchart illustrating a method for generating medical consultation responses based on multimodal fusion, as provided in some embodiments of this application.

[0057] Figure 3 This is a flowchart of a dynamic routing Top-k gating selection algorithm provided in some embodiments of this application;

[0058] Figure 4 This is a schematic diagram of the structure of a medical consultation response generation method based on multimodal fusion provided in some embodiments of this application. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0061] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0062] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0063] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0064] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).

[0065] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.

[0066] Due to the complexity of medical scenarios, large-scale medical models still face challenges in handling mathematical reasoning and medical entity recognition tasks. For example, when dealing with drug dosage calculations or symbolic logic problems, models exhibit high calculation error rates, weak ability to interpret image information, and insufficient knowledge generalization capabilities. Furthermore, existing models mostly employ single-modal processing flows, failing to effectively integrate multimodal information such as text and images. These issues result in low-quality medical consultation responses.

[0067] Please refer to Figure 1 To address the problem of low-quality medical consultation responses generated by large medical models in existing technologies, this application provides a method for generating medical consultation responses based on multimodal fusion, comprising steps S101 to S104, specifically:

[0068] S101: Receive multimodal data input by the user, and select multiple first intelligent agent experts for feature extraction and multiple second intelligent agent experts for reasoning analysis from a preset pool of intelligent agent experts based on the multimodal data.

[0069] Furthermore, in some embodiments of this application, the multimodal data includes: text data and image data; the image data includes: DICOM (Digital Imaging and Communications in Medicine) images and general medical images.

[0070] Furthermore, in some embodiments of this application, reference is made to... Figure 3 The flowchart shown illustrates that upon receiving multimodal data, a Task object is created for the user currently inputting the multimodal data, and a unique ID is assigned. Specifically, this includes assigning a unique ID to each medical consultation response generation task, while also associating it with the session context (such as historical dialogue data) under the same task. The above process can adopt the object modeling methods commonly used in software engineering, and this application does not specifically limit the technical means used in this process.

[0071] Furthermore, in some embodiments of this application, the association of the session context under the same task includes: integrating background data such as patient basic information and historical diagnostic results as the session context under the same task. The above process can be based on the context association technology of database query. This application does not specifically limit the technical means used in this process.

[0072] Furthermore, in some embodiments of this application, upon receiving multimodal data, the method further includes: encoding and preprocessing the multimodal data, specifically: generating semantic vectors by combining text data with pre-trained models such as Contrastive Language-Image Pre-training (CLIP); and extracting visual features from image data using pre-trained models such as CLIP. CLIP is an open-source multimodal pre-training model.

[0073] Through the above process, the preprocessed multimodal features are unified into a 768-dimensional vector, laying the foundation for subsequent joint encoding.

[0074] Furthermore, in some embodiments of this application, the step of selecting multiple first agent experts for feature extraction and multiple second agent experts for inference analysis from a preset pool of agent experts based on the multimodal data includes:

[0075] The multimodal data includes: text data and image data;

[0076] The first intelligent agent expert includes at least: a visual intelligent agent expert and a text intelligent agent expert; the second intelligent agent expert includes at least: the text intelligent agent expert.

[0077] When the image data contains a DICOM image, the DICOM agent expert is used as the first agent expert;

[0078] Extract user demand keywords from the text data, and select the corresponding reasoning agent expert as the second agent expert based on the user demand keywords.

[0079] Furthermore, in some embodiments of this application, the reasoning agent expert includes: a diagnostic agent expert and a medical knowledge agent expert.

[0080] Furthermore, in some embodiments of this application, the preset intelligent agent experts include, but are not limited to: various text intelligent agent experts for natural language processing, visual intelligent agent experts for image data processing, DICOM intelligent agent experts for parsing DICOM files, diagnostic intelligent agent experts for medical reasoning analysis, and medical knowledge intelligent agent experts. This application does not impose specific limitations on the preset intelligent agent experts in the method; corresponding intelligent agent experts can be embedded according to the user's actual needs, exhibiting good scalability.

[0081] Furthermore, in some embodiments of this application, the preset intelligent agent expert is a segmented medical large model. For example, if the total number of parameters in the medical large model is 34B, 3B of the parameters are used to perform data processing tasks for the diagnostic intelligent agent expert, 2.5B of the parameters are used to perform data processing tasks for the medical knowledge intelligent agent expert, 4B of the parameters are used to perform data processing tasks for the visual intelligent agent expert, 5B of the parameters are used to perform data processing tasks for the text intelligent agent expert, and so on. In other words, all the intelligent agent experts can be understood as a combination of all the intelligent agent experts into a single medical large model.

[0082] Furthermore, in some embodiments of this application, the first and second intelligent agent experts can be of the same type. For example, during feature extraction, the text intelligent agent expert in the first intelligent agent expert performs basic operations such as word segmentation and semantic vectorization of the text data, while the text intelligent agent expert in the second intelligent agent expert performs in-depth semantic parsing and text generation based on the feature data provided by each first intelligent agent expert and the inference analysis results output by other second intelligent agent experts. That is, although they are all text intelligent agent experts, different natural language processing models are needed to achieve tasks at different stages.

[0083] To more clearly illustrate the process of selecting multiple first-agent experts and multiple second-agent experts, and Figure 2 The process of selecting an intelligent agent expert in step two of the diagram is shown below. Next, we will combine... Figure 3The dynamic routing Top-k gating selection algorithm flow shown below will be further explained:

[0084] First, the first-level intelligent agent experts need to include basic visual and text intelligent agent experts to parse image and text data from the user-input multimodal data. Second, based on whether the multimodal data contains DICOM images, it is determined whether a DICOM intelligent agent expert needs to be selected from the first-level intelligent agent experts to parse DICOM images in the multimodal data. Further, the text data in the user-input multimodal data is analyzed to extract text keywords (such as user demand keywords like diagnosis, medical, or explanation). Based on the corresponding text keywords, a second-level intelligent agent expert is determined for subsequent inference analysis. For example, if the text data includes medical keywords, a medical knowledge intelligent agent expert is added to the second-level intelligent agent expert list; if the text data includes diagnosis keywords, a diagnosis intelligent agent expert is added to the second-level intelligent agent expert list. Each time an intelligent agent expert is added, a corresponding weight coefficient is set for each intelligent agent expert according to the task complexity and urgency. Through reinforcement learning, the weight coefficients of each intelligent agent expert are optimized by combining their historical performance.

[0085] Furthermore, in some embodiments of this application, the intelligent agent experts select Top-k experts through a sparse gating network by using a dynamic routing mechanism, specifically as follows:

[0086]

[0087] g i (x) = softmax(W) g x+b g )

[0088] The following sparse regularization term is also introduced:

[0089]

[0090] Where Top-k represents the k highest-weighted agent experts in the system; y represents the result after fusing the outputs of the k highest-weighted agent experts; g i (x) is the gating weight function of the i-th agent expert; x is the input data; E i (x) represents the output of the i-th agent expert; W g b is the weight matrix of the gated network; g is the bias vector of the gated network; softmax is the normalized exponential function; λ represents the sparse regularization term introduced in the large-scale medical model; λ is the sparsity control parameter, a larger value tends to activate fewer experts; N is the total number of agent experts in the large-scale medical model; ||g i(x)‖ represents the L0 norm, indicating that the number of activated agent experts is ≤3, meaning that at most three agent experts are activated for each task.

[0091] As can be seen from the above embodiments, this application improves the medical big data model's ability to process multimodal data such as text and images by including visual intelligent agent experts and text intelligent agent experts in the first intelligent agent expert; it ensures the medical big data model's text generation capability by adding text intelligent agent experts to the second intelligent agent expert; in addition, the DICOM intelligent agent expert is activated only when DICOM images are present in the multimodal data, and the user's needs are confirmed from the text data and the inference intelligent agent expert used in this medical consultation response generation task is reasonably selected, thereby avoiding invalid intelligent agent expert call operations and improving the computational efficiency and resource utilization of the medical big data model.

[0092] S102: The first intelligent agent expert extracts medical data features and digital symbol features from the multimodal data, and uses a latent attention caching mechanism to jointly encode the medical data features and digital symbol features to obtain a multimodal feature vector.

[0093] Furthermore, in some embodiments of this application, the step of extracting medical data features and digital symbol features from the multimodal data through the first intelligent agent expert includes:

[0094] The visual intelligent agent expert extracts medical image semantic features from the image data;

[0095] The text-based intelligent agent expert extracts query semantic features from the text data;

[0096] When a DICOM image is present in the image data, the DICOM intelligent agent expert extracts medical metadata features from the DICOM image.

[0097] By concatenating the medical image semantic features, the query semantic features, and the medical metadata features, the medical data features are obtained.

[0098] Quantitative medical parameters are extracted from the medical data features to obtain the digital symbol features.

[0099] Furthermore, in some embodiments of this application, the text agent expert incorporates a retrieval strategy that integrates Retrieval-Augmented Generation (RAG) knowledge graphs. Specifically, after a user inputs text data, the text agent expert processes the text data during the retrieval phase. At this time, the text agent expert utilizes the powerful semantic association capabilities of the knowledge graph to deeply analyze medical-related queries in the user-input text data, mapping them to corresponding entities and relationships in the knowledge graph. For example, when a user inquires about the symptoms of a certain disease, the system will accurately retrieve all symptom-related guidelines based on the association between diseases and symptoms in the knowledge graph. The retrieved relevant information from the medical knowledge graph is used as context input into the generative model. Combined with the general language understanding and generation capabilities of the generative model trained on a large-scale corpus using RAG technology, an accurate, detailed, and medically professional answer is generated for the user's query. Meanwhile, we continuously optimize the retrieval and generation process, adjusting retrieval strategies and generation model parameters based on dynamic updates to the medical knowledge graph and user feedback, to ensure timely and accurate integration of the latest medical knowledge and provide users with higher-quality medical information retrieval and answer services.

[0100] Furthermore, in some embodiments of this application, the medical data features include: medical image semantic features extracted from image data, query semantic features extracted from text data, and medical metadata features extracted from DICOM images when DICOM images are present in the image data; the digital symbol features include: numerical calculation features (such as dose calculation, HU value, pixel spacing and other quantitative parameters) parsed from the query semantic features, and medical measurement data (such as lesion size, density value, signal strength and other numerical medical indicators) extracted from medical metadata features and medical image semantic features.

[0101] Furthermore, in some embodiments of this application, the step of jointly encoding the medical data features and digital symbol features through a latent attention caching mechanism to obtain a multimodal feature vector includes: the calculation formula for the joint encoding is specifically as follows:

[0102] c = MLP(Attn(H,M))

[0103]

[0104] Where c is the multimodal feature vector; MLP is the multilayer perceptron operator; Attn is the latent attention caching mechanism operator; H is the medical data feature, H∈R d M represents the symbolic feature of numbers, M∈R d R dLet W be a d-dimensional real vector space; softmax is a normalized exponential function; W Q W K and W V The parameters are learnable, and W Q W K and W V ∈R d×d R d×d Let be a d×d dimensional real vector space, where d = 768.

[0105] It is understandable that the bimodal joint encoding process in step three is also based on the above formula, where Attn(H,M) is the cross-modal attention calculation.

[0106] As can be seen from the above embodiments, this application receives text data and image data, and extracts relevant medical data features and mathematical symbol features from the text data and image data through different intelligent agent experts, so as to facilitate subsequent joint feature encoding, realize the deep fusion of multimodal information of text and image data, and improve the comprehensive performance of the medical big model in complex scenarios; furthermore, in the feature extraction process, the corresponding intelligent agent expert is activated according to the data type included in the multimodal data, which can effectively improve the computational efficiency of the medical big model.

[0107] S103: Input the output of the first intelligent agent expert and the multimodal feature vector to the corresponding second intelligent agent expert, and call each second intelligent agent expert in a preset calling order to obtain the output of each second intelligent agent expert; wherein, the second intelligent agent experts are connected in the preset calling order.

[0108] Furthermore, in some embodiments of this application, the step of inputting the output of the first intelligent agent expert and the multimodal feature vector to the corresponding second intelligent agent expert, and calling each of the second intelligent agent experts according to a preset calling order, includes:

[0109] The reasoning agent experts include: diagnostic agent experts and medical knowledge agent experts;

[0110] The multimodal feature vector is input to each of the inference agent experts among the plurality of second agent experts;

[0111] If the diagnostic intelligent agent expert is among the plurality of second intelligent agent experts, then the semantic features of the medical image are input to the diagnostic intelligent agent expert.

[0112] If the medical knowledge intelligent agent expert is among the plurality of second intelligent agent experts and the image data contains a DICOM image, then the medical metadata features are input to the medical knowledge intelligent agent expert.

[0113] The outputs of each of the reasoning agent experts among the plurality of second agent experts are input into the text agent expert to obtain a comprehensive processing result.

[0114] Preferably, in some embodiments of this application, reference is made to Figure 2 The flowchart shown in step four illustrates the multi-agent expert parallel processing process. There is also a second preset calling order between the first and second agent experts, specifically:

[0115] First, the first priority call analyzes the user-input image data using the visual and DICOM agents within the first intelligent agent expert group to obtain the analysis results. The visual agent expert processes non-DICOM image data, while the DICOM agent expert processes DICOM images (this can be understood as the first priority call being used simultaneously to obtain multimodal feature vectors and for acquiring input data for subsequent second intelligent agent experts). The second priority call involves the medical knowledge agent expert within the second intelligent agent expert group receiving the multimodal feature vectors and the output of the DICOM agent expert within the first intelligent agent expert group; the diagnostic agent expert within the second intelligent agent expert group also receives the multimodal feature vectors and the output of the visual agent expert within the first intelligent agent expert group. The third priority call involves the text expert integrating the multimodal feature vectors, the outputs from the first and second priority calls, and the output of each intelligent agent expert to obtain the comprehensive processing result. In other words, the dependencies between the first and second intelligent agent experts are as follows: the visual agent expert connects to the diagnostic agent expert, the DICOM agent expert connects to the medical knowledge agent expert, and all four intelligent agent experts (visual, DICOM, diagnostic, and medical knowledge) are connected to the text agent expert.

[0116] As can be seen from the above embodiments, this application improves the computational efficiency and resource utilization of the large medical model by rationally dividing the task of generating medical consultation responses among intelligent agent experts with different capabilities. At the same time, by transmitting information among intelligent agent experts, information sharing among different intelligent agent experts is realized, giving full play to the capabilities of each intelligent agent expert and improving the quality of the final generated medical consultation responses. In addition, during reasoning analysis, the corresponding reasoning intelligent agent expert is activated as needed according to user requirements, saving the computational resources of the large medical model.

[0117] S104: Verify the consistency between the medical knowledge and mathematical logic outputs of each of the first and second intelligent agents. When all verifications pass, aggregate the outputs of each of the first and second intelligent agents to generate a medical consultation response.

[0118] Furthermore, in some embodiments of this application, the step of verifying the consistency between the medical knowledge and mathematical logic output by each of the first and second intelligent agent experts includes:

[0119] The consistency of medical knowledge output by each first intelligent agent expert and each second intelligent agent expert is verified by using a pre-set medical knowledge graph.

[0120] The mathematical logic consistency of the outputs of each first intelligent agent expert and each second intelligent agent expert is verified by a preset symbolic logic engine.

[0121] Calculate the confidence level of the outputs of each first agent expert and each second agent expert, and verify whether each confidence level meets a preset threshold.

[0122] Furthermore, in some embodiments of this application, the verification of the consistency between the medical knowledge and mathematical logic output by each of the second intelligent agent experts can also be:

[0123] The use of terminology is verified by knowledge graph intelligent agents, and consistency checks are performed against a pre-set disease knowledge base to achieve consistency verification of medical knowledge; the diagnostic intelligent agents verify the diagnostic logic and check the rationality of symptom-disease matching to achieve consistency of mathematical logic.

[0124] Furthermore, in some embodiments of this application, reference is made to... Figure 2 Step five of the illustrated process, verifying the consistency of medical knowledge and mathematical logic in the outputs of each of the second intelligent agents, can also be achieved as follows: Based on multiple preset knowledge rules included in the clinical knowledge graph, the results output by each intelligent agent are checked against guidelines to ensure consistency, such as incorrect terminology usage or inaccurate anatomical descriptions; a symbolic logic engine (such as the Z3 solver) is invoked to verify the mathematical consistency of formula derivations in the outputs of each intelligent agent, such as incorrect dosage calculations or abnormal numerical ranges; finally, a confidence score is included in the output of each intelligent agent, and further quality verification is performed by determining whether the confidence score meets a preset threshold, such as setting the preset threshold to 0.3. When any verification by any intelligent agent fails, a backup intelligent agent is activated, and an intelligent agent is selected through intelligent routing.

[0125] Furthermore, in some embodiments of this application, the consistency verification of medical knowledge includes: verifying the medical discovery descriptions of visual agent experts; verifying the terminology used by medical knowledge agent experts; verifying the diagnostic logic of diagnostic agent experts; and verifying the interpretation of medical parameters by DICOM agent experts.

[0126] Furthermore, in some embodiments of this application, the consistency verification of the mathematical logic includes: verifying the mathematical correctness of the dose calculation formula in the output of each agent expert; verifying the rationality of the HU value, pixel spacing, and other values ​​in the output of each agent expert; and verifying the logical consistency of the measurement data in the output of each agent expert.

[0127] Furthermore, in some embodiments of this application, after verifying the consistency between the medical knowledge and mathematical logic output by each of the first and second intelligent agent experts, the method further includes:

[0128] If any verification fails, a backup agent expert is added to the preset agent experts, and multiple first agent experts and multiple second agent experts are reselected from the preset agent experts to regenerate the medical consultation response.

[0129] Furthermore, in some embodiments of this application, when any agent expert fails any verification, an error log is recorded, then a backup expert is activated and the process is repeated, followed by downgrade processing. Downgrade processing refers to a simplified processing strategy adopted when agent expert verification fails. Triggering conditions for downgrade processing include: failure of medical knowledge consistency verification, such as incorrect terminology or inaccurate anatomical description; failure of mathematical logic consistency verification, such as incorrect dosage calculation or abnormal numerical range; agent expert response timeout, such as processing time exceeding a preset threshold; and agent expert output confidence being too low, such as the agent expert output confidence being less than 0.3.

[0130] As can be seen from the above embodiments, when verification fails, this application enhances the system's ability to cope with complex or abnormal data by automatically adding backup intelligent agents and reorganizing expert processes, thereby ensuring the quality of the final report.

[0131] Furthermore, in some embodiments of this application, reference is made to... Figure 2 Step six of the process shown, which involves aggregating the outputs of each of the first and second intelligent agents to generate a medical consultation response, includes: selecting a leading intelligent agent expert from among the intelligent agents, aggregating the outputs of each intelligent agent expert, generating a structured medical consultation response, verifying the quality of the medical consultation response, and updating the historical records.

[0132] Furthermore, in some embodiments of this application, after aggregating the outputs of each of the second intelligent agent experts to generate a medical consultation response, the method further includes:

[0133] If the medical consultation response also needs to generate medical images, the medical images are generated through an improved convolutional neural network; wherein, the improved convolutional neural network is obtained by parsing the textual conditional constraints from the medical consultation response through a conditional diffusion model, and then embedding the textual conditional constraints into a preset U-Net convolutional neural network through an attention mechanism.

[0134] Furthermore, in some embodiments of this application, reference is made to... Figure 2 Step seven, as shown, involves the conditional diffusion model, which includes: designing an embedded conditional diffusion model based on the visual understanding capabilities of a large medical model; adding textual conditional constraints to the U-Net diffusion model architecture based on an attention mechanism to enable textual understanding; and ultimately achieving the model's text-to-image multimodal generative capability. The conditional diffusion model further includes: medical-specific conditional encoding; medical image constraint mechanisms; and special processing embedding methods, i.e., embedding medical domain knowledge in each attention layer of the U-Net. The medical domain knowledge, i.e., the textual conditional constraints, includes: anatomical priors, such as ensuring the generated images conform to human anatomy; pathological manifestation constraints, such as the generated pathological features must conform to medical cognition; and image quality standards, such as the contrast and resolution of the generated images meeting clinical requirements.

[0135] In summary, the medical consultation response generation method based on multimodal fusion provided in this application has the following advantages compared to existing technologies: Firstly, by using multimodal data, the intelligent agent experts used in this medical consultation response generation task are rationally selected, thereby avoiding invalid intelligent agent expert invocation operations and improving the computational efficiency and resource utilization of the large medical model. Secondly, medical data features and digital symbol features are extracted from the multimodal data, and the medical data features and digital symbol features are jointly encoded through a latent attention caching mechanism to achieve deep fusion of medical data and mathematical reasoning, improving the comprehensive performance of the large medical model in complex scenarios and verifying the consistency between the medical knowledge and mathematical logic output by the second intelligent agent expert, ensuring the quality of the final generated medical consultation response. Thirdly, by sequentially invoking the second intelligent agent experts and connecting them according to the invocation order, the second intelligent agent experts called later can receive the output of the second intelligent agent experts called earlier, improving the degree of information sharing and the rationality of information transmission routes between different intelligent agent experts during information transmission, fully leveraging the capabilities of each second intelligent agent expert, and improving the quality of the final generated medical consultation response.

[0136] Based on the above embodiments of the medical consultation response generation method based on multimodal fusion, this application also provides a method for obtaining the training dataset of a large medical model implementing the above embodiments of the medical consultation response generation method based on multimodal fusion, specifically including:

[0137] First, set the mixing ratio of different types of data in the training dataset. Specifically, set the proportion of clinical guideline data to 45%, the proportion of medical literature data to 30%, the proportion of mathematical reasoning data to 15%, and the proportion of patient dialogues to 10%.

[0138] The uses of different types of data in the training process of medical large-scale models are defined as follows: Clinical guideline data is used in the pre-training phase to provide the model with validated standard treatment procedures and diagnostic suggestions within the medical field, helping the model establish a basic medical knowledge framework; medical literature data is used in the domain adaptation phase to enable the model to adapt to the professional expressions and latest knowledge in specific medical fields through the latest medical research results and papers, improving its understanding of professional medical content; mathematical reasoning data is used in the capability enhancement phase to train the model to handle complex calculations, dosage estimation, statistical analysis, and other content involved in the medical field, improving the model's logical reasoning and numerical calculation capabilities; and mathematical reasoning data is used in the interactive training phase to help the model learn how to communicate with users in a natural, professional, and easy-to-understand way by simulating doctor-patient dialogue scenarios, improving the model's interactive capabilities and humanized expression in real-world application scenarios.

[0139] Based on the aforementioned embodiments of the medical consultation response generation method and training dataset acquisition method based on multimodal fusion, this application also provides a method for training the aforementioned large-scale medical model. Specifically, different weights are assigned to different types of data, so that the large-scale medical model will focus more on data types with higher weights during training, thereby better learning and adapting to the characteristics and patterns of these data types. At the same time, data types with lower weights also have some influence, but it is relatively small, thus helping the large-scale medical model achieve a balance among different data types.

[0140] Furthermore, in some embodiments of this application, after the medical big data model has been trained, the method further includes: fine-tuning the medical big data model, specifically: firstly, fine-tuning on the MIMIC-Ⅲ (Medical Information Mart for IntensiveCare) and multi-center collected medical image report multimodal datasets to enable it to have mathematical reasoning capabilities, and simultaneously fine-tuning on the mathematical reasoning dataset to improve its symbolic logic accuracy; testing the dynamic routing mechanism of the medical big data model, specifically: testing the dynamic routing effect of the Top-k gated network to ensure that input data can be correctly allocated to relevant intelligent agent experts; testing the bimodal joint encoding capability of the medical big data model, specifically: testing the bimodal joint encoding effect of the potential attention caching mechanism of the medical big data model to ensure that medical data features and mathematical symbolic features can be effectively integrated; and testing the conditional diffusion model reinforcement learning capability of the medical big data model, specifically: testing the image generation capability of the medical big data model, and further optimizing the quality of generated images based on reinforcement learning through an expert scoring mechanism.

[0141] Based on the above embodiments of the medical consultation response generation method, training dataset acquisition method, and medical large model training method, this application also provides a method for deploying the above-mentioned medical large model, including:

[0142] (1) Backend Architecture Design: API (Application Programming Interface) Routing Settings. Multiple API routes are defined using FastAPI to handle different types of task requests, such as medical question answering, drug dosage calculation, and image description. Each route corresponds to a specific functional module, ensuring that requests can quickly locate the appropriate processing program. Upon backend startup, a pre-trained medical model is loaded, and based on the expert module division in the implementation steps, different parts of the medical model are assigned to corresponding intelligent agents. Simultaneously, the background task function of FastAPI is used to periodically check the update status of the medical model, ensuring that the model is always up-to-date. The logic judgment module integrates automatic analysis of the user's input, determining its medical domain and question type. Through predefined rules and keyword matching, combined with natural language processing technology, the intelligent agent to be invoked is determined. For example, if the user input involves drug dosage calculation, the medical knowledge intelligent agent and the symbolic logic intelligent agent are activated.

[0143] (2) Front-end Interaction Design: The front-end adopts a simple and intuitive question-and-answer dialogue interface. Users can enter medical questions or related descriptions in the input boxes. The interface displays historical dialogue records in real time and supports multi-turn dialogue, facilitating continuous interaction between users and the system. The front-end and back-end transmit data via HTTPS (Hypertext Transfer Protocol Secure) to ensure the security and privacy of user data. Simultaneously, user input data undergoes appropriate preprocessing and validation to prevent malicious attacks and the submission of invalid data.

[0144] (3) Front-end and back-end interaction design: The back-end logic judgment module adjusts the expert module used in real time based on the user's input. By calling the FastAPI API interface, the user's question is passed to the corresponding expert module combination for processing, improving operating efficiency and avoiding resource waste. It also supports the forwarding history function, allowing users to forward the history of multi-turn conversations to other users or save it locally. The back-end is responsible for storing the history, using a database for persistent data management to ensure the integrity and traceability of the history.

[0145] (4) Performance Optimization and Scaling Design: Leveraging the asynchronous nature of FastAPI, time-consuming operations such as medical large-scale model inference and data processing are handled asynchronously, improving system response speed and concurrency performance. When processing user requests, the backend can handle multiple tasks simultaneously, avoiding excessive user waiting time. To cope with high-concurrency user access, multiple FastAPI backend service instances are deployed, and traffic is distributed through a load balancer. When user requests increase, new service instances can be easily added, achieving horizontal scaling of the system.

[0146] As can be seen from the above embodiments, the medical big data model provided in this application adopts a modular, multi-expert collaborative architecture. The system deploying the medical big data model consists of a multimodal input parsing layer, an expert module fusion layer, a task-driven routing layer, and a unified output layer, supporting dynamic sparse routing and on-demand activation of intelligent agent experts. By integrating multi-agent expert and reasoning capabilities, the system can process multimodal medical data from radiological images, pathological slides, ophthalmological and dermatological images, and reselect multiple first-agent experts and multiple second-agent experts from preset intelligent agent experts to interface with unstructured text information such as clinical data, test reports, and medical literature, achieving cross-modal semantic understanding and knowledge fusion.

[0147] like Figure 4As shown, based on the above-described method embodiments, this application provides a medical consultation response generation device based on multimodal fusion, including: an intelligent agent expert selection module 201, a joint encoding module 202, an intelligent agent expert invocation module 203, and a verification module 204.

[0148] Further, in some embodiments of this application, the intelligent agent expert selection module 201 is used to receive multimodal data input by the user, and select multiple first intelligent agent experts for feature extraction and multiple second intelligent agent experts for reasoning analysis from a preset pool of intelligent agent experts based on the multimodal data; the joint encoding module 202 is used to extract medical data features and digital symbol features from the multimodal data through the first intelligent agent experts, and jointly encode the medical data features and digital symbol features through a latent attention caching mechanism to obtain a multimodal feature vector; the intelligent agent expert invocation module 203 is used to input the output of the first intelligent agent expert and the multimodal feature vector to the corresponding second intelligent agent expert, and invoke each second intelligent agent expert according to a preset invocation order to obtain the output of each second intelligent agent expert; wherein, the second intelligent agent experts are connected according to the preset invocation order; the verification module 204 is used to verify the consistency of medical knowledge and mathematical logic in the outputs of each first intelligent agent expert and each second intelligent agent expert respectively, and when all verifications pass, aggregate the outputs of each first intelligent agent expert and each second intelligent agent expert to generate a medical consultation response.

[0149] Furthermore, in some embodiments of this application, the agent expert selection module 201 is used to select multiple first agent experts for feature extraction and multiple second agent experts for inference analysis from a preset set of agent experts based on the multimodal data. This includes: the multimodal data includes text data and image data; the first agent experts include at least visual agent experts and text agent experts; the second agent experts include at least the text agent experts; when the image data contains DICOM images, the DICOM agent expert is selected as the first agent expert; user demand keywords are extracted from the text data, and the corresponding inference agent expert is selected as the second agent expert based on the user demand keywords.

[0150] Furthermore, in some embodiments of this application, the joint encoding module 202 is used to extract medical data features and digital symbol features from the multimodal data through the first intelligent agent expert, including: extracting medical image semantic features from the image data through the visual intelligent agent expert; extracting query semantic features from the text data through the text intelligent agent expert; when there is a DICOM image in the image data, extracting medical metadata features from the DICOM image through the DICOM intelligent agent expert; concatenating the medical image semantic features, the query semantic features, and the medical metadata features to obtain the medical data features; and extracting quantitative medical parameters from the medical data features to obtain the digital symbol features.

[0151] Further, in some embodiments of this application, the intelligent agent expert invocation module 203 is used to input the output of the first intelligent agent expert and the multimodal feature vector to the corresponding second intelligent agent expert, and to invoke each of the second intelligent agent experts according to a preset invocation order, including: the reasoning intelligent agent experts include: diagnostic intelligent agent experts and medical knowledge intelligent agent experts; inputting the multimodal feature vector to each of the reasoning intelligent agent experts among the plurality of second intelligent agent experts; if there is a diagnostic intelligent agent expert among the plurality of second intelligent agent experts, then the medical image semantic features are also input to the diagnostic intelligent agent expert; if there is a medical knowledge intelligent agent expert among the plurality of second intelligent agent experts and the image data contains a DICOM image, then the medical metadata features are also input to the medical knowledge intelligent agent expert; inputting the output of each of the reasoning intelligent agent experts among the plurality of second intelligent agent experts to the text intelligent agent expert to obtain a comprehensive processing result.

[0152] Furthermore, in some embodiments of this application, the verification module 204 is used to verify the consistency of medical knowledge and mathematical logic output by each of the first intelligent agent experts and each of the second intelligent agent experts, including: verifying the consistency of medical knowledge output by each of the first intelligent agent experts and each of the second intelligent agent experts through a preset medical knowledge graph; verifying the consistency of mathematical logic output by each of the first intelligent agent experts and each of the second intelligent agent experts through a preset symbolic logic engine; calculating the confidence level of the output by each of the first intelligent agent experts and each of the second intelligent agent experts, and verifying whether each confidence level meets a preset threshold.

[0153] Furthermore, in some embodiments of this application, the verification module 204, after verifying the consistency of the medical knowledge and mathematical logic output by each of the first intelligent agent experts and each of the second intelligent agent experts, further includes: if any verification fails, adding a backup intelligent agent expert to a preset intelligent agent expert list, and reselecting multiple first intelligent agent experts and multiple second intelligent agent experts from the preset intelligent agent experts list to regenerate the medical consultation response.

[0154] Furthermore, in some embodiments of this application, after aggregating the outputs of each of the first intelligent agent experts and each of the second intelligent agent experts to generate a medical consultation response, the method further includes: if the medical consultation response also needs to generate a medical image, then the medical image is generated through an improved convolutional neural network; wherein, the improved convolutional neural network is obtained by parsing the textual conditional constraints from the medical consultation response through a conditional diffusion model, and then embedding the textual conditional constraints into a preset U-Net convolutional neural network through an attention mechanism.

[0155] In summary, the medical consultation response generation device based on multimodal fusion provided in this application has the following advantages compared to the prior art: By using multimodal data, it rationally selects the intelligent agent experts used in this medical consultation response generation task, thereby avoiding invalid intelligent agent expert invocation operations and improving the computational efficiency and resource utilization of the large medical model; furthermore, it extracts medical data features and digital symbol features from the multimodal data respectively, and jointly encodes the medical data features and digital symbol features through a latent attention caching mechanism, realizing deep fusion of medical data and mathematical reasoning, improving the comprehensive performance of the large medical model in complex scenarios, and verifying the consistency between the medical knowledge and mathematical logic output by the second intelligent agent expert, ensuring the quality of the final generated medical consultation response; by sequentially calling the second intelligent agent experts and connecting them according to the calling order, the second intelligent agent expert called later can receive the output of the second intelligent agent expert called earlier, improving the degree of information sharing and the rationality of information transmission routes between different intelligent agent experts during information transmission, fully leveraging the capabilities of each second intelligent agent expert, and improving the quality of the final generated medical consultation response.

[0156] It is understood that the above-described device embodiments correspond to the method embodiments of this application, and can realize the medical consultation response generation method based on multimodal fusion provided by any of the above-described method embodiments of this application.

[0157] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can specifically be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0158] Based on the above embodiments of the medical consultation response generation method based on multimodal fusion, another embodiment of this application provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the medical consultation response generation method based on multimodal fusion of any embodiment of this application.

[0159] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete this application. The one or more module units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.

[0160] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0161] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.

[0162] Based on the above-described method embodiments, another embodiment of this application provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the medical consultation response generation method based on multimodal fusion as described in any of the above-described method embodiments of this application.

[0163] The modules / units integrated in the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

Claims

1. A method for generating medical consultation responses based on multimodal fusion, characterized in that, include: Receive multimodal data input by the user, and select multiple first intelligent agent experts for feature extraction and multiple second intelligent agent experts for reasoning analysis from a preset pool of intelligent agent experts based on the multimodal data. The first intelligent agent expert extracts medical data features and digital symbol features from the multimodal data, and uses a latent attention caching mechanism to jointly encode the medical data features and digital symbol features to obtain a multimodal feature vector. The output of the first intelligent agent expert and the multimodal feature vector are input to the corresponding second intelligent agent expert, and each second intelligent agent expert is called in a preset calling order to obtain the output of each second intelligent agent expert; wherein, the second intelligent agent experts are connected in the preset calling order. The consistency between the medical knowledge and mathematical logic outputs of each first intelligent agent expert and each second intelligent agent expert is verified separately. When all verifications pass, the outputs of each first intelligent agent expert and each second intelligent agent expert are aggregated to generate a medical consultation response.

2. The method for generating medical consultation responses based on multimodal fusion as described in claim 1, characterized in that, The step of selecting multiple first agent experts for feature extraction and multiple second agent experts for inference analysis from a preset pool of agent experts based on the multimodal data includes: The multimodal data includes: text data and image data; The first intelligent agent expert includes at least: a visual intelligent agent expert and a text intelligent agent expert; the second intelligent agent expert includes at least: the text intelligent agent expert. When the image data contains a DICOM image, the DICOM agent expert is used as the first agent expert. Extract user demand keywords from the text data, and select the corresponding reasoning agent expert as the second agent expert based on the user demand keywords.

3. The method for generating medical consultation responses based on multimodal fusion as described in claim 2, characterized in that, The extraction of medical data features and digital symbol features from the multimodal data by the first intelligent agent expert includes: The visual intelligent agent expert extracts medical image semantic features from the image data; The text-based intelligent agent expert extracts query semantic features from the text data; When a DICOM image is present in the image data, the DICOM intelligent agent expert extracts medical metadata features from the DICOM image. By concatenating the medical image semantic features, the query semantic features, and the medical metadata features, the medical data features are obtained. Quantitative medical parameters are extracted from the medical data features to obtain the digital symbol features.

4. The method for generating medical consultation responses based on multimodal fusion as described in claim 3, characterized in that, The step of inputting the output of the first intelligent agent expert and the multimodal feature vector to the corresponding second intelligent agent expert, and calling each second intelligent agent expert according to a preset calling order, includes: The reasoning agent experts include: diagnostic agent experts and medical knowledge agent experts; The multimodal feature vector is input to each of the inference agent experts among the plurality of second agent experts; If the diagnostic intelligent agent expert is among the plurality of second intelligent agent experts, then the semantic features of the medical image are input into the diagnostic intelligent agent expert. If the medical knowledge intelligent agent expert is among the plurality of second intelligent agent experts and the image data contains a DICOM image, then the medical metadata features are input to the medical knowledge intelligent agent expert. The outputs of each of the reasoning agent experts among the plurality of second agent experts are input into the text agent expert to obtain a comprehensive processing result.

5. The method for generating medical consultation responses based on multimodal fusion as described in claim 1, characterized in that, The step of verifying the consistency between the medical knowledge and mathematical logic output by each of the first and second intelligent agent experts includes: The consistency of medical knowledge output by each first intelligent agent expert and each second intelligent agent expert is verified by using a pre-set medical knowledge graph. The mathematical logic consistency of the outputs of each first intelligent agent expert and each second intelligent agent expert is verified by a preset symbolic logic engine. Calculate the confidence level of the outputs of each first agent expert and each second agent expert, and verify whether each confidence level meets a preset threshold.

6. The method for generating medical consultation responses based on multimodal fusion as described in claim 1, characterized in that, After verifying the consistency between the medical knowledge and mathematical logic output by each of the first and second intelligent agent experts, the process further includes: If any verification fails, a backup agent expert is added to the preset agent experts, and multiple first agent experts and multiple second agent experts are reselected from the preset agent experts to regenerate the medical consultation response.

7. A method for generating medical consultation responses based on multimodal fusion as described in any one of claims 1 to 6, characterized in that, After aggregating the outputs of each of the first and second intelligent agent experts to generate a medical consultation response, the process also includes: If the medical consultation response also needs to generate medical images, the medical images are generated through an improved convolutional neural network; wherein, the improved convolutional neural network is obtained by parsing the textual conditional constraints from the medical consultation response through a conditional diffusion model, and then embedding the textual conditional constraints into a preset U-Net convolutional neural network through an attention mechanism.

8. A medical consultation response generation device based on multimodal fusion, characterized in that, include: The module includes an agent expert selection module, a joint encoding module, an agent expert invocation module, and a verification module. The intelligent agent expert selection module is used to receive multimodal data input by the user, and select multiple first intelligent agent experts for feature extraction and multiple second intelligent agent experts for reasoning analysis from a preset pool of intelligent agent experts based on the multimodal data. The joint encoding module is used to extract medical data features and digital symbol features from the multimodal data through the first intelligent agent expert, and to jointly encode the medical data features and digital symbol features through a latent attention caching mechanism to obtain a multimodal feature vector. The intelligent agent expert invocation module is used to input the output of the first intelligent agent expert and the multimodal feature vector to the corresponding second intelligent agent expert, and invoke each second intelligent agent expert according to a preset invocation order to obtain the output of each second intelligent agent expert; wherein, the second intelligent agent experts are connected to each other according to the preset invocation order; The verification module is used to verify the consistency of the medical knowledge and mathematical logic output by each of the first intelligent agent experts and each of the second intelligent agent experts. When all verifications pass, the outputs of each of the first intelligent agent experts and each of the second intelligent agent experts are aggregated to generate a medical consultation response.

9. A medical consultation response generation device based on multimodal fusion as described in claim 8, characterized in that, The agent expert selection module is used to select, based on the multimodal data, multiple first agent experts for feature extraction and multiple second agent experts for inference analysis from a preset pool of agent experts, including: The first intelligent agent expert includes at least: a visual intelligent agent expert and a text intelligent agent expert; the second intelligent agent expert includes at least: the text intelligent agent expert. When the image data contains a DICOM image, the DICOM agent expert is used as the first agent expert. Extract user demand keywords from the text data, and select the corresponding reasoning agent expert as the second agent expert based on the user demand keywords.

10. A terminal device, characterized in that, The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements a method for generating medical consultation responses based on multimodal fusion as described in any one of claims 1 to 7.