Method and device for assisting inquiry based on multi-modal data, medium, equipment and product

By partitioning and multi-modal graphs of multiple characteristics of traditional Chinese medicine, and combining with network fusion models, the data island problem in traditional Chinese medicine diagnosis is solved, high-precision intelligent consultation results are achieved, and reliability suggestions are provided for traditional Chinese medicine diagnosis.

CN120376066APending Publication Date: 2025-07-25BEIJING YANHUANG SIAN CHAY MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510430601.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In traditional Chinese medicine diagnosis, independent analysis of tongue, face and pulse data lacks multimodal correlation modeling, resulting in a unilateralized syndrome judgment. It is difficult for existing models to model the complex topological relationship between tongue coating color, pulse waveform, and facial areas. The black box model cannot be combined with the traditional Chinese medicine's "holistic view" theory, and the diagnosis efficiency and accuracy are low.

Method used

By partitioning the multi-category features of the target object, a multi-modal graph is constructed, and a hybrid architectural model of the graph attention network and the graph convolution network is used to process the feature information, and the consultation results are generated, and correlation analysis is performed in combination with the organ generation and restraint relationship to improve the accuracy of the consultation.

Benefits of technology

The joint analysis of multiple types of characteristics is realized, the accuracy of consultation results is improved, and the auxiliary suggestions for reliability is provided to doctors, combining a dual diagnosis model of data-driven and traditional Chinese medicine theory verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376066A_ABST
    Figure CN120376066A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent inquiry, and particularly provides an auxiliary inquiry method and device based on multi-modal data, a medium, equipment and a product, and the method can comprise the steps: obtaining feature information corresponding to multiple types of features of a target object, and the multiple types of features comprise a tongue image, a face, a pulse image or a palm; partitioning each type of features in the plurality of types of features, and obtaining a plurality of feature point locations corresponding to each type of features; connecting a plurality of feature point locations corresponding to each type of features according to a viscera sinkage-restraint relationship to obtain a multi-modal graph; and inputting the multi-modal diagram and the feature information into a pre-trained network fusion model to obtain an inquiry result of the target object, the inquiry result including a syndrome reasoning result and a feature thermodynamic diagram. According to the embodiment of the invention, a reliable auxiliary inquiry result can be provided for the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent medical inquiries. Specifically, it relates to a method, device, medium, equipment, and product for assisting medical inquiries based on multi-modal data. Background Art

[0002] With the continuous development of intelligent traditional Chinese medicine, more and more users tend to use traditional Chinese medicine-related knowledge to analyze their own situations. However, due to the complexity of traditional Chinese medicine theory knowledge, usually traditional Chinese medicine practitioners or users rely on accumulating their own experience to conduct traditional Chinese medicine inquiries on the target object or themselves. However, the efficiency and accuracy of this way of medical inquiry cannot be guaranteed. Moreover, in traditional Chinese medicine diagnosis, tongue, facial, and pulse data are analyzed independently, lacking multi-modal correlation modeling, resulting in one-sided syndrome judgment.

[0003] Therefore, how to provide a technical solution for a method of assisting medical inquiries based on multi-modal data with relatively high accuracy has become a technical problem that urgently needs to be solved. Summary of the Invention

[0004] Some embodiments of this application aim to provide a method, device, medium, equipment, and product for assisting medical inquiries based on multi-modal data. Through the technical solutions of the embodiments of this application, intelligent medical inquiries can be performed on the target object, the accuracy of the medical inquiry results can be improved, and relatively reliable auxiliary suggestions can be provided for traditional Chinese medicine practitioners or the target object themselves.

[0005] In a first aspect, some embodiments of this application provide a method for assisting medical inquiries based on multi-modal data, including: obtaining feature information corresponding to multiple types of features of the target object, where the multiple types of features include: tongue image, face, pulse condition, or palm; partitioning each type of feature among the multiple types of features to obtain multiple feature points corresponding to each type of feature; connecting the multiple feature points corresponding to each type of feature according to the relationship of generation and restriction among zang-fu organs to obtain a multi-modal graph; inputting the multi-modal graph and the feature information into a pre-trained network fusion model to obtain the medical inquiry result of the target object, where the medical inquiry result includes: syndrome inference result and feature heat map, where the syndrome inference result represents the syndrome type and symptom information of the target object, and the feature heat map is used to display at least one feature point corresponding to the syndrome inference result.

[0006] Some embodiments of this application construct a multi-modal graph by partitioning multiple types of features of the target object and combining the relationship of generation and restriction among zang-fu organs, and then obtain the medical inquiry result of the target object by combining the network fusion model and the feature information. Some embodiments of this application can jointly analyze multiple types of features through the multi-modal graph constructed from the multi-modal feature data of the target object, and then realize intelligent medical inquiries in combination with the network fusion model, improving the accuracy of intelligent medical inquiries and providing reliable auxiliary reference suggestions for doctors and the target object.

[0007] In some embodiments, obtaining the feature information corresponding to multiple types of features of the target object includes: controlling a feature acquisition device to perform image acquisition on the multiple types of features of the target object to obtain the tongue image, facial image, pulse condition information, or palm image in the feature information.

[0008] In some embodiments of the present application, a feature acquisition device is used to perform feature acquisition on a target object to obtain corresponding feature information, providing data support for subsequent intelligent medical inquiries.

[0009] In some embodiments, connecting multiple feature points corresponding to each type of feature according to the generating and restraining relationships among zang-fu organs to obtain a multimodal graph includes: obtaining the internal organs in the human body associated with each feature point; using each feature point as a single node, and connecting each feature point according to the association relationships among the internal organs in the generating and restraining relationships among zang-fu organs to obtain the multimodal graph.

[0010] In some embodiments of the present application, the feature points are connected to construct a multimodal graph through the internal organs managed by each feature point and the generating and restraining relationships among zang-fu organs, providing data support for subsequent intelligent medical inquiries.

[0011] In some embodiments, the pre-trained network fusion model includes: a graph attention network module and a graph convolutional network module.

[0012] In some embodiments of the present application, different models are used to form a network fusion model, providing model support for intelligent medical inquiries.

[0013] In some embodiments, inputting the multimodal graph and the feature information into the pre-trained network fusion model to obtain the medical inquiry result of the target object includes: using the graph attention network module to analyze the feature information and assign corresponding edge weights to each edge in the multimodal graph; using the graph convolutional network module to generate the medical inquiry result according to the feature information and the corresponding edge weights assigned to each edge.

[0014] In some embodiments of the present application, the graph attention network module and the graph convolutional network module are used to process the input multimodal graph and feature information respectively to determine the medical inquiry result, realizing intelligent medical inquiries.

[0015] In some embodiments, before inputting the multimodal graph and the feature information into the pre-trained network fusion model, the method further includes: obtaining a training data set, where the training data set includes: multiple sample modal graphs corresponding to multiple types of features of multiple target objects and syndrome sample data corresponding to each sample modal graph in the multiple sample modal graphs; using the training data set to train an initial fusion model to obtain the network fusion model.

[0016] In some embodiments of the present application, the initial fusion model is trained with a training dataset to obtain a network fusion model, providing model support for intelligent medical consultation.

[0017] In a second aspect, some embodiments of the present application provide a device for assisting medical consultation based on multi-modal data, including: a feature acquisition module for acquiring feature information corresponding to multiple types of features of a target object, where the multiple types of features include: tongue image, face, pulse condition, or palm; a partitioning module for partitioning each type of feature among the multiple types of features to obtain multiple feature points corresponding to each type of feature; a graph construction module for connecting the multiple feature points corresponding to each type of feature according to the relationship of the generation and restriction of zang-fu organs to obtain a multi-modal graph; and a medical consultation prediction module for inputting the multi-modal graph and the feature information into a pre-trained network fusion model to obtain the medical consultation result of the target object, where the medical consultation result includes: a syndrome inference result and a feature heat map, where the syndrome inference result represents the syndrome type and symptom information of the target object, and the feature heat map is used to display at least one feature point corresponding to the syndrome inference result.

[0018] In a third aspect, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any one of the embodiments of the first aspect can be implemented.

[0019] In a fourth aspect, some embodiments of the present application provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the method described in any one of the embodiments of the first aspect can be implemented.

[0020] In a fifth aspect, some embodiments of the present application provide a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the method described in any one of the embodiments of the first aspect can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of some embodiments of the present application, the drawings required to be used in some embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a system diagram of assisting medical consultation based on multi-modal data provided by some embodiments of the present application;

[0023] Figure 2One of the method flowcharts for assisted medical interview based on multi-modal data provided by some embodiments of the present application;

[0024] Figure 3 Another method flowchart for assisted medical interview based on multi-modal data provided by some embodiments of the present application;

[0025] Figure 4 Block diagram of the device for assisted medical interview based on multi-modal data provided by some embodiments of the present application;

[0026] Figure 5 Schematic diagram of an electronic device provided by some embodiments of the present application. Detailed implementation manners

[0027] Next, the technical solutions in some embodiments of the present application will be described in conjunction with the accompanying drawings in some embodiments of the present application.

[0028] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for differential description and cannot be construed as indicating or implying relative importance.

[0029] In the related art, in traditional Chinese medicine diagnosis, tongue, face, and pulse data are analyzed independently, resulting in a situation of data islands and lacking multi-modal association modeling, which leads to one-sided syndrome judgment. Existing CNN / RNN models are difficult to model the complex topological relationships among tongue coating colors, pulse waveforms, and facial regions, and have great limitations in feature extraction. Moreover, existing black-box models cannot be combined with the traditional Chinese medicine theory of "holistic view" and are difficult to generate diagnostic bases that conform to doctors' cognition.

[0030] In view of this, some embodiments of the present application provide a method for assisted medical interview based on multi-modal data. In this method, multiple types of features of the target object can be partitioned, then a multi-modal graph is constructed with each feature point as a node, and finally a network fusion model is used to process the multi-modal graph and feature information to generate an interview result. Some embodiments of the present application combine the data of multiple types of features for correlation analysis, making the syndrome judgment comprehensive. Moreover, in this implementation process, the features are correlated in combination with the relationship of the generation and restriction of zang-fu organs, improving the accuracy of the interview result and providing reliable auxiliary suggestions for doctors.

[0031] Next, the overall composition structure of the system for assisted medical interview based on multi-modal data provided by some embodiments of the present application will be exemplarily elaborated. Figure 1

[0032] Figure 1 ​​As shown in the figure, some embodiments of the present application provide a system diagram for assisted medical consultation based on multimodal data. The system for assisted medical consultation based on multimodal data includes: a feature acquisition device 100 and a terminal device 200. Among them, the feature acquisition device 100 can be an integrated device that collects at least two features of the tongue, face, pulse, and palm. According to the voice prompt of the feature acquisition device 100, the target object can independently perform on-demand acquisition of the tongue, face, pulse, and palm. After the acquisition is completed, the feature acquisition device 100 can send the acquired feature information to the terminal device 200. After the terminal device 200 obtains the feature information, it is necessary to partition multiple types of features corresponding to the feature information to obtain multiple feature points; and then, in combination with the relationship of mutual generation and restriction among zang-fu organs, connect the multiple feature points to generate a multimodal map. Finally, input the multimodal map and the feature information into the network fusion model to output the medical consultation result.

[0033] In some embodiments of the present application, a trained network fusion model is pre-deployed in the terminal device 200. The terminal device 200 can be a computer device or a server device, etc., and the embodiments of the present application do not make specific limitations here.

[0034] In some other embodiments of the present application, if the terminal device 200 has the function of collecting the tongue, face, pulse, and palm of the target object, it may not be connected to the feature acquisition device 100 at this time. Specifically, it can be selected according to the actual situation, and the embodiments of the present application do not make specific limitations here.

[0035] The following Figure 2 exemplarily elaborates the implementation process of assisted medical consultation based on multimodal data executed by the terminal device 200 provided by some embodiments of the present application.

[0036] Please refer to the Figure 2 , Figure 2 which is a flowchart of a method for assisted medical consultation based on multimodal data provided by some embodiments of the present application. The method for assisted medical consultation based on multimodal data may include:

[0037] S210, obtaining feature information corresponding to multiple types of features of a target object, where the multiple types of features include: tongue image, face, pulse condition, or palm.

[0038] For example, in some embodiments of the present application, the terminal device 200 needs to obtain feature images (as a specific example of feature information) corresponding to at least two of the four types of features of the target object: tongue image, face, pulse condition, and palm (abbreviated as tongue, face, pulse, and palm).

[0039] In some embodiments of the present application, S210 may include: controlling the feature acquisition device to perform image acquisition on the multiple types of features of the target object to obtain the tongue image, face image, pulse condition information, or palm image in the feature information.

[0040] For example, in some embodiments of the present application, the tongue image camera on the feature acquisition device can obtain high-definition images of the front and underside of the tongue of the target object. After determining that the tongue image acquisition is completed, the camera and the facial thermal imager are started: the facial image and the facial temperature distribution are recorded, and this feature is related to the visceral functions. After confirming that the facial image acquisition is completed, the target object is prompted to place the wrist at the designated position, and the pulse sensor is started: the pulse information at different positions of the radial artery (i.e., the pulse condition information) is measured. In addition, the camera can be started to collect the palm image of the target object.

[0041] S220. Partition each type of feature in the multiple types of features to obtain multiple feature points corresponding to each type of feature.

[0042] For example, in some embodiments of the present application, the tongue surface pulse is taken as an example for illustration. The tongue image is divided into multiple functional areas, such as the tip of the tongue, the middle of the tongue, and the root of the tongue (as a specific example of the feature point). Correspondingly, the facial image corresponds to the five-zang organs area in traditional Chinese medicine, such as the forehead, the wing of the nose, and the cheeks (as a specific example of the feature point). Pulse condition: the cun, guan, and chi positions of the radial artery (as a specific example of the feature point). Each feature point is a node in the subsequent constructed multi-modal graph.

[0043] For another example, the tongue image can be divided into feature points such as the tongue body, the tongue coating, and the cracks (RGB + texture feature extraction). The facial nodes can include: the complexion and temperature of areas such as the forehead, the nose, and the cheeks (LAB color space + infrared thermal imaging). The pulse nodes can include: the wave speed and rhythm of floating, middle, and deep pulse taking (wavelet transform features). It should be noted that the partitioning method can be flexibly selected according to the actual scenario, and the embodiments of the present application are not limited thereto.

[0044] S230. Connect the multiple feature points corresponding to each type of feature according to the generating and restraining relationships among the zang-fu organs to obtain a multi-modal graph.

[0045] For example, in some embodiments of the present application, by taking each feature point as a node, the generating and restraining relationships among the zang-fu organs generate corresponding edges, thereby constructing a multi-modal graph. For example, based on the zang-fu organ - body surface mapping relationship in "Huangdi Neijing" (such as "the heart opens into the tongue" and "the liver shows on the eyes"), prior knowledge edges are constructed to obtain a multi-modal graph.

[0046] In some embodiments of the present application, S220 may include: obtaining the internal organs in the human body associated with each feature point; taking each feature point as a single node, and connecting each feature point according to the association relationship among the internal organs in the generating and restraining relationships of the zang-fu organs to obtain the multi-modal graph.

[0047] For example, in some embodiments of the present application, the tip of the tongue corresponds to the heart, the middle of the tongue corresponds to the spleen and stomach, and the root of the tongue corresponds to the kidney. Correspondingly, facial images correspond to the five internal organs areas in traditional Chinese medicine. For example, the forehead corresponds to the heart, the alae nasi corresponds to the spleen, and the cheeks correspond to the liver and lungs. Pulse condition: The cun, guan, and chi positions of the radial artery respectively correspond to the heart, liver, and kidney. Based on the generating and restraining relationships among the internal organs, connections are established for the above-mentioned characteristic points, such as heart-lung, spleen-liver, etc.

[0048] S240, input the multi-modal graph and the feature information into a pre-trained network fusion model to obtain the consultation result of the target object. The consultation result includes: a syndrome inference result and a feature heat map. The syndrome inference result characterizes the syndrome type and symptom information of the target object, and the feature heat map is used to display at least one characteristic point corresponding to the syndrome inference result. The pre-trained network fusion model includes: a graph attention network module and a graph convolutional network module.

[0049] For example, in some embodiments of the present application, the network fusion model is a hybrid architecture of GAT (i.e., the graph attention network module) and GCN (i.e., the graph convolutional network module). Input the multi-modal graph and the feature image constructed above into the network fusion model, and the corresponding consultation result can be output.

[0050] In some embodiments of the present application, S240 may include: using the graph attention network module to analyze the feature information and assign corresponding edge weights to each edge in the multi-modal graph; using the graph convolutional network module to generate the consultation result according to the feature information and the corresponding edge weights assigned to each edge.

[0051] For example, in some embodiments of the present application, GAT can capture the dynamic weights between nodes in the multi-modal graph and strengthen the important relationships between feature information. GCN can utilize the global structure information to extract deep features, so as to output a consultation result with high accuracy. Specifically, GAT respectively processes the original features of the tongue, face, and pulse (as a specific example of feature information). Based on the original features, dynamic edge weights are injected through edge_attr to realize the adaptive adjustment of the "tongue-pulse-face" association intensity. For example, edge weights are generated based on the multi-modal fusion features to replace the fixed adjacency matrix, adapting to the dynamic association characteristics of traditional Chinese medicine syndrome types. Calculate the weight coefficients of the tongue, face, and pulse: α i = softmax(W^T h i ), where h iLet \(X\) be the modal feature vectors and \(W\) be the learnable parameters. The GCN can select the feature information and regions corresponding to the higher values from the edge weights (i.e., the tongue surface veins), determine the symptom information of the target object, and identify the syndrome type to which the target object belongs, such as qi deficiency, blood stasis, etc. The edge weights obtained through graph attention show the key regions. For example, abnormal conditions at the tip of the tongue may be related to heart fire. At this time, the tip of the tongue can be marked on the feature heat map (as a specific example of a key region). For example, the key region of the tongue surface veins can be deduced by the Grad-CAM algorithm. In addition, edges with higher edge weights can be extracted from the graph network to obtain the syndrome reasoning results, such as "red tongue → liver fire hyperactivity → taut and rapid pulse".

[0052] In some embodiments of the present application, before performing S240, the method for auxiliary medical history taking based on multimodal data may further include: obtaining a training data set, where the training data set includes: multiple sample modal graphs corresponding to the multiple types of features of multiple target objects and syndrome sample data corresponding to each sample modal graph in the multiple sample modal graphs; using the training data set to train an initial fusion model to obtain the network fusion model.

[0053] For example, in some embodiments of the present application, a network fusion model is obtained by training an initial fusion model (such as an initial GAT+GCN). For example, first, the historical features of the target object are collected and sorted out to obtain multiple sample modal graphs and corresponding syndrome sample data (such as symptom information, syndrome type, the symptom association relationship between feature points and internal organs, etc.). The initial GAT+GCN is trained with the integrated training data set to obtain a network fusion model that meets the requirements. Specifically, the training process includes: an independent encoding layer: processing the historical features of the tongue, face, and pulse respectively (compatible with the distribution differences of different sensor data). Cross-modal GAT: injecting dynamic edge weights through edge_attr to adaptively adjust the association intensity of "tongue-pulse-face". The GCN includes TemporalGCNConv, that is, aggregating the historical features of the target object within a time window to capture the chronic evolution law of the tongue surface vein features; a window sliding mechanism: setting window_size = 3 to balance long-term and short-term dependencies and avoid the gradient problem of RNN-like models. It can be understood that the process of training this model is similar to the conventional training process, and the description is omitted here.

[0054] As can be seen from some embodiments of the present application above, the present application transforms the viscera-body surface mapping relationship (i.e., the viscera generating and restricting relationship) into the prior edges of the graph network, solving the problem of randomly initializing the edges of traditional GNNs. The adaptive weight assignment based on the attention mechanism GAT is superior to fixed-weight splicing (for example, the weight of pulse condition data is increased by 30% in cold syndrome). Through the heat map and the reasoning path, a dual diagnosis mode of "data-driven + theoretical verification" is realized, providing reliable reference opinions for doctors.

[0055] The following is an illustrative explanation of the process of determining syndrome types using the correlation between the face and the internal organs in traditional Chinese medicine.

[0056] The facial thermal imager uses non-contact infrared thermal imaging technology to obtain facial temperature distribution and associate it with the functions of the internal organs in traditional Chinese medicine. The entire process includes four core steps: data acquisition, image processing, feature extraction, and verification analysis:

[0057] 1) Data collection

[0058] Facial thermal imagers use infrared sensors to measure the temperature distribution on the skin surface to form a thermal map. Uncooled infrared detectors (such as microbolometers) are used, and the resolution is usually 320×240 or higher to ensure clear details; the temperature measurement accuracy is ±0.1°C to capture small temperature differences. Scanning method: The subject maintains a neutral expression and sits quietly in a constant temperature room (25°C±1°C) for 5 to 10 minutes to adapt to the ambient temperature. Shoot at a fixed distance (usually 30 to 50 cm), and perform a complete scan from the front + 45° side. Time series capture: Collect multiple frames of images per second, remove abnormal temperature points, and obtain steady-state thermal distribution.

[0059] 2) Image processing

[0060] Since individual baseline temperatures are different, the whole face temperature is first normalized to achieve temperature normalization. The normalization formula is as follows:

[0061]

[0062] Where T is the original temperature, T min , T max are the minimum and maximum temperatures of the individual.

[0063] Facial area division (based on TCM viscera theory): forehead area (heart, brain), nose tip / nose wings (spleen and stomach), left and right cheekbones (lungs), mandible (kidneys), temples (liver and gallbladder); color mapping: different temperature areas use pseudo-color (such as red = high temperature, blue = low temperature) for visualization analysis.

[0064] 3) Key feature extraction

[0065] Extract temperature features related to organ functions from thermal images. Average temperature: analyze the overall temperature level of a certain area (e.g., a high forehead temperature may correspond to "heart fire"). Temperature difference analysis: calculate the temperature difference between different areas: forehead-zygomatic bone (heart-lung), nose tip-zygomatic bone (spleen, stomach-lung); dynamic thermal map: compare the temperature changes of patients in different states (quiet vs. after drinking hot water) to analyze the body's regulatory ability.

[0066] 4) Result verification

[0067] Internal verification within an individual: The subject repeats the test (for example, once in the morning and once in the afternoon), and the consistency is calculated. By comparing the facial temperature distribution at rest and after exercise, verify whether the temperature response pattern is stable. Or conduct clinical verification: Select healthy and diseased populations (such as spleen-stomach deficiency-cold, heart-fire hyperactivity), statistically analyze the differences in heat maps, and evaluate the correlation between temperature characteristics and traditional Chinese medicine syndromes, so that the fusion model can output accurate medical consultation results.

[0068] The following combines the appended Figure 3 Exemplarily elaborate on the specific process of assisted medical consultation based on multi-modal data provided by some embodiments of the present application.

[0069] Please refer to the appended Figure 3 , Figure 3 which is a flowchart of a method for assisted medical consultation based on multi-modal data provided by some embodiments of the present application.

[0070] The above process will be elaborated below by way of example.

[0071] S310, Use a feature acquisition device to collect images of the multiple types of features of the target object to obtain feature information.

[0072] S320, Partition each type of feature among the multiple types of features to obtain multiple feature points corresponding to each type of feature.

[0073] S330, Obtain the internal organs in the human body associated with each feature point.

[0074] S340, Take each feature point as a single node, and connect each feature point according to the association relationship between internal organs in the relationship of generation and restriction of zang-fu organs to obtain a multi-modal graph.

[0075] S350, Input the multi-modal graph and feature information into a pre-trained network fusion model to obtain the medical consultation result of the target object.

[0076] It can be understood that the specific implementation process of S310 to S350 can refer to the method embodiments provided above. To avoid repetition, the detailed description is appropriately omitted here.

[0077] Please refer to Figure 4 , Figure 4 which shows a block diagram of the composition of a device for assisted medical consultation based on multi-modal data provided by some embodiments of the present application. It should be understood that this device for assisted medical consultation based on multi-modal data corresponds to the above method embodiments and can execute each step involved in the above method embodiments. The specific functions of this device for assisted medical consultation based on multi-modal data can be seen in the above description. To avoid repetition, the detailed description is appropriately omitted here.

[0078] Figure 4The device for assisting diagnosis based on multimodal data includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the device for assisting diagnosis based on multimodal data, and the device for assisting diagnosis based on multimodal data includes: a feature acquisition module 410, used to obtain feature information corresponding to multiple types of features of the target object, wherein the multiple types of features include: tongue image, face, pulse or palm; a partitioning module 420, used to partition each type of feature in the multiple types of features to obtain multiple feature points corresponding to each type of feature; a graph construction module 430, used to connect the multiple feature points corresponding to each type of feature according to the relationship between the viscera and organs to obtain a multimodal graph; a diagnosis prediction module 440, used to input the multimodal graph and the feature information into a pre-trained network fusion model to obtain the diagnosis result of the target object, wherein the diagnosis result includes: a syndrome reasoning result and a feature heat map, wherein the syndrome reasoning result represents the syndrome type and symptom information of the target object, and the feature heat map is used to display at least one feature point corresponding to the syndrome reasoning result.

[0079] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.

[0080] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the operations of the method corresponding to any of the above methods provided in the above embodiments.

[0081] Some embodiments of the present application further provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operations corresponding to any of the above methods provided in the above embodiments.

[0082] like Figure 5 As shown, some embodiments of the present application provide an electronic device 500, which includes: a memory 510, a processor 520, and a computer program stored in the memory 510 and executable on the processor 520, wherein the processor 520 can implement a method as described in any of the above embodiments when reading the program from the memory 510 through a bus 530 and executing the program.

[0083] Processor 520 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 520 can be a microprocessor.

[0084] The memory 510 can be used to store instructions executed by the processor 520 or data related to the instruction execution process. These instructions and / or data can include code for implementing some or all of the functions of one or more modules described in the embodiments of the present application. The processor 520 in the embodiments of the present disclosure can be used to execute the instructions in the memory 510 to implement the methods shown above. The memory 510 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.

[0085] The above are only the embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0086] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

[0087] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

Claims

1. A method for assisted medical history taking based on multi-modal data, characterized in that, Including: Obtaining feature information corresponding to multiple types of features of a target object, where the multiple types of features include: tongue image, face, pulse condition, or palm; Partitioning each type of feature among the multiple types of features to obtain multiple feature points corresponding to each type of feature; Connecting the multiple feature points corresponding to each type of feature according to the relationship of generation and restriction among zang-fu organs to obtain a multi-modal graph; Inputting the multi-modal graph and the feature information into a pre-trained network fusion model to obtain the medical consultation result of the target object, where the medical consultation result includes: syndrome reasoning result and feature heat map, where the syndrome reasoning result represents the syndrome type and symptom information of the target object, and the feature heat map is used to display at least one feature point corresponding to the syndrome reasoning result.

2. The method according to claim 1, wherein The obtaining feature information corresponding to multiple types of features of a target object includes: Controlling a feature acquisition device to perform image acquisition on the multiple types of features of the target object to obtain a tongue image, a face image, pulse condition information, or a palm image in the feature information.

3. The method according to claim 1 or 2, characterized in that The connecting the multiple feature points corresponding to each type of feature according to the relationship of generation and restriction among zang-fu organs to obtain a multi-modal graph includes: Obtaining the internal organs in the human body associated with each feature point; Taking each feature point as a single node and connecting each feature point according to the association relationship among internal organs in the relationship of generation and restriction among zang-fu organs to obtain the multi-modal graph.

4. The method according to claim 1 or 2, characterized in that, The pre-trained network fusion model includes: a graph attention network module and a graph convolutional network module.

5. The method according to claim 4, wherein The inputting the multi-modal graph and the feature information into a pre-trained network fusion model to obtain the medical consultation result of the target object includes: Analyzing the feature information by using the graph attention network module to assign corresponding edge weights to each edge in the multi-modal graph; Using the graph convolutional network module to generate the medical consultation result according to the feature information and the corresponding edge weights assigned to each edge.

6. The method according to claim 1 or 2, characterized in that, Before the inputting the multi-modal graph and the feature information into a pre-trained network fusion model, the method further includes: Obtaining a training data set, where the training data set includes: multiple sample modal graphs corresponding to the multiple types of features of multiple target objects and syndrome sample data corresponding to each sample modal graph in the multiple sample modal graphs; Training an initial fusion model by using the training data set to obtain the network fusion model.

7. A device for assisting medical inquiries based on multimodal data, characterized in that, Including: A feature acquisition module, configured to obtain feature information corresponding to multiple types of features of a target object, where the multiple types of features include: tongue image, face, pulse condition, or palm; A partitioning module, configured to partition each type of feature among the multiple types of features to obtain multiple feature points corresponding to each type of feature; A graph construction module, configured to connect the multiple feature points corresponding to each type of feature according to the relationship of generation and restriction among zang-fu organs to obtain a multi-modal graph; The inquiry prediction module is configured to input the multi-modal graph and the feature information into a pre-trained network fusion model to obtain the inquiry result of the target object, where the inquiry result includes: a syndrome inference result and a feature heat map, and the syndrome inference result represents the syndrome type and symptom information of the target object, and the feature heat map is used to display at least one feature point corresponding to the syndrome inference result.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, where the computer program, when run by a processor, executes the method according to any one of claims 1-6.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor, where the computer program, when run by the processor, executes the method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program, where the computer program, when run by a processor, executes the method according to any one of claims 1-6.

Citation Information

Cited By

  • AI health analysis method and system based on tongue image and pulse wave and related equipment

    CN121730769A

  • AI health analysis method and system based on tongue appearance and pulse wave and related equipment

    CN121730769B