Intelligent hospital guide method, model training method and device, equipment and storage medium

By combining the medical image data and text data, multimodal feature fusion technology is used to identify disease types and provide personalized guidance based on medical resource databases, the shortcomings of existing intelligent guidance technology in terms of accuracy and personalization are solved, and more efficient and accurate guidance services are achieved.

CN120089343APending Publication Date: 2025-06-03PING AN HEALTH INSURANCE CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510246086.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing intelligent diagnosis technology is not accurate when conducting initial diagnosis of patients, resulting in low accuracy of the guidance recommendations provided and the inability to provide personalized guidance services.

Method used

By combining the user's medical image data and medical text data, multimodal feature fusion technology is used to extract image features and text semantic features, input a preset disease recognition model to obtain disease type, and then the guide information is determined based on the medical resource database.

Benefits of technology

It improves the comprehensiveness and accuracy of the diagnosis and treatment, provides users with personalized diagnosis and treatment services, and enhances the accuracy of the patient's health status and the effectiveness of medical advice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089343A_ABST
    Figure CN120089343A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and medical information, and provides an intelligent hospital guide method, a model training method and device, equipment and a storage medium, and the intelligent hospital guide method comprises the steps: obtaining the doctor-seeing image data and the doctor-seeing text data of a user; based on a first feature extraction network, extracting image features corresponding to the doctor-seeing image data; based on a second feature extraction network, extracting text semantic features corresponding to the treatment text data; fusing the image features and the text semantic features to obtain multi-modal features; inputting the multi-modal features into a preset disease recognition model to obtain a disease type; and based on a preset medical resource database, determining hospital guide information according to the disease type. According to the method, by combining the doctor-seeing image data and the doctor-seeing text data of the user, the comprehensiveness and accuracy of doctor guide can be improved, and the corresponding personalized doctor guide service can be provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and medical information technology, and particularly to an intelligent diagnosis guidance method, a model training method, a device, a device, and a storage medium. Background Art

[0002] Intelligent diagnosis guidance is a technology that assists medical staff in guiding patients to seek medical treatment. After a patient inputs or selects symptoms, it can provide the patient with medical treatment suggestions. For example, it is recommended to go to the internal medicine consultation room for medical treatment.

[0003] Although the current intelligent diagnosis guidance technology provides preliminary diagnosis guidance services for patients by means of artificial intelligence and natural language processing technology, the accuracy is not high when making a preliminary diagnosis of patients, resulting in low accuracy of the provided diagnosis guidance suggestions and the inability to provide personalized diagnosis guidance services. Summary of the Invention

[0004] The main purpose of this application is to provide an intelligent diagnosis guidance method, a model training method, a device, a device, and a storage medium. By combining the user's medical treatment image data and medical treatment text data, it can not only improve the comprehensiveness and accuracy of diagnosis guidance, but also provide corresponding personalized diagnosis guidance services for users.

[0005] In a first aspect, this application provides an intelligent diagnosis guidance method, including:

[0006] Obtain the user's medical treatment image data and medical treatment text data;

[0007] Based on a first feature extraction network, extract the image features corresponding to the medical treatment image data;

[0008] Based on a second feature extraction network, extract the text semantic features corresponding to the medical treatment text data;

[0009] Fuse the image features and the text semantic features to obtain multi-modal features;

[0010] Input the multi-modal features into a preset disease recognition model to obtain the disease type;

[0011] Based on a preset medical resource database, determine the diagnosis guidance information according to the disease type.

[0012] In a second aspect, this application also provides a training method for a disease recognition model, including:

[0013] Obtain a plurality of training samples, where the training samples include medical treatment image data and medical treatment text data, and the disease type labels corresponding to the medical treatment image data and the medical treatment text data;

[0014] Extract the image features corresponding to the medical image data based on the first feature extraction network;

[0015] Extract the text semantic features corresponding to the medical text data based on the second feature extraction network;

[0016] Fuse the image features and the text semantic features to obtain multimodal features;

[0017] Input the multimodal features into the disease recognition model to obtain the disease type;

[0018] Based on a preset loss function, determine the loss value according to the disease type and the disease type label;

[0019] Adjust the model parameters of the disease recognition model according to the loss value.

[0020] In a third aspect, the present application also provides an intelligent medical guidance device, including:

[0021] A first acquisition module, configured to acquire the medical image data and medical text data of a user;

[0022] A first feature extraction module, configured to extract the image features corresponding to the medical image data based on the first feature extraction network;

[0023] A second feature extraction module, configured to extract the text semantic features corresponding to the medical text data based on the second feature extraction network;

[0024] A feature fusion module, configured to fuse the image features and the text semantic features to obtain multimodal features;

[0025] An identification module, configured to input the multimodal features into a preset disease recognition model to obtain the disease type;

[0026] A medical guidance module, configured to determine medical guidance information according to the disease type based on a preset medical resource database.

[0027] In a fourth aspect, the present application also provides a computer device, where the computer device includes a memory and a processor;

[0028] The memory is used to store a computer program;

[0029] The processor is configured to execute the computer program and, when executing the computer program, implement the intelligent medical guidance method and the training method of the disease recognition model as described above.

[0030] Fifth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the intelligent diagnosis guidance method and the steps of the training method of the disease recognition model as described above are implemented.

[0031] The present application provides an intelligent diagnosis guidance method, a model training method, a device, a device and a storage medium. Among them, the intelligent diagnosis guidance method includes: obtaining the medical image data and medical text data of a user; extracting the image features corresponding to the medical image data based on a first feature extraction network; extracting the text semantic features corresponding to the medical text data based on a second feature extraction network; fusing the image features and the text semantic features to obtain multi-modal features; inputting the multi-modal features into a preset disease recognition model to obtain a disease type; and determining diagnosis guidance information according to the disease type based on a preset medical resource database. The present application inputs the multi-modal features that integrate the medical image data and medical text data of the user into a preset disease recognition model to obtain the corresponding disease type; and then determines the diagnosis guidance information corresponding to the disease type of the user based on the preset medical resource database; by combining the medical image data and medical text data of the user, the comprehensiveness and accuracy of the diagnosis guidance can be improved, and corresponding personalized diagnosis guidance services can be provided for the user. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 It is a schematic flowchart of an intelligent diagnosis guidance method provided by an embodiment of the present application;

[0034] Figure 2 It is a schematic connection diagram of a server and a terminal device provided by an embodiment of the present application;

[0035] Figure 3 It is a schematic framework diagram of intelligent diagnosis guidance provided by an embodiment of the present application;

[0036] Figure 4 It is a schematic flowchart of a training method of a disease recognition model provided by an embodiment of the present application;

[0037] Figure 5 It is a schematic block diagram of an intelligent diagnosis guidance device provided by an embodiment of the present application;

[0038] Figure 6A schematic block diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0040] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.

[0041] An embodiment of the present application provides an intelligent medical guidance method, a model training method, a device, a device, and a storage medium. Among them, the intelligent medical guidance method can be applied to a terminal device, and the terminal device can be a device such as a mobile phone, a tablet computer, a notebook computer, or a desktop computer. It can also be applied to a server, and the server can be a single server or a server cluster, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0042] Next, some implementation manners of the present application will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0043] Please refer to Figure 1 , Figure 1 A schematic flowchart of an intelligent medical guidance method provided by an embodiment of the present application. It should be noted that the intelligent medical guidance method provided by the embodiment of the present application can be used for a terminal device, and of course, it can also be used for a server.

[0044] As Figure 2 shown, the intelligent medical guidance method is applied to a server, and the server is communicatively connected to a terminal device. Through the server, the medical guidance information obtained by the intelligent medical guidance method can be sent to the terminal device. Of course, it is not limited to this, and no limitation is made here.

[0045] In specific implementation, the terminal device includes, but is not limited to, any one of a mobile phone, a tablet computer, a notebook computer, and a desktop computer; the server can be a single server, a server cluster, or a cloud server providing cloud computing services.

[0046] As Figure 1 shown, the intelligent diagnosis guidance method includes steps S101 to S106.

[0047] Step S101: Obtain the user's medical image data and medical text data.

[0048] In the embodiments of the present application, both the medical image data and the medical text data can include the image data and medical data corresponding to traditional Chinese medicine, as well as the image data and medical data corresponding to Western medicine. For example, the image data corresponding to traditional Chinese medicine can include the user's tongue image, facial image and other medical image data; the medical text data corresponding to traditional Chinese medicine can include the symptom content described by the user related to traditional Chinese medicine, such as dreaminess, dry mouth, bitter taste, etc. The image data corresponding to Western medicine can include the user's X-ray image, CT image, abdominal color Doppler ultrasound image and other medical images; the medical data corresponding to Western medicine can include the symptom content described by the user related to Western medicine, such as body temperature, runny nose, chest tightness, etc. By obtaining the user's medical image data and medical text data in multiple aspects, more accurate diagnosis guidance services can be provided for the user.

[0049] As Figure 3 shown, the embodiments of the present application can obtain the user's medical image data and medical text data uploaded through the human-computer interaction interface. Furthermore, the embodiments of the present application can also set a speech-to-text window on the human-computer interaction interface, so that when the user uses the speech-to-text function, they can confirm whether the symptom content they describe is accurate before uploading, improving the accuracy of data acquisition. And by providing a visual interface for the user to operate, it can not only improve the user experience, but also improve the intelligence of diagnosis guidance.

[0050] Step S102: Based on the first feature extraction network, extract the image features corresponding to the medical image data.

[0051] It can be understood that after obtaining the user's medical image data in the embodiments of the present application, it is necessary to preprocess the medical image data first, for example, reduce or eliminate Gaussian noise in the image; adjust the image size, contrast and saturation; normalize, etc., to improve the clarity of the image and reduce the computational complexity of subsequent processing.

[0052] For example, the embodiments of the present application can use a convolutional neural network to extract the image features corresponding to the medical image data. It should be noted that a convolutional neural network is a deep learning model for processing image data, which can include a convolutional layer, a pooling layer and a fully connected layer. Through the convolutional layer, local features such as edges and textures of the medical image data can be extracted; through the pooling layer, the dimension of the local features corresponding to the medical image data can be reduced to reduce the computational complexity; through the fully connected layer, the extracted image features can be used for disease classification.

[0053] Specifically, the preprocessed medical visit image data (such as tongue image, facial image, X-ray image, CT image, etc.) is input into the convolutional layer, and local features of the medical visit image data are extracted through multiple convolutional layers. For example, the first-layer convolutional kernel can extract low-level features such as edges and textures of the medical visit image data, and the deep convolutional kernel can extract more advanced features, such as the lesion area. The extracted image features are input into the pooling layer to reduce the dimension of the image features using max pooling or average pooling. Finally, the image features extracted by the convolutional layer are flattened and then input into the fully connected layer.

[0054] Taking the extraction of the tongue image as an example, the tongue coating color, tongue body shape, etc. in the tongue image can be extracted through a convolutional neural network. Taking the facial image as an example, the eyes, lip color, skin condition, etc. in the facial image can be extracted through a convolutional neural network. Taking the X-ray image or CT image as an example, the lung shadow and fracture line in the X-ray image or CT image can be extracted through a convolutional neural network.

[0055] Step S103: Based on the second feature extraction network, extract the text semantic features corresponding to the medical visit text data.

[0056] Exemplarily, after obtaining the medical visit text data input by the user, it is necessary to first perform preprocessing such as cleaning, word segmentation, stop word removal, and part-of-speech tagging on the medical visit text data, so as to more accurately extract the key information about the specific symptoms of the user in the medical visit text data, such as the pain time, degree, etc.

[0057] For example, the embodiments of the present application can extract the text semantic features corresponding to the medical visit text data based on a Transformer model (such as the BERT pre-trained model). Specifically, the words in the vocabulary obtained after preprocessing the medical visit text data (that is, each word corresponds to a unique index) are input into the word embedding layer of the model to convert the words into dense vectors; then, the dense vectors are encoded into text semantic features through the encoding layer of the model; then, the encoded text semantic features are subjected to max pooling or average pooling through the pooling layer to obtain text semantic features of a fixed length.

[0058] Step S104: Fuse the image features and the text semantic features to obtain multi-modal features.

[0059] It can be understood that the methods for fusing the image features and the text semantic features can include concatenation fusion, weighted fusion, multiplication fusion, addition fusion, tensor fusion, and attention mechanism fusion, etc.

[0060] Among them, splicing and fusion refers to directly splicing the image features and the text semantic features into a new feature vector. Weighted fusion refers to assigning different weights according to the importance of each feature, and then adding the weighted features. Multiplication fusion refers to multiplying the pixels at the corresponding positions of the feature maps of different modalities to obtain a comprehensive feature map. Addition fusion refers to adding the pixels at the corresponding positions of the feature maps of different modalities to obtain a comprehensive feature map. Tensor fusion refers to converting the feature maps of different modalities into tensors, and then using operations such as tensor multiplication and dot product for fusion. The attention mechanism refers to weighting the feature maps corresponding to the image features and the feature maps corresponding to the text semantic features through the attention mechanism, and assigning a greater weight to the feature map with a larger weight, so as to more effectively fuse the features of different modalities.

[0061] In the embodiment of the present application, by combining the image features corresponding to the medical visit image data and the text semantic features corresponding to the medical visit text data, a richer and more comprehensive feature representation can be obtained, thereby improving the recognition and understanding of the model.

[0062] Step S105: Input the multi-modal features into a preset disease recognition model to obtain the disease type.

[0063] Exemplarily, the disease recognition model in the embodiment of the present application can be a multi-modal deep learning model, which can perform disease recognition based on multi-modal data. For example, the multi-modal deep learning model can fuse tongue image, facial image, X-ray image and the text content of the user's described symptoms to identify the type of lung disease (e.g., emphysema - 90%, pulmonary tuberculosis - 10%); it can also fuse electrocardiogram, facial image and text content to identify the type of heart disease (e.g., coronary heart disease - 85%).

[0064] It can be understood that the embodiment of the present application can input the multi-modal features obtained by fusing different types of medical visit image data and medical visit text data into a preset disease recognition model to comprehensively analyze the user's health status through the disease recognition model; in addition, by combining the medical visit image data in traditional Chinese medicine (such as tongue image) and the medical visit image data in Western medicine (such as X-ray image, electrocardiogram, etc.), the accuracy of disease recognition can be improved and more accurate medical advice can be provided for the user.

[0065] Step S106: Based on a preset medical resource database, determine the guiding diagnosis information according to the disease type.

[0066] Exemplarily, the medical resource database in the embodiments of the present application may include the user's attribute information (name, age, address, contact information, etc.), historical medical treatment data, medical institution information (basic information of traditional Chinese medicine and Western medicine hospitals and clinics), medical equipment (information such as equipment name, model, quantity, usage status, maintenance records, etc.), medical staff (work experience, professional qualifications, etc. of traditional Chinese medicine and Western medicine doctors and nurses), and drug resources (traditional Chinese medicine and Western medicine). Among them, as Figure 3 shown, if it is a new user, the attribute information and historical medical treatment data input by the new user can be obtained through the human-computer interaction interface.

[0067] Specifically, after the disease recognition model in the embodiments of the present application recognizes the disease type, it can, based on the preset medical resource database, integrate data in the aspects of traditional Chinese medicine and Western medicine, and accurately provide the user with referral information such as professional hospitals and doctors for treating this disease type.

[0068] In another embodiment, the medical resource database in the embodiments of the present application may include guidance data on body conditioning published by well-known traditional Chinese medicine doctors and Western medicine doctors in different specialties. Based on this, a personal health management plan can also be provided for the user based on the preset medical resource database.

[0069] The intelligent referral method provided in the above embodiments includes: obtaining the user's medical treatment image data and medical treatment text data; extracting the image features corresponding to the medical treatment image data based on the first feature extraction network; extracting the text semantic features corresponding to the medical treatment text data based on the second feature extraction network; fusing the image features and the text semantic features to obtain multi-modal features; inputting the multi-modal features into a preset disease recognition model to obtain the disease type; and determining the referral information based on the preset medical resource database according to the disease type. The present application inputs the multi-modal features that integrate the user's medical treatment image data and medical treatment text data into a preset disease recognition model to obtain the corresponding disease type; and then, based on the preset medical resource database, determines the referral information corresponding to the user's disease type; by combining the user's medical treatment image data and medical treatment text data, it can not only improve the comprehensiveness and accuracy of the referral, but also provide the user with corresponding personalized referral services.

[0070] In an exemplary implementation manner, step S104 may include steps S1041 to S1043.

[0071] Step S1041: Align the dimensions of the image features and the text semantic features based on the fully connected layer.

[0072] Step S1042: Normalize the image features and the text semantic features.

[0073] Step S1043: Based on the attention mechanism, fuse the image features and text semantic features to obtain multi-modal features.

[0074] Specifically, in the embodiments of the present application, the image features and text semantic features can be first mapped to a space of the same dimension based on a fully connected layer to align the dimensions of the image features and text semantic features; then, the image features and text semantic features are normalized to eliminate the differences between different features; then, based on the attention mechanism, the weights of the image features and text semantic features are dynamically adjusted. For example, let the text semantic features focus on the important regions in the image features, or let the image features focus on the keywords in the text semantic features, so as to capture the fine-grained interaction information between the two modalities of image features and text semantic features, thereby obtaining richer and more accurate multi-modal features.

[0075] In an exemplary embodiment, the medical image data includes tongue image and facial image, and the medical text data includes first text data; step S105 may specifically include: inputting the multi-modal features corresponding to the tongue image, facial image, and first text data into a preset disease recognition model to obtain the traditional Chinese medicine disease type.

[0076] It can be understood that the first text data in the embodiments of the present application includes text data related to traditional Chinese medicine, such as dreaminess, dry mouth, bitter taste in the mouth, sallow complexion, etc. Inputting the tongue image, facial image, and text data related to traditional Chinese medicine into a preset disease recognition model (i.e., a multi-modal deep learning model) can obtain the traditional Chinese medicine disease type. For example, if the image features corresponding to the tongue image and facial image uploaded by the user are redder tongue coating, thick yellow tongue coating, and multiple red marks on the face; the text semantic features corresponding to the first text data include recent cough, thick yellow sputum, and painful swallowing; the traditional Chinese medicine disease type, such as wind-heat cold, can be identified through the disease recognition model.

[0077] In another embodiment, after obtaining the gastroscope image uploaded by the user and the second text content related to Western medicine input by the user (such as upper abdominal pain for several days, vomiting, etc.); after extracting the image features corresponding to the gastroscope image (such as ulcer, polyp, etc.) and the text semantic features corresponding to the second text content, the Western medicine disease type, such as chronic gastritis, can be identified through the disease recognition model.

[0078] Through the disease recognition model in the embodiments of the present application, the traditional Chinese medicine disease type or Western medicine disease type can be identified according to the obtained medical image data and medical text data, so that a more comprehensive and accurate disease type can be identified.

[0079] In an exemplary embodiment, the disease type includes traditional Chinese medicine disease type and Western medicine disease type; step S106 may include step S601A to step S603A.

[0080] Step S601A: Based on a preset medical resource database, establish the corresponding relationship between traditional Chinese medicine (TCM) disease types and Western medicine (WM) disease types.

[0081] Step S602A: Determine at least one disease according to the corresponding relationship between TCM disease types and WM disease types.

[0082] Step S603A: Determine the triage information according to the priority levels corresponding to at least one disease.

[0083] Exemplarily, the medical resource database preset in the embodiments of the present application stores medical data related to the field of traditional Chinese medicine and medical data related to the field of Western medicine, and establishes the corresponding relationship between TCM disease types and WM disease types. For example, TCM disease types may include wind-cold common cold, wind-heat common cold, spleen-stomach weakness, lung yin deficiency, etc. WM disease types may include cold, chronic gastritis, irritable bowel syndrome, functional dyspepsia, pneumonia, chronic bronchitis, etc. Among them, wind-cold common cold and wind-heat common cold in traditional Chinese medicine generally correspond to cold in Western medicine. Therefore, the corresponding relationship between wind-cold common cold, wind-heat common cold and cold can be established; spleen-stomach weakness in traditional Chinese medicine generally corresponds to chronic gastritis, irritable bowel syndrome, functional dyspepsia, etc. in Western medicine. Therefore, the corresponding relationship between spleen-stomach weakness and chronic gastritis, irritable bowel syndrome and functional dyspepsia can be established; lung yin deficiency in traditional Chinese medicine generally corresponds to pneumonia, chronic bronchitis, etc. in Western medicine. Therefore, the corresponding relationship between lung yin deficiency and pneumonia, chronic bronchitis can be established.

[0084] The embodiments of the present application can determine at least one disease according to the corresponding relationship between TCM disease types and WM disease types, so as to more accurately determine the triage information corresponding to the disease according to the priority level of the disease, such as the priority level corresponding to the severity of the disease, so as to provide a more intelligent and comprehensive triage service for users. Specifically, if the identified TCM disease types are spleen-stomach weakness and lung yin deficiency, and the WM disease types are chronic gastritis, functional dyspepsia and chronic bronchitis, it can be determined that the user has dyspepsia and chronic bronchitis, and chronic bronchitis is more severe than functional dyspepsia. Therefore, the triage information corresponding to chronic bronchitis can be determined first, so that the user can receive professional treatment as soon as possible.

[0085] In an exemplary embodiment, the medical resource database includes a doctor evaluation database of each hospital; step S106 may include steps S601B to S603B.

[0086] Step S601B: Obtain the attribute information and historical medical record data of the user.

[0087] Step S602B: Determine the first target hospital according to the disease type, attribute information and historical medical record data.

[0088] Step S603B: Based on the doctor evaluation database, determine a first target doctor from the first target hospital, where the score of the first target doctor is higher than the scores of other doctors in the department corresponding to the disease type in the first target hospital.

[0089] Exemplarily, the medical resource data in the embodiments of the present application may include doctor evaluation databases of each hospital. For example, the doctor evaluation database may include the evaluation scores of doctors in each department of each hospital. Among them, the evaluation criteria for the evaluation scores may include years of practice, professional skills, professional titles, and comprehensiveness (such as patient scores, etc.).

[0090] It can be understood that for a new user, it is necessary to obtain the attribute information (name, gender, age, address, etc.) and historical medical treatment data input by the new user on the human-computer interaction interface; for an existing user, the attribute information and historical medical treatment data of the user can be directly called. The embodiments of the present application can screen out the hospitals within a preset range from the user's address according to the obtained attribute information and historical medical treatment data of the user; then screen out the first target hospital that has been relatively successful in treating this disease type from these multiple hospitals, and then, based on the doctor evaluation database, query the scores of each doctor in the department corresponding to this disease type in the first target hospital, so as to recommend the first target doctor with a higher score to the user. For example, it is recognized that the user's disease type is thyroid disease (such as hyperthyroidism), and further screening is performed among multiple hospitals within 5 km or 10 km from the user's address to screen out the first target hospital that has been relatively successful in treating thyroid disease, and query the scores of each doctor in the endocrinology department of the first target hospital that treats thyroid disease, so as to recommend the first target doctor with a higher score to the user. The embodiments of the present application can recommend suitable hospitals and doctors to the user according to the specific situation of the user, so as to realize personalized medical guidance services.

[0091] In an exemplary embodiment, the method may further include step B1 and step B2.

[0092] Step B1: Obtain the medical treatment feedback data and medical treatment evaluation data of the user.

[0093] Step B2: Determine a second target hospital and a second target doctor according to the medical treatment feedback data and medical treatment evaluation data.

[0094] It can be understood that the medical treatment feedback data in the embodiments of the present application may include the user's evaluation of the accuracy of the medical guidance service (identified disease type, recommended first target hospital, first target doctor). The medical treatment evaluation data may include the scores of the user for the first target hospital and the first target doctor visited.

[0095] Exemplarily, the embodiments of the present application can dynamically optimize the navigation system based on the intelligent navigation method according to the user's medical treatment feedback data and medical treatment evaluation data. Specifically, the recognition accuracy of disease types can be improved according to the user's accuracy evaluation of the navigation service. For example, the user can be guided to input more text data about symptom descriptions in the form of questions. The second target hospital and the second target doctor can also be determined according to the user's medical treatment evaluation data.

[0096] It can be understood that the second target hospital and the first target hospital may be the same or different; the second target doctor and the first target doctor may be the same or different. For example, if the user gives a high score to the previously recommended first target hospital and / or the first target doctor, the user can continue to be recommended. If the user gives a low score to the previously recommended first target hospital and / or the first target doctor, another hospital can be recommended to the user, or other doctors in the same hospital and department can be recommended.

[0097] The embodiments of the present application improve the accuracy of recognizing disease types and adjust the navigation information suitable for users according to the user's medical treatment feedback data and medical treatment evaluation data, which can further improve the user experience.

[0098] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of the training method of the disease recognition model provided by the embodiments of the present application. The training method of the disease recognition model includes steps S201 to S207.

[0099] Step S201: Obtain a plurality of training samples, where the training samples include medical treatment image data and medical treatment text data, and the disease type labels corresponding to the medical treatment image data and the medical treatment text data.

[0100] It can be understood that the embodiments of the present application can collect a large number of training samples on the premise of obtaining legal authorization and not involving user privacy. Each training sample can include medical treatment image data and medical treatment text data, and the disease type labels corresponding to the medical treatment image data and the medical treatment text data.

[0101] Step S202: Based on the first feature extraction network, extract the image features corresponding to the medical treatment image data.

[0102] Step S203: Based on the second feature extraction network, extract the text semantic features corresponding to the medical treatment text data.

[0103] Step S204: Fuse the image features and the text semantic features to obtain multi-modal features.

[0104] Step S205: Input the multi-modal features into the disease recognition model to obtain the disease type.

[0105] It should be noted that the relevant discussions in steps S202 to S205 can refer to the relevant embodiments of the foregoing intelligent medical guidance method, which will not be elaborated here.

[0106] Step S206: Based on a preset loss function, determine a loss value according to the disease type and the disease type label.

[0107] Step S207: Adjust the model parameters of the disease recognition model according to the loss value.

[0108] It can be understood that the preset loss function is mainly in the training stage of the disease recognition model. After each batch of training samples are input into the disease recognition model, the predicted value (i.e., the recognized disease type) can be output through forward propagation, and then the difference value between the predicted value and the true value (i.e., the disease type label), that is, the loss value, can be calculated according to the loss function. After obtaining the loss value, the disease recognition model can update each model parameter through backpropagation to reduce the loss between the true value and the predicted value, so that the predicted value output by the disease recognition model approaches the true value, thereby achieving the purpose of learning.

[0109] Please refer to Figure 5 , Figure 5 , which is a schematic block diagram of an intelligent medical guidance device provided by an embodiment of the present application. This intelligent medical guidance device can be configured in a server or a terminal device and is used to execute the foregoing intelligent medical guidance method.

[0110] As Figure 5 shown, the intelligent medical guidance device includes: a first acquisition module 110, a first feature extraction module 120, a second feature extraction module 130, a feature fusion module 140, an identification module 150, and a medical guidance module 160.

[0111] The first acquisition module 110 is used to acquire the user's medical visit image data and medical visit text data.

[0112] The first feature extraction module 120 is used to extract the image features corresponding to the medical visit image data based on the first feature extraction network.

[0113] The second feature extraction module 130 is used to extract the text semantic features corresponding to the medical visit text data based on the second feature extraction network.

[0114] The feature fusion module 140 is used to fuse the image features and the text semantic features to obtain multi-modal features.

[0115] The identification module 150 is used to input the multi-modal features into a preset disease recognition model to obtain the disease type.

[0116] The triage module 160 is used to determine triage information based on a preset medical resource database according to the disease type.

[0117] In an exemplary embodiment, the visit image data includes tongue image and facial image, and the visit text data includes the first text data; the recognition module 150 can specifically be used to input the multi-modal features corresponding to the tongue image, facial image, and the first text data into a preset disease recognition model to obtain the traditional Chinese medicine disease type.

[0118] In an exemplary embodiment, the disease type includes traditional Chinese medicine disease type and Western medicine disease type; the triage module 160 can include a relationship establishment sub-module, a disease determination sub-module, and a triage determination sub-module.

[0119] The relationship establishment sub-module is used to establish the corresponding relationship between the traditional Chinese medicine disease type and the Western medicine disease type based on a preset medical resource database.

[0120] The disease determination sub-module is used to determine at least one disease according to the corresponding relationship between the traditional Chinese medicine disease type and the Western medicine disease type.

[0121] The triage determination sub-module is used to determine triage information according to the priority corresponding to at least one disease.

[0122] In an exemplary embodiment, the medical resource database includes the doctor evaluation database of each hospital; the triage module 160 can include a first acquisition sub-module, a first target determination sub-module, and a second target determination sub-module.

[0123] The first acquisition sub-module is used to acquire the attribute information and historical visit data of the user.

[0124] The first target determination sub-module is used to determine the first target hospital according to the disease type, attribute information, and historical visit data.

[0125] The second target determination sub-module is used to determine the first target doctor from the first target hospital based on the doctor evaluation database, and the score of the first target doctor is higher than the scores of other doctors in the department corresponding to the disease type in the first target hospital.

[0126] In an exemplary embodiment, the device can further include a second acquisition module and a third target determination sub-module.

[0127] The second acquisition module is used to acquire the visit feedback data and visit evaluation data of the user.

[0128] The third target determination sub-module is used to determine the second target hospital and the second target doctor according to the visit feedback data and visit evaluation data.

[0129] In an exemplary embodiment, the feature fusion module 140 may include a dimension alignment sub-module, a normalization sub-module, and a fusion sub-module.

[0130] The dimension alignment sub-module is configured to perform dimension alignment on the image features and the text semantic features based on a fully connected layer.

[0131] The normalization sub-module is configured to normalize the image features and the text semantic features.

[0132] The fusion sub-module is configured to fuse the image features and the text semantic features based on an attention mechanism to obtain multi-modal features.

[0133] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0134] The method of the present application can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0135] Exemplarily, the above method and device can be implemented in the form of a computer program, and the computer program can run on a computer device.

[0136] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a server or a terminal device.

[0137] As Figure 6 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a storage medium and an internal memory.

[0138] The storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to execute the steps of any intelligent diagnosis guidance method and the steps of the training method of the disease recognition model.

[0139] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0140] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, it can cause the processor to execute the steps of any intelligent diagnosis guidance method and the steps of the training method of the disease recognition model.

[0141] The network interface is used for network communication, such as sending assigned tasks, etc.

[0142] Those skilled in the art can understand that Figure 6 the structure shown in

[0143] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0144] Among them, in one embodiment, the processor is used to execute the computer program and can implement the following steps when executing the computer program:

[0145] Obtain the user's medical image data and medical text data;

[0146] Based on the first feature extraction network, extract the image features corresponding to the medical image data;

[0147] Based on the second feature extraction network, extract the text semantic features corresponding to the medical text data;

[0148] Fuse the image features and text semantic features to obtain multimodal features;

[0149] Input the multimodal features into a preset disease recognition model to obtain the disease type;

[0150] Based on a preset medical resource database, determine the guiding diagnosis information according to the disease type.

[0151] Correspondingly, the processor is used to execute the computer program and can also implement the following steps when executing the computer program:

[0152] Obtain multiple training samples, where the training samples include medical visit image data and medical visit text data, as well as the disease type labels corresponding to the medical visit image data and medical visit text data;

[0153] Based on the first feature extraction network, extract the image features corresponding to the medical visit image data;

[0154] Based on the second feature extraction network, extract the text semantic features corresponding to the medical visit text data;

[0155] Fuse the image features and text semantic features to obtain multimodal features;

[0156] Input the multimodal features into the disease recognition model to obtain the disease type;

[0157] Based on a preset loss function, determine the loss value according to the disease type and the disease type label;

[0158] According to the loss value, adjust the model parameters of the disease recognition model.

[0159] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described intelligent guiding diagnosis can refer to the corresponding process in the embodiments of the foregoing intelligent guiding diagnosis method, and the specific training process of the above-described disease recognition model can refer to the corresponding process in the embodiments of the foregoing disease recognition model training method, which will not be elaborated here.

[0160] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps can be implemented:

[0161] Obtain the user's medical visit image data and medical visit text data;

[0162] Based on the first feature extraction network, extract the image features corresponding to the medical visit image data;

[0163] Based on the second feature extraction network, extract the text semantic features corresponding to the medical visit text data;

[0164] Fuse the image features and text semantic features to obtain multi-modal features;

[0165] Input the multi-modal features into a preset disease recognition model to obtain the disease type;

[0166] Based on a preset medical resource database, determine the guiding diagnosis information according to the disease type.

[0167] Correspondingly, when the computer program is executed by a processor, the following steps can also be implemented:

[0168] Obtain multiple training samples, where the training samples include medical visit image data and medical visit text data, as well as the disease type labels corresponding to the medical visit image data and medical visit text data;

[0169] Based on the first feature extraction network, extract the image features corresponding to the medical visit image data;

[0170] Based on the second feature extraction network, extract the text semantic features corresponding to the medical visit text data;

[0171] Fuse the image features and text semantic features to obtain multi-modal features;

[0172] Input the multi-modal features into the disease recognition model to obtain the disease type;

[0173] Based on a preset loss function, determine the loss value according to the disease type and the disease type label;

[0174] According to the loss value, adjust the model parameters of the disease recognition model.

[0175] Among them, the computer-readable storage medium can be the internal storage unit of the computer device in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0176] It should be noted that the functions or steps that the above computer-readable storage medium can achieve can be correspondingly referred to the embodiments of the foregoing intelligent guiding diagnosis method and the embodiments of the training method of the disease recognition model.

[0177] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0178] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0179] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. An intelligent diagnosis guidance method, characterized in that: include: Obtaining the user's medical image data and medical text data; Extracting image features corresponding to the medical image data based on a first feature extraction network; Extracting text semantic features corresponding to the medical consultation text data based on a second feature extraction network; Fusing the image features with the text semantic features to obtain multimodal features; Inputting the multimodal features into a preset disease recognition model to obtain the disease type; Based on a preset medical resource database, the diagnosis guidance information is determined according to the disease type.

2. The intelligent diagnosis guidance method according to claim 1, characterized in that: The medical consultation image data includes a tongue image and a facial image, and the medical consultation text data includes first text data; the multimodal features are input into a preset disease recognition model to obtain a disease type, including: The multimodal features corresponding to the tongue image, the facial image and the first text data are input into a preset disease recognition model to obtain a traditional Chinese medicine disease type.

3. The intelligent diagnosis guidance method according to claim 1, characterized in that: The disease types include traditional Chinese medicine disease types and Western medicine disease types; the method of determining the diagnosis guidance information based on the preset medical resource database according to the disease types includes: Based on a preset medical resource database, establishing a correspondence between the TCM disease type and the Western medicine disease type; Determining at least one disease according to the correspondence between the TCM disease type and the Western medicine disease type; The diagnosis guidance information is determined according to the priority corresponding to the at least one disease.

4. The intelligent diagnosis guidance method according to claim 1, characterized in that: The medical resource database includes a doctor evaluation database of each hospital; the medical resource database based on the preset medical resource database determines the guidance information according to the disease type, including: Obtain user attribute information and historical medical data; Determine a first target hospital according to the disease type, the attribute information and the historical medical data; Based on the doctor evaluation database, a first target doctor is determined from the first target hospital, and the score of the first target doctor is higher than the scores of other doctors in the department corresponding to the disease type in the first target hospital.

5. The intelligent diagnosis guidance method according to claim 4, characterized in that: The method also includes: Obtain users' medical consultation feedback data and medical consultation evaluation data; A second target hospital and a second target doctor are determined according to the medical consultation feedback data and the medical consultation evaluation data.

6. The intelligent diagnosis guidance method according to any one of claims 1 to 5, characterized in that: The fusing of the image features and the text semantic features to obtain multimodal features includes: Based on a fully connected layer, dimensionally aligning the image features and the text semantic features; Normalizing the image features and the text semantic features; Based on the attention mechanism, the image features and the text semantic features are fused to obtain multimodal features.

7. A method for training a disease recognition model, characterized in that: include: Acquire a plurality of training samples, wherein the training samples include medical consultation image data and medical consultation text data, and disease type labels corresponding to the medical consultation image data and the medical consultation text data; Extracting image features corresponding to the medical image data based on a first feature extraction network; Extracting text semantic features corresponding to the medical consultation text data based on a second feature extraction network; Fusing the image features with the text semantic features to obtain multimodal features; Inputting the multimodal features into a disease recognition model to obtain a disease type; Based on a preset loss function, determining a loss value according to the disease type and the disease type label; According to the loss value, the model parameters of the disease identification model are adjusted.

8. An intelligent medical guidance device, characterized in that: include: The first acquisition module is used to acquire the user's medical consultation image data and medical consultation text data; A first feature extraction module, used for extracting image features corresponding to the medical image data based on a first feature extraction network; A second feature extraction module, used for extracting text semantic features corresponding to the medical consultation text data based on a second feature extraction network; A feature fusion module, used for fusing the image features with the text semantic features to obtain multimodal features; An identification module, used for inputting the multimodal features into a preset disease identification model to obtain a disease type; The diagnosis guidance module is used to determine the diagnosis guidance information according to the disease type based on a preset medical resource database.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the intelligent diagnosis guidance method as described in any one of claims 1 to 6 and the disease recognition model training method as described in claim 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent diagnosis guidance method according to any one of claims 1 to 6 and the steps of the disease recognition model training method according to claim 7 are implemented.

Citation Information

Cited By

  • Intelligent hospital guide method and system, intelligent terminal and storage medium

    CN120932836A

  • Intelligent pre-inspection management method based on AI

    CN120954662A

  • Tongue image classification method and system based on multi-modal large model, and terminal

    CN121010568A

  • Fatty liver and coronary heart disease prediction method, device and program product based on tongue diagnosis

    CN121054256A

  • A method, device, and program product for predicting fatty liver disease combined with coronary heart disease based on tongue diagnosis

    CN121054256B