Auxiliary triage method and device, electronic equipment and storage medium

By using a triage model trained in stages and combining image and text information, the richness and reliability of the training set of the triage model are improved, solving the problem of insufficient triage accuracy in existing technologies and realizing more efficient departmental triage and generation of diagnostic and treatment plans.

CN120412966APending Publication Date: 2025-08-01CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410149953.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing medical triage methods rely on deep learning models, which results in insufficient triage accuracy and fails to effectively improve triage efficiency and accuracy.

Method used

A triage model with phased training is adopted. First, the image model and the text model are trained, and then their output is used as the training set for the feedforward network model. Combined with the patient's condition description images and text, the accurate features of the diseased parts of patients in each department are extracted, reducing interference from non-disease features and improving the richness and reliability of the training set.

Benefits of technology

By training with a rich training set and an efficient feedforward network model, the reliability and accuracy of the triage model are improved, ensuring the accuracy and efficiency of the triage results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412966A_ABST
    Figure CN120412966A_ABST
Patent Text Reader

Abstract

The invention discloses an auxiliary triage method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining disease information of a patient; inputting the disease information into a triage model, wherein the triage model at least comprises an image model, a text model and a forward network model which are sequentially trained; the training set of the image model comprises illness state description images of patients seeing a doctor in at least one department, the training set of the text model comprises illness state description texts of the patients seeing the doctor, and the training set of the forward network model at least comprises output of the image model and output of the text model; according to the output of the triage model, the triage result of the patient is determined, and the triage result is used for indicating the to-be-treated department of the patient. According to the invention, the triage accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to, but is not limited to, the field of digital medical technologies, and in particular, to an auxiliary triage method, device, electronic device, and storage medium. Background Art

[0002] With the continuous improvement of people's living conditions, the demand for health is also growing stronger. In recent years, the number of emergency patients in major hospitals has increased sharply, especially in hospitals leading the industry. The corresponding problems faced also include: patients lack medical and health knowledge and do not know which department to visit, which further exacerbates the pressure of medical triage.

[0003] In order to improve triage efficiency, with the rise of artificial intelligence, intelligent medical triage has begun to develop to achieve faster and more accurate disease judgment and give reasonable suggestions. For example, a proposed triage method is a triage method based on deep learning. Using deep learning methods usually takes the patient's condition description as the basis for guiding diagnosis, which leads to insufficient reliability of the triage model, and thus the triage accuracy of the triage model needs to be improved. Summary of the Invention

[0004] In view of this, this application provides an auxiliary triage method, device, electronic device, and storage medium to improve triage accuracy.

[0005] The technical solution of this application is implemented as follows:

[0006] On the one hand, an embodiment of this application provides an auxiliary triage method, and the method includes: obtaining the patient's illness information, where the illness information at least includes the patient's condition description text and the patient's condition description image, and the patient's condition description image at least includes the image of the patient's affected part; inputting the illness information into a triage model, where the triage model at least includes an image model, a text model, and a forward network model that are sequentially trained; the training set of the image model includes the condition description images of the patients who have visited at least one department, the training set of the text model includes the condition description texts of the patients who have visited, and the training set of the forward network model at least includes the output of the image model and the output of the text model; determining the triage result of the patient according to the output of the triage model, and the triage result is used to indicate the department where the patient should visit.

[0007] On the other hand, an embodiment of the present application provides an auxiliary triage device, and the device includes: an acquisition and generation module, configured to acquire and generate patient information and input it into the auxiliary diagnosis module; an auxiliary diagnosis module, configured to obtain a triage result and / or a diagnosis and treatment plan by using the patient information, and is further configured to store the patient information, the triage result, and the diagnosis and treatment plan; a navigation module: configured to obtain the patient's location information and the location information of the department to be visited, so as to generate navigation information from the patient's location to the department to be visited.

[0008] In yet another aspect, an embodiment of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0009] In still another aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements some or all of the steps in the above method.

[0010] In still another aspect, an embodiment of the present application provides a computer program, including computer-readable code, and when the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.

[0011] In still another aspect, an embodiment of the present application provides a computer program product, where the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, it implements some or all of the steps in the above method.

[0012] In the embodiment of the present application, when performing department triage, the triage model used adopts a phased training method. First, an image model and a text model are trained, and then the outputs of the trained image model and the trained text model are used as the training set of the forward network model to train the forward network model. Since the training set of the forward network model is the accurate features of the diseased parts of patients in each department extracted by the trained image model and the trained text model, without interference information of non-disease features, therefore, the training of the forward network model in the triage model provided by the embodiment of the present application is more efficient. Further, the training set of the forward network can not only come from the diseased information of the patients who have already visited the doctor, but also from the outputs of the trained image model and the trained text model. Therefore, the source of the training set of the forward network model is richer, which helps to ensure that the number of samples in the training set is sufficient, thereby improving the reliability of the trained forward network model, and further increasing the reliability of the triage model, so that when using the triage model for triage, the triage result is more accurate.

[0013] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solutions of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with this application and, together with the specification, are used to explain the technical solutions of this application.

[0015] Figure 1 is a schematic flowchart of an implementation process of an auxiliary triage method provided by an embodiment of this application;

[0016] Figure 2 is a schematic flowchart of an implementation process of obtaining text descriptions in an image of a diseased part provided by an embodiment of this application;

[0017] Figure 3 is a schematic flowchart of an implementation process of obtaining depth information of an image of a diseased part provided by an embodiment of this application;

[0018] Figure 4 is a schematic structural diagram of an image depth model provided by an embodiment of this application;

[0019] Figure 5 is a schematic flowchart of an implementation process of extracting a segmented image provided by an embodiment of this application;

[0020] Figure 6 is a schematic flowchart of an implementation process of extracting angular features provided by an embodiment of this application;

[0021] Figure 7 is a schematic diagram of human key points provided by an embodiment of this application;

[0022] Figure 8 is a schematic diagram of knee angle and ankle angle provided by an embodiment of this application;

[0023] Figure 9 is a schematic diagram of body tilt angle provided by an embodiment of this application;

[0024] Figure 10 is a schematic diagram of pelvic tilt angle provided by an embodiment of this application;

[0025] Figure 11 is a schematic flowchart of an implementation process of obtaining expression category information provided by an embodiment of this application;

[0026] Figure 12 is a schematic diagram of a triage method provided by an embodiment of this application;

[0027] Figure 13 is a schematic diagram of a method for department triage and making a diagnosis and treatment plan provided by an embodiment of this application;

[0028] Figure 14 is a schematic flow diagram of department classification provided by an embodiment of the present application;

[0029] Figure 15 is a schematic diagram of a method for obtaining vital sign data provided by an embodiment of the present application;

[0030] Figure 16 a schematic diagram of the composition structure of an auxiliary triage device provided by an embodiment of the present application;

[0031] Figure 17 is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0032] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the embodiments of the present application will be further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be construed as limitations on the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the embodiments of the present application.

[0033] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0034] The terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the embodiments of the present application belong. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the embodiments of the present application.

[0036] An embodiment of the present application provides an auxiliary triage method, which can be executed by a processor of a computer device. Herein, the computer device may refer to a device with data processing capabilities such as a server, a laptop computer, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable device), etc. An auxiliary triage method provided by an embodiment of the present application, as Figure 1 shown, the method includes the following steps S101 to step S103:

[0037] Step 101, obtaining the patient's disease information;

[0038] The disease information includes at least a text describing the patient's disease condition and an image describing the patient's disease condition, and the image describing the patient's disease condition includes at least an image of the patient's diseased part.

[0039] Here, since the medical information includes at least a text description of the patient's condition and an image description of the patient's condition, the types of medical information include not only text information but also image information. Therefore, compared with medical information that only contains a single text information or image information, the types of medical information in the embodiment of the present application are more comprehensive. Since the total amount of medical information in the embodiment of the present application is the sum of the number of two types of medical information, the amount of medical information in the embodiment of the present application is more sufficient than that containing only one type of medical information.

[0040] For example, the patient's condition description text may be the condition description text input by the patient. Of course, the condition description text may also be input by other personnel to assist the patient.

[0041] For example, these condition description texts may be directly input by the patient or his / her family, or may be read from an electronic medical record.

[0042] For example, the input method for inputting the text of the condition description may include: text input, voice input converted to text, translation of text or voice input and then converting it to text, etc.

[0043] Step 102: input the disease information into the triage model;

[0044] The triage model includes at least an image model, a text model and a forward network model that are trained in sequence. The training set of the image model includes images of condition descriptions of patients who have visited at least one department. The training set of the text model includes text descriptions of the condition of the patients who have visited the department. At least a part of the text descriptions of the condition of the patients who have visited the department is obtained by text extraction from the images of the condition descriptions of the patients who have visited the department. The training set of the forward network model includes at least the output of the image model and the output of the text model.

[0045] Step 103: determining the triage result of the patient according to the output of the triage model;

[0046] The triage result is used to indicate the department to be visited by the patient.

[0047] Exemplarily, the training set of the text model is the medical condition description text of the patients who have visited the doctor. The medical condition description text of the patients who have visited the doctor may include: the medical condition description text directly collected from the electronic medical records of the patients who have visited the doctor and the medical condition description text obtained by extracting text from the medical condition description images of the patients who have visited the doctor.

[0048] In the embodiments of the present application, the types of disease information are more comprehensive and the disease information is more sufficient. Therefore, when the triage model performs triage, the number of disease features extracted is more, and the basis for triage is more sufficient, thereby improving the accuracy of the triage result.

[0049] In the embodiments of the present application, when performing department triage, the triage model used adopts a phased training method. First, the image model and the text model are trained, and then the outputs of the trained image model and the trained text model are used as the training set of the forward network model to train the forward network model. Since the training set of the forward network model is the accurate features of the diseased parts of the patients in each department extracted by the trained image model and the trained text model, without interference information of non-disease features, therefore, the training of the forward network model in the triage model provided by the embodiments of the present application is more efficient. Moreover, the training set of the forward network can come not only from the disease information of the patients who have visited the doctor, but also from the outputs of the trained image model and the trained text model. Therefore, the source of the training set of the forward network model is richer, which helps to ensure that the number of samples in the training set is sufficient, thereby improving the reliability of the trained forward network model, and further increasing the reliability of the triage model, so that when using this triage model for triage, the triage result is more accurate.

[0050] In some embodiments, after step 101 of obtaining the disease information of the patient, the method may further include: generating an electronic medical record according to the disease information of the patient; synchronizing the electronic medical record to the electronic medical record database. In the embodiments of the present application, after obtaining the disease information of the patient, the electronic medical record is updated in real time so that the doctor who diagnoses and treats the patient can timely obtain the latest medical record of the patient.

[0051] In some embodiments, refer to Figure 2 , the obtaining of the disease information of the patient in step 101 may include steps 201 to 203, where:

[0052] Step 201, input the diseased part image of the patient into the image generation text model.

[0053] Among them, the image generation text model is used to convert image information into text information. The image generation text model includes: multi-modal recursive neural network (RNN), Showand Tell (a deep learning-based image annotation algorithm), Top-Down Bottom-Up Attention (including top-down and bottom-up attention), ClipCap, etc. Among them, Show and Tell realizes the conversion from image to natural language by mapping the image and its corresponding description text into the same vector space. This top-down attention is actively driven by the system itself. This driving force may come from the system spontaneously or may be a high-order signal generated under the stimulation of other tasks; while this bottom-up attention is driven by the outside of the system, and the driving force is a passive trigger generated under the stimulation of external things.

[0054] Step 202: Obtain the text description information corresponding to the patient according to the output of the image generation text model.

[0055] Here, in the embodiment of the present application, the step of extracting the text information of the diseased part image through the image generation text model is added. Compared with only extracting the disease-related information in the diseased part image through a single image feature extraction model, the embodiment of the present application can discover more disease-related information in the diseased part image.

[0056] Exemplarily, in step 202, it may further include: sending the output of the image generation text model to the patient's client; obtaining the adjusted text input by the patient through the patient's client, where the adjusted text is the text after the output of the image generation text model is adjusted by the patient; and using the adjusted text as the text description information corresponding to the patient.

[0057] Exemplarily, in step 202, after the patient or their accompanying person receives the output of the image generation text model, they can also modify the output of the image generation text model in combination with the specific condition of the patient's illness to form an adjusted text that is more in line with the patient's illness condition.

[0058] Step 203: Add the text description information corresponding to the patient to the patient's condition description text.

[0059] The solution provided by the embodiment of the present application discovers more disease-related information in the diseased part image to a greater extent, and the obtained disease information is more sufficient, so as to obtain a more accurate department triage and disease diagnosis and treatment plan based on the sufficient disease information.

[0060] In some embodiments, the obtaining of the patient's illness information in step 101 may further include: obtaining the patient's vital sign information, where the vital sign information is used to characterize the patient's vital signs; and adding the patient's vital sign information to the illness description text. The vital sign information may include information such as heart rate, blood pressure, blood sugar, sleep condition, body fat, etc.

[0061] Among them, the portable medical device may be a wearable smart watch, a smart body fat scale, a smart blood pressure monitor, a smart sleep monitor, etc.

[0062] Technology is becoming a part of people's lives and even their bodies. Portable medical devices are helping people regulate some physical conditions and achieve the best operation of the body. Patients can monitor their physical health conditions through portable medical devices and monitor changes or abnormalities in important vital sign parameters. The embodiments of the present application obtain vital sign data on various portable medical devices and use it as part of the illness information, broadening the types of the patient's illness information and helping to more comprehensively understand the patient's condition.

[0063] In some embodiments, the method further includes: adding at least one of the patient's past medical history and the patient's family medical history to the patient's illness description text.

[0064] For some diseases, especially major diseases, which are usually accompanied by complications, the solution provided by the embodiments of the present application requires determining whether the patient's illness manifestation is a complication caused by a major disease in the pre-onset stage. Since many major diseases are hereditary, the illness description text in the illness information of the embodiments of the present application covers the patient's past medical history and the patient's family medical history. When the patient's illness manifestation is a complication caused by a major disease, it can increase the accuracy of guiding diagnosis and treatment plans and avoid delaying the treatment opportunity for major diseases.

[0065] In some embodiments, referring to Figure 3 , the obtaining of the patient's illness information in step 101 may further include steps 301 to 303:

[0066] Step 301, inputting the image of the patient's diseased part into an image depth model;

[0067] Among them, the image depth model is used to extract the depth information of the image of the patient's diseased part.

[0068] Exemplarily, the image depth model can be a deep learning model, specifically including: convolutional neural network (CNN), generative adversarial network (GAN), deep boltzmann machine (DBM), etc.

[0069] Step 302: Obtain the depth image corresponding to the patient according to the output of the image depth model.

[0070] Step 303: Add the depth image corresponding to the patient to the disease description image of the patient.

[0071] In some cases, such as when being cut or having chapped skin, etc., it is necessary to analyze the depth of the wound. In the embodiments of the present application, the image depth information is added to the disease description image to ensure that accurate department guidance and diagnosis and treatment plans can be given even when image depth information is required.

[0072] In some embodiments, the above image depth model at least includes: a first feature extraction module and a second feature extraction module. The first feature extraction module is used to extract features from the image of the diseased part of the patient according to a first granularity, and the second feature extraction module is used to extract features from the image of the diseased part of the patient according to a second granularity, and the first granularity is greater than the second granularity;

[0073] The second feature extraction module includes M layers of processing units, and the output of the M layers of processing units is the output of the image depth model;

[0074] The input of the i-th layer of processing units in the M layers of processing units is the fused image corresponding to the (i - 1)-th layer of processing units in the M layers of processing units, and the fused image is an image obtained by fusing the output of the first feature extraction module and the output of the (i - 1)-th layer of processing units.

[0075] Wherein, M is an integer greater than or equal to 2, and i is a positive integer less than or equal to M.

[0076] Exemplarily, referring to Figure 4 , the image depth model can include a first feature extraction model and a second feature extraction module, wherein: the first feature extraction model is the Coarse branch, and the second feature extraction module is the Fine branch;

[0077] The Coarse branch includes the first coarse processing unit Coarse1 to the seventh coarse processing unit Coarse7, where: The first coarse processing unit Coarse1, after performing a convolution operation with a convolution kernel of 11×11 and a stride of 4 on the input original image, then performing a pooling operation with a pooling kernel of 2×2, finally outputs the first-layer coarse extraction feature map Coarse1. The second coarse processing unit Coarse2, after performing a convolution operation with a convolution kernel of 5×5 on the first-layer coarse extraction feature map Coarse1, then performing a pooling operation with a pooling kernel of 2×2, finally outputs the second-layer coarse extraction feature map Coarse2. The third coarse processing unit Coarse3, after performing a convolution operation with a convolution kernel of 3×3 on the second-layer coarse extraction feature map Coarse2, outputs the third-layer coarse extraction feature map Coarse3. The fourth coarse processing unit Coarse4, after performing a convolution operation with a convolution kernel of 3×3 on the third-layer coarse extraction feature map Coarse3, outputs the fourth-layer coarse extraction feature map Coarse4. The fifth coarse processing unit Coarse5, after performing a convolution operation with a convolution kernel of 3×3 on the fourth-layer coarse extraction feature map Coarse4, outputs the fifth-layer coarse extraction feature map Coarse5. The sixth coarse processing unit Coarse6, after performing a full convolution operation on the fifth-layer coarse extraction feature map Coarse5, outputs the sixth-layer coarse extraction feature map Coarse6. The seventh coarse processing unit Coarse7, after performing a full convolution operation on the sixth-layer coarse extraction feature map Coarse6, outputs the seventh-layer coarse extraction feature map Coarse7.

[0078] The Fine branch includes the first fine processing unit Fine1 to the fifth fine processing unit Fine5, where: The first fine processing unit Fine1, after performing a convolution operation with a convolution kernel of 9×9 and a stride of 2 on the input original image, then performing a pooling operation with a pooling kernel of 2×2, finally outputs the first-layer fine extraction feature map Fine1. The second fine processing unit Fine2, this layer concatenates (i.e., feature combination operation) the seventh-layer coarse extraction feature map Coarse7 and the first-layer fine extraction feature map Fine1 to obtain the second-layer fine extraction feature map Fine2. The third fine processing unit Fine3, after performing a convolution operation with a convolution kernel of 5×5 on the second-layer fine extraction feature map Fine2, obtains the third-layer fine extraction feature map Fine3. The fourth fine processing unit Fine4, after performing a convolution operation with a convolution kernel of 5×5 on the third-layer fine extraction feature map Fine2, obtains the fourth-layer fine extraction feature map Fine4. The fifth fine processing unit Fine5, this layer uses a Refined Lee filter to perform noise reduction processing on the fourth-layer fine extraction feature map Fine4 to obtain the final depth image.

[0079] In some embodiments, referring to Figure 5 , the obtaining of the patient's disease information in step 101 may further include steps 501 to 503:

[0080] Step 501, input the image of the diseased part of the patient into an image segmentation model;

[0081] The image segmentation model is used to perform image segmentation on the image of the diseased part of the patient to identify the diseased part of the patient.

[0082] Exemplarily, the image segmentation model can be the segmentanything model (SAM) under SuperMap Software, mean-shift, watershed algorithm, simple linear iterative clustering (SLIC), conditional random fields (CRF), U-Net, DeepLab, Pyramid Scene Parsing Network (PSPNet), etc. Among them, the mean-shift method is an image segmentation algorithm based on the color space distribution. The output of this algorithm is a "color-separated" image after color filtering, whose color will become gradually changing and the fine texture will become smooth. The watershed algorithm is a relatively basic mathematical morphology segmentation algorithm. The superpixel segmentation algorithm is a simple and effective image segmentation algorithm. U-Net is an improved fully convolutional neural network structure, named because its structure looks like the letter U when drawn, and is applied to the semantic segmentation of medical images. DeepLab is an image segmentation method that combines a deep convolutional neural network and a probabilistic graph model. The Pyramid Scene Parsing Network is a deep convolutional neural network model for image semantic segmentation.

[0083] Step 502, obtain the segmented image corresponding to the patient according to the output of the image segmentation model.

[0084] Exemplarily, step 502 may further include: sending the output of the image segmentation model to the patient's client; obtaining the adjusted image input by the patient through the patient's client, where the adjusted image is the image after the output of the image segmentation model is adjusted by the patient; using the adjusted image as the segmented image corresponding to the patient.

[0085] Here, after the patient or relevant personnel receive the output of the image segmentation model, they can make more refined adjustments to the output of the image segmentation model according to the actual situation of the patient's diseased area, and select a diseased area that is more in line with the actual diseased area.

[0086] Step 503: Add the segmented image corresponding to the patient to the patient's condition description image.

[0087] In some embodiments, when symptoms such as redness, swelling, and bruising appear at the affected area, the affected area has a distinct boundary, and the patient's affected part image often includes non - affected areas. In the embodiments of the present application, the segmentation model is used to segment the affected part image, removing the non - affected areas in the affected part image and only retaining the affected area, which can avoid the interference information of areas unrelated to the disease and reduce the efficiency and accuracy of subsequent triage and diagnosis.

[0088] In some embodiments, referring to Figure 6 , the obtaining of the patient's disease information in step 101 above may include steps 601 to 603:

[0089] Step 601: Collect at least one frame of the patient's image during exercise;

[0090] The at least one frame of image includes the human body image of at least one angle of the patient.

[0091] Step 602: Detect the human body image and determine the angle feature corresponding to the patient;

[0092] The angle feature is used to represent the angles of at least one human body key part of the patient.

[0093] Exemplarily, the method for obtaining the angles of human body key parts may include: 1) identifying the human body posture, 2) setting 17 human body key points, and 3) calculating the human body key angles. Among them,

[0094] 1) Identifying the human body posture:

[0095] Record the front video and two side videos of the patient walking from the front, left, and right of the patient respectively, and upload them to the server. Through image detection algorithms such as R - CNN, YOLO, and Single Shot Multibox Detector (SSD), detect the pedestrians in the frame by frame, and crop the image with the outer rectangular frame of the pedestrian as the boundary. Among them, the full name of R - CNN is Region - CNN, which is the first algorithm to successfully apply deep learning to object detection. R - CNN is based on algorithms such as convolutional neural network, linear regression, and support vector machine to achieve object detection technology. The full name of YOLO is you only look once, which means that the category and position of the objects in the image can be recognized only by browsing once.

[0096] From the cropped image, detect the above-mentioned labeled human skeletal key points based on any one of the models such as OpenPose, PoseNet, and DeepPose. Among them, OpenPose is an open-source real-time multi-person pose estimation library developed by Carnegie Mellon University. It estimates the human pose by analyzing the key points of the human body in images or videos, identifies various parts of the body, and infers the pose information of the human body. PoseNet is a human pose analysis model that can identify the parts of the human body in pictures and then describe the human pose with 17 reference points. DeepPose performs human pose estimation through a deep neural network.

[0097] 2) Set 17 human key points:

[0098] Such as Figure 7 , the human key points are respectively located at the left wrist, right wrist, left elbow, right elbow, left shoulder, right shoulder, neck, left waist, right waist, left knee, right knee, left ankle, right ankle, left heel, right heel, left toe, and right toe.

[0099] See Figures 8 to 10 , according to these key points, a total of the following 4 human angles are generated:

[0100] Knee angle: The knee angle can be used to reflect the situation of genu valgum or valgus foot, manifested as the calf being unable to straighten and bending outward. This gait is very distinctive, looking clumsy, with the knees together and the ankles valgus. In addition, the knee angle can also be used to describe the "O" shape legs caused by walking with an in-toe gait.

[0101] Ankle angle: The ankle angle is mainly used to describe the situation of walking with both feet. If both feet walk on tiptoe, it may be related to muscle tension or may be caused by damage to the spine or brain. At the moment when the heel touches the ground while walking, the knees should be kept straight. If not, it means that the mobility of the patella or the extension ability of the hip may be limited, and there is a risk of knee joint injury.

[0102] Body tilt angle (the angle between the line connecting the left and right waists and the neck): The body tilt angle is a description of the overall body posture. Some experts have found that when walking, the shoulders lean forward and the body leans forward, which may be a signal of gastrointestinal diseases. It may be suffering from chronic gastritis, gastric ulcer, or duodenal diseases.

[0103] Pelvic tilt angle (the tilt angle of the line connecting the left and right waists): The pelvic tilt angle is a description of the pelvic health condition. If both the left and right feet cross the body midline, it increases the rotation range of the pelvis and the lower back. An excessive waist rotation range is likely to cause lumbar muscle strain and joint degeneration, and may ultimately lead to a herniated disc.

[0104] 3) Calculate the human key angles:

[0105] In the video recorded from the left side, min(left knee angle), max(left knee angle), min(left ankle angle), and max(left ankle angle) are analyzed, and the corresponding frames are recorded as key frames;

[0106] In the video recorded from the right side, min(right knee angle), max(right knee angle), min(right ankle angle), and max(right ankle angle) are analyzed;

[0107] In the video recorded from the front, min(body tilt angle), max(body tilt angle), min(pelvic tilt angle), and max(pelvic tilt angle) are analyzed;

[0108] There are a total of 12 key part angle information for each patient. Taking the calculation process of the left knee angle cosθ at the t-th frame as an example, the calculation process for each angle is as follows: t 左膝 The calculation process for each angle is as follows:

[0109] Left knee angle cosθ 左膝 It is calculated from key points such as the left waist key point (x 左腰 , y 左腰 ), the left knee key point (x 左膝 , y 左膝 ), and the left heel key point (x 左脚跟 , y 左脚跟 ). For example, it is calculated using the following expressions (1) to (3):

[0110]

[0111] min(left knee angle) = min(cosθ 左膝 ) (2)

[0112] max(left knee angle) = max(cosθ 左膝 ) (3)

[0113] where x1 = x 左膝 - x 左腰 ; y1 = y 左膝 - y 左腰 ; x2 = x 左膝 - x 左脚跟 ; y2 = y 左膝 - y 左脚跟 .

[0114] Each frame of the current video corresponds to a cosθ 左膝 , cosθ 左膝The minimum and maximum values in the sequence are respectively taken as min (left knee angle) and max (left knee angle), which are the characteristic values of the left knee angle. In addition, the angle video frames corresponding to the minimum and maximum values are key frames and are stored in the database for the doctor to view later to assist in diagnosis.

[0115] Step 603: Add the angle features corresponding to the patient to the patient's disease information.

[0116] When diseases such as fractures, arthritis, and lumbar disc herniation occur, the angles of the corresponding parts of the human body will play an important role in the triage and diagnosis of diseases. Therefore, in the embodiments of the present application, the angles of the key parts of the human body are also added to the disease information to ensure that when diseases such as fractures, arthritis, and lumbar disc herniation occur, accurate triage can be carried out and a diagnosis and treatment plan can be given.

[0117] In some embodiments, referring to Figure 11 The obtaining of the patient's disease information in step 101 may further include steps 1101 to 1103, where:

[0118] Step 1101: Collect at least one frame of the patient's face image.

[0119] Exemplarily, a face image can be collected based on a recorded frontal video or a captured frontal photo.

[0120] Step 1102: Perform expression category recognition on the at least one frame of face image to obtain the patient's expression category information; the expression category information is used to represent the patient's emotion.

[0121] Exemplarily, face image expression recognition can be performed according to MTCNN (multi-task cascaded convolutional networks). MTCNN is a multi-task cascaded convolutional neural network based on deep learning and is applicable to face detection at multiple scales.

[0122] Exemplarily, after detecting a face, face features are extracted based on a deep learning model and the face is classified. For example, facial expressions are divided into 6 categories, including: Category 1: Happy; Category 2: Sad; Category 3: Angry; Category 4: Fear / Surprise; Category 5: Disgust; Category 6: Neutral.

[0123] Exemplarily, the facial expressions in each video frame can be classified. Among them, the frames where the facial expressions are in categories 1-5 are all key frames, and the key frames will be saved for the doctor to view later to assist in the diagnosis of the condition. The voting statistics can also be performed on the expression sequence results, and the category with the highest number of votes is the expression analysis result corresponding to the current video, that is, the expression category information.

[0124] Step 1103, add the expression category information of the patient to the disease information of the patient.

[0125] The patient's emotion can reflect the severity of the disease, the degree of pain, discomfort, etc. By extracting the patient's facial expression, it can further assist in the diagnosis. Especially for diseases without obvious external wounds, it is more necessary to rely on the patient's emotion to assist in judging the severity of the condition. Therefore, the disease information in the embodiment of the present application includes the expression category to assist doctors in disease diagnosis when necessary.

[0126] In some embodiments, after step 103, according to the output of the triage model, after determining the triage result of the patient, the method may further include: obtaining multiple department information according to the output of the triage model, where the department information is used to indicate the departments where the patient can seek medical treatment; sorting the departments where the patient can seek medical treatment according to the multiple department information to obtain a department sequence, and the department sequence is used to indicate the departments where the patient is to seek medical treatment.

[0127] Exemplarily, the department sequence can be obtained according to the probabilities of the patient being assigned to each department. The department sequence can be a sequence including all departments, or a sequence only including the top several ranked departments.

[0128] Some diseases may be related to multiple departments, and the relevance to multiple departments is almost the same. At this time, it is not comprehensive enough to only give one best department for the patient to seek medical treatment. Therefore, the embodiment of the present application provides a department sequence for the patient to seek medical treatment, so that the patient can select the best department for seeking medical treatment according to the relevance of each department to the disease, the number of patients in each department, etc.

[0129] In some embodiments, the method may further include: feeding back the department sequence to the corresponding doctor of the selected department. After the patient obtains the triage result, the patient will select the department for seeking medical treatment according to the triage result, but some patients may go to the wrong department. Therefore, in the embodiment of the present application, the department sequence is fed back to the corresponding doctor of the selected department. When the patient goes to the wrong department, the doctor can find out in time and assist the patient to find the appropriate department for seeking medical treatment, thus avoiding wasting more energy and time of doctors and patients due to misdiagnosis.

[0130] In some embodiments, after determining the triage result of the patient in step 103, the method further includes: generating an electronic medical record according to the patient's disease information and the patient's triage result; synchronizing the electronic medical record to the electronic medical record database.

[0131] In the embodiments of the present application, the disease information and triage results are synchronously sent to the electronic medical record database in a timely manner, so that the doctor at the registration desk or the doctor in the department where the patient seeks medical treatment can accurately guide the patient in a timely manner according to the triage results when it is found that the patient has chosen the wrong consulting room, without wasting more time and energy on diagnosis and medical treatment.

[0132] In some embodiments, after step 103, the method may further include:

[0133] Input the patient's condition description text and at least one prompt into the text generation pre-trained model, where the prompt is used to indicate the items included in the patient's diagnosis and treatment plan.

[0134] Exemplarily, the input of the text generation pre-trained model may be "The patient's condition is as follows: 'Text (condition description text)', and 'Please give Text (prompt)'."

[0135] Exemplarily, the prompt, that is, prompt, as the name implies, means "prompt". Simply put, a prompt is an instruction given to the text generation pre-trained model. For example, prompt 1: The possible disease names of the patient; prompt 2: The diagnosis results including the severity of the condition and the mortality rate; prompt 3: The probability of the patient being hospitalized; prompt 4: The treatment plan for the patient; prompt 5: The medication plan including the names of drugs, available drugs, drug prices, and drug combinations; prompt 6: The diet recommendations, diet precautions, and recipes.

[0136] Exemplarily, when the prompts in the input are the above prompts 1 to 6, the output of the text generation pre-trained model, that is, the diagnosis and treatment plan, includes:

[0137] The specific content of the possible disease names of the patient;

[0138] The specific content of the diagnosis results including the severity of the condition and the mortality rate;

[0139] The specific content of the probability of the patient being hospitalized;

[0140] The specific content of the treatment plan for the patient;

[0141] The specific content of the medication plan including the names of drugs, available drugs, drug prices, and drug combinations;

[0142] The specific content of the diet recommendations, diet precautions, and recipes;

[0143] Obtain the diagnosis and treatment plan of the patient output by the text generation pre-trained model.

[0144] By using the method described in the embodiment of the present application, a diagnosis and treatment plan can be generated based on the condition description text in the medical information, so as to assist doctors in diagnosing and treating patients more quickly and accurately.

[0145] In some embodiments, after obtaining the patient's diagnosis and treatment plan, the method may further include: generating the patient's electronic medical record based on the patient's medical information, the patient's triage results, and the patient's diagnosis and treatment plan; and synchronizing the patient's electronic medical record to an electronic medical record database. In this embodiment of the present application, after generating the diagnosis and treatment plan, all information about the patient's visit is promptly updated to the electronic medical record database, making it convenient for patients and doctors to access and use it at any time.

[0146] In some embodiments, after step 103, the method may further include: obtaining patient location information of the patient's location and department location information of the department to be treated.

[0147] For example, the patient's position can be located while the video or image of the patient is being captured by the camera, or the patient's position can be located by the location of the client.

[0148] For example, the path between the patient's location and the location of the department to be treated can be indicated by the Beidou system, GPS, or a hospital map built into the navigation module.

[0149] Many times, patients are not familiar with the distribution of various departments in the hospital. The auxiliary diagnosis method of the embodiment of the present application can also navigate the patients so that they can reach the department to be treated accurately and quickly.

[0150] In some embodiments, before step 101, the method may further include steps 1201 to 1205, wherein:

[0151] Step 1201 : Obtain a first training set, where the first training set includes images describing the condition of the patient who has been treated and text describing the condition of the patient who has been treated.

[0152] Step 1202 : Using the images describing the condition of the patients who have been treated as training sets, the initial image model is trained until the initial image model converges to obtain the image model.

[0153] Step 1203: After obtaining the image model, the text describing the condition of the patient who has been treated is used as a training set to train the initial text model until the initial text model converges to obtain the text model.

[0154] Step 1204, after obtaining the text model, use at least the output of the initial image model and the output of the initial text model as a training set to train the initial forward network model until the initial forward network model converges to obtain the forward network model.

[0155] Step 1205, obtain the triage model based on the image model, the text model, and the forward network model.

[0156] In the embodiments of the present application, the triage model is trained in stages. First, the image model and the text model are trained, and then the output of the trained image model and the output of the trained text model are used as the training set of the forward network model to train the forward network model. Because the training set of the forward network model includes the accurate features of the diseased parts of patients in each department extracted by the trained image model and the trained text model, and there is no interference information unrelated to the disease in the accurate features of the diseased parts of patients in each department. Therefore, using the output of the trained image model and the output of the trained text model as the training set of the forward network model can make the training of the forward network model more efficient, and also make the source of the training set of the forward network model richer, thereby improving the reliability of the forward network model, and further making the department classification result of the triage model more accurate.

[0157] In some embodiments, in step 1201, the obtaining of the first training set may further include: obtaining the diseased part image of the already-seen patient; inputting the diseased part image of the already-seen patient into the image segmentation model to identify the diseased part of the already-seen patient; obtaining the segmented image corresponding to the already-seen patient according to the output of the image segmentation model; adding the segmented image corresponding to the already-seen patient to the condition description image of the already-seen patient.

[0158] In some embodiments, in step 1204, after obtaining the text model, before using at least the output of the initial image model and the output of the initial text model as a training set to train the initial forward network model, the method may further include:

[0159] Obtaining at least one of the expression category information of the already-seen patient and the angle feature corresponding to the already-seen patient, where the expression category information of the already-seen patient is used to represent the emotion of the already-seen patient; the angle feature corresponding to the already-seen patient is used to represent the angle of at least one human key part of the already-seen patient.

[0160] In some embodiments, in step 1204, after obtaining the text model, when using at least the output of the initial image model and the output of the initial text model as a training set to train the initial forward network model, the method may further include:

[0161] Use at least one of the expression category information of the visited patient and the angular feature corresponding to the visited patient, the output of the initial image model, and the output of the initial text model as the training set to train the initial forward network model.

[0162] To facilitate better understanding of the training process of the triage model by those skilled in the art, a specific example of triage model training is given below:

[0163] The first stage is to train the image model:

[0164] 11) Obtain a large number of patient images from each department in the hospital's electronic medical record library or other medical platforms and use them as the training set of the image model. The label of each image is the name of this department.

[0165] 12) Based on Darknet as the basic network, this image model extracts high-level image features of the image, and this high-level image feature is represented by a one-dimensional feature of size 1000.

[0166] 13) Use a fully connected layer with an input of 1000 and an output of the number of departments to further extract features.

[0167] 14) Use the softmax layer to normalize the output of the fully connected layer to the interval [0,1]. Each normalized value represents the probability of being recommended for this department.

[0168] 15) Represent the one-hot of the department data as the label data, calculate the cross-entropy loss between the label data and the softmax output data, continuously adjust the network weights during the training process to reduce this loss value. When the reduction of the loss value approaches stability or is less than a certain preset value, the training ends.

[0169] The second stage is to train the text model:

[0170] 21) During the training process of the second stage, fix the weights of the image model, and the input is the text description of the condition, past medical history, and family history (such as: multiple red and swollen areas on the legs).

[0171] 22) Based on the BERT (bidirectional encoder representation from transformers) model, extract text features, and the text features are one-dimensional vectors of size embedding size.

[0172] 23) After connecting the text feature vector and the image feature vector output by the image model, use it as the comprehensive feature.

[0173] 24) Use a fully connected layer based on the comprehensive features to further extract high-level semantics. The input of this connection layer is a one-dimensional vector with a size of embedding size + 3 * 1000, and the number of output nodes is equal to the number of departments. Each node corresponds to a department, and the value of the node is the probability of being predicted to belong to that department.

[0174] 25) Use the one-hot representation of the department as the label data. The cross-entropy loss between the label data and the output of the softmax layer is the loss of the entire network. Continuously optimizing the loss and reducing the loss value is the training process. When the decrease in the loss value approaches stability or is less than a certain preset value, the training ends.

[0175] The third stage is to train the forward network model:

[0176] 31) Fix the parameters of the image model (Darknet) and the parameters of the text model (BERT). The output of the text model is a one-dimensional vector with a size of embedding size, and the output of the image model is a one-dimensional vector with a size of 3 * 1000. Add an angle feature with a size of 14 and an emotion feature with a size of 1. Connect the output of the above text model, the output of the image model, the angle feature, and the emotion feature to form an output feature with a feature size of embedding size + 3 * 1000 + 14 + 1.

[0177] 32) Input the output feature in 1) into a fully connected layer with a size of embedding size + 3 * 1000 + 14 + 1 to obtain the fully connected feature.

[0178] 33) Input the fully connected feature in 2) into a fully connected layer with the number of nodes equal to the number of departments, and set a softmax layer to normalize the output of the fully connected layer to the interval [0, 1].

[0179] 34) The cross-entropy loss between the label data and the output of the softmax layer is the loss of the entire classification model network. Continuously optimizing the loss and reducing the loss value is the training process. When the decrease in the loss value approaches stability or is less than a certain preset value, the training ends.

[0180] To enable those skilled in the art to understand the embodiments of the present application as a whole, the following gives an example of the embodiments of the present application.

[0181] This technical solution takes multi - type and multi - dimensional patient information as input, and infers auxiliary diagnosis information through a deep learning model. That is, the input is the current patient's condition description, family medical history, past medical history, and images of the affected area. Through the images of the affected area, corresponding depth information, segmentation information, key part angle information, etc. are inferred. These features are fused to recommend departments for the patient and provide a map navigation for the department location. At the same time, it will predict diagnosis information such as the patient's disease type, mortality rate, hospitalization rate, treatment plan, medication plan, diet recommendation, etc., to assist doctors in medical diagnosis.

[0182] The neural network learning model designed in this solution refers to the high - level features of multiple inputs as the basis for auxiliary diagnosis. Since some current diseases, especially major diseases, are usually accompanied by complications, fusing the features of family medical history and past medical history can further determine whether there are complications caused by major diseases in the early stage, thereby increasing the accuracy of guiding diagnosis and avoiding delaying the treatment time of major diseases. This solution also combines the depth image and segmentation image corresponding to the original image, making the auxiliary diagnosis result more accurate.

[0183] The method of using the triage model provided in the embodiments of this application is as Figure 12 shown. The triage model in the embodiments of this application can be built into the intelligent triage server:

[0184] After patient 1207 arrives at the hospital, they perform identity verification and login by swiping their ID card / medical card (identity authentication 1206); after logging in, the server can quickly retrieve the patient's historical case information (including historical symptoms, historical departments visited, historical doctor diagnoses, historical medications, etc.); the patient or their family member inputs the current condition by voice (voice input of condition 1204) and automatically uploads it to the server; the affected area is captured by a dedicated camera (capture of affected area 1205), and at the same time, the camera can locate the patient's position to facilitate the next department navigation; the captured image is automatically uploaded to the intelligent triage server 1200. The depth information of the image is automatically estimated based on the image of the affected area, and the depth image is also uploaded to the intelligent triage server 1200 synchronously. If symptoms such as redness, bruising, etc. appear in the affected area, the affected area image can be automatically segmented through image segmentation technology such as the Segment Anything Model (SAM), and the segmented image is fed back to the user. The user can manually adjust and confirm the segmentation boundary, and the confirmed segmented image is synchronized to the intelligent triage server 1200.

[0185] In the embodiments of this application, department navigation and the processing of images of the affected area are also added.

[0186] In the embodiments of the present application, the description of the condition is supplemented according to the image of the affected area. Based on image-to-text technologies such as ClipCap, a medical description text of the patient's current condition is generated according to the image of the affected area. This part of the text will also be used as part of the description of the patient's current condition and synchronized to the patient's electronic medical record database 1201.

[0187] In the embodiments of the present application, according to the texturized patient condition, medical history and other electronic medical record data, and in turn combined with multiple prompts as the input for a text generation pre-trained model (such as GPT), text 1203 is output for the doctor. The text 1203 output for the doctor includes: the patient's diagnosis results (disease type, mortality rate, hospitalization rate, etc.), medication suggestions (treatment plan), diet suggestions and other information. When the patient confirms the attending doctor, text 1202 is output for the patient. The text 1202 output for the patient includes: wound area confirmation, department recommendation and map navigation.

[0188] The implementation process of a triage method provided by the embodiments of the present application is as follows:

[0189] 1) The inputs are two types of data:

[0190] One type is text-based descriptions, which include the description of the current patient's condition, family medical history, and past medical history. These text data can be dictated by the patient or the patient's family member and generated by speech-to-text, or directly input by the patient or their family member in text form, or extracted from the electronic medical record.

[0191] The other type is medical images and videos, which can be recorded by the patient himself / herself using a mobile phone or other devices with camera sensors.

[0192] 2) Extract more features of the image of the affected area:

[0193] Based on the deep learning model, the depth information of the image of the affected area is extracted to generate a depth image; based on the deep learning model, the image is segmented to obtain the segmented image of the affected area; based on the image-to-text model, the image of the affected area is input to generate the corresponding description text of the condition. The above depth information, segmented image, and description text of the condition are used as supplements to the original condition description.

[0194] 3) Perform department triage and give a diagnosis and treatment plan:

[0195] See Figure 13, high-level features of the diseased part image, depth image, and segmentation image are extracted based on a convolutional neural network model; by analyzing the recorded video, the key part angle and expression category are obtained, these features are fused, and based on the fused features, department recommendation is performed. The department with the highest probability value is selected as the best recommended department and feedback to the patient, and the ranking information of all departments from high to low probability is pushed to the corresponding doctor of the selected department to provide certain help for the doctor's diagnosis and reduce the occurrence of misdiagnosis.

[0196] In some embodiments, through a text generation pre-training model, the diseased part image, depth image, segmentation image, key part angle, expression category, disease description text, family medical history, and past medical history are input, and combined with prompt information, the disease type, mortality rate, hospitalization rate, treatment plan, medication plan, diet recommendation, etc. of the current patient are obtained.

[0197] In some embodiments, the steps of image segmentation of the diseased part image by an image segmentation model may include: if symptoms such as local skin redness, bruising, and dark purple appear in the patient, the diseased part image can be automatically segmented by image segmentation technology, and the segmented image and contour information are fed back to the user. The user can manually adjust and confirm the segmented contour such as moving and stretching. The confirmed segmented image will be synchronized to the electronic case database as patient information.

[0198] In some embodiments, the steps of generating a disease description text based on the diseased part image may include: based on image-to-text technologies such as ClipCap, using the diseased part image as input, the output is the text description corresponding to the disease, and the text description is fed back to the user. The user can modify and confirm the text. This part of the text will be added to the patient's disease description as part of the patient's electronic case content and will also be synchronized to the electronic case database.

[0199] In some embodiments, the steps of extracting image depth information from the diseased part image by an image depth model may include: obtaining the specific depth from a single original picture is equivalent to inferring the three-dimensional space from a two-dimensional image to obtain the relative depth information of the diseased part image. This solution intends to adopt a similar Coarse network structure, with the input being the original image of the diseased part and the output being its corresponding depth information map. Coarse uses AlexNet as the backbone network architecture, and the backbone network adopted in this solution is not limited to AlexNet and can also be VGG, ResNet, etc.

[0200] In some embodiments, the image depth information extraction model includes a Coarse branch and a Fine branch. Among them, the Coarse branch includes: a first coarse extraction layer, which performs a convolution operation with a convolution kernel of 11×11 and a stride of 4 on the input original picture, then performs a pooling operation with a pooling kernel of 2×2, and finally outputs the first-layer coarse extraction feature map Coarse1; a second coarse extraction layer, which performs a convolution operation with a convolution kernel of 5×5 on the first-layer coarse extraction feature map Coarse1, then performs a pooling operation with a pooling kernel of 2×2, and finally outputs the second-layer coarse extraction feature map Coarse2; a third coarse extraction layer, which performs a convolution operation with a convolution kernel of 3×3 on the second-layer coarse extraction feature map Coarse2 and outputs the third-layer coarse extraction feature map Coarse3; a fourth coarse extraction layer, which performs a convolution operation with a convolution kernel of 3×3 on the third-layer coarse extraction feature map Coarse3 and outputs the fourth-layer coarse extraction feature map Coarse4; a fifth coarse extraction layer, which performs a convolution operation with a convolution kernel of 3×3 on the fourth-layer coarse extraction feature map Coarse4 and outputs the fifth-layer coarse extraction feature map Coarse5; a sixth coarse extraction layer, which performs a full convolution operation on the fifth-layer coarse extraction feature map Coarse5 and outputs the sixth-layer coarse extraction feature map Coarse6; a seventh coarse extraction layer, which performs a full convolution operation on the sixth-layer coarse extraction feature map Coarse6 and outputs the seventh-layer coarse extraction feature map Coarse7.

[0201] Furthermore, the Fine branch includes: a first fine extraction layer which performs a convolution operation with a convolution kernel of 9×9 and a stride of 2 on the input original image, then performs a pooling operation with a pooling kernel of 2×2, and finally outputs the first fine extraction feature map Fine1; a second fine extraction layer which concatenates (a feature combination operation) the seventh coarse extraction feature map Coarse7 and the first fine extraction feature map Fine1 to obtain the second fine extraction feature map Fine2; a third fine extraction layer which performs a convolution operation with a convolution kernel of 5×5 on the second fine extraction feature map Fine2 to obtain the third fine extraction feature map Fine3; a fourth fine extraction layer which performs a convolution operation with a convolution kernel of 5×5 on the third fine extraction feature map Fine2 to obtain the fourth fine extraction feature map Fine4; a fifth fine extraction layer which uses a Refined Lee filter to perform noise reduction processing on the fourth fine extraction feature map Fine4 to obtain the final depth image. Among them, the stride in the convolution operation refers to the step size of the convolution kernel moving on the image, and the size of the stride directly affects the result of the convolution operation and the size of the feature map. Concatenate is generally used to combine features, fuse the features extracted by multiple convolutional feature extraction frameworks, or fuse the information of the output layer. This combination is actually a combination of dimensions.

[0202] In some embodiments, the steps of detecting a human body image to determine the angular features of a patient may include: First, it is planned to set that the total number of human body key points is 17 points, namely the left wrist, right wrist, left elbow, right elbow, left shoulder, right shoulder, neck, left waist, right waist, left knee, right knee, left ankle, right ankle, left heel, right heel, left toe, and right toe. A total of 4 types of angles are generated based on these key points, namely the knee angle, ankle angle, body tilt angle, and pelvic tilt angle.

[0203] Among them, the knee angle may refer to: the angle formed by the line connecting from the right waist through the right knee to the right ankle, and the angle formed by the line connecting from the left waist through the left knee to the left ankle. The knee angle can be used to reflect the situation of knee valgus or valgus foot, manifested as the calf being unable to straighten and bending outward. This gait is very distinctive, looking clumsy, with the knees together and the ankles valgus. In addition, the knee angle can also be used to describe the "O" - shaped legs caused by walking in an in - toed posture.

[0204] The ankle angle can refer to the angle formed by the line connecting the right knee, through the right heel, to the right toe, and the angle formed by the line connecting the left knee, through the left heel, to the left toe. The ankle angle is mainly used to describe the walking situation of the feet. If a person walks on tiptoe, it may be related to muscle tension or may be caused by damage to the spine or brain. At the moment when the heel touches the ground during walking, the knee should be kept straight. Otherwise, it may mean that the mobility of the patella or the extension ability of the hip is limited, and there is a risk of knee joint injury.

[0205] The body tilt angle can refer to the angle between the line connecting the left and right waists and the neck. The body tilt angle is a description of the overall body posture. Some experts have found that when walking, the shoulders lean forward and the body leans forward, which may be a signal of gastrointestinal diseases. It may be suffering from chronic gastritis, gastric ulcer or duodenal diseases.

[0206] The pelvic tilt angle can refer to the tilt angle of the line connecting the left and right waists. The pelvic tilt angle is a description of the pelvic health condition. If both feet cross the body midline, it increases the rotation range of the pelvis and the lower back. An excessive waist rotation range is likely to cause lumbar muscle strain and joint degeneration, and may eventually lead to a herniated disc.

[0207] Secondly, calculate the angles of the above key parts of the human body. In some embodiments, record the frontal video and two side videos of the patient walking from the front, left, and right of the patient respectively, and upload them to the server. Through image detection algorithms such as YOLO, detect the pedestrians in the frame by frame, and crop the image with the outer rectangular frame of the pedestrian as the boundary. Then, from the cropped image, detect the above-mentioned marked human bone key points based on OpenPose. Among them, OpenPose is an open-source real-time multi-person pose estimation library developed by Carnegie Mellon University. It estimates the human pose by analyzing the key points of the human body in the image or video, identifies each part of the body, and infers the pose information of the human body. PoseNet is a human pose analysis model that can identify the parts of the human body in the picture and then describe the human pose with 17 reference points;

[0208] Analyze min(left knee angle), max(left knee angle), min(left ankle angle), and max(left ankle angle) from the video recorded from the left, and record the corresponding frame as the key frame.

[0209] Analyze min(right knee angle), max(right knee angle), min(right ankle angle), and max(right ankle angle) from the video recorded from the right.

[0210] Analyze min(body tilt angle), max(body tilt angle), min(pelvic tilt angle), and max(pelvic tilt angle) from the video recorded from the front.

[0211] There are a total of 12 pieces of angle information for the key parts of each patient. The method of calculating the cosine of the left knee angle θ in the t-th frame as described above can be used to calculate them respectively, so as to obtain the angle features of the above-mentioned key parts. In addition, the angle video frames corresponding to the minimum value and the maximum value are key frames and are stored in the database for the doctor to view later to assist in diagnosis. t 左膝 The method of calculating the cosine of the left knee angle θ in the t-th frame as described above can be used to calculate them respectively, so as to obtain the angle features of the above-mentioned key parts. In addition, the angle video frames corresponding to the minimum value and the maximum value are key frames and are stored in the database for the doctor to view later to assist in diagnosis.

[0212] In some embodiments, the step of performing expression category recognition on at least one frame of face image to obtain the expression category information of the patient may include:

[0213] Based on the recorded video of walking forward, using MTCNN (Multi-task Cascaded Convolutional Networks), which is a multi-task cascaded convolutional neural network based on deep learning and is suitable for face detection at multiple scales. After detecting the face, extract face features based on the deep learning model and classify the face. There are 6 categories of face expressions in total, namely Category 1: Happy, Category 2: Sad, Category 3: Angry, Category 4: Fear / Surprise, Category 5: Disgust, Category 6: Neutral. Classify each video frame. The frames where the face expressions of categories 1-5 are located are all key frames, and the key frames will be saved to the server for the doctor to view later to assist in the diagnosis of the condition.

[0214] Conduct a vote count on the expression sequence results, and the category with the highest number of votes is the expression analysis result corresponding to the current video.

[0215] In some embodiments, the method for recommending departments is as follows:

[0216] Based on other electronic case information such as the original image, depth image, segmentation image, current condition, past medical history, and family medical history, predict the best department. And automatically recommend the route from the current location to the recommended department.

[0217] As Figure 14 shown, the input is four types of data: text, image, key part angle, and emotion category. The text data includes three types of data: condition description, family medical history, and past medical history; the images are three types: the original image, segmentation image, and depth image of the patient's diseased part. The original image is obtained through a camera sensor, the segmentation image is the output of an image segmentation model, and the depth image is the output of a depth detection model. The output of department classification is the probability of waiting to see a doctor in each department.

[0218] The one-dimensional vector of size embedding_size extracted based on BERT is used as the text feature, the one-dimensional vector of size 1000 extracted based on Darknet is used as the image feature, the angle feature is a one-dimensional vector of size 14, and the emotion feature is a one-dimensional feature of size 1. After connecting the four types of features, a comprehensive feature of size (embedding_size + 1000 * 3 + 14 + 1) is obtained. Then, a fully connected layer with the number of input nodes (embedding_size + 3 * 1000 + 14 + 1) and the number of output nodes equal to the number of departments is connected. The output of the last softmax layer is the probability of each department, that is, the final classification result is obtained. The classification result is the probability of being recommended to each department, and the department corresponding to the maximum probability value is the best recommended department.

[0219] In some embodiments, the method of department navigation is as follows:

[0220] The server will preset the hospital map. The device where the camera sensor for taking pictures has a built-in positioning device. According to the recommended department and the location of the patient, map navigation is performed, and the patient can easily find the corresponding department through the navigation information.

[0221] In addition, the image, voice information, and recommended department of the current patient will form a new electronic medical record and be stored in the server to facilitate the next diagnosis of the patient's condition.

[0222] In some embodiments, the training process of the triage model in the embodiments of the present application is as follows:

[0223] Phase 1: Train the image model.

[0224] A large number of patient images of each department are obtained from the hospital's electronic medical record library or other medical platforms. The label of each image is the name of this department. Based on Darknet as the basic network, high-level image features of the images are extracted, and this feature is represented by a one-dimensional feature of size 1000. Then, a fully connected layer with an input of 1000 and an output of the number of departments is connected to further extract features. The last layer is a softmax layer, which normalizes the output of the fully connected layer to the [0, 1] interval, and each value represents the probability of being recommended to this department. The one-hot representation of the department data is used as the label data. Calculate the cross-entropy loss between the label data and the softmax output data. The training process is a process of continuously adjusting the network weights to reduce the loss value.

[0225] In the embodiments of the present application, the training processes for the diseased part image, segmentation image, and depth image are similar and will not be elaborated here.

[0226] Phase 2: Train the text model.

[0227] After the first stage is completed, the image model can correctly extract the features of the patient's image.

[0228] During the training process of the second stage, fix the weights of this image model. The inputs are the disease description text, past medical history, and family medical history (e.g., multiple red and swollen areas on the legs). Based on the BERT model, extract the text features. The text features are one-dimensional vectors of size embedding size. After concatenating the text feature vector and the image feature vector as the comprehensive feature, it is followed by a fully connected layer to further extract high-level semantics. The input of this connection layer is a one-dimensional vector of size embedding size + 3 * 1000, and the number of output nodes is equal to the number of departments. Each node corresponds to a department. The value of the node is the probability of being predicted to belong to that department. The one-hot representation of the department is used as the label. The cross-entropy loss between the label data and the output of the softmax layer is the loss of the entire network. Continuously optimizing the loss and reducing the loss value is the training process. When the decrease in the loss value approaches stability or is less than a certain preset value, the training ends.

[0229] Train the forward network model in stage 3.

[0230] After the second stage is completed, fix the parameters of the image model (Darknet) and the parameters of the text model (Bert). The output of the text model is a one-dimensional model of size embedding size, and the output of the image model is a one-dimensional vector of size 3 * 1000. Append an angle feature of size 14 and an emotion feature of size 1. The connected feature size is embedding size + 3 * 1000 + 14 + 1, followed by a fully connected layer of the same size and a fully connected layer with the number of nodes equal to the number of departments and a softmax layer. The cross-entropy loss between the label data and the output of the softmax layer is the loss of the entire network. Continuously optimizing the loss and reducing the loss value is the training process. When the decrease in the loss value approaches stability or is less than a certain preset value, the training ends.

[0231] In some embodiments, the method for generating other auxiliary diagnosis and treatment plans is as follows:

[0232] Connect the patient's condition and medical history, and sequentially combine multiple suffix prompts as the input of the text generation pre-training model, so that the text output obtained by reasoning is a medical diagnosis or treatment.

[0233] Among them, the input of the text generation pre-training model is: "The patient's condition is as follows: Text(condition), Text(medical history), please give Text(prompt)".

[0234] The output of the text generation pre-training model is: Text (the solution corresponding to the prompt). Exemplarily, the prompt can include, but is not limited to, the following situations:

[0235] Prompt1: The possible disease names of the patient.

[0236] Prompt2: The diagnosis results including the severity of the condition and the mortality rate.

[0237] Prompt3: The probability of the patient being hospitalized.

[0238] Prompt4: The treatment plan for the patient.

[0239] Prompt5: The medication plan including the drug names, available medications, drug prices, and drug combinations.

[0240] Prompt6: Including diet recommendations, diet precautions, and recipes.

[0241] When the patient confirms the attending doctor, the above recommended information can be synchronized to the doctor, and then a diagnosis and treatment plan can be generated. The specific process can be referred to Figures 5 to 11 in the embodiments, and the embodiments of the present application will not elaborate on this.

[0242] In some embodiments, technology is becoming a part of people's lives and even their bodies. Wearable medical devices are helping people adjust some physical conditions and achieve the best operation of the body. Patients can monitor their physical health conditions through wearable devices, monitor important vital sign parameters, detect changes or abnormalities, and current wearable devices play a key role in preventing many common diseases. For example, in the case of hypertension, many cases of hypertension do not cause obvious symptoms, and over time, it can lead to stroke, heart failure, atrial fibrillation or other dangerous diseases. Vital sign data such as heart rate, sleep condition, blood pressure, and blood sugar can all be obtained through smart wearable devices.

[0243] Such as Figure 15 shown, important vital sign data such as heart rate, blood pressure, and blood sugar are monitored through smart wearable devices. These devices will be synchronized to user terminals such as computers, mobile phones, tablets and other portable devices. When the wounded person requests rescue guidance by phone through the terminal device, after obtaining the permission to read the vital sign data of the wounded person, the vital sign data can be synchronized to the rescue guidance service through 5G high-speed transmission. After texturing the structured data of heart rate, blood pressure, blood sugar, etc. (such as text description: the heart rate is 80 beats per minute), it is used as the input of the triage model and the text generation pre-training large model in this application.

[0244] Based on the foregoing embodiments, an embodiment of the present application provides an auxiliary triage device, which includes each module included and each unit included in each module, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0245] Figure 16 It is a schematic diagram of the composition structure of an auxiliary triage device provided by an embodiment of the present application. As Figure 16 shown, the auxiliary triage device 1600 includes:

[0246] A first acquisition module 1601, configured to acquire the patient's illness information, where the illness information at least includes the patient's illness description text and the patient's illness description image, and the patient's illness description image at least includes the image of the patient's affected part;

[0247] A triage module 1602, configured to input the illness information into a triage model, where the triage model at least includes an image model, a text model, and a forward network model that are sequentially trained; the training set of the image model includes the illness description images of the patients who have visited at least one department, the training set of the text model includes the illness description texts of the patients who have visited, and at least a part of the illness description texts of the patients who have visited is obtained by text extraction from the illness description images of the patients who have visited, and the training set of the forward network model at least includes the output of the image model and the output of the text model;

[0248] A determination module 1603, configured to determine the triage result of the patient according to the output of the triage model, where the triage result is used to indicate the department where the patient needs to visit.

[0249] In some embodiments, the first acquisition module includes: a generation unit, configured to input the image of the patient's affected part into an image generation text model, and the image generation text model is used to convert image information into text information; a first acquisition unit, configured to obtain the text description information corresponding to the patient according to the output of the image generation text model; a first addition unit, configured to add the text description information corresponding to the patient to the patient's illness description text.

[0250] In some embodiments, a first obtaining unit is configured to send the output of an image-to-text model to the patient client of the patient; obtain the adjusted text input by the patient through the patient client, where the adjusted text is the text obtained by the patient's adjustment of the output of the image-to-text model; and use the adjusted text as the text description information corresponding to the patient.

[0251] In some embodiments, a first obtaining module is configured to perform at least one of the following: obtain the vital sign information of the patient, where the vital sign information is used to characterize the vital signs of the patient; add the vital sign information of the patient to the text describing the condition; and add at least one of the patient's past medical history and family history to the text describing the patient's condition.

[0252] In some embodiments, an obtaining module is configured to input the image of the diseased part of the patient into an image depth model, where the image depth model is used to extract the depth information of the image of the diseased part of the patient; obtain the depth image corresponding to the patient according to the output of the image depth model; and add the depth image corresponding to the patient to the image describing the patient's condition.

[0253] In some embodiments, the depth information model at least includes: a first feature extraction module and a second feature extraction module. The first feature extraction module is configured to perform feature extraction on the image of the diseased part of the patient according to a first granularity, and the second feature extraction module is configured to perform feature extraction on the image of the diseased part of the patient according to a second granularity, where the first granularity is greater than the second granularity; the second feature extraction module includes M processing units, and the output of the M processing units is the output of the image depth model; the input of the i-th processing unit in the M processing units is the fused image corresponding to the (i - 1)-th processing unit in the M processing units, and the fused image is the image obtained by fusing the output of the first feature extraction module and the output of the (i - 1)-th processing unit; where M is an integer greater than or equal to 2, and i is a positive integer less than or equal to M.

[0254] In some embodiments, the obtaining module includes: an input unit configured to input the image of the diseased part of the patient into an image segmentation model, where the image segmentation model is used to perform image segmentation on the image of the diseased part of the patient to identify the diseased part of the patient; a second obtaining unit configured to obtain the segmented image corresponding to the patient according to the output of the image segmentation model; and a second adding unit configured to add the segmented image corresponding to the patient to the image describing the patient's condition.

[0255] Wherein, the second obtaining unit is configured to: send the output of the image segmentation model to the patient client of the patient; obtain the adjusted image input by the patient through the patient client, where the adjusted image is the image obtained by the patient's adjustment of the output of the image segmentation model; and use the adjusted image as the segmented image corresponding to the patient.

[0256] In some embodiments, a first acquisition module is configured to: collect at least one frame of images of a patient during exercise, where the at least one frame of images includes human body images of the patient from at least one angle; detect the human body images to determine angle features corresponding to the patient, where the angle features are used to represent the angles of at least one key human body part of the patient; and add the angle features corresponding to the patient to the patient's disease information.

[0257] In some embodiments, a first acquisition module is configured to: collect at least one frame of face images of a patient; perform expression category recognition on the at least one frame of face images to obtain expression category information of the patient, where the expression category information is used to represent the patient's mood; and add the expression category information of the patient to the patient's disease information.

[0258] In some embodiments, a determination module is configured to: obtain multiple department information according to the output of a triage model, where the department information is used to indicate the departments where the patient can seek medical treatment; and sort the departments where medical treatment can be sought according to the multiple department information to obtain a department sequence, where the department sequence is used to indicate the departments where the patient is to seek medical treatment.

[0259] In some embodiments, the apparatus further includes: a first generation module configured to input a patient's condition description text and at least one prompt word into a text generation pre-trained model, where the prompt word is used to indicate items included in the patient's diagnosis and treatment plan; and a second acquisition module configured to obtain the patient's diagnosis and treatment plan output by the text generation pre-trained model.

[0260] In some embodiments, the above-mentioned apparatus further includes: a third acquisition module configured to acquire patient location information of the location where the patient is located and department location information of the location where the department to be visited is located; and a second generation module configured to generate clinic navigation information according to the patient location information and the department location information, where the clinic navigation information is used to indicate the path between the location where the patient is located and the location where the department to be visited is located.

[0261] In some embodiments, the above device further includes: a fourth acquisition module, configured to acquire a first training set before the first acquisition module acquires the patient's disease information, where the first training set includes the disease description images of the patients who have already visited the doctor and the disease description texts of the patients who have already visited the doctor; a first training module, configured to use the disease description images of the patients who have already visited the doctor as a training set to train an initial image model until the initial image model converges, so as to obtain an image model; a second training module, configured to, after obtaining the image model, use the disease description texts of the patients who have already visited the doctor as a training set to train an initial text model until the initial text model converges, so as to obtain a text model; a third training module, configured to, after obtaining the text model, use at least the output of the initial image model and the output of the initial text model as a training set to train an initial forward network model until the initial forward network model converges, so as to obtain a forward network model; a triage model establishment module, configured to obtain a triage model based on the image model, the text model, and the forward network model.

[0262] In some embodiments, the fourth acquisition module further includes: a sample collection unit, configured to acquire the disease site images of the patients who have already visited the doctor; a disease site recognition unit, configured to input the disease site images of the patients who have already visited the doctor into an image segmentation model to identify the disease sites of the patients who have already visited the doctor; a third acquisition unit, configured to acquire the segmented images corresponding to the patients who have already visited the doctor according to the output of the image segmentation model; a third addition unit, configured to add the segmented images corresponding to the patients who have already visited the doctor to the disease description images of the patients who have already visited the doctor.

[0263] In some embodiments, the fourth acquisition module further includes: a depth information extraction unit, configured to input the disease site images of the patients who have already visited the doctor into an image depth model, where the image depth model is used to extract the depth information of the disease site images of the patients who have already visited the doctor; a fourth acquisition unit, configured to acquire the depth images corresponding to the patients who have already visited the doctor according to the output of the image depth model; a fourth addition unit, configured to add the depth images corresponding to the patients who have already visited the doctor to the disease description images of the patients who have already visited the doctor.

[0264] In some embodiments, the third training module includes: a fifth acquisition unit, configured to acquire at least one of the expression category information of the patients who have already visited the doctor and the angle feature corresponding to the patients who have already visited the doctor before using at least the output of the initial image model and the output of the initial text model as a training set to train the initial forward network model, where the expression category information of the patients who have already visited the doctor is used to represent the emotions of the patients who have already visited the doctor; the angle feature corresponding to the patients who have already visited the doctor is used to represent the angles of at least one human key part of the patients who have already visited the doctor.

[0265] In some embodiments, a third training module is configured to use at least one of the expression category information of the patients who have received medical treatment and the angular features corresponding to the patients who have received medical treatment, the output of the initial image model, and the output of the initial text model as a training set to train the initial forward network model.

[0266] The description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the embodiments of the present application, please refer to the description of the method embodiments of the embodiments of the present application for understanding.

[0267] It should be noted that in the embodiments of the present application, if the above-mentioned auxiliary triage method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any arbitrary combination of hardware, software, and firmware.

[0268] The embodiments of the present application provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0269] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.

[0270] The embodiments of the present application provide a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.

[0271] An embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. This computer program product can be specifically implemented in a way of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0272] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities or similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the device, storage medium, computer program, and computer program product embodiments of the embodiments of the present application, please refer to the descriptions of the method embodiments of the embodiments of the present application for understanding.

[0273] It should be noted that Figure 17 is a schematic diagram of a hardware entity of a computer device in an embodiment of the present application. As Figure 17 shown, the hardware entity of the computer device 1700 includes: a processor 1701, a communication interface 1702, and a memory 1703, where:

[0274] The processor 1701 generally controls the overall operation of the computer device 1700.

[0275] The communication interface 1702 can enable the computer device to communicate with other terminals or servers through a network.

[0276] The memory 1703 is configured to store instructions and applications executable by the processor 1701, and can also cache data to be processed or already processed by the processor 1701 and each module in the computer device 1700 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). Data transmission can be performed between the processor 1701, the communication interface 1702, and the memory 1703 through a bus 1704.

[0277] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above steps / processes do not mean the order of execution. The order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The sequence numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0278] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.

[0279] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings or communication connections between the components shown or discussed can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical or other forms.

[0280] The units described as separate components above may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0281] In addition, each functional unit in the embodiments of the present application can be all integrated in a processing unit, or each unit can be separately taken as a unit, or two or more units can be integrated in one unit; the above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0282] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs that can store program codes.

[0283] Alternatively, if the above integrated units of the present application are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the related art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in various embodiments of the present application. And the foregoing storage medium includes: various media such as removable storage devices, ROM, magnetic disks, or optical discs that can store program codes.

[0284] As described above, the above are only the implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application.

Claims

1. An auxiliary triage method, characterized in that, Including: Obtaining the patient's disease information, where the disease information at least includes the patient's disease description text and the patient's disease description image, and the patient's disease description image at least includes the image of the diseased part of the patient; Inputting the disease information into a triage model, where the triage model at least includes a trained image model, a text model, and a forward network model; the training set of the image model includes the disease description images of the patients who have visited at least one department, the training set of the text model includes the disease description texts of the patients who have visited, and the training set of the forward network model at least includes the output of the image model and the output of the text model; Determining the triage result of the patient according to the output of the triage model, where the triage result is used to indicate the department where the patient is to be seen.

2. The method according to claim 1, characterized in that, The obtaining of the patient's disease information includes: Inputting the image of the diseased part of the patient into an image generation text model, where the image generation text model is used to convert image information into text information; Obtaining the text description information corresponding to the patient according to the output of the image generation text model; Adding the text description information corresponding to the patient to the patient's disease description text.

3. The method according to claim 2, characterized in that, The obtaining of the text description information corresponding to the patient according to the output of the image generation text model includes: Sending the output of the image generation text model to the patient's client; Obtaining the adjusted text input by the patient through the patient's client, where the adjusted text is the text after the output of the image generation text model is adjusted by the patient; Taking the adjusted text as the text description information corresponding to the patient.

4. The method according to claim 1, characterized in that, The obtaining of the patient's disease information includes: Obtaining the vital sign information of the patient, where the vital sign information is used to characterize the vital signs of the patient; Adding the vital sign information of the patient to the disease description text; Adding at least one of the patient's past medical history and the patient's family medical history to the patient's disease description text.

5. The method according to claim 1, wherein The obtaining of the patient's disease information includes: Inputting the image of the diseased part of the patient into an image depth model, where the image depth model is used to extract the depth information of the image of the diseased part of the patient; Obtaining the depth image corresponding to the patient according to the output of the image depth model; Adding the depth image corresponding to the patient to the patient's disease description image.

6. The method according to claim 5, wherein The image depth information model at least includes: a first feature extraction module and a second feature extraction module. The first feature extraction module is used to extract features of the image of the diseased part of the patient according to a first granularity, and the second feature extraction module is used to extract features of the image of the diseased part of the patient according to a second granularity, where the first granularity is greater than the second granularity; The second feature extraction module includes M processing units, and the output of the M processing units is the output of the image depth model. The input of the $i$-th processing unit in the $M$-layer processing unit is the fused image corresponding to the $(i - 1)$-th processing unit in the $M$-layer processing unit, and the fused image is an image obtained by fusing the output of the first feature extraction module and the output of the $(i - 1)$-th processing unit; where $M$ is an integer greater than or equal to 2, and $i$ is a positive integer less than or equal to $M$.

7. The method according to claim 1, wherein The obtaining of the patient's disease information includes: Inputting the image of the patient's diseased part into an image segmentation model, where the image segmentation model is used to perform image segmentation on the image of the patient's diseased part to identify the patient's diseased part; Obtaining the segmented image corresponding to the patient according to the output of the image segmentation model; Adding the segmented image corresponding to the patient to the patient's condition description image.

8. The method according to claim 7, characterized in that, The obtaining of the segmented image corresponding to the patient according to the output of the image segmentation model includes: Sending the output of the image segmentation model to the patient's client; Obtaining the adjusted image input by the patient through the patient's client, where the adjusted image is an image obtained by the patient adjusting the output of the image segmentation model; Taking the adjusted image as the segmented image corresponding to the patient.

9. The method according to any one of claims 1 to 8, characterized in that, The obtaining of the patient's disease information includes: Collecting at least one frame of image of the patient during exercise, where the at least one frame of image includes human body images of at least one angle of the patient; Detecting the human body images to determine the angle features corresponding to the patient, where the angle features are used to represent the angles of at least one human body key part of the patient; Adding the angle features corresponding to the patient to the patient's disease information.

10. The method according to any one of claims 1 to 8, characterized in that, The obtaining of the patient's disease information includes: Collecting at least one frame of face image of the patient; Performing expression category recognition on the at least one frame of face image to obtain the expression category information of the patient, where the expression category information is used to represent the patient's emotion; Adding the expression category information of the patient to the patient's disease information.

11. The method according to any one of claims 1 to 8, characterized in that The determining of the triage result of the patient according to the output of the triage model includes: Obtaining multiple department information according to the output of the triage model, where the department information is used to indicate the departments where the patient can seek medical treatment; Sorting the departments where medical treatment can be sought according to the multiple department information to obtain a department sequence, where the department sequence is used to indicate the departments where the patient is to seek medical treatment.

12. The method according to any one of claims 1 to 8, characterized in that, After obtaining the patient's disease information, the method further includes: Inputting the patient's condition description text and at least one prompt word into a text generation pre-trained model, where the prompt word is used to indicate the items included in the patient's diagnosis and treatment plan; Obtaining the patient's diagnosis and treatment plan output by the text generation pre-trained model.

13. The method according to any one of claims 1 to 8, characterized in that After obtaining the triage result, the method further includes: Obtaining the patient location information of the location where the patient is located and the department location information of the location where the to-be-visited department is located; Generate a consulting room navigation information according to the patient position information and the department position information, where the consulting room navigation information is used to indicate a path between the position where the patient is located and the position of the department to be visited.

14. The method according to any one of claims 1 to 8, characterized in that, Before obtaining the patient's illness information, the method further includes: Obtain a first training set, where the first training set includes the disease description images of the patients who have received medical treatment and the disease description texts of the patients who have received medical treatment; Use the disease description images of the patients who have received medical treatment as a training set to train an initial image model until the initial image model converges to obtain the image model; After obtaining the image model, use the disease description texts of the patients who have received medical treatment as a training set to train an initial text model until the initial text model converges to obtain the text model; After obtaining the text model, use at least the output of the initial image model and the output of the initial text model as a training set to train an initial forward network model until the initial forward network model converges to obtain the forward network model; Obtain the triage model based on the image model, the text model, and the forward network model.

15. The method according to claim 14, wherein The obtaining of the first training set includes: Obtain the disease part image of the patient who has received medical treatment; Input the disease part image of the patient who has received medical treatment into the image segmentation model to identify the disease part of the patient who has received medical treatment; Obtain the segmented image corresponding to the patient who has received medical treatment according to the output of the image segmentation model; Add the segmented image corresponding to the patient who has received medical treatment to the disease description image of the patient who has received medical treatment.

16. The method according to claim 14, wherein The obtaining of the first training set includes: Obtain the disease part image of the patient who has received medical treatment; Input the disease part image of the patient who has received medical treatment into an image depth model, where the image depth model is used to extract the depth information of the disease part image of the patient who has received medical treatment; Obtain the depth image corresponding to the patient who has received medical treatment according to the output of the image depth model; Add the depth image corresponding to the patient who has received medical treatment to the disease description image of the patient who has received medical treatment.

17. The method according to claim 14, wherein Before using at least the output of the initial image model and the output of the initial text model as a training set to train an initial forward network model, the method further includes: Obtain at least one of the expression category information of the patient who has received medical treatment and the angle feature corresponding to the patient who has received medical treatment, where the expression category information of the patient who has received medical treatment is used to represent the emotion of the patient who has received medical treatment; the angle feature corresponding to the patient who has received medical treatment is used to represent the angles of at least one human key part of the patient who has received medical treatment; The using at least the output of the initial image model and the output of the initial text model as a training set to train an initial forward network model includes: Use at least one of the expression category information of the patient who has received medical treatment and the angle feature corresponding to the patient who has received medical treatment, the output of the initial image model, and the output of the initial text model as a training set to train the initial forward network model.

18. An auxiliary triage device, characterized in that, Includes: A first acquisition module, configured to acquire the disease information of a patient, where the disease information at least includes the text description of the patient's condition and the image description of the patient's condition, and the image description of the patient's condition at least includes the image of the diseased part of the patient; A triage module, configured to input the disease information into a triage model, where the triage model at least includes an image model, a text model, and a forward network model that are trained in sequence; the training set of the image model includes the image descriptions of the conditions of the patients who have visited at least one department, the training set of the text model includes the text descriptions of the conditions of the patients who have visited, and at least a part of the text descriptions of the conditions of the patients who have visited is obtained by text extraction from the image descriptions of the conditions of the patients who have visited, and the training set of the forward network model at least includes the output of the image model and the output of the text model; A determination module, configured to determine the triage result of the patient according to the output of the triage model, where the triage result is used to indicate the department to which the patient should go for treatment.

19. An electronic device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the computer program, the steps in the method according to any one of claims 1 to 17 are implemented.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 17 are implemented.

21. A computer program product, the computer program product comprising a non-transitory computer-readable storage medium storing a computer program, characterized in that, When the computer program is read and executed by a computer, the steps in the method according to any one of claims 1 to 17 are implemented.