Traditional Chinese medicine heart failure intelligent diagnosis and treatment system and method based on machine vision

Through the intelligent diagnosis and treatment system of heart failure in traditional Chinese medicine based on machine vision, the intelligent tongue diagnosis instrument, facial acquisition camera and infrared thermal imager are used to collect data, combined with multimodal fusion deep learning model, the existing problems of low efficiency and difficulty in diagnosis are solved, early screening and personalized diagnosis are achieved, and diagnostic accuracy and convenience are improved.

CN120284204APending Publication Date: 2025-07-11JINAN UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510435631.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing heart failure disease detection system requires doctors to study the examination data of each patient. The diagnosis efficiency is low, and the patient needs to go to the hospital for data collection multiple times. The diagnosis results are difficult to standardize, and there are strong human experience differences.

Method used

Using a Chinese medicine heart failure intelligent diagnosis and treatment system based on machine vision, the patient's tongue image, facial acquisition camera and infrared thermal imager is used to collect the patient's tongue image, facial and infrared heat map data through intelligent tongue diagnostic instruments, facial acquisition cameras and infrared thermal imagers. The patient's tongue image, facial and infrared heat maps are extracted using image segmentation algorithms and convolutional neural networks, and features are combined with multimodal fusion attention mechanisms and deep learning models to realize the fusion of multi-source data of tongue image, facial and infrared heat maps, generate heart failure diagnosis vectors, and perform early screening and Chinese medicine syndrome identification.

Benefits of technology

Early screening of heart failure and identification of traditional Chinese medicine syndromes has been achieved, the accuracy and efficiency of diagnosis has been improved, the differences in human experience have been reduced, personalized traditional Chinese medicine treatment recommendations have been provided, the sensitivity and convenience of diagnosis has been improved, and telemedicine and dynamic disease monitoring have been supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120284204A_ABST
    Figure CN120284204A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent medical treatment and traditional Chinese medicine diagnosis and treatment, in particular to a traditional Chinese medicine heart failure intelligent diagnosis and treatment system and method based on machine vision, and the diagnosis and treatment method comprises the steps: collecting a tongue image and a face image of a patient and an infrared thermogram of a corresponding part, and carrying out the preprocessing of the collected tongue image, face image and infrared thermogram; features in the tongue image, the face image and the infrared thermal image are quantified into tongue image feature vectors, face feature vectors and temperature feature vectors by using different algorithms, a multi-modal fusion deep learning model is established, and the feature vectors are input into the multi-modal fusion deep learning model for training; and inputting the collected tongue image, the face image and the infrared thermal image into a multi-modal fusion deep learning model, outputting a diagnosis result, and carrying out visual display. The diagnosis and treatment system is applied to the diagnosis and treatment method, the system can objectively quantify traditional Chinese medicine diagnosis information, the sensitivity of early recognition is remarkably improved, and the accuracy and efficiency of diagnosis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent medicine and traditional Chinese medicine diagnosis, and particularly relates to a traditional Chinese medicine heart failure intelligent diagnosis and treatment system and method based on machine vision. Background Art

[0002] Heart failure, abbreviated as HF, refers to the syndrome that occurs when the systolic and diastolic functions of the heart are impaired, and the venous return blood volume cannot be fully discharged from the heart, resulting in blood stasis in the venous system and insufficient blood perfusion in the arterial system, thus causing a cardiac circulatory disorder syndrome. It is a common and critical cardiovascular disease. Early detection and intervention are crucial for improving the prognosis of HF. In traditional Chinese medicine diagnosis, traditional methods such as observing the tongue and face have important value in clinical practice. However, due to strong subjectivity and experience dependence, the diagnostic results are difficult to standardize, and it is difficult to ensure the consistency of examination data among different doctors. Especially in the diagnosis and treatment of complex chronic diseases such as HF, if only relying on traditional experience, it may be difficult to capture early lesion information in a timely and accurate manner.

[0003] With the rapid development of artificial intelligence and machine vision, objectively quantifying features such as tongue images and facial colors through high-precision image acquisition and deep learning models has become an important direction for the modernization of traditional Chinese medicine. The emergence of new technologies such as infrared thermal imaging and multispectral imaging has made it possible to analyze diseases from a multi-modal perspective, and it is possible to comprehensively and quantitatively capture the changes in the overall functions emphasized in traditional Chinese medicine. However, there is currently a lack of an integrated system that integrates tongue images, facial features, infrared thermal maps, and spectral information and is oriented towards the diagnosis and treatment of heart failure in traditional Chinese medicine. Moreover, existing heart failure disease detection systems require doctors to study the examination data of each patient to determine whether they are heart failure patients, and then give diagnostic suggestions according to their specific conditions. This process will waste a lot of time of patients and doctors, with low diagnostic efficiency, and patients need to go to the hospital multiple times for data collection, which is rather inconvenient. Summary of the Invention

[0004] The purpose of the present invention is to provide a traditional Chinese medicine heart failure intelligent diagnosis and treatment system and method based on machine vision to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A traditional Chinese medicine heart failure intelligent diagnosis and treatment method based on machine vision, including the following steps:

[0006] S1: Use an intelligent tongue diagnosis instrument to collect the tongue image features of a patient to obtain the tongue image of the patient, use a facial acquisition camera to collect the facial features of the patient to obtain the facial image of the patient, and use an infrared thermal imager to collect the infrared thermal map of the whole body or a part of the patient to obtain the infrared thermal map of the corresponding part of the patient;

[0007] S2: Preprocess the collected tongue image, facial image, and infrared thermal image;

[0008] S3: Extract the tongue body from the preprocessed tongue image through an image segmentation algorithm, use a convolutional neural network to quantify the tongue texture and coating features of the tongue body into a tongue image feature vector W1, use a face recognition system to perform skin color analysis and key point localization on the preprocessed facial image, extract the skin color abnormality features of the key points and quantify them into a facial feature vector W2, search for the thermal abnormality area in the preprocessed infrared thermal image, and extract the temperature of the thermal abnormality area to generate a temperature feature vector W3;

[0009] S4: Adopt a multi-modal fusion attention mechanism to fuse the tongue image feature vector W1, facial feature vector W2, and temperature feature vector W3 to generate a multi-modal fusion feature vector F;

[0010] S5: Introduce the heart failure syndrome set, and quantify and encode the heart failure syndrome set to generate a heart failure fixed vector W. Re-fuse the heart failure fixed vector W with the multi-modal fusion feature F to construct a heart failure diagnosis vector A, and calculate the heart failure risk probability and heart failure syndrome categories;

[0011] S6: Establish a multi-modal fusion deep learning model, input the collected tongue image feature vector W1, facial feature vector W2, and temperature feature vector W3 into the multi-modal fusion deep learning model for training, and iteratively optimize the weight vectors of the tongue image feature vector W1, facial feature vector W2, and temperature feature vector W3;

[0012] S7: Input the collected tongue image, facial image, and infrared thermal image into the multi-modal fusion deep learning model, predict the heart failure risk probability, classify the heart failure syndrome categories, and perform visual display at the same time.

[0013] Further, the S3 includes the following steps: precisely separate the tongue body from the background in the preprocessed tongue image through an image segmentation algorithm, detect the crack distribution features, tooth mark distribution features, tongue texture features, and coating features on the tongue body based on a convolutional neural network and encode them. The crack distribution features are quantified into the number of cracks and crack length, the tooth mark distribution features are quantified into the number of tooth marks and tooth mark depth, the tongue texture features are quantified into the color threshold of the tongue texture, and the coating features are quantified into the thickness of the coating and the color threshold of the coating. Normalize the quantified crack distribution features, tooth mark distribution features, tongue texture features, and coating features to generate a tongue image feature vector W1.

[0014] Further, S3 further includes the following steps: The face recognition system locates the key point positions of the eyelids, lips or cheekbones in the preprocessed facial image through the MTCNN or Facenet algorithm and encodes them, extracts abnormal indicators such as dullness, flushing or cyanosis in the key point positions through skin color analysis, and extracts the color thresholds of the abnormal indicators for normalization processing to generate the facial feature vector W2.

[0015] Further, S3 further includes the following steps: Extract the thermal anomaly features in the preprocessed infrared thermal image in the areas of the forehead, nose tip, extremities or chest and abdomen through the heat map analysis network and encode them, and extract the temperature gradient or abnormal heat area range of the corresponding parts for normalization processing to generate the temperature feature vector W3.

[0016] Further, the multimodal fusion feature vector F = α1·W1 + α2·W2 + α3·W3, where α1, α2, and α3 are the weights of the tongue image feature vector, facial feature vector, and temperature feature vector respectively, and α1, α2, and α3 are adaptively obtained through the shared Softmax attention network mechanism.

[0017] Further, S5 includes the following steps: The heart failure diagnosis vector A = Concat(F, W), the heart failure syndrome category P1 = Softmax(Wf·A + b), and the heart failure tendency probability P2 = Sigmoid(Wf·A + b), where Wf is the weight parameter for heart failure discrimination and b is the heart failure judgment error factor.

[0018] Further, the iterative optimization in S6 includes the following steps:

[0019] S61: Initialize the tongue image feature vector W1, facial feature vector W2, and temperature feature vector W3, as well as the corresponding weights α1, α2, and α3.

[0020] S62: Construct a comprehensive loss function L, where the loss function L includes a classification loss Lc, a modality consistency loss Lm, and a regularization term Lr, that is: L = Lc + λ1Lm + λ2Lr, and λ1, λ2 are preset regularization coefficients.

[0021] S63: Use the stochastic gradient descent method to optimize the network parameters of the multimodal fusion deep learning model.

[0022] S64: Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the network parameters.

[0023] S65: Update the values of the network parameters according to the gradient information and the optimization algorithm.

[0024] Repeat S64 and S65 until the weight vector reaches the maximum number of iterations or the loss function converges.

[0025] A traditional Chinese medicine heart failure intelligent diagnosis and treatment system based on machine vision, which is applied to any one of the described traditional Chinese medicine heart failure intelligent diagnosis and treatment methods based on machine vision, includes a front-end data acquisition module, a back-end AI analysis module, and a clinical terminal module respectively connected to a central control module;

[0026] Central control module: Used to control each module to work properly through the host;

[0027] The front-end data acquisition module includes an intelligent tongue diagnosis instrument, an infrared thermal imager, and a facial acquisition camera, which are used to obtain multi-modal images in real time;

[0028] The back-end AI analysis module includes a multi-modal feature fusion model and other algorithms. The back-end AI analysis module runs on a local server or in the cloud, completes image parsing and diagnostic reasoning, and gives a diagnostic result;

[0029] The clinical terminal module includes a doctor workstation or a mobile tablet, which realizes the visual display of medical record management, diagnostic results, and treatment suggestions, and docks with the hospital system.

[0030] Furthermore, a high-definition camera and a ring-shaped LED light source are set in the intelligent tongue diagnosis instrument. The resolution of the high-definition camera is ≥2 million pixels. The color temperature of the ring-shaped LED light source is 5500K±200K, and the color rendering index is ≥90. The standardized light source maintains an illuminance of about 2000~3000lx in the tongue surface area. The configuration of the facial acquisition camera is the same as that of the high-definition camera.

[0031] Furthermore, the infrared detector in the infrared thermal imager uses an uncooled microbolometer, with a wavelength band of 8~14μm, a resolution of 320×240, 640×480, or 1024×768, and a temperature measurement accuracy of ±0.5℃.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] 1. The traditional Chinese medicine heart failure intelligent diagnosis and treatment system based on machine vision provided by the present invention integrates multi-source data such as tongue images, facial images, and infrared thermal images into heart failure diagnosis and treatment. Through comprehensive analysis of the patient's multi-source data, it realizes early screening of heart failure and identification of traditional Chinese medicine syndromes, and gives personalized traditional Chinese medicine treatment suggestions. This system can objectively quantify traditional Chinese medicine diagnosis information, and the physiological and pathological characteristics from multiple angles corroborate each other, reducing the differences in human experience, significantly improving the sensitivity of early recognition, and improving the accuracy and efficiency of diagnosis. At the same time, it is of great significance for the course tracking and dynamic adjustment of treatment plans for heart failure patients in clinical practice.

[0034] 2. The intelligent diagnosis and treatment method of traditional Chinese medicine for heart failure based on machine vision provided by the present invention conducts joint modeling of tongue image, facial image and infrared thermal image. The three-modal features complement each other, more comprehensively depicting the typical traditional Chinese medicine manifestations such as deficiency, cold, stasis, and water under the syndrome of heart failure. Through the attention mechanism, dynamic adjustment of modal weights is realized, enabling different patients to adapt to different modal focuses, improving the accuracy of personalized diagnosis. Embedding traditional Chinese medicine syndromes in the model enhances the traditional Chinese medicine semantic ability and interpretability of model reasoning.

[0035] 3. The intelligent diagnosis and treatment method of traditional Chinese medicine for heart failure based on machine vision provided by the present invention. In the outpatient clinic or specialist consulting room, a doctor can guide the patient to complete the collection of images such as the tongue surface and face. The data is uploaded to the AI platform in real time, and a diagnosis report and traditional Chinese medicine prescription suggestions can be returned within seconds for the doctor to make a comprehensive decision. The diagnosis results and prescription suggestions can be automatically recorded in the electronic medical record, and subsequent efficacy tracking and follow-up can be carried out. For patients diagnosed with heart failure, it is also possible to regularly collect their tongue surface and face follow-up images, dynamically monitor the condition and prompt possible recurrence or transformation. Moreover, patients can use a simple camera device at home to upload images, and doctors can view the AI analysis results online and remotely guide the adjustment of prescriptions or conduct rehabilitation management, greatly improving the diagnosis efficiency and enhancing the convenience of patients' medical treatment.

[0036] 4. The intelligent diagnosis and treatment method of traditional Chinese medicine for heart failure based on machine vision provided by the present invention. When the model determines that a patient has a specific traditional Chinese medicine syndrome, it can automatically retrieve the corresponding plan in the knowledge base and make personalized adjustments in combination with information such as the patient's age, constitution, and concomitant diseases. The diagnosis and treatment suggestions list the analysis results of key tongue surface and face features and the theoretical basis of traditional Chinese medicine, enhancing the interpretability of the plan and the trust of doctors. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic flow chart of the diagnosis and treatment method of the present invention;

[0038] Figure 2 It is a schematic structural diagram of the diagnosis and treatment system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] In the description of the following invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper", "lower", "left", "right", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation. The term "connection" only represents the connection between devices and has no special meaning.

[0041] In addition, the technical fields and installation methods involved in the embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0042] Specific embodiments: Please refer to Figure 1 - Figure 2 , a traditional Chinese medicine heart failure intelligent diagnosis and treatment method based on machine vision, including the following steps:

[0043] S1: Use an intelligent tongue diagnosis instrument to collect the tongue image features of the patient to obtain the tongue image of the patient, use a facial collection camera to collect the facial features of the patient to obtain the facial image of the patient, and use an infrared thermal imager to collect the infrared thermal map of the whole body or part of the patient to obtain the infrared thermal map of the corresponding part of the patient;

[0044] S2: Preprocess the collected tongue image, facial image, and infrared thermal map.

[0045] S3: Extract the tongue body from the preprocessed tongue image through an image segmentation algorithm, use a convolutional neural network to quantify the tongue texture and tongue coating features of the tongue body into a tongue image feature vector W1, use a face recognition system to perform skin color analysis and key point localization on the preprocessed facial image, extract the skin color abnormal features of the key points and quantify them into a facial feature vector W2, and find the thermal abnormal area in the preprocessed infrared thermal map, and extract the temperature of the thermal abnormal area to generate a temperature feature vector W3;

[0046] S4: Use a multi-modal fusion attention mechanism to fuse the tongue image feature vector W1, facial feature vector W2, and temperature feature vector W3 to generate a multi-modal fusion feature vector F;

[0047] S5: Introduce a heart failure syndrome set, quantify and encode the heart failure syndrome set to generate a heart failure fixed vector W, fuse the heart failure fixed vector W with the multi-modal fusion feature F again to construct a heart failure diagnosis vector A, and calculate the heart failure risk probability and heart failure syndrome category;

[0048] S6: Establish a multi-modal fusion deep learning model, input the collected tongue image feature vector W1, facial feature vector W2, and temperature feature vector W3 into the multi-modal fusion deep learning model for training, and iteratively optimize the weight vectors of the tongue image feature vector W1, facial feature vector W2, and temperature feature vector W3;

[0049] S7: Input the collected tongue image, facial image, and infrared thermal image into a multi-modal fusion deep learning model to predict the risk probability of heart failure, classify the syndromes of heart failure, and simultaneously perform visual display.

[0050] Further, S1 includes the following steps: After the patient inserts the tongue surface into the collection port of the intelligent tongue diagnosis instrument, the device automatically takes multiple images, performs image quality detection and illumination correction to obtain the patient's tongue image. The facial image is collected by a facial collection camera. During collection, the patient needs to keep the expression relaxed and look directly at the facial collection camera to obtain the patient's facial image, so as to obtain information such as complexion, skin texture, and edema. Use an infrared thermal imager to collect the infrared thermal image of the patient's whole body or chest and abdomen, face to obtain the infrared thermal image of the corresponding part of the patient, which is used to analyze and observe the temperature distribution of blood vessel endings, facial heat areas, etc., and capture the common peripheral circulation abnormalities in heart failure patients.

[0051] Further, S2 includes the following steps: Use image processing methods or neural network preprocessing methods to denoise, crop, standardize, and color correct the collected tongue image, facial image, and infrared thermal image, which is convenient for the subsequent model to identify the fine structure and components of the tongue tissue and improve the model's recognition of potential lesions in the tongue image.

[0052] Further, S3 includes the following steps: Through an image segmentation algorithm, such as the improved Deeplabv3+, accurately separate the tongue body area in the tongue image or the facial skin area in the facial image. Based on a convolutional neural network, detect the crack distribution characteristics, tooth mark distribution characteristics, tongue texture characteristics, and tongue coating characteristics on the tongue body and encode them. The crack distribution characteristics are quantified as the number of cracks and the crack length. The tooth mark distribution characteristics are quantified as the number of tooth marks and the depth of tooth marks. The tongue texture characteristics are quantified as the color threshold of the tongue texture. The tongue coating characteristics are quantified as the thickness of the tongue coating and the color threshold of the tongue coating. Normalize the quantified crack distribution characteristics, tooth mark distribution characteristics, tongue texture characteristics, and tongue coating characteristics to generate the tongue image feature vector W1;

[0053] The specific steps for quantifying the tongue image feature vector W1 are as follows:

[0054] Tongue texture color or tongue coating color: Extract the pixels in the tongue body area, calculate the average RGB value of the area pixels or take the average value after converting the area pixels to the Lab color space. You can also further extract the standard deviation or hue, saturation and other indicators of the tongue body area pixels;

[0055] Tongue coating thickness: Use color contrast, depth estimation or use a CNN to extract the tongue coating feature map, and then estimate the tongue coating thickness level through image brightness / texture complexity. According to the thickness level: thick / thin / none, it can be quantified as 0 / 1 / 2;

[0056] Cracks: Extract high-frequency feature maps based on CNN, such as using the intermediate layer of ResNet or Grad-CAM, or combining image edge detection algorithms (Canny / Sobel) with morphological operations to quantify crack images into: the number of cracks, the total length of cracks, the average width of cracks, and the crack direction distribution. Among them, the crack direction distribution can be made into a direction distribution histogram;

[0057] Tooth marks: Identify the serrated structure based on the tongue edge contour, and use the CNN convolutional layer to detect the high-frequency changes of the serrated structure on the tongue edge, and quantify the tooth marks into: the number of tooth marks, the depth of each tooth mark, or the length / width ratio of the tooth mark area;

[0058] Normalization processing of tongue image features: Normalize the tongue color, tongue coating color, tongue coating thickness, the number of cracks, the total length of cracks, the average width of cracks, the crack direction distribution, the number of tooth marks, the depth of each tooth mark, or the length / width ratio of the tooth mark area to the interval [0,1] to generate the tongue image feature vector W1.

[0059] Furthermore, the face recognition system locates and encodes the key point positions of the eyelids, lips, or cheekbones in the preprocessed face image through the MTCNN or Facenet algorithm, extracts the abnormal indicators containing dullness, flushing, or cyanosis in the key point positions through skin color analysis, and extracts the color thresholds of the abnormal indicators for normalization processing to generate the face feature vector W2;

[0060] The specific quantification steps of the face feature vector W2 are as follows:

[0061] Face key area location: Use the Facenet algorithm to extract the key points in the face image, such as the center position and bounding box of the eyelid, lip, and cheekbone areas, crop each key area image and standardize the size for subsequent feature extraction;

[0062] Skin color and abnormal color analysis: Convert the key area pixels from the RGB space to the Lab or HSV color space and extract: L value: used to detect dullness, paleness, etc.; a value: used to detect flushing; b value: used for jaundice indication, where the L value is the brightness, the a value is the red-green axis, and the b value is the yellow-blue axis;

[0063] Abnormal indicator extraction: Mean color value (Mean), standard deviation (Std), and color deviation degree;

[0064] Color abnormality normalization processing: Normalize the L value, a value, b value, mean color value, standard deviation, and color deviation degree to the interval [0,1] to finally generate the face feature vector W2.

[0065] Further, the thermal anomaly features in the forehead, nose tip, extremities, or chest and abdomen regions of the preprocessed infrared thermal image are extracted by the thermal map analysis network and encoded, and the temperature gradient or the range of the abnormal thermal area of the corresponding part is extracted for normalization to generate the temperature feature vector W3;

[0066] The specific quantization steps of the temperature feature vector W3 are as follows:

[0067] Region extraction and segmentation: The thermal map analysis network is used to extract typical regions such as the forehead, nose tip, extremities, chest and abdomen. The following temperature indicators are extracted for each region: average temperature Tavg, maximum temperature Tmax, and minimum temperature Tmin, and the temperature gradient is calculated: ΔT = Tcenter - Tperiphery;

[0068] Thermal anomaly detection: Using a threshold or a learning method, determine whether there is a phenomenon of "hot spot" or "uneven heat and cold", calculate the area ratio or the number of abnormal regions. If the proportion of the abnormal thermal area > 10%, it is determined as abnormal;

[0069] Normalization of the abnormal thermal area: The average temperature Tavg, maximum temperature Tmax, minimum temperature Tmin, temperature gradient, and the area ratio or the number of abnormal regions of each region are normalized to the [0, 1] interval to generate the temperature feature vector W3.

[0070] Through the joint modeling of tongue image, facial image and infrared thermal image, the three-modal features complement each other, more comprehensively depicting the typical TCM manifestations such as deficiency, cold, stasis, and fluid retention under heart failure syndromes. The dynamic adjustment of the modal weights is realized through the attention mechanism, enabling different patients to adapt to different modal focuses, improving the accuracy of personalized diagnosis. Embedding TCM syndromes in the model enhances the TCM semantic ability and interpretability of the model inference.

[0071] Further, the multi-modal fusion feature vector F = α1·W1 + α2·W2 + α3·W3, where α1, α2, and α3 are the weights of the tongue image feature vector, facial image feature vector, and temperature feature vector respectively, and α1, α2, and α3 are adaptively obtained through the shared Softmax attention network mechanism.

[0072] Further, S5 includes the following steps: the heart failure diagnosis vector A = Concat(F, W), the heart failure syndrome category P1 = Softmax(Wf·A + b), the heart failure tendency probability P2 = Sigmoid(Wf·A + b), where Wf is the weight parameter for heart failure discrimination, b is the heart failure judgment error factor, the Softmax function is used to calculate the probability distribution of multiple heart failure syndrome categories, the output result of the Softmax function is a probability vector, and the sum of the probability vectors of all heart failure syndrome categories is 1. The Sigmoid function is used to calculate the tendency probability, and the output result of the Sigmoid function is a standard value between [0, 1], representing the risk tendency probability of the current sample having heart failure.

[0073] Further, the iterative optimization in S6 includes the following steps:

[0074] S61: Initialize the tongue image feature vector W1, the facial feature vector W2, the temperature feature vector W3, and the corresponding weights α1, α2, α3.

[0075] S62: Construct the comprehensive loss function L. The loss function L includes the classification loss Lc, the modality consistency loss Lm, and the regularization term Lr, i.e., L = Lc + λ1Lm + λ2Lr, where λ1 and λ2 are preset regularization coefficients.

[0076] S63: Use the stochastic gradient descent method to optimize the network parameters of the multi-modal fusion deep learning model.

[0077] S64: Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the network parameters.

[0078] S65: Update the values of the network parameters according to the gradient information and the optimization algorithm.

[0079] Repeat S64 and S65 until the maximum number of iterations is reached or the loss function converges.

[0080] Furthermore, S7 includes the following steps: Input the collected tongue image, facial image, and infrared thermal map into a multi-modal fusion deep learning model. Through the attention mechanism or weighted fusion strategy, the model automatically focuses on key feature combinations, such as "pale and swollen tongue body + dull complexion + abnormal local heat area", etc., to identify potential heart failure risks and classify traditional Chinese medicine syndromes. During the identification process, the system compares and performs weighted operations on the abnormal features of the patient's tongue image, complexion, and thermal map according to common traditional Chinese medicine syndromes of heart failure, such as heart qi deficiency, heart yang deficiency, blood stasis, and internal retention of fluid. It automatically discriminates the syndrome type and gives the corresponding syndrome tendency. If early heart failure risk is identified, the system will issue a warning prompt and visually mark the abnormal areas in the image, such as the purple and dull area on the tongue surface, the locally red or temperature-abnormal area on the face, etc. Finally, a diagnostic report is output. The diagnostic report includes the heart failure risk score or probability, the visualized feature thermal map, and the traditional Chinese medicine syndrome categories related to heart failure, such as heart qi deficiency, heart yang deficiency, internal retention of fluid, and blood stasis in the heart. At the same time, a diagnostic explanation is given, such as "pale and swollen tongue, dull complexion, indicating yang deficiency, and it is necessary to warm and tonify the heart yang", "pale and swollen tongue + white and slippery tongue coating + dull complexion, with a tendency of heart yang deficiency", etc. And the system has corresponding rules and weight tables built in, quantifying and accumulating the collected multi-modal features, and finally outputting the most likely 1-2 main syndrome types according to the scores of multiple syndromes. For example, if the model determines that "heart yang deficiency" is in the patient's diagnosis result, it can be shown in the report: "The tongue color is relatively light, the depth of the tooth marks on the tongue edge > 2 pixels, and the average facial temperature is 1.0°C lower than the normal baseline, indicating yang deficiency and cold excess, and it is advisable to warm yang and dispel cold".

[0081] Furthermore, S7 also includes the following steps: There is a traditional Chinese medicine clinical knowledge base set in the multi-modal fusion deep learning model. In the traditional Chinese medicine clinical knowledge base, there are classic prescriptions, traditional Chinese medicines, and acupuncture treatment plans corresponding to the syndrome categories of heart failure. When the model determines that the patient has a specific traditional Chinese medicine syndrome according to the diagnostic report, it automatically retrieves the corresponding plan in the knowledge base, and combines information such as the patient's age and constitution to add, subtract, or replace the prescription, lists the corresponding prescription recommendations in the diagnostic report, and at the same time lists the reasons for the recommendations in the diagnostic report, such as "The patient's tongue image shows that cold-dampness is relatively severe, so the prescription mainly for warming yang and promoting diuresis is used", etc., to enhance the doctor's understanding and trust in the AI-recommended plan.

[0082] Furthermore, S7 also includes the following steps: The diagnostic report and treatment suggestions are finally visually displayed on the doctor's workstation or mobile tablet. In actual application, in the outpatient clinic or specialist consulting room, the doctor can guide the patient to complete image acquisition of the tongue surface, face, etc. The data is uploaded to the AI platform in real time, and the diagnostic report and traditional Chinese medicine prescription suggestions can be returned within seconds for the doctor's comprehensive decision-making. The diagnostic results and prescription suggestions can be automatically recorded in the electronic medical record, and the curative effect can be tracked and followed up subsequently. For the diagnosed heart failure patients, the system can regularly collect their follow-up images of the tongue surface and face, dynamically monitor the condition and prompt possible recurrence or transformation. In the telemedicine scenario, the patient uses a simple version of the camera device at home to upload images, and the doctor can view the AI analysis results online and remotely guide the adjustment of the prescription or perform rehabilitation management, which greatly improves the diagnostic efficiency and enhances the convenience of the patient's medical treatment.

[0083] A traditional Chinese medicine heart failure intelligent diagnosis and treatment system based on machine vision is applied to any one of the above-mentioned traditional Chinese medicine heart failure intelligent diagnosis and treatment methods based on machine vision, and includes a front-end data acquisition module, a back-end AI analysis module, and a clinical terminal module that are respectively connected to the central control module;

[0084] Central control module: used to control each module to work properly through the host;

[0085] The front-end data acquisition module includes an intelligent tongue diagnosis instrument, an infrared thermal imager, and a face acquisition camera, and is used to obtain multi-modal images in real time;

[0086] The back-end AI analysis module includes a multi-modal feature fusion model and other algorithms. The back-end AI analysis module runs on the local server or in the cloud, and is responsible for uniformly encoding and fusing heterogeneous medical data from sources such as tongue image, face image, and infrared thermal image, realizing accurate identification of traditional Chinese medicine syndromes related to heart failure, and giving diagnostic results;

[0087] The clinical terminal module includes a doctor's workstation or a mobile tablet, realizes the visualization display of medical record management, diagnostic results, and treatment suggestions, and docks with the hospital system.

[0088] Furthermore, the intelligent tongue diagnosis instrument is provided with a high-definition camera and a ring-shaped LED light source that can simulate sunlight. The resolution of the high-definition camera is ≥2 million pixels to ensure clear capture of details such as tongue surface texture, cracks, and tongue coating thickness by the camera. The color temperature of the ring-shaped LED light source is 5500K±200K, and the color rendering index is ≥90, which can reduce the interference of ambient light while ensuring the true restoration of the tongue color. The ring-shaped LED light source maintains an illuminance of about 2000~3000lx in the tongue surface area, and a diffuser or soft light structure is used to avoid obvious shadows in the captured image. The configuration of the face acquisition camera is the same as that of the high-definition camera.

[0089] Further, the infrared detector in the infrared thermal imager uses an uncooled microbolometer, with a wavelength band of 8 - 14μm, a resolution of 320×240, 640×480 or 1024×768, and a temperature measurement accuracy of ±0.5°C.

[0090] Through comprehensive analysis of data such as the tongue image, facial features, and infrared thermal image of patients, the system realizes early screening for heart failure and identification of traditional Chinese medicine syndromes, and gives personalized traditional Chinese medicine treatment suggestions. The system can objectively quantify traditional Chinese medicine diagnostic information, with physiological and pathological features from multiple angles corroborating each other, reducing differences in human experience, significantly improving the sensitivity of early identification, and enhancing the accuracy and efficiency of diagnosis. At the same time, it is of great significance for the course tracking and dynamic adjustment of treatment plans for heart failure patients in clinical practice.

Claims

1. An intelligent diagnosis and treatment method for heart failure in traditional Chinese medicine based on machine vision, characterized in that: It includes the following steps: S1: Use an intelligent tongue diagnosis instrument to collect the tongue image features of the patient to obtain the tongue image of the patient, use a facial collection camera to collect the facial features of the patient to obtain the facial image of the patient, and use an infrared thermal imager to collect the infrared thermal map of the whole body or part of the patient to obtain the infrared thermal map of the corresponding part of the patient; S2: Preprocess the collected tongue image, facial image and infrared thermal map; S3: Extract the tongue body in the preprocessed tongue image through an image segmentation algorithm, use a convolutional neural network to quantify the tongue texture and tongue coating features of the tongue body into a tongue image feature vector W1, use a face recognition system to perform skin color analysis and key point localization on the preprocessed facial image, extract the skin color abnormality features of the key points and quantify them into a facial feature vector W2, find the thermal abnormal area in the preprocessed infrared thermal map, and extract the temperature of the thermal abnormal area to generate a temperature feature vector W3; S4: Use a multi-modal fusion attention mechanism to fuse the tongue image feature vector W1, facial feature vector W2 and temperature feature vector W3 to generate a multi-modal fusion feature vector F; S5: Introduce a heart failure syndrome set, and quantify and encode the heart failure syndrome set to generate a heart failure fixed vector W. Re-fuse the heart failure fixed vector W and the multi-modal fusion feature vector F to construct a heart failure diagnosis vector A, and calculate the heart failure risk probability and heart failure syndrome category; S6: Establish a multi-modal fusion deep learning model, input the collected tongue image feature vector W1, facial feature vector W2 and temperature feature vector W3 into the multi-modal fusion deep learning model for training, and iteratively optimize the weight vectors of the tongue image feature vector W1, facial feature vector W2 and temperature feature vector W3; S7: Input the collected tongue image, facial image and infrared thermal map into the multi-modal fusion deep learning model, predict the heart failure risk probability, classify the heart failure syndrome category, and perform visual display at the same time.

2. The intelligent diagnosis and treatment method for heart failure of traditional Chinese medicine based on machine vision according to claim 1, characterized in that: The S3 includes the following steps: Precisely separate the tongue body in the preprocessed tongue image from the background through an image segmentation algorithm, detect the crack distribution features, tooth mark distribution features, tongue texture features and tongue coating features on the tongue body based on a convolutional neural network and encode them. The crack distribution features are quantified into the number of cracks and the crack length, the tooth mark distribution features are quantified into the number of tooth marks and the tooth mark depth, the tongue texture features are quantified into the color threshold of the tongue texture, and the tongue coating features are quantified into the thickness of the tongue coating and the color threshold of the tongue coating. Normalize the quantified crack distribution features, tooth mark distribution features, tongue texture features and tongue coating features to generate a tongue image feature vector W1.

3. The intelligent diagnosis and treatment method for heart failure of traditional Chinese medicine based on machine vision according to claim 2, characterized in that: The S3 also includes the following steps: The face recognition system locates and encodes the key point positions of the eyelids, lips or cheekbones in the preprocessed facial image through the MTCNN or Facenet algorithm, extracts the abnormal indicators containing dullness, flushing or cyanosis in the key point positions through skin color analysis, and extracts the color threshold of the abnormal indicators for normalization to generate a facial feature vector W2.

4. The intelligent diagnosis and treatment method for heart failure of traditional Chinese medicine based on machine vision according to claim 3, characterized in that: S3 further includes the following steps: extracting and encoding the thermal anomaly features in the forehead, tip of the nose, extremities or chest and abdomen regions of the preprocessed infrared thermal image through a heat map analysis network, and extracting the temperature gradient or abnormal heat area range of the corresponding part for normalization processing to generate a temperature feature vector W3.

5. The intelligent diagnosis and treatment method for heart failure of traditional Chinese medicine based on machine vision according to claim 1, characterized in that: The multi-modal fusion feature vector F = α1·W1 + α2·W2 + α3·W3, where α1, α2, and α3 are the weights of the tongue image feature vector, facial feature vector, and temperature feature vector respectively, and α1, α2, and α3 are adaptively obtained through a shared Softmax attention network mechanism.

6. The intelligent diagnosis and treatment method for heart failure of traditional Chinese medicine based on machine vision according to claim 1, characterized in that: S5 includes the following steps: the heart failure diagnosis vector A = Concat(F, W), the heart failure syndrome category P1 = Softmax(Wf@A + b), and the heart failure tendency probability P2 = Sigmoid(Wf@A + b), where Wf is the weight parameter for heart failure discrimination and b is the heart failure judgment error factor.

7. The intelligent diagnosis and treatment method for heart failure of traditional Chinese medicine based on machine vision according to claim 1, characterized in that: The iterative optimization in S6 includes the following steps: S61: Initialize the tongue image feature vector W1, facial feature vector W2, temperature feature vector W3, and the corresponding weights α1, α2, α3. S62: Construct a comprehensive loss function L, where the loss function L includes a classification loss Lc, a modality consistency loss Lm, and a regularization term Lr, i.e., L = Lc + λ1Lm + λ2Lr, and λ1, λ2 are preset regularization coefficients. S63: Optimize the network parameters of the multi-modal fusion deep learning model using the stochastic gradient descent method. S64: Calculate the gradient of the loss function with respect to the network parameters using the backpropagation algorithm. S65: Update the values of the network parameters according to the gradient information and the optimization algorithm. Repeat S64 and S65 until the maximum number of iterations is reached or the loss function converges.

8. A traditional Chinese medicine heart failure intelligent diagnosis and treatment system based on machine vision, which is applied to a traditional Chinese medicine heart failure intelligent diagnosis and treatment method according to any one of claims 1 to 7, and is characterized in that: It includes a front-end data acquisition module, a back-end AI analysis module, and a clinical terminal module respectively connected to the central control module. Central control module: used to control each module to work properly through the host. The front-end data acquisition module includes an intelligent tongue diagnosis instrument, an infrared thermal imager, and a facial acquisition camera, and is used to obtain multi-modal images in real time. The back-end AI analysis module includes a multi-modal feature fusion model and other algorithms. The back-end AI analysis module runs on a local server or in the cloud, completes image parsing and diagnostic reasoning, and gives a diagnostic result. The clinical terminal module includes a doctor workstation or a mobile tablet, realizes the visual display of medical record management, diagnostic results, and treatment suggestions, and docks with the hospital system.

9. The intelligent diagnosis and treatment system for heart failure of traditional Chinese medicine based on machine vision according to claim 8, characterized in that: The intelligent tongue diagnosis instrument is equipped with a high-definition camera and a ring-shaped LED light source. The resolution of the high-definition camera is ≥ 2 million pixels. The color temperature of the ring-shaped LED light source is 5500K ± 200K, and the color rendering index is ≥ 90. The standardized light source maintains an illuminance of about 2000 - 3000 lx in the tongue surface area. The configuration of the facial acquisition camera is the same as that of the high-definition camera.

10. The intelligent diagnosis and treatment system for heart failure of traditional Chinese medicine based on machine vision according to claim 9, characterized in that: The infrared detector in the infrared thermal imager uses an uncooled microbolometer, with a wavelength band of 8 - 14 μm, a resolution of 320×240, 640×480 or 1024×768, and a temperature measurement accuracy of ±0.5°C.

Citation Information

Cited By

  • Diagnostic system for arthritis cold and heat symptoms

    CN121054201A

  • Diagnostic system for arthritic cold and heat syndrome

    CN121054201B