A method, device, and program product for predicting fatty liver disease combined with coronary heart disease based on tongue diagnosis

CN121054256BActive Publication Date: 2026-09-11THE FIFTH MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511219005.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-09-11
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

然而,现有研究仍面临三大瓶颈:一是遗传标志物的种族异质性导致风险预测模型普适性不足;二是缺乏针对合并疾病等复杂代谢异常人群的特异性生物标志物;三是肝脏与心血管系统的双向调控机制尚未完全阐明,限制了个体化干预策略的开发

Benefits of technology

[0053] 1. Research on the characteristics of tongue diagnosis in traditional Chinese medicine for fatty liver combined with coronary heart disease revealed that multi-dimensional characteristics such as facial appearance, tongue quality, tongue coating, and sublingual features are important diagnostic indicators for fatty liver combined with coronary heart disease. These multi-dimensional characteristics can be used to effectively predict fatty liver combined with coronary heart disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121054256B_ABST
    Figure CN121054256B_ABST
Patent Text Reader

Abstract

The application relates to the field of intelligent medical treatment, in particular to a prediction method, equipment and program product for fatty liver combined with coronary heart disease based on tongue diagnosis. The method comprises the following steps: S1, acquiring a tongue diagnosis image of a fatty liver patient; S2, inputting the tongue diagnosis image into a tongue diagnosis model to obtain a prediction result of whether the fatty liver patient is combined with coronary heart disease or not; wherein the tongue diagnosis image sequentially passes through an input layer, a convolution module, N tongue diagnosis modules, a pooling layer, a full connection layer and an output layer of the tongue diagnosis model to obtain the prediction result, N is a natural number greater than 1, the tongue diagnosis module comprises a feature extraction module, a local attention layer and a global attention layer, and an output feature vector of the convolution module is sequentially input into the feature extraction module, the local attention layer and the global attention layer of the tongue diagnosis module to obtain an output feature vector of the tongue diagnosis module. The application can effectively identify fatty liver combined with coronary heart disease and has good clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent healthcare, specifically to a method, device, program product, and computer-readable storage medium for predicting fatty liver combined with coronary heart disease based on tongue diagnosis. Background Technology

[0002] Fatty liver disease and coronary heart disease (CHD), as two typical manifestations of metabolic syndrome, have become hot topics in the interdisciplinary field of cardiovascular and hepatology in recent years. Epidemiological data show that patients with non-alcoholic fatty liver disease (NAFLD) have a 2-3 times higher risk of developing CHD than the general population, and this association is particularly significant in individuals with insulin resistance. Current research has moved beyond simple clinical observation to the molecular mechanism level, discovering that PNPLA3 gene polymorphism, chronic low-grade inflammation (manifested as elevated markers such as NLR and PLR), and lipid metabolism disorders constitute a common pathophysiological basis. A multinational cohort study published in *Nature Metabolism* in 2024 further confirmed that for every 5% increase in liver fat content, the coronary artery calcification score increases by 12.7%, providing quantitative evidence for the construction of imaging-based joint prediction models. However, existing research still faces three major bottlenecks: first, the racial heterogeneity of genetic markers leads to insufficient universality of risk prediction models; second, there is a lack of specific biomarkers for complex metabolic abnormalities such as comorbid diseases; and third, the bidirectional regulatory mechanisms of the liver and cardiovascular system have not been fully elucidated, limiting the development of individualized intervention strategies. Summary of the Invention

[0003] To address the aforementioned problems, this invention provides a method for predicting fatty liver disease complicated with coronary heart disease based on tongue diagnosis. Through prediction, patients with comorbidities can receive timely multidisciplinary treatment, integrating cardiology into fatty liver treatment to create a treatment plan tailored to the patient's specific condition, thus achieving comorbidity management. This method has significant clinical value. The prediction method specifically includes:

[0004] S1. Obtain tongue diagnosis images of patients with fatty liver;

[0005] S2. Input the tongue diagnosis image into the tongue diagnosis model to predict whether the patient with fatty liver has coronary heart disease or not;

[0006] The tongue diagnosis image is sequentially passed through the input layer, convolutional module, N tongue diagnosis modules, pooling layer, fully connected layer, and output layer of the tongue diagnosis model to obtain the prediction result. N is a natural number greater than 1. The tongue diagnosis module includes a feature extraction module, a local attention layer, and a global attention layer. The output feature vector of the convolutional module is input into the tongue diagnosis module and sequentially passes through the feature extraction module, the local attention layer, and the global attention layer to obtain the output feature vector of the tongue diagnosis module.

[0007] The tongue diagnosis module includes a feature extraction module, a local attention layer, and a global attention layer, each of which contains residual connections. The output feature vector of the convolution module is input to the feature extraction module to obtain tongue diagnosis features. The tongue diagnosis features and the output features of the convolution module are fused through residual connections to obtain a first fused feature. The first fused feature is input to the local attention layer to obtain local tongue diagnosis features. The local tongue diagnosis features and the first fused feature are fused through residual connections to obtain a second fused feature. The second fused feature is input to the global attention layer to obtain global tongue diagnosis features. The global tongue diagnosis features and the second fused feature are fused through residual connections to output the output feature vector of the tongue diagnosis module.

[0008] Optionally, the feature extraction module includes a convolutional layer, a depthwise separable convolutional layer, and a lightweight channel attention layer. The output vector of the convolutional module is passed through the convolutional layer, the depthwise separable convolutional layer, and the lightweight channel attention layer in sequence to obtain tongue diagnosis features.

[0009] Optionally, the local attention layer is Block Attention, which divides the tongue diagnosis features into non-overlapping local windows, and each local window calculates the local tongue diagnosis features through a self-attention mechanism;

[0010] Optionally, the global attention layer is Grid Attention, which captures the global feature relationship between local tongue diagnosis features to obtain global tongue diagnosis features;

[0011] Optionally, the convolution module includes L convolutional layers, where L is a natural number greater than 1. The tongue diagnosis image is input through the input layer and then convolved through the L convolutional layers to obtain the output vector of the convolution module.

[0012] Optionally, the tongue diagnosis image is a tongue surface image;

[0013] Optionally, the method further includes facial images, acquiring facial images of patients with fatty liver, and inputting the facial images and tongue diagnosis images into a tongue diagnosis model for prediction to obtain prediction results of whether patients with fatty liver have coronary heart disease or not.

[0014] Optionally, the method further includes facial images, obtaining facial images of patients with fatty liver, and S2 is replaced by: inputting the facial images and tongue diagnosis images into the prediction model of fatty liver combined with coronary heart disease obtained by the following method for constructing a prediction model of fatty liver combined with coronary heart disease to predict whether the patient with fatty liver has coronary heart disease or not.

[0015] Optionally, the facial image includes a facial region image and an ear region image, and the tongue diagnosis image includes a tongue surface region image and a sublingual vein region image; the tongue surface region image, the sublingual vein region image, the facial region image, and the ear region image are input into the prediction model for fatty liver combined with coronary heart disease obtained by the following prediction model construction method for fatty liver combined with coronary heart disease to obtain the prediction result of whether the fatty liver patient has coronary heart disease or not.

[0016] S1 is replaced by: acquiring the second tongue diagnosis image, facial image, and clinical data of patients with fatty liver;

[0017] The S2 is replaced by: the second tongue diagnosis image and facial image are sequentially passed through an image embedding layer and an image coding layer to obtain image features; the clinical data are sequentially passed through a text embedding layer and a text coding layer to obtain text features; the image features and text features are passed through an integration layer to obtain fusion features; and the fusion features are passed through a classification prediction layer to predict whether a patient with fatty liver has coronary heart disease or not.

[0018] The second tongue diagnosis image includes an image of the tongue surface area and an image of the sublingual area;

[0019] Optionally, the image encoding layer and the text encoding layer are encoded using a pre-trained model to obtain image features or text features;

[0020] Optionally, the pre-trained model of the image coding layer adopts one or more of the following: ViT-B-32, UniFormer, Swing Transformer, PLIP, CLIP;

[0021] Optionally, the pre-trained model of the text encoding layer adopts one or more of the following: bert-base-chinese, ELECTRA-zh, RoBERTa-zh, ALBERT-zh, DistilBERT-zh, ERNIE-3.0;

[0022] Optionally, the image features and text features are input into the fusion layer, concatenated, and then subjected to nonlinear transformation and feature learning through a neural network. Finally, a classification prediction layer is used to make a prediction result.

[0023] Optionally, the clinical data includes one or more of the following: medical history, gender, age, and blood indicators;

[0024] Optionally, the blood indicators include one or more of the following: red blood cell distribution width, monocytes, blood glucose, triglycerides, albumin / globulin ratio, total bilirubin, low-density lipoprotein, mean platelet volume, apolipoprotein A1, total cholesterol, prealbumin, and alkaline phosphatase.

[0025] S1 is replaced by: obtaining facial features, tongue texture, tongue coating, and sublingual data of patients with fatty liver;

[0026] S2 is replaced by: predicting whether a patient with fatty liver has coronary heart disease based on the facial features, tongue texture, tongue coating, and sublingual data.

[0027] The facial features include one or more of the following: nasal folds, left eye color, right eye color, left ear folds, right ear folds, left ear color, right ear color, lip color, main color, and gloss.

[0028] Optionally, the tongue body includes one or more of the following: age, thickness, color, tongue tip color-Y, tongue tip color-Cr, tongue tip color-Cb, tongue root color-Y, tongue root color-Cr, tongue root color-Cb, tongue edge color-Cr, tongue middle color-R, tongue middle color-Cr, and tongue middle color-Cb.

[0029] Optionally, the tongue coating includes one or more of the following: tongue color, tongue coating color-ASM, tongue coating second-order color moment BGR, and tongue coating correlation coefficient COR;

[0030] Optionally, the sublingual vein includes one or more of the following: sublingual vein shape, sublingual vein color, sublingual-Cr, sublingual-Cb, and sublingual-Y;

[0031] Optionally, the method further includes body fluids and constitution, obtaining body fluids and constitution data, and making predictions based on the facial appearance, tongue texture, tongue coating, sublingual data, body fluids, and constitution to obtain prediction results of whether fatty liver patients have coronary heart disease or not.

[0032] Optionally, the method further includes obtaining clinical information, and predicting whether a patient with fatty liver has coronary heart disease based on the facial features, tongue texture, tongue coating, sublingual data, and clinical information. The clinical information includes one or more of the following: age, diabetes, hypertension, hepatitis B, and smoking history.

[0033] The prediction is made using one or more of the following algorithms: SVM, LightGBM, MLP, Random Forest, XGBoost, GBM, and AdaBoost.

[0034] The purpose of this invention is to provide a method for constructing a predictive model for fatty liver combined with coronary heart disease, comprising:

[0035] Acquire patient tongue and facial image datasets;

[0036] The tongue image data and facial image dataset are respectively input into a parallel visual encoding module for feature extraction to obtain tongue image features and facial image features; the visual encoding module is a neural network formed by combining different neural network layers;

[0037] The tongue and facial features are input into a feature adaptive modulator module for feature fusion to obtain modulated fused features; the feature adaptive modulator module is a neural network formed by combining an expert gated network and a dimensionality reduction network;

[0038] The modulation and fusion features are input into the classification layer for classification to obtain the predicted classification result;

[0039] The predicted classification results are compared with the true classification labels to calculate the loss. The above steps are repeated until the loss function remains unchanged, thus obtaining the prediction model for fatty liver combined with coronary heart disease.

[0040] Optionally, the visual encoding module sequentially includes an input layer, a backbone module, a main body block, a head module, and an output layer; the backbone module includes L cascaded convolutional layers for high-dimensional feature extraction; the main body block includes S cascaded sub-modules for global and local feature extraction, wherein the sub-modules are network modules composed of convolutional layers and residual connections of different layers; the head module includes a global average pooling layer for feature dimensionality reduction.

[0041] Optionally, the sub-module sequentially includes a first convolutional module, a restricted receptive field attention module, and an unrestricted receptive field attention module, with the modules connected in series and residual connections also included between the modules; features are input to the next sub-module or the head module after passing through the first convolutional module, the restricted receptive field attention module, and the unrestricted receptive field attention module in the sub-module.

[0042] Optionally, the restricted receptive field attention module extracts local features through a preset fixed receptive field, and the unrestricted receptive field attention module extracts global features through a dynamic receptive field, wherein the global features include features of different regions and / or features of the same region at different time points.

[0043] Optionally, the preset fixed receptive field is to limit the computational receptive field of each position by spatial constraints, and to form a local region with the current position as the center and a radius of r, and to obtain local features by attention calculation on the features of the local region;

[0044] Optionally, the dynamic receptive field is a region centered on the current position with a radius of R. The region decays linearly with distance. When calculating attention features, the attention weights are weighted by a dynamic distance mask to obtain global features.

[0045] Optionally, the first convolutional module includes cascaded depthwise separable convolution and lightweight channel attention, and high-dimensional features are obtained by integrating depthwise separable convolution and lightweight channel attention to extract image detail features.

[0046] Optionally, the feature adaptive modulator module includes a hybrid expert module, a fusion layer, and a dimensionality reduction layer. The hybrid expert module performs gating weight calculation on tongue image features and facial image features to obtain tongue image features and facial image features with different weights. The tongue image features and facial image features with different weights are fused through the fusion layer to obtain gated fused features. The gated fused features are then dimensionality reduced by the dimensionality reduction layer to output modulated fused features.

[0047] Optionally, the tongue image data includes images of the tongue surface region and the sublingual vein region, and the facial image data includes images of the face region and the ear region. The images of the tongue surface region, the sublingual vein region, the face region, and the ear region are respectively input to a parallel visual encoding module for feature extraction to obtain tongue surface features, sublingual vein features, facial features, and ear features. The tongue surface features, sublingual vein features, facial features, and ear features are then input to a feature adaptive modulator module for feature fusion to obtain modulated fusion features.

[0048] Optionally, the tongue surface features, sublingual vein features, facial features, and ear features are input to the feature adaptive modulator module. First, the weights of the corresponding features are calculated by the hybrid expert module through gating weight calculation. Then, the tongue surface features, sublingual vein features, facial features, ear features, and their corresponding weights are weighted and fused to obtain the gated fused features. The gated fused features are then reduced in dimensionality by a dimensionality reduction layer and output as modulated fused features.

[0049] The purpose of this invention is to provide a computer program product that includes a computer program or instructions, which are executed by a processor to implement the above-described method for predicting fatty liver combined with coronary heart disease based on tongue diagnosis.

[0050] The purpose of this invention is to provide a computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, wherein the computer program or instructions are executed by the processor to implement the above-described method for predicting fatty liver combined with coronary heart disease based on tongue diagnosis.

[0051] The purpose of this invention is to provide a computer-readable storage medium having a computer program or instructions stored thereon, which is executed by a processor to implement the above-described method for predicting fatty liver combined with coronary heart disease based on tongue diagnosis.

[0052] Advantages of this invention:

[0053] 1. Research on the characteristics of tongue diagnosis in traditional Chinese medicine for fatty liver combined with coronary heart disease revealed that multi-dimensional characteristics such as facial appearance, tongue quality, tongue coating, and sublingual features are important diagnostic indicators for fatty liver combined with coronary heart disease. These multi-dimensional characteristics can be used to effectively predict fatty liver combined with coronary heart disease.

[0054] 2. For fatty liver combined with coronary heart disease, this invention proposes a tongue diagnosis model and a multimodal model for disease prediction. The tongue diagnosis model learns local features of the tongue surface, the relationships between local features, and global features to achieve high-accuracy prediction. Furthermore, the multimodal model, by fusing relevant parameters from traditional Chinese medicine tongue diagnosis and Western medicine clinical information, predicts whether fatty liver patients will have coronary heart disease. Results demonstrate that the proposed solution has high accuracy, providing an important approach for predicting fatty liver combined with coronary heart disease and possessing significant clinical application value.

[0055] 3. This invention is a prediction method for patients with multiple diseases. Compared to existing comorbidity prediction methods, this invention predicts whether other diseases are present based on a known disease, rather than simply ranking the number of diseases from case data. Compared to conventional disease prediction, the patient population in this invention is individuals with fatty liver disease, not the general population. Individuals with fatty liver disease exhibit significant differences in physiological symptoms, facial appearance, and tongue appearance compared to the general population. Determining whether another disease is present under these different baseline conditions significantly increases the difficulty. This invention predicts whether coronary heart disease is present in individuals with fatty liver disease, which is more in line with actual clinical situations and provides a predictive technical solution.

[0056] 4. For fatty liver combined with coronary heart disease, this invention also proposes a prediction model constructed using both restricted and unrestricted receptive field attention mechanisms. Restricted receptive field attention guides the model to concentrate computational resources on local regions within an image or sequence. This design enables the model to more efficiently and accurately capture and analyze the interdependencies between object features within a small area. The unrestricted receptive field attention mechanism breaks the limitations of local scope, extending its receptive field to the entire global sequence. This allows the model to comprehensively capture and model long-distance dependencies in the data, thereby better understanding cross-regional or cross-time point information associations, resulting in an accurate and stable prediction model. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This is a schematic diagram of the prediction method for fatty liver combined with coronary heart disease based on tongue diagnosis provided in an embodiment of the present invention.

[0059] Figure 2 A schematic diagram of a tongue diagnosis-based predictive system for fatty liver complicated with coronary heart disease provided in an embodiment of the present invention;

[0060] Figure 3 A schematic diagram of a tongue diagnosis-based predictive device for fatty liver combined with coronary heart disease provided in an embodiment of the present invention;

[0061] Figure 4 The TongueViT network structure diagram provided in the embodiments of the present invention;

[0062] Figure 5 Training and accuracy of the TongueViT recognition model for coronary heart disease and tongue appearance features provided in this embodiment of the invention;

[0063] Figure 6 The confusion matrix of the three coronary heart disease classification models provided in this embodiment of the invention for the tongue image validation set;

[0064] Figure 7 The thermal response of the TongueViT coronary heart disease recognition model based on tongue image features provided in this embodiment of the invention;

[0065] Figure 8 Training and accuracy of the TongueViT recognition model for coronary heart disease and facial features provided in this embodiment of the invention;

[0066] Figure 9 The thermal response of the TongueViT coronary heart disease recognition model based on facial features provided in this embodiment of the invention;

[0067] Figure 10 This is a network structure diagram of the AI-CAD prediction model provided in an embodiment of the present invention;

[0068] Figure 11a This is a hierarchical clustering heatmap of discrete characteristics of a coronary heart disease positive population provided in an embodiment of the present invention; Figure 11b This is a hierarchical clustering heatmap of continuous characteristics of a coronary heart disease positive population provided in an embodiment of the present invention;

[0069] Figure 12 The training iteration results of the AI-CAD prediction model provided in the embodiments of the present invention;

[0070] Figure 13 The confusion matrices of the three coronary heart disease classification models provided in this embodiment of the invention for the validation set;

[0071] Figure 14 The ROC and PR curves of the three coronary heart disease classification models provided in this embodiment of the invention for the validation set;

[0072] Figure 15 The thermal response of the tongue surface to the AI-CAD coronary artery disease recognition model provided in this embodiment of the invention;

[0073] Figure 16The thermal response of the sublingual region to the AI-CAD coronary artery disease recognition model provided in this embodiment of the invention;

[0074] Figure 17 The thermal response of the face to the AI-CAD coronary heart disease recognition model provided in this embodiment of the invention;

[0075] Figure 18 The thermal response of the left and right ears to the AI-CAD coronary heart disease recognition model provided in this embodiment of the invention;

[0076] Figure 19 The TongueMMM network structure diagram provided in the embodiments of the present invention;

[0077] Figure 20 Training and accuracy of the coronary artery disease and multimodal Tongue MMM recognition model provided in this embodiment of the invention;

[0078] Figure 21 The contribution of the top ten discrete features provided in the embodiments of the present invention;

[0079] Figure 22 The contribution of the top ten consecutive features provided in the embodiments of the present invention;

[0080] Figure 23 The contribution of the top ten blood indicators provided in the embodiments of the present invention. Detailed Implementation

[0081] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0082] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as S101, S102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0083] Figure 1 A schematic diagram of the method for predicting fatty liver complicated with coronary heart disease based on tongue diagnosis provided in this embodiment of the invention, specifically including:

[0084] S1: Obtain tongue diagnosis images of patients with fatty liver;

[0085] In one embodiment, the tongue diagnosis image is a tongue surface image.

[0086] In one embodiment, due to the uneven distribution of coronary heart disease (CHD) patient data, the image data of CHD patients was enhanced in various ways during model training (noise enhancement, brightness enhancement, Gaussian blur enhancement), resulting in a total of 2489 images (864 CHD cases and 1625 non-CHD cases). The enhanced dataset was then divided approximately according to a training set / validation set ratio of 100 / 5, resulting in a training set of 2363 cases (804 CHD cases and 1559 non-CHD cases) and a validation set of 2363 cases (60 CHD cases and 66 non-CHD cases). Classic classification networks ResNet50 and GoogleNet were selected as controls.

[0087] S2: Input the tongue diagnosis image into the tongue diagnosis model to predict whether the patient with fatty liver has coronary heart disease or not.

[0088] The tongue diagnosis image is sequentially passed through the input layer, convolutional module, N tongue diagnosis modules, pooling layer, fully connected layer, and output layer of the tongue diagnosis model to obtain the prediction result. N is a natural number greater than 1. The tongue diagnosis module includes a feature extraction module, a local attention layer, and a global attention layer. The output feature vector of the convolutional module is input into the tongue diagnosis module and sequentially passes through the feature extraction module, the local attention layer, and the global attention layer to obtain the output feature vector of the tongue diagnosis module.

[0089] In one embodiment, the tongue diagnosis module further includes residual connections in the feature extraction module, local attention layer, and global attention layer. The output feature vector of the convolution module is input to the feature extraction module to obtain tongue diagnosis features. The tongue diagnosis features and the output features of the convolution module are fused through residual connections to obtain a first fused feature. The first fused feature is input to the local attention layer to obtain local tongue diagnosis features. The local tongue diagnosis features and the first fused feature are fused through residual connections to obtain a second fused feature. The second fused feature is input to the global attention layer to obtain global tongue diagnosis features. The global tongue diagnosis features and the second fused feature are fused through residual connections to output the output feature vector of the tongue diagnosis module.

[0090] Optionally, the feature extraction module includes a convolutional layer, a depthwise separable convolutional layer, and a lightweight channel attention layer. The output vector of the convolutional module is passed through the convolutional layer, the depthwise separable convolutional layer, and the lightweight channel attention layer in sequence to obtain tongue diagnosis features.

[0091] In one embodiment, the local attention layer is Block Attention, which divides the tongue diagnosis features into non-overlapping local windows, and each local window calculates the local tongue diagnosis features through a self-attention mechanism.

[0092] In one embodiment, the global attention layer is Grid Attention, which captures the global feature relationship between local tongue diagnosis features to obtain global tongue diagnosis features.

[0093] In one embodiment, the convolution module includes L convolutional layers, where L is a natural number greater than 1. The tongue diagnosis image is input through the input layer and then convolved through the L convolutional layers to obtain the output vector of the convolution module.

[0094] In one embodiment, the method further includes facial images, acquiring facial images of patients with fatty liver, and inputting the facial images and tongue diagnosis images into a tongue diagnosis model to predict whether the patients with fatty liver have coronary heart disease.

[0095] In one specific embodiment, the TongueVit network model architecture diagram is as follows: Figure 4 As shown, this model combines the advantages of Convolutional Neural Networks (CNNs) and Transformers to construct an efficient image classification model. The model consists of multiple stacked TongueViT Blocks, each integrating three key modules: MBConv, BlockAttention, and GridAttention. The MBConv module utilizes depthwise separable convolutions (DSVs) and a lightweight channel attention mechanism (Squeeze-and-Excitation, SE) to effectively improve local feature extraction capabilities with low computational cost. The BlockAttention module divides the feature map into multiple non-overlapping local windows and applies a self-attention mechanism within each window, focusing on capturing detailed feature dependencies within local regions. The GridAttention module employs a global self-attention mechanism to capture global feature relationships. The network adopts a hierarchical structure, progressively extracting multi-level features from fine-grained to coarse-grained through multiple stages. Downsampling operations are used between each stage to reduce the spatial resolution of the feature map while expanding the channel dimension to enhance feature expressiveness. Finally, the model outputs classification results through global average pooling layers and fully connected layers.

[0096] In one specific embodiment, a validation set of 126 tongue surface cases (60 with coronary artery disease and 66 without coronary artery disease) was used to verify the recognition results after inverse calculation using three different trained models:

[0097]

[0098] Figure 5The loss curves shown are the loss curves of three training coronary heart disease classification models (TongueViT, ResNet50, and GoogLeNet). The figure shows that after 3000 iterations, the loss becomes relatively flat and continues to flatten in subsequent training, indicating that the model's fitting ability has reached its optimal state.

[0099] The ROC curve plot shows the false positive rate (the proportion of incorrect predictions made by the model in all negative samples, also known as the false alarm rate) on the horizontal axis and the true positive rate (the proportion of correct predictions made by the model in all positive samples, also known as the sensitivity) on the vertical axis. The positive samples are tongue images of patients with coronary heart disease. In the figure, the ROC curve of TongueViT is closer to the upper left corner of the axis, indicating that the model performance is relatively better.

[0100] In the PR curve graph, the horizontal axis represents recall (the proportion of samples that were actually positive but were correctly predicted as positive), and the vertical axis represents precision (the number of samples that were predicted as positive and were also actually positive divided by the total number of samples predicted as positive). Positive examples represent the tongue images of patients with coronary artery disease. In the graph, the PR curve for TongueViT is closer to the upper right corner of the axis, indicating that this model achieves a better trade-off between precision and recall, resulting in better performance. The confusion matrices of the three coronary artery disease classification models for the tongue image validation set are shown below. Figure 6 As shown.

[0101] In one specific embodiment, Figure 7 This study presents a comparative analysis of the tongue thermal response areas in two groups: those without coronary heart disease (CHD) and those with CHD, using a tongue thermal model for CHD. The results show that the non-CHD group exhibited a generally lower level of tongue thermal response, while the CHD group displayed a significant characteristic thermal distribution pattern, specifically enhanced thermal response in the thick coating area at the base of the tongue, the fissure area, the tip of the tongue, and the lateral sides of the tongue. This differential thermal distribution provides an important reference for the tongue-based auxiliary diagnosis of CHD. The TongueViT model, validated on 126 cases of tongue image test data in a clinical validation set, achieved an area under the curve (AUC) of 0.877, an accuracy of 86.50%, a precision of 95.7%, a specificity of 97.00%, and a recall of 75.00%. Among the three models, TongueViT performed best. Thermal response mapping revealed that this tongue image AI model primarily focused on key areas such as the thick coating, fissures, the tip of the tongue, and the lateral sides of the tongue during feature recognition. This closely matches the correlation analysis results of tongue appearance characteristics in coronary heart disease, specifically: 1) tongue coating characteristics; 2) tongue tip and tongue color; 3) tongue texture (age / tenderness); and 4) identification of features such as teeth marks. The regional attention of this model is consistent with the feature distribution patterns in traditional Chinese medicine tongue diagnosis theory.

[0102] In one specific embodiment, a facial model was constructed from face image data. The area under the curve (AUC) obtained by the TongueViT model in the process of validating the recognition results of 126 face image test data in the clinical validation set was 0.907, the accuracy was 88.90%, the precision was 94.20%, the specificity was 95.50%, and the recall was 81.70%, indicating that the model performed well. Figure 8 The loss curves shown are those of three training models for coronary artery disease classification (TongueViT, ResNet50, and GoogLeNet). The graphs show that after more than 4000 iterations, the loss curves become relatively flat and continue to flatten in subsequent training, indicating that the model's fitting ability has reached its optimal state. In the ROC curve graph, TongueViT's ROC curve is closer to the upper left corner of the coordinate system, indicating relatively better model performance. In the PR curve graph, TongueViT's PR curve is closer to the upper right corner of the coordinate system, indicating that this model has a better trade-off between precision and recall, resulting in better performance.

[0103] Figure 9 The figure illustrates a comparative analysis of the frontal thermal response areas in a facial model of patients with and without coronary artery disease (CAD). The results show that the non-CAD group exhibited a generally lower level of frontal thermal response, while the CAD group displayed a significant characteristic thermal distribution pattern, specifically enhanced thermal response in the nasolabial folds, around the lips, around the eyes, and on the cheeks. This difference in thermal distribution provides important reference for the auxiliary diagnosis of CAD through facial examination.

[0104] In one embodiment, the method further includes facial images, acquiring facial images of patients with fatty liver, and S2 is replaced by: inputting the facial images and tongue diagnosis images into the prediction model for fatty liver combined with coronary heart disease obtained by the prediction model construction method for fatty liver combined with coronary heart disease to predict whether the patient with fatty liver has coronary heart disease or not.

[0105] Optionally, the facial image data includes facial region images and ear region images, and the tongue diagnosis data includes tongue surface region images and sublingual vein region images; the tongue surface region images and sublingual vein region images, facial region images and ear region images are input into the prediction model for fatty liver combined with coronary heart disease obtained by the prediction model construction method for fatty liver combined with coronary heart disease to obtain the prediction result of whether the fatty liver patient has coronary heart disease or not.

[0106] In one embodiment, the method for constructing a predictive model for fatty liver combined with coronary heart disease is as follows:

[0107] Acquire patient tongue and facial image datasets;

[0108] The tongue image data and facial image dataset are respectively input into a parallel visual encoding module for feature extraction to obtain tongue image features and facial image features; the visual encoding module is a neural network formed by combining different neural network layers;

[0109] The tongue and facial features are input into a feature adaptive modulator module for feature fusion to obtain modulated fused features; the feature adaptive modulator module is a neural network formed by combining an expert gated network and a dimensionality reduction network;

[0110] The modulation and fusion features are input into the classification layer for classification to obtain the predicted classification result;

[0111] The predicted classification result is compared with the true classification label to calculate the loss. The above steps are repeated until the loss function remains unchanged, thus obtaining the prediction model for fatty liver combined with coronary heart disease.

[0112] In one embodiment, the visual encoding module sequentially includes an input layer, a backbone module, a main body block, a head module, and an output layer; the backbone module includes L cascaded convolutional layers for high-dimensional feature extraction; the main body block includes S cascaded sub-modules for global and local feature extraction, wherein the sub-modules are network modules composed of convolutional layers and residual connections of different layers; the head module includes a global average pooling layer for feature dimensionality reduction.

[0113] Optionally, the sub-module sequentially includes a first convolutional module, a restricted receptive field attention module, and an unrestricted receptive field attention module, with the modules connected in series and residual connections also included between the modules; features are input to the next sub-module or the head module after passing through the first convolutional module, the restricted receptive field attention module, and the unrestricted receptive field attention module in the sub-module.

[0114] In one embodiment, the restricted receptive field attention module extracts local features by using a preset fixed receptive field, while the unrestricted receptive field attention module extracts global features by using a dynamic receptive field to extract long-distance dependency features. The global features include features from different regions and / or features from the same region at different time points.

[0115] In one embodiment, the preset fixed receptive field is to limit the computational receptive field of each position by spatial constraints, forming a local region with the current position as the center and a radius of r, and performing attention calculation on the features of the local region to obtain local features.

[0116] In one embodiment, the dynamic receptive field is a region centered on the current position with a radius of R. The region decays linearly based on distance. When calculating attention features, the attention weights are weighted by a dynamic distance mask to obtain global features.

[0117] In one embodiment, the first convolutional module includes cascaded depthwise separable convolution and lightweight channel attention, and high-dimensional features are obtained by integrating depthwise separable convolution and lightweight channel attention to extract image detail features.

[0118] In one embodiment, the feature adaptive modulator module includes a hybrid expert module, a fusion layer, and a dimensionality reduction layer. The hybrid expert module performs gating weight calculation on tongue and facial features to obtain tongue and facial features with different weights. The tongue and facial features with different weights are fused through the fusion layer to obtain gated fused features. The gated fused features are then dimensionality reduced by the dimensionality reduction layer to output modulated fused features.

[0119] In one embodiment, the tongue image data includes a tongue surface region and a sublingual vein region, and the facial image data includes a face region and an ear region. The tongue surface region, sublingual vein region, face region, and ear region are respectively input to a parallel visual encoding module for feature extraction to obtain tongue surface features, sublingual vein features, facial features, and ear features. The tongue surface features, sublingual vein features, facial features, and ear features are then input to a feature adaptive modulator module for feature fusion to obtain modulated fusion features.

[0120] Optionally, the tongue surface features, sublingual vein features, facial features, and ear features are input to the feature adaptive modulator module. First, the weights of the corresponding features are calculated by the hybrid expert module through gating weight calculation. Then, the tongue surface features, sublingual vein features, facial features, ear features, and their corresponding weights are weighted and fused to obtain the gated fused features. The gated fused features are then reduced in dimensionality by a dimensionality reduction layer and output as modulated fused features.

[0121] In one specific embodiment, the network structure of the predictive model for fatty liver combined with coronary heart disease (AI-CAD) is as follows: Figure 10 As shown, using segmented images of the patient's tongue, sublingual veins, face, and ears as input to the network model, the probability of the patient having coronary heart disease can be predicted after passing through the TongueVit (Vision Transformer, ViT) visual encoding module and the Feature Adaptive Modulator (FAM) module. Specifically, the tongue and sublingual veins are input into the same visual encoding module, as are the facial and ear regions. The outputs of the two visual encoding modules are fed in parallel to the FAM. In the AI-CAD visual encoder, the backbone network uses a TongueVit network structure with a restricted and unrestricted receptive field hybrid attention mechanism to extract local and global color, texture, and morphological features of the image. [The text then repeats the description of using different visual encoding modules.] The network branch processes two tongue images. The network branches process two face images, and the weight parameters of the two branches are independent. This allows each branch to focus on specific image features, while the weights of each branch are shared across the two images within the same group. , Output a set of visually encoded high-dimensional feature vectors These feature vector groups from different parts are synchronously input into the FAM module. The FAM performs differential weighting on each feature group, fusing the multidimensional information of the tongue surface into a high-dimensional feature vector. Then, the data is processed by a multilayer perceptron (MLP) to reduce the dimensionality and output the final probability of coronary heart disease.

[0122] In one specific embodiment, the TongueViT model is designed with three core modules: Stem, Block, and Head. This network architecture follows a hierarchical design principle, progressively extracting hierarchical features from the input data, from fine-grained to coarse-grained, through a multi-stage representation learning process. At each stage of feature extraction, the model performs downsampling, effectively reducing the spatial resolution of the feature maps while expanding the channel dimension to enhance the expressive power of the features. Finally, a global average pooling layer integrates all learned feature information and condenses it into a compact feature vector. The main structure of the model is as follows:

[0123] 1) Stem: As a pre-feature extraction module, it adopts a two-layer convolutional structure with an input dimension of [448, 448, 3] and an output dimension of [224, 224, 64].

[0124] 2) Block: This is the core module of the model, integrating three key sub-modules:

[0125] ① Lightweight Bottleneck Convolution Module (MBConv): This module integrates depthwise separable convolution (DConv) and a lightweight channel attention mechanism (Squeeze-and-Excitation, SE). It effectively improves the ability to extract detailed image features at a lower computational cost.

[0126] ② Limited Field-of-View (LFV) Attention Mechanism: This mechanism guides the model to concentrate computational resources on local regions within an image or sequence by setting a limited receptive field. This design enables the model to more efficiently and accurately capture and analyze the interdependencies between object features within a small area, guiding the model to focus on local texture features (such as cracks, pinpoints, and peeling) in the tongue image that have diagnostic value for coronary heart disease. Through this feature-guided attention mechanism, the model can more accurately capture key tongue image representations related to coronary heart disease, thereby improving disease prediction capabilities.

[0127] ③ Non-Limited Field-of-View (NLFV) Attention: Unlike LFV Attention, this mechanism breaks the limitations of local scope, extending its receptive field to the entire global sequence. This allows the model to comprehensively capture and model long-distance dependencies in the data, thereby better understanding cross-regional or cross-time point information associations. It guides the model to accurately capture color features closely related to coronary heart disease diagnosis in tongue appearance (such as tongue color, coating color, lip color, ear color, and other key color indicators), effectively improving the accuracy and reliability of disease prediction.

[0128] Both LFV Attention and NLFV Attention can effectively suppress feature responses in non-critical regions, thereby optimizing the model's attention distribution. The complementary combination of these two mechanisms enhances the model's ability to model at different spatial scales.

[0129] 3) The Head module is extremely concise, containing only a single global average pooling layer (GAP), outputting a feature vector of length 784. The dimensions of each module are shown in Table 1:

[0130] Table 1 Feature Dimensions in the Main Modules of TongueViT

[0131]

[0132] In one specific embodiment, a hierarchical clustering algorithm is used to cluster the feature column attribute values ​​of a coronary heart disease positive population, displaying the clustering relationships between data points (or features), such as... Figure 11a and Figure 11b As shown, the feature clustering results are as follows:

[0133] a. Discrete features: Tongue color, coating color, ear color, lip color, etc. These features cluster into a branch, indicating that these features may collectively reflect certain physiological or pathological characteristics to some extent. Similar groupings include plumpness / thinness, pinpoints / spikes, peeling, and cracks.

[0134] b. The color values ​​of the tongue edges, root, and tip, as well as the coating color, are each clustered into one branch, while the color moments and texture features of the tongue image are each clustered into independent branches. These feature groupings suggest that each category of features may collectively reflect specific physiological or pathological characteristics to some extent.

[0135] In summary, the texture features (petechiae, peeling, cracks, ear folds) and color features (tongue color, coating color, ear color, lip color) of the tongue surface have a positive effect on the prediction of coronary heart disease models. Fully considering these features in the design of the model architecture can effectively improve the performance of coronary heart disease prediction models.

[0136] In one specific embodiment, the FAM feature adaptive modulator module sets four independent expert gating weights for the four sets of feature vectors that need to be fused. Each gating weight is specifically responsible for controlling the fusion ratio of the corresponding feature group and adaptively fusing them into a unified representation, which is ultimately used for classification or other downstream tasks.

[0137] ① The features of the four sets of images—tongue surface, under the tongue, front view, and left and right ears—are as follows: , , , ;

[0138] ② Weight generation method: ,in {1,2,3,4} and These are learnable weight and bias vectors. The activation function is used. Modulation weights are generated for the corresponding graphs based on the four sets of feature inputs. , , , ;

[0139] ③ Weighted fusion yields the following features: .

[0140] ④ Using MLP for dimensionality reduction, the final output is: .

[0141] In one specific embodiment, this invention introduces two mechanisms: Restricted Receptive Field Attention (LFV Attention) and Unrestricted Receptive Field Attention (NLFV Attention). The computational processes of both are similar, with the core difference lying only in the generation of the receptive field mask. This design effectively improves the model's ability to model spatial locality and long-range dependencies, while balancing computational efficiency and feature robustness.

[0142] 1) The restricted receptive field attention mechanism introduces spatial constraints to limit the computational receptive field at each location. This mechanism initializes a mask matrix and, when calculating attention, only considers points within a local region centered at the current point with a radius of R, ignoring points outside the region. This reduces computational overhead and enhances the ability to extract local features, effectively preserving spatial locality.

[0143] 2) The unrestricted receptive field attention mechanism is not limited by a fixed spatial range. It uses dynamic distance masks to weight attention weights, cleverly balancing the capture of global long-range dependencies with the filtering of irrelevant background interference. This design not only enables the model to focus on long-distance correlations between targets at any location in the image, but also effectively suppresses weak correlations between regions through geometric distance constraints.

[0144] The specific implementation steps of the two attention mechanisms are as follows:

[0145] (1) Input feature map preprocessing:

[0146] Input feature map Reshaped into a two-dimensional matrix form (where These are the height, width, and number of channels of the input feature, respectively, and H=W, which facilitates subsequent calculations.

[0147] The feature map is obtained: , .

[0148] (2) Feature projection:

[0149] enter The feature map generates a matrix of query (Q), key (K), and value (V) through linear projection.

[0150] , , ,in , , This is a learnable weight matrix.

[0151] (3) Mask generation:

[0152] ① Restricted receptive field mask generation:

[0153] For spatial coordinates At each location, a corresponding mask matrix can be generated. The mask value is in spatial coordinates The radius R of the origin is 1 within a local region and 0 outside this range, i.e., the spatial coordinates are... mask value at Satisfy the following formula:

[0154]

[0155] in ,and .

[0156] ② Unrestricted receptive field mask generation:

[0157] For spatial coordinates At each location, a corresponding mask matrix can be generated. The mask value is in spatial coordinates Within a circular region with radius R centered at the origin, the value decreases linearly, approaching 1 as the distance increases and approaching 0 as the distance increases. This is reflected in the spatial coordinates. mask value at Satisfy the following formula:

[0158]

[0159] in ,and .

[0160] (4) Attention calculation:

[0161] The calculation process is as follows:

[0162] ① For spatial coordinates Correlation matrix between features and other locations:

[0163] , .

[0164] ② Apply the softmax function to the correlation matrix so that the sum of the attention weight matrices is 1;

[0165]

[0166] ③ Apply spatial coordinates In a restricted receptive field attention mechanism, the mask at a point will set the correlation outside the R region centered on that point to zero, while the correlation within the R region remains unchanged. In an unrestricted receptive field attention mechanism, the correlation between the other points and that point will decrease linearly from near to far.

[0167]

[0168] ④ Use attention weights to perform a weighted summation on the Value(V) matrix to generate the attention output. :

[0169]

[0170] (5) Feature reconstruction and output:

[0171] The output matrix is ​​reshaped into a spatial format, and the channel dimensions are reduced from [previous format] by a linear transformation. Map back This yields the final output feature map.

[0172] .

[0173] In one specific implementation, the model is trained as follows:

[0174] Dataset: The data of patients with coronary heart disease includes segmentation images of the tongue, sublingual region, frontal view, left ear, and right ear. Among them, there are 864 cases of coronary heart disease and 1625 cases of non-coronary heart disease. Due to the uneven distribution of the data of patients with coronary heart disease, the image data of patients with coronary heart disease were enhanced in various ways (noise enhancement, brightness enhancement, Gaussian blur enhancement) during model training. A total of 2489 images were enhanced. The enhanced dataset was divided according to an approximate ratio of training set / validation set = 100 / 5, resulting in a training set of 2363 cases (804 cases of coronary heart disease and 1559 cases of non-coronary heart disease) and a validation set of 126 cases (60 cases of coronary heart disease and 66 cases of non-coronary heart disease).

[0175] (1) Data augmentation processing methods

[0176] There are three data augmentation methods, and 0-3 of them will be randomly selected and combined.

[0177] ① Noise amplification:

[0178] Random noise was added to the image pixel values ​​to simulate interference factors in real-world scenarios and improve the model's robustness. Salt-and-pepper noise was used in the experiment, with the noise ratio p randomly selected within the range of [0.01, 0.03], meaning that pixel values ​​at ratio p were randomly replaced with the maximum value of the image pixel values.

[0179] ② Increased brightness and contrast:

[0180] By adjusting the brightness, contrast, and saturation of the image, visual changes under different lighting conditions are simulated. The brightness factor is set to be randomly selected within the range of [0.5, 0.9].

[0181] ③ Gaussian blur enhancement:

[0182] Convolving the image with a Gaussian filter simulates optical or motion blur effects, enhancing the model's ability to recognize blurred targets. The Gaussian kernel size is set to... Add random, slight blurring to the image data.

[0183] Training: Network input preprocessing methods:

[0184] ① All images are now 448 pixels in size:

[0185] Fill the image into a square shape by using the shortest side of the image to fill with black pixels, then compress it proportionally. size.

[0186] ② Image data normalization:

[0187] Divide the original value of the pixel (usually an integer between 0 and 255) by 255 to convert it to a floating-point number between [0.0, 1.0].

[0188] ③ Image data digitization:

[0189] Move the center point (mean) of the data to 0. This helps the optimizer converge faster and more stably. That is, subtract the mean [0.485, 0.456, 0.406] from the normalized image data.

[0190] ④ Image data standardization:

[0191] The numerical range of each feature (channel) is scaled to a similar scale, i.e., the centered image data is divided by the standard deviation [0.229, 0.224, 0.225].

[0192] (2) Training parameters:

[0193] The Adam optimizer is used, and the weight decay is set to... Momentum parameters , The batch size was set to 64, and training was performed using two NVIDIA RTX 5090 GPUs. The total number of training epochs was 300, and the initial learning rate was set to... It employs distributed decay, where the learning rate decreases to 10% of its original value every 50 epochs, meaning the learning rate is multiplied by 0.1 every 50 epochs.

[0194] (3) Loss function:

[0195] During training, the image data of each patient group predicts the corresponding coronary heart disease outcome, i.e., the category is coronary heart disease or non-coronary heart disease. The model's judgment of coronary heart disease is optimized by minimizing the difference between the prediction and the target. The classification loss is defined as follows:

[0196]

[0197] in For the input image data, Total number of categories One-hot encoding of the true category of image data. This set of image data belongs to the category The predicted probability.

[0198] Training results: After 300 rounds of iterative training, the loss curve tends to flatten out, as shown below. Figure 12 As shown, this indicates that the model has converged and its performance is relatively stable.

[0199] Validation set testing:

[0200] To verify the effectiveness of the AI-CAD network, the core network TongueViT was replaced with the mainstream baseline models ResNet50 and GoogLeNet. A comparative experiment was conducted using a validation set of 126 tongue samples (including 60 cases of coronary artery disease and 66 cases of non-coronary artery disease). The three models were used to perform inference predictions on the validation set samples, and the recognition results of each model are summarized below:

[0201]

[0202] Comparing the three models, the TongueViT core network performed best. The TongueViT model achieved an area under the curve (AUC) of 0.912 and an accuracy of 90.50% during the validation of the clinical validation set of 126 cases. Figure 13 , Figure 14 As shown.

[0203] In one specific embodiment, a model interpretability experiment was conducted: the interpretability of medical AI is crucial for understanding model decisions and improving credibility. This paper calculates the gradient of the coronary heart disease category output with respect to the convolutional layer feature map, and then generates an intuitive coronary heart disease heatmap by weighted summation of the gradient and the feature map, providing a visual explanation for model decisions, such as... Figure 15 As shown, in patients with fatty liver and coronary heart disease, the tongue thermogram exhibits highly specific distribution characteristics: significantly enhanced thermal signals are observed in the thick coating area at the base of the tongue, the fissure area on the tongue surface, the tip of the tongue, and the sides of the tongue body. This specific thermal distribution lays an important foundation for the auxiliary diagnosis of coronary heart disease using tongue diagnosis. Figure 16 The results showed that sublingual thermography exhibited significantly specific distribution characteristics in patients with fatty liver and coronary heart disease. Specifically, the thermal response of the sublingual vein region was significantly enhanced. This specific thermal distribution provides important evidence for the auxiliary diagnosis of coronary heart disease through tongue examination. Figure 17The results showed that facial thermography exhibited significantly specific distribution characteristics in patients with fatty liver and coronary heart disease. Specifically, the thermal response was significantly enhanced in the nasal folds, perioral region, perioral region, and bilateral cheek areas. This specific thermal distribution provides important evidence for the auxiliary diagnosis of coronary heart disease through facial examination. Figure 18 The results showed that in patients with fatty liver and coronary heart disease, the thermograms of the left and right ears exhibited significantly specific distribution characteristics. Specifically, the thermal response in the ear crease region was significantly enhanced. This specific thermal distribution provides important evidence for the auxiliary diagnosis of coronary heart disease through facial examination.

[0204] In one specific embodiment, to verify the positive effect of multi-input scenarios (i.e., simultaneous input of images of the subject's tongue, sublingual region, frontal view, and left and right ears) on the efficacy of the coronary heart disease prediction model, an ablation experiment was conducted on the input images. The experiment compared the performance of the multi-input model and the single-input model, specifically including five models: using only tongue images, using only sublingual images, using only frontal images, using only left and right ear images, and using multi-dimensional images. The recognition results for various inputs are statistically analyzed as follows:

[0205]

[0206] The experimental results above show that using only sublingual or left / right ear images contributes little to the coronary artery disease prediction model; however, using only frontal or tongue images results in good prediction performance; and using multiple input images achieves the highest accuracy and AUC value. Therefore, applying multiple input images to the coronary artery disease prediction model is effective.

[0207] In one specific embodiment, an ablation experiment of the network attention module:

[0208] To verify the effectiveness of the LFV and NLFV attention mechanisms, this study compares and analyzes the model performance under different attention configurations, specifically including five model architecture configurations: no attention mechanism, traditional cross-attention mechanism with only LFV attention, only NLFV attention, and LFV+NLFV. The recognition results for each configuration are as follows:

[0209]

[0210] The "-" indicates that no attention mechanism is used.

[0211] The experimental results show that the model performs worst when no attention mechanism is used. Furthermore, using Cross attention, LFV attention, or NLFV attention alone limits model performance because they cannot simultaneously capture texture and color features. In contrast, the synergistic combination of LFV and NLFV attention mechanisms demonstrates significant advantages: LFV attention enhances the detailed modeling of local texture features, while NLFV attention improves the macroscopic perception of global color features. This multi-scale feature fusion mechanism complementarily enhances the model's representational capabilities across different spatial dimensions, ultimately leading to a significant improvement in the accuracy of tongue image recognition for coronary heart disease.

[0212] In one embodiment, S1 is replaced by: acquiring a second tongue diagnosis image, facial image, and clinical data of a patient with fatty liver;

[0213] The S2 is replaced by: the second tongue diagnosis image and facial image are sequentially passed through an image embedding layer and an image coding layer to obtain image features; the clinical data are sequentially passed through a text embedding layer and a text coding layer to obtain text features; the image features and text features are passed through an integration layer to obtain fusion features; and the fusion features are passed through a classification prediction layer to predict whether a patient with fatty liver has coronary heart disease or not.

[0214] In one embodiment, the second tongue diagnosis image includes a tongue surface image and a sublingual image.

[0215] Optionally, the image encoder and text encoding layer encode image features or text features using a pre-trained model.

[0216] Optionally, the pre-trained model of the image encoder may be one or more of the following: ViT-B-32, UniFormer, Swing Transformer, PLIP, CLIP.

[0217] Optionally, the pre-trained model of the text encoder may be one or more of the following: bert-base-chinese, ELECTRA-zh, RoBERTa-zh, ALBERT-zh, DistilBERT-zh, ERNIE-3.0.

[0218] Optionally, the image features and text features are input into the fusion layer, concatenated, and then subjected to nonlinear transformation and feature learning through a neural network. Finally, a classification prediction layer is used to make a prediction result.

[0219] Optionally, the clinical information includes one or more of the following: medical history, gender, age, and blood indicators.

[0220] Optionally, the blood indicators include one or more of the following: red blood cell distribution width, monocytes, blood glucose, triglycerides, albumin / globulin ratio, total bilirubin, low-density lipoprotein, mean platelet volume, apolipoprotein A1, total cholesterol, prealbumin, and alkaline phosphatase.

[0221] In one specific embodiment, the Tongue Diagnosis Multimodal Model (TongueMMM) is used, such as... Figure 19 The result of the multimodal TCM visual language model, which combines tongue surface, sublingual, facial images and natural language, is shown. It consists of three parts: Tongue visual encoder, medical terminology encoder and Tongue modality fusion learner. (1) Tongue visual encoder: This encoder is based on the ViT-B-32 pre-trained model and is optimized through transfer learning strategy. It processes multi-source visual data such as tongue surface, sublingual and facial images at the same time, and uses the self-attention mechanism of Transformer to capture valuable medical attributes of local image details. (2) Tongue text encoder: It uses Chinese word library to encode the patient's gender, age, medical history, physical indicators and other data by word embedding. After the text encoder extracts effective feature data, it uses bert-base-chinese pre-trained model transfer learning to learn the ability to understand the text features of medical terms. (3) Tongue modality fusion learning: By combining the visual encoding features of tongue surface, sublingual and facial images with the natural language features such as the patient's physical indicators, a unified multimodal feature vector is constructed. The MLP network is used to perform nonlinear transformation and feature learning on the fused feature vector, and finally outputs a feature representation that can be used for disease classification.

[0222] In one specific embodiment, three images—tongue surface, sublingual image, and frontal image—along with textual data such as medical history, gender, and age, were used as input. A validation set of 126 tongue surface images and patient information (60 cases of coronary heart disease and 66 cases of non-coronary heart disease) was used to verify the recognition results through back-calculation using a trained multimodal model. The TongueMMM model, trained using three types of image data (tongue surface, sublingual image, and facial image) and patient-related textual data, achieved an accuracy of 90.5%, a precision of 96.2%, and a specificity of 97%. Compared to single-modal facial visual models and tongue image visual models, the TongueMMM model showed improved performance, thanks to the richer and more comprehensive information provided by multimodal data, effectively compensating for the limitations of single-modal data. The loss curve of TongueMMM training after 500 epochs shows that after 180 epochs, the loss leveled off and remained relatively flat in subsequent training, indicating that the model's fitting ability reached its optimal state. Figure 20In the ROC curve graph, TongueMMM's ROC curve is close to the upper left corner of the coordinate axis, indicating good model performance and a slight improvement over the single-modal vision model. In the PR curve graph, TongueMMM's PR curve is close to the upper right corner of the coordinate axis, indicating good model performance and a slight improvement over the single-modal vision model. The confusion matrix of the coronary artery disease prediction model is the validation set.

[0223] In another embodiment, three images—tongue surface, sublingual area, and frontal view—are used as input, along with textual data on medical history, gender, age, and blood indicators. The blood indicators selected include (red blood cell distribution width, monocyte percentage, blood glucose, triglycerides, albumin / globulin ratio, total bilirubin, low-density lipoprotein, mean platelet volume, apolipoprotein A1, total cholesterol, prealbumin, and alkaline phosphatase), which are converted into four values: normal, low, high, and default, according to relevant standards. The TongueMMM model achieved an accuracy of 91.3%, precision of 96.2%, and specificity of 97%. Compared to using only gender, age, and medical history as input, this model incorporates examination indicators, resulting in improved accuracy. Therefore, examination indicators contribute to the improvement of model performance.

[0224] In one embodiment, S1 is replaced by: acquiring facial features, tongue texture, tongue coating, and sublingual data of a patient with fatty liver.

[0225] S2 is replaced by: predicting whether a patient with fatty liver has coronary heart disease based on the facial features, tongue texture, tongue coating, and sublingual data.

[0226] In one embodiment, the facial features include one or more of the following: nasal folds, left eye color, right eye color, left ear folds, right ear folds, left ear color, right ear color, lip color, main color, and gloss.

[0227] Optionally, the tongue body includes one or more of the following: age, thickness, color, tongue tip color-Y, tongue tip color-Cr, tongue tip color-Cb, tongue root color-Y, tongue root color-Cr, tongue root color-Cb, tongue edge color-Cr, tongue middle color-R, tongue middle color-Cr, and tongue middle color-Cb.

[0228] Optionally, the tongue coating includes one or more of the following: tongue color, tongue coating color-ASM, tongue coating second-order color moment BGR, and tongue coating correlation coefficient COR;

[0229] Optionally, the sublingual vein includes one or more of the following: sublingual vein shape, sublingual vein color, sublingual-Cr, sublingual-Cb, and sublingual-Y;

[0230] Optionally, the method further includes body fluids and constitution, obtaining body fluids and constitution data, and making predictions based on the facial appearance, tongue texture, tongue coating, sublingual data, body fluids, and constitution to obtain prediction results of whether fatty liver patients have coronary heart disease or not.

[0231] Optionally, the method further includes blood indicators, which are used to predict whether a patient with fatty liver has coronary heart disease based on facial features, tongue texture, tongue coating, sublingual data, body fluids, constitution, and blood indicators. The blood indicators include one or more of the following: red blood cell count, hemoglobin, platelet count, blood protein, alanine aminotransferase, blood glucose, aspartate aminotransferase, gamma-glutamyl transferase, alkaline phosphatase, and adenosine deaminase.

[0232] Optionally, the method further includes obtaining clinical information, and predicting whether a patient with fatty liver has coronary heart disease based on the facial features, tongue texture, tongue coating, sublingual data, and clinical information. The clinical information includes one or more of the following: age, diabetes, hypertension, hepatitis B, and smoking history.

[0233] In one embodiment, the prediction is made using one or more of the following algorithms: SVM, LightGBM, MLP, Random Forest, XGBoost, GBM, and AdaBoost.

[0234] In one specific embodiment, the numerical features (the values ​​of the features are numerical): age, number of petechiae, height ratio of collateral vein 1, width ratio of collateral vein 1, area ratio of collateral vein 1, height ratio of collateral vein 2, width ratio of collateral vein 2, area ratio of collateral vein 2, root-coating color, root-tongue color, whole tongue-coating color, whole tongue-crimson, sublingual-Y, sublingual-Cr, sublingual-Cb, tongue coating-first order color moment-R, tongue coating-second order color moment-G, and tongue coating-third order color moment-G.

[0235] Categorical characteristics (characteristic values ​​are of two or more types): gender, smoking history, diabetes, hypertension (diabetes and hypertension: data from comorbidities), hepatitis B, tongue color, tongue coating color, saliva, sublingual abnormalities, sublingual pulse color, sublingual pulse shape, constitution, peeling, thickness, no coating, greasy, fat or thin, old or young, prickles, cracks, liver stagnation lines, teeth marks, petechiae, abnormal tongue shape, eye color (identifying red and swollen eyes, yellow eyes), complexion (identifying a complexion that is subtly reddish-yellow, bright and moist, etc.). (Dark, withered, or exposed skin color), main color (identifying red, yellow, white, black, and bluish skin), ear creases (identifying ear creases and mild, moderate, and severe), ear color (identifying white, red, yellow, and bluish-purple ear colors), lip color (identifying white, red, dark red, bluish-purple, black, or spotted), dark circles under the eyes (identifying mild and severe dark circles), eye expression (identifying bright and dull eyes), nose creases (identifying nose creases and mild, moderate, and severe), and luster (identifying normal luster, slight luster, and no luster).

[0236] Tongue color classification: 6 types (pale white tongue, pale tongue, pale red tongue, red tongue, crimson tongue, bluish-purple tongue), supporting 3 organ region analyses;

[0237] Tongue shape detection: 9 types (fat, thin, old, young, prickles, ecchymosis, petechiae, cracks, teeth marks); Coating color identification: 6 types (white coating, light yellow coating, yellow coating, scorched yellow coating, gray-black coating, scorched black coating); Coating texture identification: thin coating, thick coating, putrid, greasy, peeling, no coating, little coating; Body fluid identification: moist, slippery, dry;

[0238] Sublingual veins: 1. Sublingual identification: normal, weak, and stagnant; 2. Degree of blood stasis: including quantitative analysis of pulse color and shape;

[0239] Data preprocessing: 1. Missing value handling: Numerical variables: filled with the mean; Categorical variables: added a missing category; Target variable: deleted rows containing missing values; 2. Data type conversion: converted the target variable from 'yes' / 'no' to 1 / 0; uniformly coded categorical variables to numeric type. Information was collected from 1913 individuals with fatty liver disease, 288 with coronary heart disease, 513 with hypertension, and 266 with diabetes.

[0240] In one embodiment, the five-part division of the tongue involves dividing the tongue into five parts: left, right, upper, lower, and middle. The algorithm steps are as follows: Take the smallest bounding rectangle of the tongue, connect the upper 1 / 5 of the left and right sides, connect the lower 1 / 5 of the left and right sides, connect the left 1 / 5 of the upper and lower sides, and connect the right 1 / 5 of the upper and lower sides. These four lines divide the smallest bounding rectangle into nine regions: A, B, C, D, E, F, G, H, and I. Region B is designated as the root of the tongue, regions A, D, and G as the left side of the tongue, regions C, F, and I as the right side of the tongue, region E as the middle of the tongue, and region H as the tip of the tongue.

[0241] The tongue coating color analysis process is as follows: Tongue regions are extracted from the TCM tongue image database, and tongue body HSV color space clustering is performed. Based on the clustering results, tongue color and coating color data cards are automatically generated. A pixel color attribute classifier X is constructed using the XGBoost machine learning algorithm, and a whole tongue color classification model s and a whole tongue coating color classification model t are also constructed using the XGBoost machine learning algorithm. The color classifier X is used to calculate the color attribute classification of all pixels on the tongue body, obtaining the number of pixels c and the color ratio f for each category. Then, the tongue color model s is used to calculate the whole tongue color, and the coating color model t is used to calculate the whole tongue coating color. According to the local feature definition requirements, local tongue color and coating color features are calculated for five regions of the tongue body.

[0242] Y, Cr, and Cb are components of the YCrCb color space.

[0243] In one specific embodiment, the overall predictive power of vital signs, facial features, tongue texture, tongue coating, and sublingual features on the target variable is quantitatively analyzed, and the importance ranking of the features is output, quantifying the contribution of each feature. A random forest model is used to extract the model's feature contribution attribute as the feature importance score.

[0244] Select the top 10 features of discrete features, such as Figure 21 As shown, the importance scores of fatty liver severity and age exceed 0.20, while the importance of hepatitis B, hypertension, coronary heart disease (family history), lip color, constitution, and diabetes decrease in that order.

[0245] Select the top 10 features of continuous features, such as Figure 22 As shown, the importance of sublingual-Cr (color mean), BMI, waist circumference, and sublingual-Cb is relatively high, exceeding 0.02. The importance of the following features decreases in the following order: root-lightness (pixel percentage), root-tongue color (pixel percentage), sublingual-Y, tip-lightness, edge-moistening color-Cb, root-moistening color-Cb, and tongue coating-roughness.

[0246] Select the top 10 features of blood characteristics. Figure 23 As shown, the importance score of the white / globulin ratio distribution width exceeds 0.35, and the importance of total bilirubin, creatine kinase, lactate dehydrogenase, total protein, blood glucose, uric acid, and aspartate aminotransferase decreases in that order.

[0247] In one specific embodiment, this invention employs a multi-stage parallel feature selection strategy, using three methods in synergistic analysis to ultimately determine the key feature set for predicting coronary heart disease. The specific selection process is as follows: First stage (univariate analysis): Based on five dimensions—physical signs, facial appearance, tongue texture, tongue coating, and sublingual tissue—the correlation between each feature and coronary heart disease is statistically tested (p < 0.05) to initially select feature variables with statistical significance. Second stage (machine learning feature importance assessment): Using a random forest model, discrete features, continuous features, and physical signs are modeled separately. The predictive contribution of each variable is quantified using the Gini importance score, retaining features with high importance. Third stage (cluster analysis): Through hierarchical clustering, the potential structure of features related to coronary heart disease is analyzed, and feature combinations with predictive value are selected. This method, through multi-dimensional parallel selection, ensures the statistical significance of features while improving the predictive ability of the model, providing a reliable feature foundation for subsequent modeling. The final characteristics used for predicting coronary heart disease are as follows: age, diabetes, hypertension, hepatitis B, tongue color, tongue coating color, smoking history, constitution, physical condition, body fluids, thickness, nasal folds, luster, eye color (right eye), eye color (left eye), ear folds (right ear), ear folds (left ear), cheek redness, sublingual veins, sublingual vein color, ear color (left ear), ear color (right ear), dominant color, lip color, gender, family history of coronary heart disease, family history of hypertension, waist circumference, BMI, sublingual Cr, sublingual Cb, sublingual Y, and apex of the tongue. Tongue color - Y, tip - tongue color - Cr, tip - tongue color - Cb, root - tongue color - Y, edge - tongue color - Cr, middle - tongue coating - R, middle - tongue coating color - Cr, middle - tongue coating color - Cb, root - tongue coating color - Cr, root - tongue coating color - Cb, tongue coating - ASM, tongue coating - second-order color moment BGR, tongue coating - correlation COR, red blood cell count, hemoglobin, platelet count, hemoglobin, alanine aminotransferase, blood glucose, aspartate aminotransferase, gamma-glutamyl transferase, alkaline phosphatase, adenosine deaminase.

[0248] In a dataset containing 1913 samples, we randomly divided the dataset into a training set of 1626 samples and a test set of 287 samples. We trained the dataset using Support Vector Machine (SVM), Lightweight Gradient Boosting Machine Learning (LightGBM), and Multilayer Perceptron (MLP) models, and evaluated their performance on the test set. The SVM model achieved an accuracy of 90.59%, precision of 82.86%, recall of 58%, specificity of 97.47%, and an AUC of 0.88. The LightGBM model achieved an accuracy of 87.80%, precision of 77.78%, recall of 42%, specificity of 97.47%, and an AUC of 0.88. The MLP model achieved an accuracy of 87.80%, precision of 67.44%, recall of 58.0%, and specificity of 94.09%, with an AUC of 0.87. Experimental results combining machine learning models (SVM, LightGBM, and MLP) showed that the coronary heart disease identification model had an AUC of approximately 0.87, an accuracy of approximately 88%, a precision of approximately 76%, a recall of approximately 53%, and a specificity of approximately 97%. This indicates that the model can effectively distinguish the negative class (those without coronary heart disease).

[0249] The present invention also discloses a computer program product or system, including a computer program that, when executed by a processor, implements the above-described steps of the method for predicting fatty liver combined with coronary heart disease based on tongue diagnosis.

[0250] Figure 2 A schematic diagram of a tongue-diagnosis-based predictive system for fatty liver complicated with coronary heart disease provided in this embodiment of the invention specifically includes:

[0251] Acquisition module: Acquires tongue images of patients with fatty liver;

[0252] Prediction module: Inputs tongue diagnosis images into the tongue diagnosis model to predict whether fatty liver patients have coronary heart disease or not;

[0253] The tongue diagnosis image is sequentially passed through the input layer, convolutional module, N tongue diagnosis modules, pooling layer, fully connected layer, and output layer of the tongue diagnosis model to obtain the prediction result. N is a natural number greater than 1. The tongue diagnosis module includes a feature extraction module, a local attention layer, and a global attention layer. The output feature vector of the convolutional module is input into the tongue diagnosis module and sequentially passes through the feature extraction module, the local attention layer, and the global attention layer to obtain the output feature vector of the tongue diagnosis module.

[0254] Figure 3 The schematic diagram of the computer device provided in this embodiment of the invention specifically includes:

[0255] A memory and a processor; the memory is used to store program instructions; the processor is used to invoke the program instructions, when any of the above-mentioned methods for predicting fatty liver combined with coronary heart disease based on tongue diagnosis are executed.

[0256] The present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, represents any of the above-described methods for predicting fatty liver combined with coronary heart disease based on tongue diagnosis.

[0257] The verification results of this verification embodiment show that assigning inherent weights to indications can improve the performance of this method compared to the default settings. Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, indirect coupling or communication connection of devices or units, and may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated; the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of this embodiment. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0258] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0259] The computer device provided by the present invention has been described in detail above. For those skilled in the art, there will be changes in the specific implementation and application scope based on the ideas of the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for constructing a prediction model of fatty liver combined with coronary heart disease, characterized in that, include: Acquire patient tongue and facial image datasets; The tongue image data includes images of the tongue surface region and the sublingual vein region, and the facial image data includes images of the face region and the ear region. The images of the tongue surface region, the sublingual vein region, the face region, and the ear region are respectively input into a parallel visual encoding module for feature extraction to obtain tongue surface features, sublingual vein features, facial features, and ear features. The tongue surface features, sublingual vein features, facial features, and ear features are input into the feature adaptive modulator module for feature fusion to obtain modulated fusion features; The visual encoding module includes, in sequence, an input layer, a backbone module, a main body block, a header module, and an output layer. The backbone module includes L cascaded convolutional layers for high-dimensional feature extraction; The main block includes S serially connected sub-modules; each sub-module sequentially includes a first convolutional module, a restricted receptive field attention module, and an unrestricted receptive field attention module. Features are input to the next sub-module or the head module after passing through the first convolutional module, the restricted receptive field attention module, and the unrestricted receptive field attention module in the sub-module. The first convolutional module includes cascaded depthwise separable convolution and lightweight channel attention. By integrating depthwise separable convolution and lightweight channel attention, high-dimensional features are extracted from image detail features. The restricted receptive field attention module extracts local features by using a preset fixed receptive field, while the unrestricted receptive field attention module extracts global features by using a dynamic receptive field to extract long-distance dependency features. The head module includes a global average pooling layer for feature dimensionality reduction; The feature adaptive modulator module includes a hybrid expert module, a fusion layer, and a dimensionality reduction layer; The tongue surface features, sublingual vein features, facial features, and ear features are input into the feature adaptive modulator module. First, the weights of the corresponding features are calculated by the hybrid expert module. Then, the weights of the tongue surface features, sublingual vein features, facial features, ear features, and corresponding weights are weighted and fused to obtain the gated fused features. The gated fused features are then reduced in dimensionality by a dimensionality reduction layer and output as modulated fused features. The modulation fusion features are input into the classification layer for classification to obtain the predicted classification result; The predicted classification result is compared with the true classification label to calculate the loss. The above steps are repeated until the loss function remains unchanged, thus obtaining the prediction model for fatty liver combined with coronary heart disease. 2.The method of claim 1, wherein, The first convolutional module, the restricted receptive field attention module, and the unrestricted receptive field attention module are connected in series, and residual connections are also included between the modules.

3. The method for constructing a predictive model for fatty liver complicated with coronary heart disease according to claim 1, characterized in that, The global features include features from different regions and / or features from the same region at different time points.

4. The method for constructing a predictive model for fatty liver complicated with coronary heart disease according to claim 1, characterized in that, The local features include texture features, color features, and morphological features.

5. The method for constructing a predictive model for fatty liver complicated with coronary heart disease according to claim 1, characterized in that, The preset fixed receptive field is a computational receptive field that is limited by spatial constraints at each position. A local region is formed with the current position as the center and a radius of r. Attention calculation is performed on the features of the local region to obtain local features.

6. The method for constructing a predictive model for fatty liver complicated with coronary heart disease according to claim 1, characterized in that, The dynamic receptive field is a region with radius R centered on the current position. The region decays linearly with distance. When calculating attention features, the attention weights are weighted by a dynamic distance mask to obtain global features.

7. A method for predicting fatty liver complicated with coronary heart disease, characterized in that, The method includes: acquiring facial images and tongue diagnosis images of a patient with fatty liver, inputting the facial images and tongue diagnosis images into a prediction model for fatty liver combined with coronary heart disease obtained by the prediction model construction method for fatty liver combined with coronary heart disease according to any one of claims 1-6, and making a prediction result to obtain whether the patient with fatty liver has coronary heart disease or not.

8. The method for predicting fatty liver complicated with coronary heart disease according to claim 7, characterized in that, The facial image includes a facial region image and an ear region image, and the tongue diagnosis image includes a tongue surface region image and a sublingual vein region image. The tongue surface region image, the sublingual vein region image, the facial region image, and the ear region image are input into the prediction model for fatty liver combined with coronary heart disease obtained by the prediction model construction method of any one of claims 1-6 to predict whether the fatty liver patient has coronary heart disease or not.

9. A computer program product comprising a computer program or instructions, characterized in that, The computer program or instructions are executed by the processor to implement the method for constructing a predictive model for fatty liver combined with coronary heart disease as described in any one of claims 1-6, or to implement the method for predicting fatty liver combined with coronary heart disease as described in any one of claims 7-8.

10. A computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, characterized in that, The computer program or instructions are executed by the processor to implement the method for constructing a predictive model for fatty liver combined with coronary heart disease as described in any one of claims 1-6, or to implement the method for predicting fatty liver combined with coronary heart disease as described in any one of claims 7-8.

11. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instructions are executed by the processor to implement the method for constructing a predictive model for fatty liver combined with coronary heart disease as described in any one of claims 1-6, or to implement the method for predicting fatty liver combined with coronary heart disease as described in any one of claims 7-8.

Citation Information

Patent Citations

  • Intelligent hospital guide method, model training method and device, equipment and storage medium

    CN120089343A