Multi-mode diabetes risk prediction method based on combination of traditional Chinese medicine and western medicine
By constructing a cohort of prediabetic patients and combining it with information from traditional Chinese medicine tongue and pulse diagnosis, a multimodal joint representation model for diabetes risk prediction was built. This solved the problem of insufficient utilization of traditional Chinese medicine information in existing technologies and achieved efficient prediction of long-term diabetes risk.
Patent Information
- Application Number
- CN202511396825.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-12-19
AI Technical Summary
Existing methods for predicting diabetes risk are unable to fully utilize information from traditional Chinese medicine diagnoses, resulting in insufficient accuracy in predicting long-term diabetes risk.
A cohort of patients with prediabetes was constructed, and medical records, tongue and pulse diagnosis information, pulse signals, retinal and fundus images, and clinical indicators were collected from TCM experts. Tongue and pulse feature extractors were pre-trained through multimodal comparative learning, and a multimodal joint representation model for diabetes risk prediction was constructed by combining retinal images and clinical indicators.
It significantly improved the predictive performance of long-term diabetes risk, enhanced the utilization value of TCM diagnostic information, and improved the accuracy and stratification of risk prediction.
Smart Images

Figure CN121171601A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of medical image processing, in particular to a multi-modal diabetes risk prediction method combining traditional Chinese medicine and Western medicine. BACKGROUND
[0002] Diabetes is a metabolic disease that seriously threatens human health, and its long-term development can lead to complications such as cardiovascular and cerebrovascular diseases, retinopathy, and kidney damage. If pre-diabetic patients do not receive timely intervention, they will gradually develop into diabetic patients or even suffer from serious complications within 5-10 years. Existing diabetes risk prediction methods mainly rely on clinical indicators and epidemiological statistical models, but they are difficult to capture individualized characteristics of patients.
[0003] Traditional Chinese medicine has unique advantages in early identification of diabetes. Information such as tongue and pulse can reflect the overall metabolism and blood state of patients, but its diagnosis relies on experience and lacks objective and quantitative tools. Meanwhile, in modern medicine, retinal images serve as a "window" to the body's microcirculation, providing a direct reflection of vascular lesions and a close relationship with the development of diabetes. Therefore, how to combine traditional Chinese medicine and Western medicine diagnosis information to build a multi-modal fusion risk prediction model and improve the long-term risk stratification prediction ability of diabetes has become an important issue in clinical and scientific research. SUMMARY
[0004] The purpose of the present application is to solve the problem that the existing method cannot fully utilize traditional Chinese medicine diagnosis information and has insufficient accuracy in predicting long-term risk of diabetes, and to provide a multi-modal diabetes risk prediction method combining traditional Chinese medicine and Western medicine.
[0005] The above invention is mainly realized by the following technical scheme:
[0006] S1, a pre-diabetic patient cohort is constructed, and medical history information of tongue and pulse, tongue image, pulse signal, retinal fundus image, and clinical indicator information of traditional Chinese medicine experts are collected. According to the 10-year follow-up results, the patients in the cohort are labeled as low-risk, medium-risk, and high-risk diabetes patients, and the steps are as follows:
[0007] (1) At baseline, patients with at least two high-risk factors for diabetes are classified as pre-diabetic patients and included in the pre-diabetic patient cohort;
[0008] (2) Traditional Chinese medicine experts record the medical history information of tongue diagnosis and pulse diagnosis of the patients in the cohort;
[0009] (3) Collect clinical indicator information including age, gender, weight, body mass index (BMI), and blood pressure, and use collection equipment to collect retinal fundus images, tongue images, and pulse signals of patients;
[0010] (4) 10 years of follow-up of the patients in the queue, according to the results of follow-up of pre-diabetic patients into pre-diabetic patients reversed to normal group, pre-diabetic patients in the follow-up process did not change group, pre-diabetic patients progress to diabetes group, and are divided into low-risk diabetes, diabetes high-risk population and diabetes risk population.
[0011] S2, using the tongue and pulse diagnosis of TCM experts medical information and collected tongue image, pulse signal to tongue feature extractor, pulse feature extractor and text information feature extractor for multi-modal contrast learning pre-training, the steps are:
[0012] (1) Obtain the tongue and pulse diagnosis of TCM experts medical information, extract the tongue and pulse diagnosis text;
[0013] (2) The tongue and pulse diagnosis text is one-to-one corresponding to the tongue image and pulse signal of the patient;
[0014] (3) respectively construct tongue feature extractor, pulse feature extractor and text information feature extractor, use tongue feature extractor, pulse feature extractor and text information feature extractor to extract tongue image feature vector, pulse signal feature vector, tongue diagnosis text feature vector, pulse diagnosis text feature vector, the process is shown in formula (1):
[0015] (1)
[0016] Among them, tongue image feature vector is represented by pulse signal feature vector is represented by tongue diagnosis text feature vector is represented by pulse diagnosis text feature vector is represented by tongue feature extractor is represented by pulse feature extractor is represented by text information feature extractor is represented by tongue image is represented by pulse signal is represented by tongue and pulse diagnosis text is represented by
[0017] (4) respectively construct tongue image feature vector and tongue diagnosis text feature vector pair, pulse signal feature vector and pulse diagnosis text feature vector pair, and conduct contrast learning training, the loss function is shown in formula (2):
[0018] (2)
[0019] Among them, total loss of contrast learning is represented by represents the contrastive learning loss of the tongue feature extractor and the text information feature extractor, represents the contrastive learning loss of the pulse feature extractor and the text information feature extractor, and respectively shown in formula (3) and formula (4):
[0020] (3)
[0021] (4)
[0022] wherein represents the contrastive learning loss of the tongue feature extractor and the text information feature extractor, N represents the number of tongue diagnosis text feature vectors, max represents the maximum function, is a constant, represents the Euclidean distance, represents the contrastive learning loss of the pulse feature extractor and the text information feature extractor, M represents the number of pulse diagnosis text feature vectors, i and j represent the indexes of the vectors, represents the tongue image feature vector with index i, represents the tongue diagnosis text feature vector with index i, represents the tongue diagnosis text feature vector with index j, represents the pulse signal feature vector with index i, represents the pulse diagnosis text feature vector with index i, represents the pulse diagnosis text feature vector with index j.
[0023] (5) Pre-training the tongue feature extractor, the pulse feature extractor and the text information feature extractor using the above loss.
[0024] S3, a multi-modal joint representation diabetes risk prediction model is constructed, which includes a tongue feature extractor, a pulse feature extractor, a text information feature extractor, a retinal image feature extractor, a clinical information encoder, a feature fusion module and a risk classifier. The multi-modal joint representation diabetes risk prediction model is trained to convergence by using tongue image, pulse signal, tongue and pulse medical record information, retinal fundus image, clinical index information and diabetes risk level label. The steps are:
[0025] (1) The tongue feature extractor, the pulse feature extractor, the text information feature extractor, the retinal image feature extractor and the clinical information encoder are used to extract the feature vectors of the tongue image, the pulse signal, the tongue and pulse diagnosis text, the retinal fundus image and the clinical index information. The process is shown in formula (5):
[0026] (5)
[0027] wherein denotes the tongue image feature vector, denotes the pulse signal feature vector, denotes the tongue diagnosis text feature vector, denotes the pulse diagnosis text feature vector, denotes the retinal fundus image feature vector, denotes the clinical indicator feature vector, denotes the tongue feature extractor, denotes the pulse feature extractor, denotes the text information feature extractor, denotes the retinal image feature extractor, denotes the clinical information encoder, denotes the tongue image, denotes the pulse signal, denotes the tongue and pulse diagnosis text, denotes the retinal fundus image, denotes the clinical indicator;
[0028] (2) The tongue image feature vector, the pulse signal feature vector, the tongue diagnosis text feature vector, the pulse diagnosis text feature vector, the retinal fundus image feature vector, and the clinical indicator feature vector are fused and input into the risk classifier to obtain a risk prediction result, and the process is shown in formula (6):
[0029] (6)
[0030] denotes the risk prediction result, denotes the risk classifier, denotes the feature fusion module, denotes the tongue image feature vector, denotes the pulse signal feature vector, denotes the tongue diagnosis text feature vector, denotes the pulse diagnosis text feature vector, denotes the retinal fundus image feature vector, denotes the clinical indicator feature vector;
[0031] (3) The risk prediction result is used to calculate a loss function, and the model parameters are updated through back propagation to train the model to convergence.
[0032] S4, using the current tongue image, pulse signal, retinal fundus image, medical record information of tongue and pulse, and clinical index information of the patient as input, using the converged multi-modal joint representation diabetes risk prediction model to predict the diabetes risk of the patient in 10 years.
[0033] Inventive Effects
[0034] The present application provides a multi-modal diabetes risk prediction method combining traditional Chinese and Western medicine. First, a pre-diabetes patient cohort is constructed, and medical record information of tongue and pulse, tongue image, pulse signal, retinal fundus image, and clinical indicators of traditional Chinese medicine experts are collected, and the patients are labeled as low risk, medium risk, and high risk combined with 10-year follow-up results. Then, multi-modal contrast learning pre-training is performed based on tongue, pulse, and diagnostic text to improve the feature extraction capabilities of the tongue feature extractor, pulse feature extractor, and text information feature extractor. A multi-modal joint representation diabetes risk prediction model is constructed again, which includes a tongue feature extractor, a pulse feature extractor, a text information feature extractor, a retinal image feature extractor, a clinical information encoder, a feature fusion module, and a risk classifier. The multi-modal joint representation diabetes risk prediction model is trained to convergence using tongue image, pulse signal, medical record information of tongue and pulse, retinal fundus image, clinical index information, and diabetes risk level label. Finally, the multi-modal data of the patient is input into the converged model to predict the diabetes risk in the next 10 years. Experiments show that this method has the following advantages: (1) multi-modal contrast learning enhances the consistency of tongue, pulse, and diagnostic text, improving the value of traditional Chinese medicine diagnostic information (2) fusion of traditional Chinese and Western medicine diagnostic features significantly improves the performance of long-term diabetes risk prediction. The present application can be widely applied to risk assessment of pre-diabetes population. BRIEF DESCRIPTION OF DRAWINGS
[0035] Figure 1 A flowchart of the multi-modal diabetes risk prediction method combining traditional Chinese and Western medicine in the present application example;
[0036] Figure 2 A flowchart of the pre-diabetes patient cohort construction in the present application example;
[0037] Figure 3 A flowchart of the tongue, pulse, and text information feature extractor training in the present application example;
[0038] Figure 4 A structure diagram of the multi-modal joint representation diabetes risk prediction model in the present application example. DETAILED DESCRIPTION Detailed implementation method one:
[0040] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0041] As shown in Figure 1 The multi-modal diabetes risk prediction method combining traditional Chinese and Western medicine comprises the following steps:
[0042] S1, a pre-diabetes patient cohort is constructed, subjects meeting the pre-diabetes diagnosis criteria are included from regional hospitals and community screening, and patient information is recorded. When collecting the tongue and pulse case information of the TCM experts, the tongue image is obtained by a tongue image collector, and the tongue and fur characteristics are recorded by the TCM experts. The pulse signal is obtained by an electronic pulse diagnosis instrument, and the pulse diagnosis record is filled by the TCM experts. The retinal fundus image and clinical index information of the patient are collected. According to the 10-year follow-up results, the patients in the cohort are labeled as low-risk diabetes patients, medium-risk diabetes patients and high-risk diabetes patients.
[0043] S2, the multi-modal contrast learning pre-training is performed on the tongue feature extractor, the pulse feature extractor and the text information feature extractor by using the case information of the tongue and pulse diagnosis of the TCM experts and the collected tongue image and pulse signal.
[0044] S3, a multi-modal joint representation diabetes risk prediction model is constructed, which comprises a tongue feature extractor, a pulse feature extractor, a text information feature extractor, a retinal image feature extractor, a clinical information encoder, a feature fusion module and a risk classifier. The multi-modal joint representation diabetes risk prediction model is trained to convergence by using the tongue image, the pulse signal, the case information of the tongue and pulse, the retinal fundus image, the clinical index information and the diabetes risk level label.
[0045] S4, the current tongue image, pulse signal, retinal fundus image, case information of the tongue and pulse, and clinical index information of the patient are input, and the converged multi-modal joint representation diabetes risk prediction model is used to predict the diabetes risk of the patient after 10 years.
[0046] The embodiments of the present application will be described in detail as follows:
[0047] The embodiments of the present application are implemented as follows.
[0048] S1, as Figure 2As shown, a pre-diabetic patient cohort is constructed, including subjects meeting the diagnostic criteria for pre-diabetes from regional hospital and community screening, and recording patient information, collecting tongue and pulse case information of TCM experts, tongue image is obtained by tongue image collector, tongue and fur characteristics are recorded by TCM experts, pulse signal is obtained by electronic pulse diagnosis instrument, and pulse diagnosis record is filled by TCM experts, retinal fundus image and clinical index information of patients are collected, according to the 10-year follow-up results, the patients in the cohort are labeled as low-risk diabetes patients, medium-risk diabetes patients and high-risk diabetes patients, the steps are as follows:
[0049] (1) Patients with at least two high-risk factors for diabetes are listed as pre-diabetic patients and enter the pre-diabetic patient cohort, the high-risk factors for diabetes include: age ≥ 40 years old, body mass index BMI ≥ 24 kg / m 2 , a first-degree relative has a history of diabetes, a lack of physical activity, a woman has a history of macrosomia or gestational diabetes, a woman has a history of polycystic ovary syndrome, a person has pachydermia, a person has a history of hypertension, a person has a history of atherosclerotic cardiovascular disease, and a person has a China diabetes risk score ≥ 25;
[0050] (2) TCM experts perform tongue diagnosis on the patients in the cohort, use a tongue image collector to obtain high-resolution images of the tongue and fur, and record information such as tongue color and thickness, fill in the tongue diagnosis medical record, perform cut diagnosis on the patients, use an electronic pulse diagnosis instrument to obtain pulse signals, and record information such as pulse position, pulse potential, and pulse shape based on the results of palpation, and form a pulse diagnosis medical record;
[0051] (3) Collect clinical index information of the patients, including age, gender, weight, body mass index BMI, blood pressure, and use an eye fundus camera to collect retinal fundus images of both eyes;
[0052] (4) Establish a follow-up mechanism for the patients in the cohort, conduct a 10-year follow-up, record the disease progression of the patients once a year, and after the follow-up ends, divide the risk level according to the health outcome of the patients: if the pre-diabetic patients reverse to the normal population group, they are the low-risk diabetes population, if the patients maintain the pre-diabetic state during the follow-up, they are the medium-risk diabetes population, and if the patients are diagnosed with diabetes during the follow-up, they are the high-risk diabetes population.
[0053] S2, as shown in Figure 3 , the case information of tongue and pulse diagnosis of TCM experts and the collected tongue image and pulse signal are used to pre-train the tongue feature extractor, the pulse feature extractor and the text information feature extractor through multi-modal contrast learning, and the steps are as follows:
[0054] (1) Acquire TCM expert tongue and pulse diagnosis medical record information, extract tongue diagnosis text and pulse diagnosis text, and perform natural language processing such as word segmentation and stop word removal before input to ensure the accuracy and consistency of semantic expression;
[0055] (2) Correspond the patient's tongue image with the tongue diagnosis text, and correspond the patient's pulse signal with the pulse diagnosis text, to ensure correct pairing of data during training and avoid semantic misplacement;
[0056] (3) Construct a tongue feature extractor, a pulse feature extractor, and a text information feature extractor, respectively. The tongue feature extractor uses ResNet-18 as the backbone network to extract deep features of the tongue image, the pulse feature extractor uses a recurrent neural network to extract time sequence features of the pulse waveform, and the text information feature extractor uses a pre-trained language model BERT to extract semantic vectors of the tongue and pulse diagnosis text. The process is shown in equation (1):
[0057] (1)
[0058] Wherein represents the tongue image feature vector, represents the pulse signal feature vector, represents the tongue diagnosis text feature vector, represents the pulse diagnosis text feature vector, represents the tongue feature extractor, represents the pulse feature extractor, represents the text information feature extractor, represents the tongue image, represents the pulse signal, represents the tongue and pulse diagnosis text.
[0059] (4) Construct the pairing of tongue image features and tongue diagnosis text features, and the pairing of pulse signal features and pulse diagnosis text features, and use a contrast learning method for training. In the training process, the positive sample pair refers to the image and the diagnosis text of the same patient, and the negative sample pair refers to the image and the diagnosis text of different patients. By minimizing the feature distance between positive sample pairs and maximizing the feature distance between negative sample pairs, it is ensured that the tongue, pulse, and diagnosis text are more consistent in semantic space. The loss function is shown in equation (2):
[0060] (2)
[0061] Wherein represents the total loss of contrast learning, represents the contrast learning loss of the tongue feature extractor and the text information feature extractor, a contrastive learning loss of the tongue appearance feature extractor and the text information feature extractor, and as shown in formulas (3) and (4) respectively:
[0062] (3)
[0063] (4)
[0064] wherein a contrastive learning loss of the tongue appearance feature extractor and the text information feature extractor, N represents the number of tongue appearance diagnosis text feature vectors, max represents a maximum function, , represents a Euclidean distance, a contrastive learning loss of the tongue appearance feature extractor and the text information feature extractor, M represents the number of pulse appearance diagnosis text feature vectors, i and j represent the indexes of vectors, represents a tongue image feature vector with index i, represents a tongue appearance diagnosis text feature vector with index i, represents a tongue appearance diagnosis text feature vector with index j, represents a pulse signal feature vector with index i, represents a pulse appearance diagnosis text feature vector with index i, represents a pulse appearance diagnosis text feature vector with index j.
[0065] (5) The tongue appearance feature extractor, the pulse appearance feature extractor and the text information feature extractor are pre-trained by using the loss function, and the pre-trained tongue appearance feature extractor, the pulse appearance feature extractor and the text information feature extractor can learn the semantic association between the tongue image, the pulse signal and the traditional Chinese medicine diagnosis text.
[0066] S3, as shown in Figure 4 A multi-modal joint representation diabetes risk prediction model is constructed, which includes a tongue appearance feature extractor, a pulse appearance feature extractor, a text information feature extractor, a retinal image feature extractor, a clinical information encoder, a feature fusion module and a risk classifier. The multi-modal joint representation diabetes risk prediction model is trained to convergence by using tongue image, pulse signal, medical record information of tongue and pulse, retinal fundus image, clinical index information and diabetes risk level label. The steps are as follows:
[0067] (1) The tongue image feature extractor, the pulse feature extractor, the text information feature extractor, the retinal image feature extractor, and the clinical information encoder are used to extract the tongue image, the pulse signal, the tongue and pulse diagnosis text, the retinal fundus image, and the clinical index information feature vectors. The retinal image feature extractor uses ResNet-18 as the backbone network, and the clinical information encoder uses a multi-layer perceptron. The process is shown in equation (5):
[0068] (5)
[0069] wherein represents the tongue image feature vector, represents the pulse signal feature vector, represents the tongue diagnosis text feature vector, represents the pulse diagnosis text feature vector, represents the retinal fundus image feature vector, represents the clinical index feature vector, represents the tongue feature extractor, represents the pulse feature extractor, represents the text information feature extractor, represents the retinal image feature extractor, represents the clinical information encoder, represents the tongue image, represents the pulse signal, represents the tongue and pulse diagnosis text, represents the retinal fundus image, represents the clinical index.
[0070] (2) The tongue image feature vector, the pulse signal feature vector, the tongue diagnosis text feature vector, the pulse diagnosis text feature vector, the retinal fundus image feature vector, and the clinical index feature vector are spliced and input into the risk classifier. The risk classifier uses a multi-layer fully connected neural network. The input layer is the fused feature vector. The middle layer uses a fully connected layer with an activation function for non-linear mapping. The output layer uses Softmax as the activation function to map the input features to the probability distribution of low risk, medium risk, and high risk of diabetes. The diabetes risk prediction result is obtained. The process is shown in equation (6):
[0071] (6)
[0072] represents the risk prediction result, represents the risk classifier, represents the splicing operation, represents the tongue image feature vector, represents a pulse signal feature vector, represents a tongue appearance diagnosis text feature vector, represents a pulse appearance diagnosis text feature vector, represents a retinal fundus image feature vector, represents a clinical index feature vector;
[0073] (3) In the training process, the difference between the risk prediction result and the true label is used to construct a loss function, and the model parameters are continuously updated through back propagation until the model converges.
[0074] S4, using the current tongue image, pulse signal, retinal fundus image, medical record information of tongue and pulse, and clinical index information of the patient as input, using the converged multi-modal joint representation diabetes risk prediction model to predict the patient's 10-year diabetes risk.
[0075] Table 1 is a comparison of the performance of the method of the present application and other methods in predicting the risk of diabetes in a pre-diabetic patient cohort, and Table 2 is a comparison of the performance of the method of the present application in different risk groups in a pre-diabetic patient cohort. The results fully demonstrate the excellent performance of the method of the present application in predicting the risk of diabetes.
[0076] Table 1 Performance comparison of the method proposed in the diabetes risk prediction task
[0077] Method Accuracy AUC AUPR Clinical indicator model 0.612 0.670 0.655 Single modality of tongue appearance 0.625 0.681 0.662 Single modality of pulse appearance 0.633 0.688 0.670 Single modality of retinal fundus image 0.642 0.693 0.676 Multi-modal fusion baseline 0.658 0.710 0.695 Method of the present invention 0.762 0.781 0.765
[0078] Table 2 Comparison of classification performance of the method of the present application in different risk groups in the diabetes risk prediction task
[0079] Risk grouping Accuracy AUC AUPR Low risk group 0.782 0.795 0.780 Medium risk group 0.761 0.782 0.764 High risk group 0.743 0.768 0.751
Claims
1. A multi-modal diabetes risk prediction method in the combination of traditional Chinese and Western medicine, characterized in that, Comprise the following steps: S1, a pre-diabetes patient cohort is constructed, and medical expert's tongue and pulse information, tongue image, pulse signal, retinal fundus image and clinical index information are collected, and according to the 10-year follow-up results, the patients in the cohort are labeled as low-risk diabetes patients, medium-risk diabetes patients and high-risk diabetes patients; S2, the medical expert's tongue and pulse diagnosis information and the collected tongue image and pulse signal are used to pre-train the tongue feature extractor, the pulse feature extractor and the text information feature extractor through multi-modal contrast learning; S3, a multi-modal joint representation diabetes risk prediction model is constructed, which comprises a tongue feature extractor, a pulse feature extractor, a text information feature extractor, a retinal image feature extractor, a clinical information encoder, a feature fusion module and a risk classifier, and the multi-modal joint representation diabetes risk prediction model is trained to convergence by using the tongue image, the pulse signal, the tongue and pulse information, the retinal fundus image, the clinical index information and the diabetes risk level label; S4, the current tongue image, pulse signal, retinal fundus image, tongue and pulse information, and clinical index information of the patient are inputted, and the converged multi-modal joint representation diabetes risk prediction model is used to predict the patient's diabetes risk in 10 years. 2.The method of claim 1, wherein, The step S1 of constructing a pre-diabetes patient cohort includes the following steps: S11, at the baseline, patients with at least two high-risk factors for diabetes are listed as pre-diabetes patients and enter the pre-diabetes patient cohort; S12, a TCM expert records the medical information of the patient's tongue diagnosis and pulse diagnosis; S13, the clinical index information of the patient including age, gender, weight, body mass index BMI and blood pressure is collected, and the retinal fundus image, tongue image and pulse signal of the patient are collected by using a collection device; S14, the patients in the cohort are followed up for 10 years, and according to the follow-up results, the pre-diabetes patients are divided into a pre-diabetes patient reversed to a normal population group, a pre-diabetes patient unchanged during follow-up group, and a pre-diabetes patient progressed to a diabetes population group, and are respectively divided into a low-risk diabetes population, a medium-risk diabetes population and a high-risk diabetes population. 3.The method of claim 1, wherein, The step S2 of pre-training the tongue feature extractor, the pulse feature extractor and the text information feature extractor through multi-modal contrast learning by using the medical expert's tongue and pulse diagnosis information and the collected tongue image and pulse signal comprises the following steps: S21, the medical expert's tongue and pulse diagnosis information is obtained, and the tongue and pulse diagnosis text is extracted; S22, the tongue and pulse diagnosis text is one-to-one corresponding to the tongue image and pulse signal of the patient; S23, respectively construct the tongue feature extractor, pulse feature extractor and text information feature extractor, use the tongue feature extractor, pulse feature extractor and text information feature extractor to extract tongue image feature vector, pulse signal feature vector, tongue diagnosis text feature vector and pulse diagnosis text feature vector respectively, the process is shown in formula (1): (1) wherein represents a tongue appearance image feature vector, represents a pulse appearance signal feature vector, represents a tongue appearance diagnosis text feature vector, represents a pulse appearance diagnosis text feature vector, represents a tongue appearance feature extractor, represents a pulse appearance feature extractor, represents a text information feature extractor, represents a tongue appearance image, represents a pulse appearance signal, represents a tongue appearance and pulse appearance diagnosis text; S24, respectively construct tongue image feature vector and tongue diagnosis text feature vector pair, pulse signal feature vector and pulse diagnosis text feature vector pair, and carry out comparative learning training, the loss function is shown in formula (2): (2) wherein denotes the total loss of the contrastive learning, denotes the contrastive learning loss of the tongue feature extractor and the text information feature extractor, denotes the contrastive learning loss of the pulse feature extractor and the text information feature extractor, and are respectively shown in formula (3) and formula (4): (3) (4) wherein denotes the contrastive learning loss of the tongue appearance feature extractor and the text information feature extractor, N denotes the number of tongue appearance diagnosis text feature vectors, and max denotes the maximum function, is a constant, denotes the Euclidean distance, denotes the contrastive learning loss of the pulse appearance feature extractor and the text information feature extractor, M denotes the number of pulse appearance diagnosis text feature vectors, i and j denotes the index of a vector, denotes the tongue appearance image feature vector with index i denotes the tongue appearance diagnosis text feature vector with index denotes the tongue appearance diagnosis text feature vector with index i denotes the pulse appearance signal feature vector with index denotes the pulse appearance diagnosis text feature vector with index j denotes the pulse appearance diagnosis text feature vector with index denotes the pulse appearance diagnosis text feature vector with index i denotes the pulse appearance diagnosis text feature vector with index denotes the pulse appearance diagnosis text feature vector with index i denotes the pulse appearance diagnosis text feature vector with index denotes the pulse appearance diagnosis text feature vector with index j S25, use the above loss to pretrain the tongue feature extractor, pulse feature extractor and text information feature extractor. 4.The method of claim 1, wherein the method is a combination of traditional Chinese and Western medicine. In step S3, a multi-modal joint representation diabetes risk prediction model is constructed, which includes tongue feature extractor, pulse feature extractor, text information feature extractor, retinal image feature extractor, clinical information encoder, feature fusion module and risk classifier, and the multi-modal joint representation diabetes risk prediction model is trained to convergence by using tongue image, pulse signal, tongue and pulse medical record information, retinal fundus image, clinical index information and diabetes risk level label, which includes the following steps: S31, use tongue feature extractor, pulse feature extractor, text information feature extractor, retinal image feature extractor and clinical information encoder to extract tongue image, pulse signal, tongue and pulse diagnosis text, retinal fundus image and clinical index information feature vector, the process is shown in formula (5): (5) wherein denotes a tongue image feature vector, denotes a pulse signal feature vector, denotes a tongue diagnosis text feature vector, denotes a pulse diagnosis text feature vector, denotes a retinal fundus image feature vector, denotes a clinical indicator feature vector, denotes a tongue feature extractor, denotes a pulse feature extractor, denotes a text information feature extractor, denotes a retinal image feature extractor, denotes a clinical information encoder, denotes a tongue image, denotes a pulse signal, denotes a tongue and pulse diagnosis text, denotes a retinal fundus image, denotes a clinical indicator; S32, fuse tongue image feature vector, pulse signal feature vector, tongue diagnosis text feature vector, pulse diagnosis text feature vector, retinal fundus image feature vector and clinical index feature vector, and input into risk classifier to get risk prediction result, the process is shown in formula (6): (6) represents a risk prediction result, represents a risk classifier, represents a feature fusion module, represents a tongue appearance image feature vector, represents a pulse appearance signal feature vector, represents a tongue appearance diagnosis text feature vector, represents a pulse appearance diagnosis text feature vector, represents a retinal fundus image feature vector, represents a clinical indicator feature vector; S33, use risk prediction result to calculate loss function, update model parameters by back propagation, and train model to convergence.