Gastrointestinal electrical signal-based diabetes and hypertension prediction system and construction method

By constructing a diabetes and hypertension prediction system based on gastrointestinal electrical signals, and utilizing machine learning algorithms and multiple gastrointestinal electrical signal indicators, the system solves the problem of early prediction of diabetes and hypertension, and enables early identification and risk prediction of individuals with potential diseases. It is suitable for community health screening.

WO2026000747A1PCT designated stage Publication Date: 2026-01-02WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/128780
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-26
Filing Date
2024-10-31
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing technologies struggle to predict diabetes and hypertension early, especially when early symptoms are mild and go unnoticed. Furthermore, current diagnostic techniques primarily target patients with existing symptoms and cannot keep pace with the trend of younger age of onset.

Method used

By collecting gastrointestinal electrical signal data, machine learning algorithms are used to construct a prediction model for the risk of developing diabetes and hypertension. By combining multiple gastrointestinal electrical signal indicators such as the percentage of postprandial coupling at the lesser curvature lead and the percentage of postprandial electrical rhythm disorder at the antral lead, a prediction system is established to identify subjects who may develop diabetes and hypertension.

Benefits of technology

It enables early identification of individuals who may develop diabetes and hypertension, improving the prospectivity and accuracy of diagnosis. It is applicable to risk prediction and screening in a wide range of populations, especially in rural areas or communities, and has good preventive and intervention effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128780_02012026_PF_FP_ABST
    Figure CN2024128780_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the field of disease prediction, and in particular, to a gastrointestinal electrical signal-based diabetes and hypertension prediction system and a construction method. The prediction system provided by the present invention comprises: a database, used for storing gastrointestinal electrical signal data, comprising the postprandial coupling percentage at the lesser curvature lead, the postprandial electrical dysrhythmia percentage at the gastric antrum lead, the preprandial dominant frequency at the lesser curvature lead, the postprandial average waveform frequency at the ascending colon lead, the postprandial normal slow wave percentage at the ascending colon lead, the postprandial coupling percentage at the descending colon lead, the postprandial average waveform frequency at the rectal lead, the preprandial average waveform frequency at the transverse colon lead, and the preprandial lead time difference at the rectal lead; and a prediction module, used for predicting a probability of the occurrence of diabetes and hypertension in a subject. Further provided in the present invention is a construction method for the prediction system. The present invention is adapted to performing diabetes and hypertension risk prediction and screening work in rural or community populations.
Need to check novelty before this filing date? Find Prior Art

Description

Prediction system and construction method of diabetes and hypertension based on gastrointestinal electrical signals

[0001] Priority application

[0002] This application claims priority to Chinese invention patent application CN 2024108342994 "Prediction system and construction method of diabetes and hypertension based on gastrointestinal electrical signals" filed on June 26, 2024, which is incorporated by reference in its entirety. TECHNICAL FIELD

[0003] The present application relates to the field of disease prediction, in particular to a prediction system and construction method of diabetes and hypertension based on gastrointestinal electrical signals. BACKGROUND

[0004] [Corrected according to Rule 26 14.11.2024] Diabetes and hypertension are two common chronic diseases that pose a significant challenge to health and healthcare systems worldwide. According to reports, the prevalence of diabetes in China in 2018 was about 12.4%, and the standardized prevalence of hypertension among adults aged 18-69 was 24.7%. Diabetes and hypertension can cause serious complications, including cardiovascular and cerebrovascular diseases, neurological dysfunction, kidney disease, and eye disease, which severely affect the quality of life of patients and increase the risk of death. Both diseases are influenced by genetics, lifestyle, environment, and other health factors, and obesity, poor eating habits, lack of exercise, high blood lipids, and smoking are common risk factors for diabetes and hypertension. According to reports, over the 10 years from 2011 to 2021, the number of diabetes patients in China increased from 90 million to 140 million, an increase of 56%; while the awareness rate of diabetes in China was only 36.7% (the proportion of people with diabetes who know they have the disease before testing). The incidence of diabetes in Chinese adults is about 11%, and nearly 36% of adults are in the pre-diabetic stage, and if no intervention is made, 1-3 out of every 4 people in the pre-diabetic stage will progress to diabetes within 10 years. For hypertension, the prevalence of hypertension among Chinese adults is 25.2%, with 270 million patients, and 2 million people die from hypertension every year. Among adult hypertension patients, more than 3 / 4 are young and middle-aged, and the incidence of the disease is growing more rapidly than in older people. In addition, the proportion of people with normal high blood pressure is increasing, which is an important source of the rapid increase in the number of people with hypertension in China and the main reason for the continued rise in the prevalence of the disease.

[0005] Early detection, effective management and appropriate treatment are crucial to prevent and reduce the negative impact of diabetes and hypertension. However, there is a certain degree of delay in the current diagnosis of diabetes and hypertension. Since diabetes and hypertension may exhibit mild symptoms in the early stage, such as mild headache, mild thirst or mild visual blurring, these symptoms are often considered as problems not worth attention, and patients may take self-treatment or ignore these symptoms until the disease progresses to a more serious stage before seeking medical help. The existing technology focuses on the object of study is limited to diabetic / hypertensive patients who have already developed serious complications, for example, the literature "Gastroparesis and functional dyspepsia patients with gastric electrical detection analysis" (Guanqi Liu. Gastroparesis and functional dyspepsia patients with gastric electrical detection analysis [J]. Chinese clinical research, 2021, (12): 1662-1664, 1669.) reported that the postprandial / preprandial power ratio and the percentage of gastric hypomotility of diabetic gastroparesis patients were significantly higher than those of diabetic functional dyspepsia patients, and the preprandial and postprandial waveform response area of diabetic gastroparesis patients was significantly larger than that of diabetic functional dyspepsia patients. It is very obvious that the existing technology is only suitable for distinguishing diabetic / hypertensive patients who have already developed serious complications (such as distinguishing diabetic gastroparesis patients from diabetic functional dyspepsia patients), and it is difficult to achieve early prediction of diabetes / hypertension, and it is even more difficult to adapt to the trend of younger age of diabetes and hypertension.

[0006] SUMMARY

[0007] In one aspect, the present application provides a diabetes and hypertension prediction system based on gastrointestinal electrical signals, characterized in that it comprises:

[0008] a database for storing data, the type of data being gastrointestinal electrical signal data, the gastrointestinal electrical signal data including gastric lead data and intestinal lead data, the gastric lead data including postprandial coupling percentage at the lesser curvature lead, postprandial electrical rhythm disorder percentage at the antral lead, preprandial dominant frequency at the lesser curvature lead, the intestinal lead data including postprandial waveform average frequency at the ascending colon lead, postprandial normal slow wave percentage at the ascending colon lead, postprandial coupling percentage at the descending colon lead, postprandial waveform average frequency at the rectal lead, preprandial waveform average frequency at the transverse colon lead and preprandial lead time difference at the rectal lead; the data including sample data from a sample population and subject data from a subject;

[0009] a data acquisition module for acquiring the data and storing the data in the database;

[0010] a model training module, the model training module uses a machine learning algorithm to train and learn the sample data, thereby determining a diabetes and hypertension risk prediction model;

[0011] The prediction module acquires the subject data through the data acquisition module and calls the diabetes and hypertension risk prediction model to analyze the subject data in order to predict the probability of the subject developing diabetes and hypertension.

[0012] In some embodiments, the sample data is divided into a training set and a validation set in a 7:3 ratio.

[0013] In some embodiments, the gastrointestinal electrical signal data are obtained simultaneously by leads located in the gastric body, antrum, lesser curvature, greater curvature, ascending colon, transverse colon, descending colon, and rectum.

[0014] In some embodiments, the system further includes a validation module for evaluating the accuracy of the diabetes and hypertension risk prediction model using the validation set, wherein the evaluation metrics include discrimination.

[0015] In some embodiments, the subject does not suffer from gastrointestinal diseases and / or gastrointestinal discomfort.

[0016] In some embodiments, the diabetes and hypertension risk prediction model is used to predict diabetes and hypertension as a whole.

[0017] In some embodiments, the predictive variables of the diabetes and hypertension risk prediction model include the percentage of postprandial coupling in the lesser curvature lead, the percentage of postprandial electrical rhythm disturbance in the antral gastric lead, the preprandial dominant frequency in the lesser curvature lead, the average postprandial waveform frequency in the ascending colon lead, the percentage of normal slow waves in the ascending colon lead, the percentage of postprandial coupling in the descending colon lead, the average postprandial waveform frequency in the rectal lead, the average preprandial waveform frequency in the transverse colon lead, and the preprandial lead time difference in the rectal lead.

[0018] On the other hand, the present invention provides a method for constructing a prediction system for diabetes and hypertension based on gastrointestinal electrical signals, characterized in that the prediction system for diabetes and hypertension based on gastrointestinal electrical signals includes a risk prediction model for the occurrence of diabetes and hypertension, and the method includes the following steps:

[0019] S1 acquires sample data from the sample population, wherein the type of sample data is gastrointestinal electrical signal data;

[0020] S2 pre-trains the sample data to select predictive variables. The selected predictive variables include the percentage of postprandial coupling in the lesser curvature lead, the percentage of postprandial electrical rhythm disorder in the antral gastric lead, the preprandial main frequency in the lesser curvature lead, the average postprandial waveform frequency in the ascending colon lead, the percentage of normal slow waves in the ascending colon lead, the percentage of postprandial coupling in the descending colon lead, the average postprandial waveform frequency in the rectal lead, the average preprandial waveform frequency in the transverse colon lead, and the preprandial lead time difference in the rectal lead.

[0021] S3 uses machine learning algorithms to train and learn the sample data based on the selected predictive variables to establish the prediction model for the risk of diabetes and hypertension, and then constructs the prediction system for diabetes and hypertension based on gastrointestinal electrical signals.

[0022] In some embodiments, the sample data is divided into a training set and a validation set in a 7:3 ratio.

[0023] In some embodiments, the pre-training includes a first round of variable selection and a second round of variable selection; the first round of variable selection includes LASSO regression analysis, and the second round of variable selection includes logistic regression analysis and stepwise regression analysis.

[0024] In some embodiments, the method further includes S4 using the validation set to evaluate the accuracy of the diabetes and hypertension risk prediction model, wherein the evaluation metrics include discrimination.

[0025] In some embodiments, the pre-training includes a ridge regression or random forest model.

[0026] Compared with the prior art, the beneficial effects of the present invention include at least the following aspects:

[0027] This invention provides a predictive system and method for diabetes and hypertension based on gastrointestinal electrical signals. Currently, clinical diagnosis of diabetes and hypertension is based on the subject's existing clinical symptoms. However, current diagnostic methods for diabetes and hypertension suffer from a degree of delay; for example, early symptoms of diabetes and hypertension are often difficult to detect or are overlooked, making targeted diagnosis challenging. Furthermore, the age of onset for diabetes and hypertension is trending younger. In other words, a large number of individuals with the potential to develop diabetes and / or hypertension (i.e., in the early stages of diabetes and / or hypertension) are currently difficult to identify. To more accurately capture changes in the gastrointestinal microenvironment in subjects potentially developing diabetes and / or hypertension, this invention creatively incorporates independent pre- and post-meal indicators at each lead. Ultimately, it retains nine predictive variables: the percentage of post-meal coupling at the lesser curvature lead, the percentage of post-meal electrical rhythm disturbances at the antral gastric lead, the pre-meal dominant frequency at the lesser curvature lead, the average post-meal waveform frequency at the ascending colon lead, the percentage of normal slow waves at the ascending colon lead, the percentage of post-meal coupling at the descending colon lead, the average post-meal waveform frequency at the rectal lead, the average pre-meal waveform frequency at the transverse colon lead, and the pre-meal lead time difference at the rectal lead. Without incorporating any clinical indicators, this invention aims to predict diabetes and hypertension as a whole, thereby distinguishing subjects without gastrointestinal diseases and / or discomfort who are potentially developing diabetes and / or hypertension from healthy individuals, demonstrating promising potential for auxiliary diagnosis.

[0028] Diabetes and hypertension are two common chronic diseases, with early symptoms often mild and easily overlooked, making them susceptible to confusion with other illnesses. This invention can effectively identify subjects potentially developing diabetes and / or hypertension, predicting and alerting them to their risk of developing these conditions. It represents a more advanced and comprehensive approach to existing diagnostic methods for diabetes and hypertension, facilitating more precise intervention and prevention. It is particularly suitable for conducting risk prediction and screening for diabetes and hypertension in large rural or community populations, and shows promising application prospects for the early prevention of these diseases. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. The elements or parts in the drawings are not necessarily drawn to scale. Obviously, the drawings described below are some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0030] Figure 1 is a schematic diagram of the placement location of the leads in an embodiment of the present invention;

[0031] Figure 2 is a schematic diagram of the nonograph of the optimal logistic regression model in Embodiment 1 of the present invention;

[0032] Figure 3 is a schematic diagram of the ROC curve of the prediction model of the test set in Embodiment 1 of the present invention;

[0033] Figure 4 is a schematic diagram of the ROC curve of the prediction model of the validation set in Embodiment 1 of the present invention;

[0034] Figure 5 is a schematic diagram of the architecture of the prediction system provided in an embodiment of the present invention;

[0035] Figure 6 is a schematic diagram of the prediction system provided in an embodiment of the present invention.

[0036] 100 is the prediction system, 102 is the data acquisition module, 104 is the model building module, 106 is the database, 108 is the model training module, 110 is the validation module, 112 is the prediction module, 202 is the first terminal, 204 is the second terminal, 206 is the network, 601 is the stomach body, 602 is the lesser curvature, 603 is the greater curvature, 604 is the antrum, 605 is the ascending colon, 606 is the transverse colon, 607 is the descending colon, and 608 is the rectum. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0038] In this document, suffixes such as "module," "part," or "unit" used to denote elements are used only for the purpose of illustrative purposes and have no specific meaning in themselves. Therefore, "module," "part," or "unit" may be used interchangeably.

[0039] In this document, the terms "upper," "lower," "inner," "outer," "front," "rear," "one end," and "the other end," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the present invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0040] In this document, "and / or" includes any and all combinations of one or more of the listed related items.

[0041] In this article, "multiple" means two or more, that is, it includes two, three, four, five, etc.

[0042] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0043] As used in this specification, the term "about" typically means + / - 5% of the value, more typically + / - 4%, more typically + / - 3%, more typically + / - 2%, even more typically + / - 1%, even more typically + / - 0.5% of the value.

[0044] In this specification, certain embodiments may be disclosed in a range-bound format. It should be understood that this "range-bound" description is merely for convenience and brevity and should not be construed as a rigid limitation on the disclosed range. Therefore, the description of a range should be considered as having specifically disclosed all possible subranges and the individual numerical values ​​within those ranges. For example, a description of the range 1-6 should be considered as having specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., and the individual numbers within those ranges, such as 1, 2, 3, 4, 5, and 6. This rule applies regardless of the breadth of the range.

[0045] Example 1: Gastrointestinal electrical signals and a predictive model for the risk of diabetes and hypertension

[0046] 1.1 Method

[0047] Subjects: Patients with diabetes and hypertension and healthy subjects were recruited from the Department of Neurology, West China Hospital of Sichuan University. They were assessed using the Mini-Mental State Examination (MMSE) and underwent electrogastrointestinal examination.

[0048] All participants voluntarily participated in the study and signed informed consent forms. All participants underwent standard diagnostic tests for diabetes and hypertension, with diagnostic criteria based on the 2022 American Diabetes Association (ADA) Standards of Diabetes Management and the 2017 American College of Cardiology (AHA) and American College of Hypertension (ACC) Joint Guidelines for Hypertension. All participants voluntarily underwent a complete pre- and post-meal electrogastroelectrogastrogram (EGEG) examination. All participants were advised to abstain from alcohol and follow a light diet for 3 days to complete an accurate EGEG test.

[0049] Diabetes diagnosis can be confirmed based on any of the following criteria: 1) Fasting blood glucose ≥126 mg / dL (7.0 mmol / L); 2) Random blood glucose ≥200 mg / dL (11.1 mmol / L), accompanied by typical symptoms of diabetes, such as polyuria, polydipsia, weight loss, and fatigue; 3) 2-hour blood glucose in an oral glucose tolerance test (OGTT) ≥200 mg / dL (11.1 mmol / L).

[0050] Diagnosis of hypertension: systolic blood pressure ≥130 mmHg or diastolic blood pressure ≥80 mmHg, and this reading must be obtained in at least two separate measurements.

[0051] Exclusion criteria include: 1) Those diagnosed with gastrointestinal diseases such as gastritis, inflammatory bowel disease, or irritable bowel syndrome within the past six months, or those with chronic diarrhea, constipation, or other gastrointestinal discomfort; 2) Those with severe heart, liver, kidney, or other major organ dysfunction, or those with metabolic diseases such as diabetes or hyperthyroidism; 3) Those with a Mini-Mental State Examination (MMSE) score of less than 24 or those who cannot effectively cooperate in completing the questionnaire and electrogastrointestinal examination; 4) Subjects who have taken any medication within one week prior to the examination.

[0052] This embodiment collected basic information of all subjects, including age and gender, diagnostic results of diabetes and hypertension, and detailed EGEG examination results.

[0053] EGEG examination: Gastrointestinal electromyographic activity signals were measured and acquired using an 8-channel gastrointestinal electrograph (XDJ-S8, Hefei Kaili Co., Hefei, China). All subjects were instructed to avoid alcohol and spicy or irritating foods for 3 days and to fast for at least 6 hours before the examination. The subjects were placed in a supine position during measurement. Four gastric electrodes (leads placed at the gastric body 601, lesser curvature 602, greater curvature 603, and antrum 604) and four intestinal electrodes (leads placed at the ascending colon 605, transverse colon 606, descending colon 607, and rectum 608) were placed on the abdominal skin (Hanjie Co. Ltd., Shanghai, China) (Figure 1). The reference electrode was located on the medial side of the right wrist, and the ground electrode was located on the medial malleolus of the right leg. During the examination, subjects were instructed to avoid any movement or talking. After 6 minutes of pre-meal EGEG recording, a postprandial functional load test was performed. After consuming approximately 200 kcal of standard food, postprandial gastrointestinal electrical signals were recorded for another 6 minutes. The placement of the leads is shown in Figure 1: Stomach body 601: 3-5 cm to the left and 1 cm above the midpoint of the line connecting the xiphoid process and the umbilicus; Stomach antrum 604: 2-4 cm to the right of the midpoint of the line connecting the xiphoid process and the umbilicus; Lesser curvature 602: 1 / 2 of the distance above the midpoint of the line connecting the xiphoid process and the umbilicus; Greater curvature 603: 1 / 2 of the distance below the midpoint of the line connecting the xiphoid process and the umbilicus. Ascending colon 605: 2-4 cm to the right of the umbilicus; Transverse colon 606: 1 cm below the umbilicus; Descending colon 607: 2-4 cm to the left of the umbilicus; Rectum 608: Below the coccyx on the back.

[0054] EGEG data processing: The EGEG sampling frequency is 1Hz and the filtering frequency is 0.008Hz-0.1Hz to filter out background noise, including heartbeat. After artifact detection, the original EGEG potential data is calculated by the software of the examination instrument, and the following parameters of the above 8 leads are derived by the software through spectrum analysis: (1) average waveform amplitude (AA); (2) average waveform frequency (AF); (3) percentage of gastrointestinal electrical rhythm disorder (RD); (4) waveform response area (RA); (5) lead time difference (TD); (6) dominant frequency (MF); (7) dominant power ratio (MPR); (8) percentage of normal slow wave (PNSW); (9) coupling percentage (CP).

[0055] Based on the above gastrointestinal electrophysiological signals, the indicators of the gastric and intestinal leads were screened and analyzed according to the lead placement location and acquisition time (before or after meals) to comprehensively reflect the electrophysiological rhythm performance of the stomach and intestines, totaling 168 indicators.

[0056] Construction of the Predictive Model: Considering the complexity of gastrointestinal electrophysiological signals, this invention employs a two-stage analysis strategy combining LASSO-based variable screening with stepwise optimal logistic regression to construct the risk prediction model. First, to obtain a subset of predictive factors, LASSO regression analysis, a regularization algorithm, is used to perform a first round of variable screening on pre- and post-meal results at various leads in the stomach and intestines. Furthermore, LASSO regression analysis undergoes 10-fold cross-validation to centralize and normalize the included variables, selecting "lambda.min" as the optimal performance. Subjects are randomly divided into training and validation sets in a 7:3 ratio. Additionally, to balance the potential impact of subjects' basic characteristics on the prediction of diabetes and hypertension risks, subjects' age and gender are included in all subsequent analyses using an inverse probability weighting method. Then, stepwise multivariate logistic regression analysis was used to merge the gastric and bowel electrophysiological results selected separately by the LASSO regression model. Further feature selection and model optimization were performed using stepwise multivariate logistic regression, and the AIC (Akaike Information Criterion) minimization principle was used for model optimization in a second round of variable selection. The statistically significant predictors (in this invention, "predictor" and "predictor variable" have the same meaning) were used to build the predictive model. The final model used odds ratio (OR) and 95% confidence interval (CI) to define the contribution of the predictors and generated a nomogram predictive model for the risk of diabetes and hypertension. It should be understood that other suitable algorithms known in the art can also be used, such as random forest methods, other regularization methods (e.g., ridge regression), neural networks, etc.

[0057] In addition, the area under the ROC curve was used to identify the quality of the nomograms for the risk of diabetes and hypertension by using data from the training and validation sets, in order to distinguish between true positives and false positives (i.e., discrimination). All analyses were performed using the R 4.1.3 packages glmnet and rms, with the significance level set to two-tailed α < 0.1.

[0058] 1.2 Results

[0059] Subject data information:

[0060] A total of 2,672 participants completed all relevant examinations, including 809 men and 1,863 women. Among them, 802 participants (297 men and 505 women) were diagnosed with chronic diseases (i.e., hypertension or diabetes). Simultaneous signal acquisition using multiple leads placed at multiple locations allows for better capture of the overall motility patterns of the stomach and intestines, thus more effectively obtaining signals that reflect the true overall state of the stomach and intestines.

[0061] This invention compared each variable among three groups: the general population, healthy individuals (non-chronic disease subjects), and subjects with chronic diseases (Table 1). Statistically significant differences were found in age and sex between subjects with and without chronic diseases (P < 0.001). Therefore, an inverse probability weighting (IPTW) approach was used to incorporate age and sex as the two basic characteristics. At the 0.05 level, differences were observed between groups in the percentage of postprandial gastric pathway electrical rhythm disturbances in the gastric body, lesser curvature, and antral leads; the percentage of postprandial intestinal pathway coupling in the lesser curvature lead; the dominant frequency of the postprandial intestinal pathway in the transverse colon lead; and the dominant frequency and coupling percentage of the preprandial gastric pathway in the lesser curvature lead.

[0062] Table 1

[0063] Screening of independent risk factors:

[0064] LASSO regression was used to independently screen variables in the gastric and intestinal pathway index datasets, retaining 6 and 12 predictor variables respectively (see Table 2). The retained predictor variables from the gastric and intestinal pathway index datasets were merged and subjected to stepwise optimal logistic regression to construct a risk assessment model based on the original EGEG index signals. After stepwise regression, 9 predictor variables were retained (see Table 3). Among these, the percentage of postprandial electrical rhythm disturbances in the antral lead, the preprandial dominant frequency in the lesser curvature lead, and the average postprandial waveform frequency in the rectal lead were positively correlated with the risk of diabetes and hypertension, while the average postprandial waveform frequency in the ascending colon lead was negatively correlated with the risk of diabetes and hypertension. A nomogram (Figure 2) was created based on the selected raw gastrointestinal electrical indices, which is helpful for personalized clinical assessment. *P<0.05 in Tables 2 and 3.

[0065] Table 2

[0066] Table 3

[0067] Using a non-zero coefficient feature variable screening method based on LASSO regression, among 168 relevant feature variables (which can also be referred to as independent variables (IVs) in this invention), the technical solution of this invention ultimately selected and retained 18 feature variables as potential predictors for the risk model of diabetes and hypertension to predict the response variable (which can also be referred to as the outcome variable (DV) in this invention) (i.e. the risk of diabetes and hypertension).

[0068] LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis is a method for shrinking and selecting feature variables in linear regression models. To obtain a subset of predictors, LASSO regression analysis imposes constraints on the model parameters, shrinking the regression coefficients of some feature variables towards zero, thereby minimizing the prediction error of the response variable. After the shrinkage process, feature variables with regression coefficients equal to zero are excluded from the model, while feature variables with non-zero regression coefficients have the strongest correlation with the response variable. In other words, the first round of feature variable screening primarily excludes feature variables whose regression coefficients are easily reduced to zero from 168 relevant feature variables, retaining 18 feature variables as predictors for the predictive model.

[0069] Furthermore, to more accurately evaluate the performance of the predictive model based on 18 predictor variables, based on the log-likelihood function (-2log-likelihood) and the type parameter of the binary dependent variable (which can be understood as whether the variable is "yes" or "no"—i.e., the target parameter to be minimized when selecting the model for cross-validation), the technical solution of this invention employs LASSO regression analysis to run 10-fold cross-validation, centralizing and normalizing the 168 included feature variables. The technical solution of this invention uses "Lambda.min" to construct a predictive model with optimal performance and highest accuracy.

[0070] Development of predictive models:

[0071] The technical solution of this invention introduces characteristic variables selected from the LASSO regression model and uses stepwise multivariate logistic regression analysis to establish a predictive model. Then, the selected characteristic variables are introduced and their statistical significance levels are analyzed. Statistically significant characteristic variables are used as predictor variables / predictor factors to establish predictive models for the risk of diabetes and hypertension.

[0072] Next, this invention uses a logistic regression model to analyze the 18 predictive variables selected in the first round of variable screening, and performs a stepwise method to select the optimal predictive variables, ultimately retaining 9 predictive variables (each predictive variable is statistically significant at the 0.05 test level). These 9 predictive variables are: the percentage of postprandial coupling at the lesser curvature lead, the percentage of postprandial electrical rhythm disturbance at the antral gastric lead, the preprandial dominant frequency at the lesser curvature lead, the average postprandial waveform frequency at the ascending colon lead, the percentage of normal slow waves at the ascending colon lead, the percentage of postprandial coupling at the descending colon lead, the average postprandial waveform frequency at the rectal lead, the average preprandial waveform frequency at the transverse colon lead, and the preprandial lead time difference at the rectal lead (Table 3).

[0073] The technical solution of this invention uses various statistical methods to test the above-mentioned nine characteristic variables, with a focus on analyzing the odds ratio (OR) of these characteristic variables. The OR value is a statistic that quantifies the strength of the association between two events, representing the ratio of the probability of the outcome occurring after exposure (in this invention, the characteristic variable being tested, hereinafter the same) to the probability of the outcome occurring without the same exposure. Specifically, the OR value in this invention can be understood as the strength of the association between diabetes and hypertension (i.e., the response variable) and exposure, representing the multiple by which the exposed person's risk of developing diabetes and hypertension (which can also be understood as disease hazard) is higher than that of the unexposed person. If the OR value of the tested characteristic variable is >1, it indicates that the risk of developing diabetes and hypertension increases due to exposure, and there is a "positive" association between this characteristic variable and diabetes and hypertension; if the OR value of the tested characteristic variable is <1, it indicates that the risk of developing diabetes and hypertension decreases due to exposure, and there is a "negative" association between this characteristic variable and diabetes and hypertension; if the OR value of the tested characteristic variable is =1, it indicates that diabetes and hypertension are not associated with this characteristic variable. The 95% confidence interval (CI) provides an estimate of the accuracy of the OR value obtained through the test. It describes how the true population value may fluctuate within the 95% confidence interval of the OR value obtained through the test, and the smaller the confidence interval, the more accurate and robust the OR value obtained through the test. Table 3 shows the odds ratios (ORs) and their 95% confidence intervals for the nine retained predictor variables, indicating that these predictor variables are all associated with diabetes and hypertension, and therefore were applied to the diabetes and hypertension risk prediction model of this invention.

[0074] Additionally, in Table 3, the regression coefficient β is the partial regression coefficient obtained through logistic regression model analysis, representing the magnitude and direction of the impact of each unit increase on the response variable (partial regression coefficients can be compared after standardization; the regression coefficient β of the tested predictor variable is an estimate). Standard error: the standard error of the regression coefficient β of the tested predictor variable, indicating the accuracy of the regression coefficient (the larger the standard error, the lower the accuracy of the tested predictor variable). Z-value: the z-statistic, which is the regression coefficient β of the tested predictor variable divided by its corresponding standard error, mainly used to determine the p-value of the tested predictor variable. P-value: the p-value corresponding to the z-statistic of the tested predictor variable (the smaller the p-value, the more important the tested predictor variable is to the response variable).

[0075] Based on the above nine predictor variables, this embodiment constructs a predictive model for the risk of developing diabetes and hypertension, and visualizes the constructed predictive model for the risk of developing diabetes and hypertension by drawing a corresponding nomogram, as shown in Figure 2.

[0076] In Figure 2, each variable (e.g., "percentage of postprandial coupling at the lesser curvature lead," "percentage of postprandial electrical rhythm disturbance at the antral gastric lead," etc.) has a scale marked on its corresponding line segment, representing the range of values ​​for that variable; the length of the line segment reflects the contribution of that variable to the outcome events (i.e., the occurrence of diabetes and hypertension). For each variable, a corresponding individual score can be obtained at the "Score" section at the top of Figure 2 for different values. After taking the values, the individual scores corresponding to all variables are added together to obtain the "Total Score." Based on the total score, the probability of developing diabetes and hypertension can be obtained at the "Predicted Probability of Diabetes and Hypertension" section at the bottom of Figure 2. As an example, in Figure 2, if a subject's total score is 170, their risk of developing diabetes and hypertension is approximately 0.3 (30%).

[0077] Validation of the prediction model:

[0078] This invention uses training and validation set data to plot corresponding receiver operating characteristic (ROC) curves to evaluate the sensitivity (also understood as true positive rate) and specificity (also understood as true negative rate) of the constructed predictive model. In Figures 3 and 4, the x-axis represents the "false positive rate," i.e., "1-specificity"; the y-axis represents the "true positive rate," i.e., "sensitivity"; the area under the ROC curve (i.e., the solid line in Figures 3 and 4) (AUC, i.e., the area enclosed by the ROC curve and the coordinate axes) analysis is used to identify the quality of the risk nomogram to distinguish true positives from false positives.

[0079] For the established risk prediction models for diabetes and hypertension, the area under the ROC curve (AUC) of the nomogram was above 0.50 (i.e., greater than the area under the dashed line): 58.36% (95% CI (confidence interval): 54.19%-62.52%) in the training set (Figure 3), and 52.29% (95% CI: 45.74%-58.84%) in the validation set (Figure 4). This indicates that the model constructed in this invention exhibits good robustness and the ability to identify patients with diabetes and hypertension.

[0080] Example 2: Using the above-mentioned diabetes and hypertension risk prediction model, predict the probability of subjects developing diabetes and hypertension.

[0081] Referring to Figure 5, which is a schematic diagram of an optional architecture of the prediction system provided in an embodiment of the present invention. To support an exemplary application, terminals (exemplarily shown as first terminal 202 and second terminal 204) are connected to the prediction system via a network. The network involved in the present invention can be a wide area network (WAN), a local area network (LAN), or a combination of both, using a wireless link to achieve data transmission. The terminals involved in the present invention can be various types of user terminals such as smartphones, tablets, and laptops. The terminals can be used to display an interface for inputting subject data and / or sample data, as well as an interface for displaying the prediction results of the prediction system.

[0082] The following describes an exemplary structure of the prediction system. In some embodiments, as shown in FIG6, the prediction system 100 may include:

[0083] Database 106 is used to store data, the type of which is gastrointestinal electrical signal data, including sample data from the sample population and subject data from the subjects;

[0084] Data acquisition module 102 is used to acquire the data and store the data in the database 106;

[0085] The model training module 108 uses machine learning algorithms to train and learn the sample data, thereby determining a predictive model for the risk of developing diabetes and hypertension.

[0086] The prediction module 112 acquires the subject data through the data acquisition module 102 and calls the determined diabetes and hypertension risk prediction model to analyze the subject data in order to predict the probability of the subject developing diabetes and hypertension.

[0087] The prediction system 100 may further include a validation module 110, which is used to evaluate the accuracy of the determined prediction model for the risk of diabetes and hypertension, and the evaluation index includes discrimination.

[0088] The database 106, the model training module 108, and the verification module 110 can be integrated into a model building module 104.

[0089] As an example, in a prediction system for diabetes and hypertension, the gastrointestinal electrical signal data specifically includes the percentage of postprandial coupling at the lesser curvature lead, the percentage of postprandial electrical rhythm disturbance at the antral lead, the preprandial dominant frequency at the lesser curvature lead, the average postprandial waveform frequency at the ascending colon lead, the percentage of normal slow waves at the ascending colon lead, the percentage of postprandial coupling at the descending colon lead, the average postprandial waveform frequency at the rectal lead, the average preprandial waveform frequency at the transverse colon lead, and the preprandial lead time difference at the rectal lead.

[0090] The following is a specific application scenario of the present invention.

[0091] Community A is conducting a community health checkup activity, collecting gastrointestinal electrocardiogram (ECG) data from a target population (e.g., all residents of Community A), and inputting the subject data through a first terminal 202. Patient B in Community A has irregular eating habits, lacks exercise, and experiences high levels of stress. During the community health checkup activity, Community A transmits the subject data via network 206 to the data acquisition module 102 of the prediction system 100. The data acquisition module 102 acquires the subject data from the subjects and stores it in database 106.

[0092] The prediction module 112 acquires the subject data and calls the established risk prediction model for diabetes and hypertension to analyze the subject data and predict the probability of the subject developing diabetes and hypertension.

[0093] As output, the prediction module 112 can generate a report indicating the risk of developing diabetes and hypertension in the subject, and transmit the prediction results to the first terminal 202 via network 206. Community A can pre-set risk thresholds for diabetes and hypertension (e.g., a probability of 30%). When the predicted risk of developing diabetes and hypertension for patient B exceeds the set risk threshold (e.g., a probability of 50%), Community A should remind patient B or their family and recommend that they visit a recommended or cooperating hospital for a thorough examination to determine if patient B has diabetes and hypertension. When patient B visits the hospital, the hospital will conduct further consultations and examinations to determine if the patient has the indicated diabetes and hypertension. The doctor can transmit patient B's diagnosis results to the prediction system 100 via the second terminal 204. Patient B's data (subject data + diagnosis results) can be used as new sample data to further train the diabetes and hypertension risk prediction model. Of course, patient B's diagnosis results can also be transmitted to the prediction system 100 via the first terminal 202; in other words, the terminals transmitting patient B's diagnosis results can be the same or different.

[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a computer terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0095] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A predictive system for diabetes and hypertension based on gastrointestinal electrical signals, characterized in that, include: A database is used to store data, the type of which is gastrointestinal electrical signal data. This data includes gastric lead data and intestinal lead data. The gastric lead data includes the percentage of postprandial coupling at the lesser curvature lead, the percentage of postprandial electrical rhythm disturbance at the antral lead, and the preprandial dominant frequency at the lesser curvature lead. The intestinal lead data includes the average postprandial waveform frequency at the ascending colon lead, the percentage of normal slow waves at the ascending colon lead, the percentage of postprandial coupling at the descending colon lead, the average postprandial waveform frequency at the rectal lead, the average preprandial waveform frequency at the transverse colon lead, and the preprandial lead time difference at the rectal lead. The data includes sample data from the sample population and subject data from the subjects. The data acquisition module is used to acquire the data and store the data in the database; The model training module uses machine learning algorithms to train and learn the sample data, thereby determining a predictive model for the risk of developing diabetes and hypertension. The prediction module acquires the subject data through the data acquisition module and calls the diabetes and hypertension risk prediction model to analyze the subject data in order to predict the probability of the subject developing diabetes and hypertension.

2. The system as described in claim 1, characterized in that, The sample data was divided into a training set and a validation set in a 7:3 ratio.

3. The system as described in claim 1, characterized in that, The gastrointestinal electrical signal data were simultaneously acquired through leads located in the stomach body, antrum, lesser curvature, greater curvature, ascending colon, transverse colon, descending colon, and rectum.

4. The system as described in claim 2, characterized in that, The system further includes a validation module, which is used to evaluate the accuracy of the diabetes and hypertension risk prediction model using the validation set, wherein the evaluation metric includes discrimination.

5. The system as described in claim 1, characterized in that, The subjects did not suffer from gastrointestinal diseases and / or gastrointestinal discomfort.

6. A method for constructing a predictive system for diabetes and hypertension based on gastrointestinal electrical signals, characterized in that, The diabetes and hypertension prediction system based on gastrointestinal electrical signals includes a risk prediction model for the occurrence of diabetes and hypertension, and the method includes the following steps: S1 acquires sample data from the sample population, wherein the type of sample data is gastrointestinal electrical signal data; S2 pre-trains the sample data to select predictive variables. The selected predictive variables include the percentage of postprandial coupling in the lesser curvature lead, the percentage of postprandial electrical rhythm disorder in the antral gastric lead, the preprandial main frequency in the lesser curvature lead, the average postprandial waveform frequency in the ascending colon lead, the percentage of normal slow waves in the ascending colon lead, the percentage of postprandial coupling in the descending colon lead, the average postprandial waveform frequency in the rectal lead, the average preprandial waveform frequency in the transverse colon lead, and the preprandial lead time difference in the rectal lead. S3 uses machine learning algorithms to train and learn the sample data based on the selected predictive variables to establish the prediction model for the risk of diabetes and hypertension, and then constructs the prediction system for diabetes and hypertension based on gastrointestinal electrical signals.

7. The method as described in claim 6, characterized in that, The sample data was divided into a training set and a validation set in a 7:3 ratio.

8. The method as described in claim 6, characterized in that, The pre-training includes a first round of variable selection and a second round of variable selection; the first round of variable selection includes LASSO regression analysis, and the second round of variable selection includes logistic regression analysis and stepwise regression analysis.

9. The method as described in claim 7, characterized in that, The method further includes S4 using the validation set to evaluate the accuracy of the diabetes and hypertension risk prediction model, wherein the evaluation metrics include discrimination.

10. The method as described in claim 6, characterized in that, The pre-training includes ridge regression or random forest models.

Citation Information

Patent Citations

  • Electrogastric signal generation method and system

    CN115998309A

  • Sleep disorder prediction system based on gastrointestinal electric signals and construction method thereof

    CN116584962A

  • Epilepsy prediction system based on gastrointestinal electric signals and construction method

    CN117898682A

  • Prediction system for diabetes and hypertension based on gastrointestinal electric signals and construction method

    CN118448058A

  • Formulation for prevention or treatment of diabetes

    WO2014137090A1