Gastrointestinal electrical signal-based epilepsy prediction system and construction method therefor
By building an epilepsy prediction system based on gastrointestinal electrical signals, using machine learning algorithms to screen out key variables, and establishing a non-invasive, low-cost epilepsy prediction model, we have solved the difficulties in epilepsy screening in existing technologies and improved the accuracy and applicability of epilepsy screening.
Patent Information
- Application Number
- PCT/CN2024/090603
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-19
- Filing Date
- 2024-04-29
- Publication Date
- 2025-09-25
AI Technical Summary
Existing epilepsy diagnosis methods mainly rely on medical history data and electroencephalogram examinations, which are difficult to be widely used in large-scale screening and are expensive, making epilepsy screening difficult to popularize.
Based on gastrointestinal electrical signal data, a model for predicting the risk of epilepsy was constructed through a machine learning algorithm. The difference indicators of gastrointestinal electrical signals before and after meals were used to screen out seven predictive variables, including the lead time difference between the stomach and intestines before meals and the main power ratio, to establish a non-invasive, low-cost epilepsy prediction system.
It improves the accuracy and versatility of epilepsy screening, is suitable for large-scale screening, reduces the economic and psychological burden on subjects, and is suitable for scenarios such as communities and physical examination centers, providing a basis for auxiliary diagnosis.
Smart Images

Figure CN2024090603_25092025_PF_FP_ABST
Abstract
Description
Epilepsy prediction system based on gastrointestinal electrical signals and its construction method
[0001] Priority application
[0002] This application claims the Chinese invention patent application CN 2024103094837 filed on March 19, 2024
[0003] The priority patent application for “Epilepsy prediction system and construction method based on gastrointestinal electrical signals” is incorporated by reference in its entirety. Technical Field
[0004] The present invention relates to the field of disease prediction, and in particular to an epilepsy prediction system based on gastrointestinal electrical signals and a construction method thereof. Background Art
[0005] Epilepsy is one of the most common paroxysmal disorders of the central nervous system, with a high incidence and more than 70 million patients worldwide.
[0006] Correct diagnosis is a prerequisite for effective treatment. Currently, the diagnosis of epilepsy relies primarily on the patient's medical history (for example, a detailed understanding of the patient's previous medical history, especially seizure history, is required during the clinical interview) as well as examinations such as electroencephalograms (EEGs) and magnetic resonance imaging (MRIs). However, these standard epilepsy diagnostic methods are difficult to apply widely in large-scale epilepsy screening (such as community screening). On the one hand, in large-scale screening, even experienced clinicians cannot determine whether a subject has epilepsy based solely on the subject's verbal description of "epileptic" seizure symptoms, given time constraints and a lack of diagnostic evidence. On the other hand, examinations such as EEGs and MRIs require specific equipment, making them difficult to use in large-scale screening. If subjects are directly asked to go to the hospital for examinations such as EEGs and MRIs, the high cost will increase the financial and psychological burden on the subjects and may also cause them to be reluctant to go to the hospital. For these reasons, large-scale epilepsy screening is difficult to popularize.
[0007] Summary of the Invention
[0008] In a first aspect, the present invention provides an epilepsy prediction system based on gastrointestinal electrical signals, characterized by comprising:
[0009] A database for storing data, wherein the data is gastrointestinal electrical signal data, the gastrointestinal electrical signal data including lead time difference of the stomach before a meal, lead time difference of the intestine before a meal, main power ratio of the intestine before a meal, normal slow wave percentage of the intestine before a meal, average waveform frequency of the stomach after a meal, lead time difference of the intestine after a meal, and main power ratio difference of the stomach before and after a meal; the data includes sample data from a sample population and subject data from a subject;
[0010] A data acquisition module, configured to acquire the data and store the data in the database;
[0011] A model training module, wherein the model training module uses a machine learning algorithm to train the sample data to determine an epilepsy risk prediction model;
[0012] A prediction module, wherein the prediction module obtains the subject data through the data acquisition module and calls the epilepsy risk prediction model to analyze the subject data to predict the probability of epilepsy occurring in the subject.
[0013] In some embodiments, the sample data is divided into a training set and a validation set in a ratio of 7:3.
[0014] In some embodiments, the gastrointestinal electrical signal data is obtained by simultaneously collecting and averaging leads located at the gastric body, gastric antrum, lesser curvature, greater curvature, ascending colon, transverse colon, descending colon, and rectum.
[0015] In some embodiments, the system further includes a validation module, which is used to evaluate the accuracy of the epilepsy risk prediction model using the validation set, and the evaluation indicators include discrimination and / or clinical practicality.
[0016] In some embodiments, the subject is an adult less than 50 years old.
[0017] In a second aspect, the present invention provides a method for constructing an epilepsy prediction system based on gastrointestinal electrical signals, characterized in that the epilepsy prediction system based on gastrointestinal electrical signals includes an epilepsy risk prediction model, and the method comprises the following steps:
[0018] S1 obtains sample data from a sample population, where the sample data is gastrointestinal electrical signal data;
[0019] S2 pre-trains the sample data to screen out predictive variables, wherein the screened predictive variables include lead time difference of the stomach before meal, lead time difference of the intestine before meal, main power ratio of the intestine before meal, normal slow wave percentage of the intestine before meal, average waveform frequency of the stomach after meal, lead time difference of the intestine after meal, and main power ratio difference of the stomach before and after meal;
[0020] S3: Based on the screened prediction variables, the sample data is trained and learned using a machine learning algorithm to establish the epilepsy risk prediction model, and then construct the epilepsy prediction system based on gastrointestinal electrical signals.
[0021] In some embodiments, the sample data is divided into a training set and a validation set in a ratio of 7:3.
[0022] In some embodiments, the pre-training includes a first round of variable screening and a second round of variable screening; the first round of variable screening includes LASSO regression analysis, and the second round of variable screening includes logistic regression analysis and stepwise regression analysis.
[0023] In some embodiments, the method further comprises S4 using the validation set to evaluate the accuracy of the epilepsy risk prediction model, wherein the evaluation index comprises discrimination and / or clinical practicality.
[0024] In some embodiments, the pre-training comprises a ridge regression or random forest model.
[0025] Compared with the prior art, the beneficial effects of the present invention include at least the following aspects:
[0026] The present invention provides an epilepsy prediction system and construction method based on gastrointestinal electrical signals. Based on the collection and calculation of 36 indicators, including the average waveform amplitude, average waveform frequency, percentage of electrical rhythm disturbance, waveform response area, lead time difference, dominant frequency, dominant power ratio, percentage of normal slow waves, and percentage of coupling, the present invention creatively designs and introduces the difference between the average waveform amplitude, average waveform frequency, percentage of electrical rhythm disturbance, waveform response area, lead time difference, dominant frequency, dominant power ratio, percentage of normal slow waves, or percentage of coupling before and after meals as new indicators of gastrointestinal electrical signals (a total of 18 indicators, referred to as difference indicators). Based on this, the present invention pre-trained (feature screening) the 54 characteristic variables in the sample data through LASSO regression analysis, logistic regression analysis and stepwise regression analysis, and finally retained "the lead time difference of the stomach before meal, the lead time difference of the intestine before meal, the main power ratio of the intestine before meal, the normal slow wave percentage of the intestine before meal, the average waveform frequency of the stomach after meal, the lead time difference of the intestine after meal and the difference in the main power ratio of the stomach before and after meal". Based on the above 7 predictive variables, an epilepsy risk prediction model was constructed. The results show that the introduction of the difference index of the stomach (or intestine) before and after meals in the present invention can reflect the difference between epilepsy patients and healthy people to a certain extent, thereby helping to improve the accuracy and versatility of the epilepsy risk prediction model.
[0027] Large-scale screening of epilepsy is of vital importance. The frequency of epileptic seizures in some epilepsy patients is not high (e.g., several times a year) and the symptoms are mild, so it is easy to ignore the symptoms of their epileptic seizures. And another part of patients, who may suffer from non-epileptic paroxysmal diseases (e.g., neurotic seizures, transient ischemic attacks), are often worried that they suffer from epilepsy, and therefore usually have a higher psychological burden. The existing diagnosis of epilepsy mainly relies on the patient's medical history data and examinations such as electroencephalograms and magnetic resonance imaging. However, for reasons such as time, human resources, and limited funds, this standard epilepsy diagnosis method cannot be applied to large-scale screening of epilepsy. The epilepsy risk prediction model based on the epilepsy prediction system and method provided by the present invention only requires the above-mentioned gastrointestinal electrical index data of the subject, is non-invasive, simple in procedure, low in price, and does not require additional examination to predict the risk of epilepsy in the subject, so the subject's acceptance and cooperation are high, which is conducive to the promotion of large-scale screening of epilepsy and provides an auxiliary diagnostic basis for the diagnosis of epilepsy.
[0028] In addition, the epilepsy prediction system provided by the present invention is not limited by the age or educational level of the subjects, nor is it affected by communication and understanding barriers during the medical consultation process or the personal experience of the doctor. Without the need to collect detailed medical history information, it can analyze and predict the subject data relatively objectively and non-invasively. It is particularly suitable for preliminary and large-scale screening of epilepsy in larger populations (such as communities and physical examination centers).
[0029] In summary, the epilepsy prediction system and method provided by the present invention not only help assist clinical evaluation, but also facilitate individualized prediction, and are applicable to a variety of application scenarios (e.g., primary medical institutions, families, hospitals, and physical examination centers) and populations. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the embodiments or the description of the prior art. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale. Obviously, the drawings described below are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without paying any creative work.
[0031] FIG1 is a schematic diagram of lead placement locations in an embodiment of the present invention;
[0032] FIG2 is a summary diagram of the baseline information and gastrointestinal electrical indices of the subjects according to Example 1 of the present invention;
[0033] FIG3 is a diagram showing the results of the second round of variable screening according to the first embodiment of the present invention;
[0034] FIG4 is a schematic diagram of a nomogram of an optimal logistic regression model according to the first embodiment of the present invention;
[0035] FIG5 is a schematic diagram of the ROC curve of the prediction model of the test set according to the first embodiment of the present invention;
[0036] FIG6 is a schematic diagram of the ROC curve of the prediction model of the validation set according to the first embodiment of the present invention;
[0037] FIG7 is a schematic diagram of a decision curve of a test set according to the first embodiment of the present invention;
[0038] FIG8 is a schematic diagram of a decision curve of a validation set according to the first embodiment of the present invention;
[0039] FIG9 is a schematic diagram of the architecture of a prediction system provided by an embodiment of the present invention;
[0040] Figure 10 is a schematic diagram of the modules of a prediction system provided by an embodiment of the present invention. 100 is the prediction system, 102 is the data acquisition module, 104 is the model construction module, 106 is the database, 108 is the model training module, 110 is the verification module, 112 is the prediction module, 202 is the first terminal, 204 is the second terminal, 206 is the network, 601 is the gastric body, 602 is the lesser curvature, 603 is the greater curvature, 604 is the gastric antrum, 605 is the ascending colon, 606 is the transverse colon, 607 is the descending colon, and 608 is the rectum. DETAILED DESCRIPTION
[0041] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0042] Herein, suffixes such as "module," "component," or "unit" used to represent elements are only used to facilitate description of the present invention and have no specific meaning. Therefore, "module," "component," or "unit" may be used interchangeably.
[0043] As used herein, terms such as "upper," "lower," "inner," "outer," "front," "back," "one end," and "the other end" indicate positions or locations based on those shown in the accompanying drawings. These terms are intended solely to facilitate and simplify the description of the present invention and are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0044] As used herein, "and / or" includes any and all combinations of one or more of the associated listed items.
[0045] Herein, "plurality" means two or more than two, ie, it includes two, three, four, five, etc.
[0046] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0047] As used in this specification, the term "about" typically means + / - 5% of the stated value, more typically + / - 4% of the stated value, more typically + / - 3% of the stated value, more typically + / - 2% of the stated value, even more typically + / - 1% of the stated value, and even more typically + / - 0.5% of the stated value.
[0048] In this specification, certain embodiments may be disclosed in a format that is within a certain range. It should be understood that such descriptions of "being within a certain range" are merely for convenience and brevity and should not be interpreted as a rigid limitation on the disclosed range. Therefore, the description of a range should be considered to have specifically disclosed all possible sub-ranges and independent numerical values within this range. For example, the description of a range of 1 to 6 should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as individual numbers within this range, such as 1, 2, 3, 4, 5, and 6. Regardless of the breadth of the range, the above rules apply.
[0049] Example 1: Gastrointestinal electrical signals and epilepsy risk prediction model
[0050] 1.1 Methods
[0051] Participants: Epilepsy patients and healthy subjects were recruited from the Department of Neurology, West China Hospital, Sichuan University, and the Mini-Mental State Examination (MMSE) and electrogastrointestinal examination were performed on them.
[0052] All subjects voluntarily participated in the study and signed informed consent. All subjects were required to be adults under 50 years of age and receive a standardized epilepsy diagnosis based on the International League Against Epilepsy (ILAE) Diagnostic Criteria for Epilepsy (2014 edition). They also voluntarily underwent a complete pre- and postprandial electrogastroenterogram (EGEG) examination. All subjects were instructed to abstain from alcohol and follow a bland diet for 3 days to complete an accurate EGEG test.
[0053] Exclusion criteria include: 1) those diagnosed with gastrointestinal diseases such as gastritis, inflammatory bowel disease, irritable bowel syndrome within six months, or those with chronic diarrhea, constipation and other gastrointestinal discomfort; 2) those with severe dysfunction of major organs such as heart, liver, and kidney, or metabolic diseases such as diabetes and hyperthyroidism; 3) those with a score of less than 24 on the Mini-Cognitive Function Examination (MMSE) or who were unable to effectively cooperate in completing the scale collection and electrogastrointestinal testing; 4) subjects who had taken any medication within 1 week before the examination were excluded.
[0054] In this example, basic information including age, gender, BMI, epilepsy diagnosis results and detailed EGEG examination results were collected from all subjects.
[0055] Gastrointestinal electromyographic (EGG) examination: Gastrointestinal electromyographic (GEM) activity signals were measured and collected using an 8-channel GEMG instrument (XDJ-S8, Hefei Kaili Co., Hefei, China). All subjects were instructed to avoid alcohol and spicy or irritating foods for 3 days and to fast for at least 6 hours before the examination. Measurements were performed in the supine position. Four gastric electrodes (leads placed at the gastric body 601, lesser curvature 602, greater curvature 603, and gastric antrum 604) and four intestinal electrodes (leads placed at the ascending colon 605, transverse colon 606, descending colon 607, and rectum 608) were placed on the abdominal skin (Hanjie Co. Ltd., Shanghai, China) (Figure 1). The reference electrode was located on the inner side of the right wrist, and the ground electrode was located at the medial malleolus of the right leg. During the examination, subjects were instructed to avoid any movement or talking. After a 6-minute pre-meal EGG recording, a meal stress test was performed. After ingesting a standard meal of approximately 200 kcal, 6-minute post-meal EGG signals were recorded. The lead placement locations are shown in Figure 1: Gastric body 601: 3 to 5 cm to the left and 1 cm upward from the midpoint of the line connecting the xiphoid process and the umbilicus; Gastric antrum 604: 2 to 4 cm to the right of the midpoint of the line connecting the xiphoid process and the umbilicus; Lesser curvature 602: 1 / 2 point upward from the midpoint of the line connecting the xiphoid process and the umbilicus; Greater curvature 603: 1 / 2 point downward from the midpoint of the line connecting the xiphoid process and the umbilicus. Ascending colon 605: 2 to 4 cm to the right, level with the umbilicus; Transverse colon 606: 1 cm below the umbilicus; Descending colon 607: 2 to 4 cm to the left, level with the umbilicus; Rectum 608: Dorsal, below the coccyx.
[0056] EGEG data processing: The EGEG sampling frequency was 1 Hz and the filtering frequency was 0.008 Hz-0.1 Hz to filter out background noise including heartbeat. After detecting artifacts, the original EGEG potential data were calculated by the software supporting the instrument and the spectrum was analyzed by the software to derive the following parameters of the above 8 leads: (1) waveform average amplitude (AA); (2) waveform average frequency (AF); (3) gastrointestinal electrical rhythm disorder percentage (RD); (4) waveform response area (RA); (5) lead time difference (TD); (6) dominant frequency (MF); (7) dominant power ratio (MPR); (8) normal slow wave percentage (PNSW); (9) coupling percentage (CP).
[0057] Based on the above-mentioned gastrointestinal electrophysiological signals, according to the electrode placement position and the acquisition time (before or after a meal), the parameter index data of the above-mentioned 8 leads before or after a meal are averaged, representing the lead signal index of the stomach before a meal, the lead signal index of the stomach after a meal, the lead signal index of the intestine before a meal, and the lead signal index of the intestine after a meal. The practice of placing multiple leads in multiple positions for simultaneous signal acquisition and then taking the average value can better capture the overall movement pattern of the stomach and intestines, so as to more effectively obtain a signal that can reflect the overall true state of the stomach and intestines. In addition, the present invention creatively designs and introduces a pre-meal and post-meal difference index, that is, by subtracting the parameter index data of the corresponding lead before a meal from the parameter index data of the above-mentioned 8 leads after a meal, and then taking the average value, representing the lead signal index of the stomach before and after a meal difference and the lead signal index of the intestine before and after a meal difference. That is, in the present invention, the electrophysiological signal of the stomach or intestine consists of three parts: before a meal, after a meal, and the difference between before and after a meal, totaling 54 indicators.
[0058] Prediction Model Construction: Considering the complexity of gastrointestinal electrophysiological signal indicators, the present invention utilizes a two-stage analysis strategy combining a LASSO-based variable screening process with stepwise optimal logistic regression to construct a risk prediction model. First, to obtain a subset of predictors, a first round of variable screening was performed using LASSO regression analysis, a regularization algorithm. Furthermore, LASSO regression analysis was run with 10-fold cross-validation to centralize and normalize the included variables, and "lambda.min" was selected as the optimal performer. Subjects were randomly divided into a training set and a validation set in a 7:3 randomization ratio. Furthermore, to balance the potential impact of subject characteristics on epilepsy risk, subject age, gender, and BMI were included in all subsequent analyses using an inverse probability weighting approach. Then, a second round of variable screening was performed using stepwise multivariate logistic regression analysis on the predictors selected from the LASSO regression model. The retained statistically significant predictors ("predictor" and "predictor variable" are used synonymously in this invention) were used to construct the prediction model. The stepwise logistic regression process uses the AIC (Akaike Information Criterion) minimization principle to optimize the model. The final model uses the odds ratio (OR) and 95% confidence interval (CI) to define the contribution of predictors and generate a nomogram prediction model for the risk of epilepsy. It should be understood that other suitable algorithms known in the art can also be used, such as random forest methods, other regularization methods (such as ridge regression), neural networks, etc.
[0059] In addition, several validation methods were used to evaluate the accuracy of the epilepsy risk prediction model using data from the training and validation sets. These included the ROC curve (area under the ROC curve) to assess the quality of the epilepsy risk nomogram, distinguishing true positives from false positives (i.e., discrimination); and decision curve analysis to determine the clinical utility of the epilepsy risk nomogram based on the net benefit at different threshold probabilities in a natural population cohort. All analyses were performed using the R 4.1.3 packages glmnet and rms, with a significance level of two-tailed alpha < 0.1.
[0060] 1.2 Results
[0061] Subject data information:
[0062] A total of 283 subjects completed all relevant examinations, including 101 males and 182 females, of whom 112 subjects were diagnosed with epilepsy (52 males and 60 females).
[0063] As shown in Figure 2, the present invention compares each variable between the overall data, healthy people (non-epileptic subjects) and epileptic subjects. The present invention found that there were statistically significant differences in age and gender between epileptic subjects and non-epileptic subjects (P<0.01). To this end, the three basic characteristic indicators of age, gender and BMI were incorporated using the inverse probability weighting (IPTW) method. Comparing the various gastrointestinal electrical indicators of each group, the present invention found that at the level of 0.1, there were inter-group differences in the lead time difference and main power ratio of the stomach before meals, the intestine before meals, and the intestine after meals. Compared with the healthy population, the lead time difference of the stomach before meals, the intestine before meals, and the intestine after meals in the epileptic subjects decreased to varying degrees. In addition, the main power ratios of the intestine before meals and the intestine after meals showed that the epileptic subjects were smaller than the healthy population. However, in the main power of the stomach before meals and the stomach after meals, the epileptic subjects were greater than the healthy population.
[0064] Screening of independent risk factors:
[0065] Using a non-zero coefficient characteristic variable screening method based on LASSO regression, the technical solution of the present invention finally selected and retained 25 characteristic variables among 54 relevant characteristic variables (which can also be referred to as independent variables (IV) in the present invention) as potential predictive variables of the epilepsy risk model to predict the response variable (which can also be referred to as outcome variable (DV) in the present invention) (i.e., the risk of epilepsy).
[0066] LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis is a shrinkage and feature variable selection method for linear regression models. To obtain a subset of predictors, LASSO regression analysis imposes constraints on model parameters, shrinking the regression coefficients of some feature variables toward zero, thereby minimizing the prediction error of the response variable. After the shrinkage process, feature variables with regression coefficients equal to zero are excluded from the model, while feature variables with non-zero regression coefficients have the strongest association with the response variable. In other words, the first round of feature variable screening mainly eliminates feature variables whose regression coefficients are easily reduced to zero among the 54 relevant feature variables, retaining 25 feature variables as predictors of the prediction model.
[0067] Furthermore, to more accurately assess the performance of the prediction model based on 25 predictor variables, the present invention employed LASSO regression analysis to perform a 10-fold cross-validation based on the log-likelihood function (-2log-likelihood) and the type parameter of the binary dependent variable (which can be understood as whether the variable is "yes / no") (i.e., the target parameter to be minimized during cross-validation model selection). This centralized and normalized the 54 feature variables involved. The present invention employed "Lambda.min" to construct a prediction model with the best performance and highest accuracy.
[0068] Development of predictive models:
[0069] The technical solution of the present invention establishes a prediction model by introducing characteristic variables selected from the LASSO regression model and using stepwise multivariate logistic regression analysis. The selected characteristic variables are then introduced and their statistical significance levels are analyzed. Statistically significant characteristic variables are used as predictors / factors to establish a prediction model for epilepsy risk.
[0070] Next, the present invention uses a logistic regression model to analyze the 25 predictor variables screened in the first round of variables, and uses a stepwise method to select the optimal predictor variable, ultimately retaining 7 predictor variables (each predictor variable is statistically significant at the 0.05 test level). These 7 predictor variables are the lead time difference of the stomach before the meal, the lead time difference of the intestine before the meal, the main power ratio of the intestine before the meal, the normal slow wave percentage of the intestine before the meal, the average frequency of the waveform of the stomach after the meal, the lead time difference of the intestine after the meal, and the difference in the main power ratio of the stomach before and after the meal (Figure 3).
[0071] The technical solution of the present invention uses a variety of statistical methods to test the above-mentioned 7 characteristic variables, among which the odds ratio (OR) of these characteristic variables is analyzed in detail. The OR value is a statistic that quantifies the strength of the association between two events, indicating the ratio of the probability of the result occurring after exposure (in the present invention, the characteristic variable being tested, the same below) to the probability of the result occurring when the same exposure is not present. The OR value in the present invention can be specifically understood as: the strength of the association between epilepsy (i.e., the response variable) and exposure, indicating that the risk of epilepsy in the exposed person (which can also be understood as the disease risk) is a multiple of that of the non-exposed person. If the OR value of the characteristic variable being tested is >1, it means that the risk of epilepsy increases due to exposure, and there is a "positive" correlation between the characteristic variable and epilepsy; if the OR value of the characteristic variable being tested is <1, it means that the risk of epilepsy decreases due to exposure, and there is a "negative" correlation between the characteristic variable and epilepsy; if the OR value of the characteristic variable being tested is =1, it means that epilepsy is not associated with the characteristic variable. The 95% confidence interval (95% Confidence Interval (CI)) provides an estimate of the accuracy of the OR value obtained by the test, which describes that the overall true value may fluctuate within the 95% confidence interval of the OR value obtained by the test, and the smaller the confidence interval, the more accurate and robust the OR value obtained by the test. Figure 3 shows the odds ratio (OR) and its 95% confidence interval of the retained 7 predictor variables, indicating that these predictor variables are all associated with epilepsy and are therefore applied to the epilepsy risk prediction model of the present invention. These 7 predictor variables are all associated with an increased risk of epilepsy, and are negatively correlated.
[0072] In addition, in Figure 3, the regression coefficient β is the partial regression coefficient related to the predictor variable and the response variable obtained through logistic regression model analysis, which indicates the magnitude and direction of the effect of each unit increase on the response variable (the partial regression coefficient can be compared after standardization, and the regression coefficient β of the tested predictor variable is an estimated value). Standard error is the standard error of the regression coefficient β of the tested predictor variable, which indicates the accuracy of the regression coefficient (the larger the standard error, the lower the accuracy of the tested predictor variable). Z value is the z statistic, which is the regression coefficient β of the tested predictor variable divided by its corresponding standard error. It is mainly used to determine the P value of the tested predictor variable. P value is the P value corresponding to the z statistic of the tested predictor variable (the smaller the P value, the more important the tested predictor variable is to the response variable).
[0073] Based on the above 7 predictive variables, this embodiment constructs an epilepsy risk prediction model, and draws a corresponding nomogram to better visualize the constructed epilepsy risk prediction model, see Figure 4.
[0074] In Figure 4, each variable (for example, "lead time difference in the stomach before meal", "lead time difference in the intestine before meal", etc.) is marked with a scale on the line segment corresponding to the variable, which represents the value range of the variable; and the length of the line segment reflects the contribution of the variable to the outcome event (i.e., the occurrence of epilepsy). Under different values, each variable can obtain a corresponding individual score at the "score" at the top of Figure 4. After taking the value, the individual scores corresponding to all variables are added together to obtain the "total score". Based on the total score, the probability of epilepsy can be obtained at the "predicted probability of epilepsy" at the bottom of Figure 4. As an example, in Figure 4, if a subject's total score is 350, his risk of developing epilepsy is approximately 0.4 (40%).
[0075] Validation of the prediction model:
[0076] The present invention uses data from the training set and validation set to plot corresponding receiver operating characteristic (ROC) curves to evaluate the sensitivity (also understood as the true positive rate) and specificity (also understood as the true negative rate) of the constructed prediction model. In Figures 5 and 6, the x-axis represents the "false positive rate," i.e., "1-specificity"; the y-axis represents the "true positive rate," i.e., "sensitivity"; the area under the ROC curve (i.e., the solid line in Figures 5 and 6) (AUC, i.e., the area under the ROC curve and the coordinate axes) is used to analyze the quality of the risk nomogram to distinguish true positives from false positives.
[0077] For the established epilepsy risk prediction model, the area under the ROC curve (AUC) of the nomogram was above 0.65 (i.e., greater than the area under the dotted line): 0.718 (95% CI (confidence interval): 0.646-0.791) in the training set (Figure 5) and 0.651 (95% CI: 0.530-0.774) in the validation set (Figure 6), indicating that the model constructed by the present invention exhibits good robustness and ability to identify patients with epilepsy. In addition, in the training set, the established epilepsy risk prediction model had an accuracy of 0.693 (95% CI: 0.691-0.696), a sensitivity of 0.671 (95% CI: 0.567-0.775), and a specificity of 0.708 (95% CI: 0.672-0.790). In the validation set, the accuracy of the established epilepsy risk prediction model was 0.667 (95% CI: 0.661-0.672), the sensitivity was 0.667 (95% CI: 0.506-0.828), and the specificity was 0.667 (95% CI: 0.537-0.790).
[0078] Furthermore, decision curve analysis (DCA) showed that in the training set, the nomogram of this embodiment could accurately predict the threshold range of epilepsy risk between 0.40 and 0.67 (Figure 7); in the validation set, the threshold range was between 0.39 and 0.61 (Figure 8). Thus, the threshold probabilities of the two datasets overlapped between 0.39 and 0.61. The overlap of the threshold probability range of the training and validation sets in the DCA curve can be understood as indicating that both the training and validation sets can demonstrate effective net benefits within this overlapping range. That is, when patients with a risk probability exceeding the threshold probability are identified as true positive patients and intervention is applied, a positive net benefit can be achieved in both the training and validation sets. In other words, patients with a predicted probability higher than 0.61 (61%) are considered to have a greater risk of epilepsy, which is a clinical reference threshold with a net benefit. This also shows that setting the risk probability threshold within this overlapping range has higher practical application value, and the magnitude of the benefit can be compared using the benefit ratio.
[0079] Example 2: Using the above epilepsy risk prediction model to predict the probability of epilepsy risk in subjects
[0080] Referring to Figure 9, Figure 9 is an optional architectural diagram of a prediction system provided by an embodiment of the present invention. In order to support an exemplary application, the terminal (exemplarily showing a first terminal 202 and a second terminal 204) is connected to the prediction system via a network. The network involved in the present invention can be a wide area network or a local area network, or a combination of the two, using a wireless link to achieve data transmission. The terminal involved in the present invention can be various types of user terminals such as smart phones, tablet computers, and laptop computers. The terminal can be used to display an interface for inputting subject data and / or sample data, as well as an interface for displaying the prediction results of the prediction system.
[0081] The following describes an exemplary structure of a prediction system. In some embodiments, as shown in FIG10 , the prediction system 100 may include:
[0082] Database 106, for storing data, wherein the type of data is gastrointestinal electrical signal data, and the data includes sample data from a sample population and subject data from a subject;
[0083] The data acquisition module 102 is used to acquire the data and store the data in the database 106;
[0084] A model training module 108 is configured to train the sample data using a machine learning algorithm to determine an epilepsy risk prediction model;
[0085] The prediction module 112 obtains the subject data through the data acquisition module 102, and calls the determined epilepsy risk prediction model to analyze the subject data to predict the probability of epilepsy occurring in the subject.
[0086] The prediction system 100 may further include a verification module 110, which is used to evaluate the accuracy of the determined epilepsy risk prediction model, where the evaluation indicators include discrimination and / or clinical practicality.
[0087] The database 106 , the model training module 108 and the verification module 110 may be integrated into a model building module 104 .
[0088] As an example, in an epilepsy prediction system, the gastrointestinal electrical signal data specifically includes the lead time difference of the stomach before a meal, the lead time difference of the intestine before a meal, the main power ratio of the intestine before a meal, the normal slow wave percentage of the intestine before a meal, the average waveform frequency of the stomach after a meal, the lead time difference of the intestine after a meal, and the difference in the main power ratio of the stomach before and after a meal.
[0089] A specific application scenario of the present invention is given below.
[0090] Community A is conducting a community health checkup activity, collecting gastrointestinal electrical signal data from a target population within the community (e.g., all residents of Community A) and inputting the subject data through first terminal 202. Patient B in Community A has experienced several symptoms similar to epileptic seizures and often worries that he or she may be a potential epileptic patient. During the community health checkup activity, Community A transmits the subject data via network 206 to data acquisition module 102 of prediction system 100. Data acquisition module 102 obtains the subject data from the subject and stores it in database 106.
[0091] The prediction module 112 obtains the subject data and calls an established epilepsy risk prediction model to analyze the subject data and predict the probability of the subject developing epilepsy.
[0092] As an output, the prediction module 112 can generate a report indicating the risk of the subject developing epilepsy, and transmit the prediction result to the first terminal 202 via the network 206. Community A can set a risk threshold for epilepsy in advance (for example, the probability of epilepsy is 40%). When the predicted risk of epilepsy for patient B exceeds the set risk threshold (for example, the probability of occurrence is 50%), community A should issue a reminder to patient B or his family and recommend that they go to a recommended or cooperating hospital for treatment and undergo a detailed examination to determine whether patient B is an epilepsy patient. When patient B goes to the hospital for treatment, the hospital will conduct further interviews and examinations on patient B to determine whether the patient suffers from the suggested epilepsy. The doctor can transmit the diagnosis result of patient B to the prediction system 100 via the second terminal 204. The data of patient B (subject data + diagnosis result) can be used as new sample data for further training of the epilepsy risk prediction model. Of course, the diagnosis result of patient B can also be transmitted to the prediction system 100 via the first terminal 202. In other words, the terminals that transmit the diagnosis result of patient B can be the same or different.
[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a computer terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0094] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. An epilepsy prediction system based on gastrointestinal electrical signals, characterized in that: include: A database for storing data, wherein the data is gastrointestinal electrical signal data, the gastrointestinal electrical signal data including lead time difference of the stomach before a meal, lead time difference of the intestine before a meal, main power ratio of the intestine before a meal, normal slow wave percentage of the intestine before a meal, waveform average frequency of the stomach after a meal, lead time difference of the intestine after a meal, and main power ratio difference of the stomach before and after a meal; the data includes sample data from a sample population and subject data from a subject; A data acquisition module, configured to acquire the data and store the data in the database; A model training module, wherein the model training module uses a machine learning algorithm to train the sample data to determine an epilepsy risk prediction model; A prediction module, wherein the prediction module obtains the subject data through the data acquisition module and calls the epilepsy risk prediction model to analyze the subject data to predict the probability of epilepsy occurring in the subject.
2. The system according to claim 1, wherein The sample data is divided into a training set and a validation set in a ratio of 7:
3.
3. The system according to claim 1, wherein The gastrointestinal electrical signal data are obtained by simultaneously collecting leads located at the gastric body, gastric antrum, lesser curvature, greater curvature, ascending colon, transverse colon, descending colon and rectum and taking the average value.
4. The system according to claim 2, wherein: The system further includes a verification module, which is used to evaluate the accuracy of the epilepsy risk prediction model using the verification set, and the evaluation indicators include discrimination and / or clinical practicality.
5. The system according to claim 1, wherein: The subjects are adults less than 50 years old.
6. A method for constructing an epilepsy prediction system based on gastrointestinal electrical signals, characterized in that: The epilepsy prediction system based on gastrointestinal electrical signals includes an epilepsy risk prediction model, and the method includes the following steps: S1 obtains sample data from a sample population, where the sample data is gastrointestinal electrical signal data; S2 pre-trains the sample data to screen out predictive variables, wherein the screened predictive variables include lead time difference of the stomach before meal, lead time difference of the intestine before meal, main power ratio of the intestine before meal, normal slow wave percentage of the intestine before meal, average waveform frequency of the stomach after meal, lead time difference of the intestine after meal, and main power ratio difference of the stomach before and after meal; S3: Based on the screened prediction variables, the sample data is trained and learned using a machine learning algorithm to establish the epilepsy risk prediction model, and then construct the epilepsy prediction system based on gastrointestinal electrical signals.
7. The method according to claim 6, wherein The sample data is divided into a training set and a validation set in a ratio of 7:
3.
8. The method according to claim 6, wherein The pre-training includes a first round of variable screening and a second round of variable screening; the first round of variable screening includes LASSO regression analysis, and the second round of variable screening includes logistic regression analysis and stepwise regression analysis.
9. The method according to claim 7, wherein The method further includes S4 using the validation set to evaluate the accuracy of the epilepsy risk prediction model, where the evaluation indicators include discrimination and / or clinical practicality.
10. The method according to claim 6, wherein The pre-training includes ridge regression or random forest models.
Citation Information
Patent Citations
Biological signal processor
CN113539472A
Epileptic seizure judgment model modeling method, epilepsy monitoring method and device
CN114431829A
Human brain abnormal discharge detection method and device, storage medium and electronic equipment
CN117503163A
Epilepsy prediction system based on gastrointestinal electric signals and construction method
CN117898682A
Method for analyzing function of the brain and other complex systems
US20100094154A1
Cited By
Target spot positioning method and device for treating epilepsy through gastrointestinal electrical stimulation and stimulation system
CN121819171A