Artificial intelligence-based ibd patient depression risk prediction method
Patent Information
- Application Number
- CN202611256308.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明的目的在于提供一种基于人工智能的IBD患者抑郁风险预测方法,该方法通过交叉递归图提取多维生理时序数据间的非线性耦合特征,并依据耦合特征与风险阈值的逐维比较实现抑郁风险等级输出,以解决现有方法忽略多生理系统非线性耦合关系、风险分级可解释性不足的问题
[0014]采用相空间重构与交叉递归图分析相结合的方式,从IBD患者肠道炎症标志物浓度序列、心率变异序列和体动加速度序列中提取动力学耦合特征。针对每对序列,先利用互信息法和虚假最近邻点法独立确定各自的最佳延迟时间和最小嵌入维数,再构建具有相同时间索引的轨道矩阵对。轨道矩阵完整保留了原序列在高维相空间中的动力学轨迹几何结构。在交叉递归图生成过程中,基于轨道矩阵对中所有成对欧氏距离均值百分之十的自适应阈值,标记矩阵对中动力学状态相似的时刻,从递归矩阵中进一步去除孤立递归点,仅保留由至少两个连续递归点组成的线段。这种处理方式能够有效滤除因噪声引起的随机状态邻近,捕捉具有确定性的非线性耦合结构。从处理后的交叉递归图中计算的递归率表征两个序列在相空间中状态重复邻近的整体密度,最长对角线长度捕捉两个系统演化轨迹长时间平行移动的稳定性,垂直递归率反映其中一个系统状态保持而另一个系统发生跃迁的间断耦合模式。这三项指标共同刻画了IBD患者体内不同生理系统之间耦合的强度、持续性和间断性特征,将原本隐含在多维时序中的非线性协同变化规律转化为可度量的数值型特征,使其能够作为抑郁风险评估的客观输入。相比于仅使用单一序列的均值或频谱特征,这种动力学耦合特征包含了跨系统交互的深层信息,能够更敏感地感知生理调控网络的早期异常偏移。
Smart Images

Figure CN122842947A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent medical auxiliary diagnostic technology, specifically to an artificial intelligence-based method for predicting the risk of depression in IBD patients. Background Technology
[0002] The incidence of depression is significantly higher in patients with inflammatory bowel disease (IBD) than in the general population. Depressive symptoms not only reduce patients' quality of life but may also exacerbate intestinal inflammation through neuroendocrine pathways, creating a vicious cycle of deterioration. Early warning of depression risk in IBD patients can help to implement timely psychological intervention and adjust comprehensive treatment plans.
[0003] Current assessments of depression risk in IBD patients primarily rely on clinical interviews, standardized psychological scales, and single-category biomarker testing. Psychological scales depend on patients' subjective recollections and self-reported states, and are subject to cultural adaptation differences and reporting bias. Single-dimensional biomarker analysis typically focuses on only one category of intestinal inflammatory markers or autonomic nervous system markers, using their mean or fluctuation range as predictive factors. This approach ignores the nonlinear characteristics of dynamic interactions between multiple physiological systems in IBD patients. Intestinal inflammatory activity, autonomic nervous system regulation, and somatic activity are not simply linearly synchronized changes, but rather involve complex coupling patterns that exhibit different tissue characteristics in healthy and depressive states. Extracting statistics from only a single signal fails to characterize the co-evolutionary patterns between multiple modalities over time, thus limiting the sensitivity and specificity of depression risk prediction.
[0004] Some studies have attempted to input multimodal data into machine learning models, but these methods often simply concatenate the features of each modality and feed them into a black-box classifier, without explicitly modeling the dynamic coupling structure between sequences at the feature construction level. Cross-recursive graphs are a classic tool for analyzing nonlinear coupling relationships between two dynamic systems, and have been explored in the field of physiological signal analysis, but have not yet been applied to predicting depression risk in IBD patients. How to extract quantitative features reflecting the strength of nonlinear coupling from multidimensional irregular sampling sequences such as intestinal inflammation markers, heart rate variability, and body acceleration, and organize them into a structured representation that can be used for individualized risk stratification, is a pressing technical problem. Furthermore, existing risk prediction models mostly directly output binary classifications or probability values, lacking transparent multidimensional evidence to support risk level determination, resulting in insufficient clinical interpretability. How to construct a risk stratification mechanism with dimensional interpretability based on multidimensional coupling features is also a problem that needs to be solved. Summary of the Invention
[0005] The purpose of this invention is to provide an artificial intelligence-based method for predicting the depression risk of IBD patients. This method extracts nonlinear coupling features between multidimensional physiological time-series data through cross-recursive graphs, and outputs the depression risk level based on the dimension-by-dimensional comparison of coupling features and risk thresholds, so as to solve the problems of existing methods ignoring the nonlinear coupling relationship of multiple physiological systems and the insufficient interpretability of risk classification.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides an artificial intelligence-based method for predicting the risk of depression in IBD patients. By jointly analyzing multi-dimensional physiological time-series data such as the concentration sequence of intestinal inflammatory markers, heart rate variability sequence, and body motion acceleration sequence of patients with inflammatory bowel disease, the method achieves accurate stratification of depression risk.
[0007] In one technical solution of the present invention, the method includes: collecting multi-dimensional physiological time-series data of IBD patients, the data including at least intestinal inflammatory marker concentration sequences, heart rate variability sequences, and body acceleration sequences. Preferably, the intestinal inflammatory markers include at least C-reactive protein and fecal calprotectin, and their concentration sequences are obtained through a hospital information system; collecting daily self-assessment symptom scores through a patient's handheld terminal, the symptom scores including abdominal pain severity, bowel movement frequency, and fatigue; and continuously collecting heart rate variability sequences and body acceleration sequences at a sampling frequency of once per minute using a wearable smart bracelet. By fusing multi-source, multi-modal data, early signs of depression can be captured from three interrelated dimensions: autonomic nervous function, motor behavior, and intestinal inflammation.
[0008] The collected multi-dimensional physiological time-series data are time-aligned and missing values are imputed to obtain an equally spaced resampled sequence set. Preferably, the high-frequency sampling time points of the wearable device are used as the baseline time axis. The concentration sequences of intestinal inflammatory markers and the self-assessment symptom score sequences are projected onto this baseline time axis through linear interpolation and cubic spline interpolation, respectively. For missing time points that still exist after interpolation, the moving average of the three time points before and after the missing time points is used to fill them. All the filled sequences are then zero-mean and unit variance scaled to obtain a standardized resampled sequence set. As a further preferred approach, for time periods with more than five consecutive missing time points, the data for the corresponding time period is marked as invalid and discarded, thereby avoiding interference from low-quality data to subsequent analysis and ensuring the reliability of coupled feature extraction.
[0009] For each pair of sequences in the resampled sequence set, phase space reconstruction is performed to obtain orbital matrix pairs. For each univariate sequence in each pair, the optimal delay time is calculated using the mutual information method, and the minimum embedding dimension is calculated using the spurious nearest neighbor method. Based on the optimal delay time and the minimum embedding dimension, each univariate sequence is mapped to an orbital matrix in a high-dimensional space. The two orbital matrices of the same pair of sequences are aligned according to their time indices to obtain orbital matrix pairs. This step can extend the original one-dimensional observation sequence to a high-dimensional state space, fully exposing the inherent coupling relationships of the system's dynamic characteristics.
[0010] Calculate the cross-recursion graph between each pair of orbital matrices and extract the recursion rate, longest diagonal length, and vertical recursion rate as coupling features. Preferably, the recursion radius threshold is set to 10% of the average Euclidean distance of all pairs of orbital matrices. For each pair of time index states in the orbital matrix pair, if the Euclidean distance between them is less than this threshold, it is marked as a recursive point at the corresponding position in the cross-recursion graph; otherwise, it is marked as a non-recursive point. Isolated points are removed from the cross-recursion graph, retaining only line segments consisting of at least two consecutive recursive points. The recursion rate is calculated as the ratio of the total number of recursive points in the cross-recursion graph to the total number of grid points. The longest diagonal length is obtained by scanning all line segments parallel to the main diagonal in the cross-recursion graph using a dynamic programming algorithm, taking the longest line segment where recursive points appear consecutively. The vertical recursion rate is calculated as the ratio of the total number of line segments consisting of at least two consecutive recursive points in the vertical direction to the total number of all vertical line segments. These coupling features quantify the similarity, determinism, and discontinuity of the dynamic behavior between the two sets of physiological timelines from different perspectives, and can sensitively reflect the changes in the dynamic coupling strength between IBD activity and depressive mood.
[0011] The coupling features from all sequence pairs are concatenated into a multidimensional feature vector. As a preferred approach, the three coupling features of the intestinal inflammatory marker concentration sequence and the heart rate variability sequence are arranged as the first group of features; the three coupling features of the intestinal inflammatory marker concentration sequence and the body acceleration sequence are arranged as the second group of features; and the three coupling features of the heart rate variability sequence and the body acceleration sequence are arranged as the third group of features. These features are then horizontally concatenated in the order of the first, second, and third groups to form the multidimensional feature vector. This feature vector integrates the nonlinear dynamic information of the three coupled pathways—gut-heart, gut-motor, and heart-motor—comprehensively characterizing the systemic physiological coupling state of IBD patients.
[0012] The method compares the multidimensional feature vector with a pre-constructed risk threshold vector dimension by dimension, counts the number of dimensions exceeding the threshold, and outputs the depression risk level based on the number of dimensions. This method does not rely on complex classifier models; it can determine the risk level simply by counting thresholds in the feature space. It has low computational cost, strong interpretability, and can be directly applied to real-time risk screening in clinical settings. This helps to identify depressive tendencies early in the course of IBD patients and take intervention measures.
[0013] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0014] A combined approach of phase space reconstruction and cross-recursive graph analysis was employed to extract kinetic coupling features from intestinal inflammatory marker concentration sequences, heart rate variability sequences, and body acceleration sequences of IBD patients. For each sequence pair, the optimal delay time and minimum embedding dimension were independently determined using mutual information and spurious nearest neighbor methods, and then orbital matrix pairs with the same time index were constructed. The orbital matrices completely preserved the kinetic trajectory geometry of the original sequences in the high-dimensional phase space. During the cross-recursive graph generation process, an adaptive threshold of 10% of the mean Euclidean distance of all pairs of orbital matrix pairs was used to mark moments of similar kinetic states in the matrix pairs. Isolated recursive points were further removed from the recursive matrix, retaining only line segments consisting of at least two consecutive recursive points. This processing method effectively filters out random state proximity caused by noise and captures deterministic nonlinear coupling structures. The recursion rate calculated from the processed cross-recursive graph characterizes the overall density of repeated state proximity between two sequences in phase space, the longest diagonal length captures the stability of long-term parallel movement of the evolution trajectories of the two systems, and the vertical recursion rate reflects the discontinuous coupling mode where one system maintains its state while the other undergoes a transition. These three indicators collectively characterize the strength, persistence, and discontinuity of coupling between different physiological systems in IBD patients, transforming the nonlinear synergistic changes originally implicit in multidimensional time series into measurable numerical features, making them an objective input for depression risk assessment. Compared to using only the mean or spectral features of a single sequence, this dynamic coupling feature contains deep information about cross-system interactions, enabling more sensitive detection of early abnormal shifts in physiological regulatory networks.
[0015] After obtaining the coupling features, the coupling features of all sequence pairs are grouped and concatenated into a multidimensional feature vector according to the inter-system relationships. Each group of features retains the system interaction meaning of its source. When determining the level of depression risk, instead of using weighted summation or distance metrics to reduce the multidimensional features to a single score, the value of each dimension of the feature vector is compared one by one with a corresponding risk threshold established in advance based on population data, and the number of dimensions exceeding the threshold is counted. Each dimension of the threshold vector corresponds to a specific physiological system coupling relationship, such as an excessive increase or decrease in synchronicity between intestinal inflammation and heart rate variability, or a long-term separation between body rhythm and autonomic nervous activity. When multiple dimensions of coupling features deviate from the normal range simultaneously, a larger number of dimensions indicates a more widespread imbalance in coordination among multiple physiological systems. Discrete risk levels are output based on the interval in which the number of dimensions falls, with each level directly associated with a traceable combination of coupling-imbalanced dimensions. This approach makes the risk level determination dimensionally transparent, allowing clinicians to clearly identify which abnormal couplings between physiological systems contributed to the current risk level, avoiding the output of untraceable single probability values from black-box models. Meanwhile, the threshold comparison mechanism does not rely on large-sample training of black-box classifiers, has relatively low dependence on sample size, and has stronger deployability under the condition that it is not easy to obtain a large number of IBD combined with depression samples. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0017] Figure 1 This is a flowchart of an artificial intelligence-based method for predicting the risk of depression in IBD patients.
[0018] Figure 2 This is a flowchart of phase space reconstruction;
[0019] Figure 3 These are curves showing the changes in standardized heart rate variability, body acceleration, and overall inflammatory activity sequences on a baseline time axis.
[0020] Figure 4 It is a curve for determining the optimal reconstruction parameters based on the mutual information method and the false nearest neighbor method;
[0021] Figure 5 It is a box plot and scatter distribution based on the coupling characteristics of the orbital matrix with the cross recursive graph. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] See Figure 1 This invention provides an artificial intelligence-based method for predicting depression risk in IBD patients, comprising: collecting multi-dimensional physiological time-series data of IBD patients, the data consisting of at least intestinal inflammatory marker concentration sequences, heart rate variability sequences, and body motion acceleration sequences; performing time alignment and missing value imputation processing on the multi-dimensional physiological time-series data to obtain an equally spaced resampled sequence set; reconstructing the phase space of each pair of sequences in the resampled sequence set to obtain orbital matrix pairs; calculating the cross recursion graph between each pair of orbital matrices and extracting the recursion rate, longest diagonal length, and vertical recursion rate as coupling features; concatenating the coupling features from all sequence pairs into a multi-dimensional feature vector; comparing the multi-dimensional feature vector with a pre-constructed risk threshold vector dimension by dimension, counting the number of dimensions exceeding the threshold, and outputting a depression risk level based on the number of dimensions.
[0024] Example 1:
[0025] In practice, the process of collecting multidimensional physiological time-series data from IBD patients includes obtaining intestinal inflammatory marker concentration sequences, heart rate variability sequences, and body acceleration sequences. The acquisition of intestinal inflammatory marker concentration sequences is performed through the hospital information system. The hospital information system stores the laboratory test results of IBD patients, including C-reactive protein (CRP) and fecal calprotectin (FCC) concentration values. The hospital information system provides a data access interface compliant with the HL7 standard. Through this interface, the system retrieves the time point and CRP concentration value for each CRP test and the time point and fecal calprotectin concentration value for each IBD patient via scheduled queries. The retrieved CRP concentration values are arranged chronologically according to their test time points to form a CRP concentration sequence. Similarly, the retrieved fecal calprotectin concentration values are arranged chronologically according to their test time points to form a fecal calprotectin concentration sequence. The intestinal inflammatory marker concentration sequences are composed of both the CRP and fecal calprotectin concentration sequences.
[0026] The daily self-assessment symptom score sequence is obtained via the patient's handheld terminal. The patient's handheld terminal has a self-assessment symptom collection application installed, which includes input controls for abdominal pain intensity, bowel movement frequency, and fatigue. The abdominal pain intensity input control is a visual analog scale. The left endpoint of the visual analog scale is marked "no pain" with a value of 0, and the right endpoint is marked "most severe pain" with a value of 10. IBD patients slide on the visual analog scale daily to select their abdominal pain intensity score. The bowel movement frequency input control is a numeric input box. IBD patients enter the total number of bowel movements for the day in this numeric input box. The fatigue score input control is another visual analog scale. The left endpoint of this other visual analog scale is marked "no fatigue" with a value of 0, and the right endpoint is marked "extreme fatigue" with a value of 10. IBD patients slide on this other visual analog scale daily to select their fatigue score. The self-report symptom collection application sends daily reminders to IBD patients at fixed times. After IBD patients input their abdominal pain severity score, bowel movement frequency score, and fatigue score, the application combines the date, abdominal pain severity score, bowel movement frequency score, and fatigue score into a single self-report symptom record and uploads it to the data center server via wireless network. The data center server stores the received self-report symptom records in chronological order, forming a self-report symptom score sequence, with each time point in the sequence corresponding to one self-report symptom record.
[0027] Heart rate variability (HRV) and body acceleration (BCA) sequences were acquired using a wearable smart bracelet. The bracelet, worn on the wrist of the IBD patient, incorporates a photoplethysmography (PPG) sensor and a triaxial accelerometer. The PPG sensor continuously acquires pulse wave signals from the wrist at a sampling frequency of 256 Hz, calculating the time interval between adjacent peaks to obtain the RR interval time series. The PPG sensor extracts all RR intervals within a one-minute time window, calculates the standard deviation of all RR intervals within the one-minute window, and uses this standard deviation as the sampled value of the HRV sequence for that time window. The wearable smart bracelet outputs one sampled value of the HRV sequence per minute, sliding the window for one minute per second to generate the HRV sequence. The triaxial accelerometer continuously acquires acceleration components in three orthogonal directions at a sampling frequency of 256 Hz, obtaining the X-axis, Y-axis, and Z-axis acceleration component time series. In a triaxial accelerometer, using a one-minute time window, the acceleration amplitude is calculated for each sampling point of the X-axis acceleration component time series, Y-axis acceleration component time series, and Z-axis acceleration component time series within the one-minute time window, as follows:
[0028]
[0029] in, Indicates the number of seconds within a one-minute time window. The acceleration amplitude at each sampling point Indicates the number of seconds within a one-minute time window. X-axis acceleration component values at each sampling point Indicates the number of seconds within a one-minute time window. Y-axis acceleration component values at each sampling point Indicates the number of seconds within a one-minute time window. Z-axis acceleration component values at each sampling point. Calculate the acceleration amplitude of all sampling points within a one-minute time window. The mean value is calculated and used as the sample value of the body acceleration sequence within that time window. The wearable smart bracelet outputs a sample value of the body acceleration sequence once per minute, with the window sliding for one minute per second to form the body acceleration sequence. The wearable smart bracelet transmits the heart rate variability sequence and body acceleration sequence in real time to the paired patient handheld terminal via Bluetooth communication protocol. The patient handheld terminal then uploads the heart rate variability sequence and body acceleration sequence to a data center server for storage via wireless network.
[0030] See Figure 3In the graph, the horizontal axis represents time in days, ranging from 0 to 7 days; the vertical axis represents the standardized value (z-score), ranging from -3 to 3. The legend indicates that the three curves represent the heart rate variability sequence (blue solid line), the body acceleration sequence (orange dashed line), and the comprehensive inflammatory activity sequence (green dotted line), all of which have been standardized with zero mean and unit variance.
[0031] The two curves of the heart rate variability sequence and the body acceleration sequence show obvious periodic oscillation characteristics, with an overall fluctuation range of ±2. Moreover, the waveform trends of the two are highly consistent, showing the synchronous dynamic change trend between the two sequences within a 7-day cycle. This indicates that there is a strong correlation and coupling relationship between heart rate variability and body acceleration during this period.
[0032] The combined inflammatory activity sequence curve differed significantly from the heart rate variability sequence and body acceleration sequence, showing a smooth and slow trend. It remained at a low level (z-score close to -1) in the time interval of 0 to 4 days, and then gradually increased in the interval of 4 to 7 days, reaching a peak close to 3, indicating that the intestinal inflammatory activity was significantly enhanced during this period.
[0033] Example 2:
[0034] In practice, the process of time alignment and missing value imputation for multi-dimensional physiological time-series data uses the high-frequency sampling time points of the wearable smart bracelet as the reference time axis. The wearable smart bracelet samples once per minute, and the reference time axis consists of a series of equally spaced timestamps, with an interval of 60 seconds between adjacent timestamps. The time point distribution of the intestinal inflammatory marker concentration sequence and the self-assessment symptom score sequence is uneven, requiring the sampling values of the intestinal inflammatory marker concentration sequence and the self-assessment symptom score sequence to be projected onto each timestamp of the reference time axis.
[0035] Both the C-reactive protein (CRP) and fecal calprotectin (PCC) concentration sequences from the intestinal inflammatory marker concentration series were projected onto a baseline time axis using linear interpolation. For a target timestamp on the baseline time axis, the preceding and following detection time points adjacent to the target timestamp were located in the CRP concentration sequence, and their CRP concentration values were obtained. The rate of change of the CRP concentration value was calculated based on the time difference between the two detection time points and the difference between the two CRP concentration values. This rate of change was multiplied by the time difference between the target timestamp and the preceding detection time point, and the result was added to the CRP concentration value of the preceding detection time point to obtain the interpolated CRP concentration for the target timestamp. The same linear interpolation operation was applied to the PCC concentration sequence to obtain the PCC concentration interpolation result for each target timestamp.
[0036] The self-reported symptom score sequence was projected onto a baseline time axis using cubic spline interpolation. The date of each day in the self-reported symptom score sequence was converted into a time point on the baseline time axis, resulting in a set of data points with time as the x-axis and the abdominal pain score, bowel movement frequency, and fatigue score from the self-reported symptom score records as the y-axis. For the set of data points corresponding to the abdominal pain score, a cubic spline interpolation function was constructed. The function value of the cubic spline interpolation function at each data point equals the abdominal pain score, and it is a cubic polynomial between adjacent data points, with continuous first and second derivatives at all data points. The target timestamp was substituted into the cubic spline interpolation function to calculate the interpolated result of the abdominal pain score corresponding to the target timestamp. Using the same cubic spline interpolation operation, cubic spline interpolation functions were constructed for the sets of data points corresponding to the bowel movement frequency and fatigue scores, respectively, and the target timestamps were substituted to obtain the interpolated results of the bowel movement frequency and fatigue scores for each target timestamp.
[0037] After linear interpolation and cubic spline interpolation are completed, interpolation results for the concentration sequences of intestinal inflammatory markers and the self-assessment symptom score sequences are obtained for all timestamps on the baseline time axis. If interpolation cannot be calculated during linear interpolation or cubic spline interpolation due to insufficient detection time points or self-assessment symptom records at the front or back ends, the corresponding timestamp is marked as a missing time point.
[0038] For any missing time points after interpolation, a moving average of three time points before and after the missing time point is used to fill them. For a missing time point on the reference time axis, three valid sampling points are determined forward and three valid sampling points backward on the reference time axis. A valid sampling point is one where valid sequence sample values are already available at that time point. The sequence sample values of the three valid sampling points before and three valid sampling points backward are obtained, resulting in six sequence sample values. , , , , , Calculate the moving average fill value. The formula is:
[0039]
[0040] in, This represents the sequence sample value of the first valid sample point in the direction preceding the missing time point. This represents the sequence sample value of the second valid sampling point in the direction preceding the missing time point. This represents the sequence sample value of the third valid sampling point in the direction preceding the missing time point. This represents the sequence sample value of the first valid sampling point in the direction following the missing time point. This represents the sequence sample value of the second valid sampling point following the missing time point. This represents the sequence sample value of the third valid sampling point following the missing time point. The calculated moving average fill value. Fill in the missing time points.
[0041] During the moving average filling process, if there are fewer than three valid sampling points in the forward or backward direction of the missing time point, the search range is expanded to five valid sampling points in both directions, and the arithmetic mean of the ten valid sampling points is used for filling. If the number of valid sampling points still cannot be met, the search is extended in the forward or backward direction to all existing valid sampling points, and the arithmetic mean of all existing valid sampling points is used for filling.
[0042] For time periods with more than five consecutive missing time points, the data for those time periods is marked as invalid and discarded. The baseline time axis is scanned from the beginning, and the number of consecutive missing time points is recorded. When the number of consecutive missing time points exceeds five, the time periods containing these five or more consecutive missing time points are marked as invalid. All interpolated results of intestinal inflammatory marker concentration sequences, self-reported symptom score sequences, heart rate variability sequences, and body acceleration sequences for all time points within the invalid time period are discarded. All records for the corresponding time points are deleted from the resampled sequence set.
[0043] After gap filling and discarding invalid time periods, zero-mean and unit variance scaling were performed on all sequences corresponding to all remaining time points on the baseline time axis. Zero-mean scaling calculates the mean of the sampled values of a single sequence across all remaining time points, subtracting the mean from each sampled value of the single sequence. Unit variance scaling calculates the standard deviation of the sampled values of a single sequence across all remaining time points, dividing each zero-meaned sampled value by the standard deviation. Zero-mean and unit variance scaling were performed independently on the C-reactive protein concentration sequence and fecal calprotectin concentration sequence (in the intestinal inflammatory marker concentration sequence), the abdominal pain score sequence, defecation frequency sequence, and fatigue score sequence (in the self-report symptom score sequence), the heart rate variability sequence, and the body acceleration sequence. After scaling was performed on all sequences, a standardized resampled sequence set was obtained.
[0044] Example 3:
[0045] In specific implementation, please refer to Figure 2The process involves reconstructing the phase space of each pair of sequences in the resampled sequence set, encompassing three steps: calculating reconstruction parameters, constructing the orbital matrix, and aligning the orbital matrix. The resampled sequence set includes standardized C-reactive protein concentration sequences, fecal calprotectin concentration sequences, abdominal pain score sequences, defecation frequency sequences, fatigue score sequences, heart rate variability sequences, and body acceleration sequences. Each sequence pair refers to selecting two different sequences from the above sets for pairing. Pairing includes pairing between intestinal inflammatory marker concentration sequences and heart rate variability sequences, pairing between intestinal inflammatory marker concentration sequences and body acceleration sequences, and pairing between heart rate variability sequences and body acceleration sequences. The two sequences within the intestinal inflammatory marker concentration sequence participate in the pairing operation independently.
[0046] For each univariate sequence in each pair of sequences, the optimal delay time is calculated using mutual information. Let the univariate sequence be... ,in Indicates the length of a univariate sequence. Represents the first digit in a univariate sequence. The sampled values corresponding to each sampling point The value range is 1 to An integer. For a candidate delay time. , univariate sequence Its delayed version Construct a joint distribution, statistically analyze the distribution intervals of univariate sequence values and delayed sequence values, divide the distribution intervals into several equal-width grids, and calculate the frequency of sample points within each grid. Calculate the joint probability based on the sample point frequencies. and marginal probability and Mutual information value The calculation formula is:
[0047]
[0048] in, Indicates the number of grid divisions. This indicates that the value of the univariate sequence falls within the first... Representative values within each grid interval This indicates that the delayed sequence value falls within the first... Representative values within each grid interval This indicates that the value of the univariate sequence falls into the first... The grid intervals and the delayed sequence values simultaneously fall into the first grid interval. The joint probability of each grid interval. This indicates that the value of the univariate sequence falls into the first... The edge probabilities of each grid interval This indicates that the delayed sequence value falls into the first... Edge probabilities of each grid interval. Candidate delay time. Starting from 1 and incrementing by 1 each time, calculate the mutual information value corresponding to each candidate delay time. The curves showing the change in mutual information value with candidate delay time are recorded, and the candidate delay time corresponding to the first local minimum of the curve is determined as the optimal delay time. The criterion for determining a local minimum is the candidate delay time... The corresponding mutual information value is less than the candidate delay time. The corresponding mutual information value, and less than the candidate delay time. The corresponding mutual information value.
[0049] The minimum embedding dimension is calculated using the spurious nearest neighbor method. Let the candidate embedding dimension be... Utilize the determined optimal delay time Reconstructing the univariate sequence into A set of orbital vectors in dimensional space, the first orbital vector in the set of orbital vectors. The orbital vectors are represented as follows: ,in The value ranges from 1 to Integers. In 3D space, using Euclidean distance as the standard, search The nearest neighbor is denoted as . , This is the index number of the nearest neighbor. Calculate... and Distance between Increase the embedding dimension to The corresponding orbital vector is and ,calculate and Distance between The criterion for identifying false nearest neighbors is to calculate the distance ratio. ,when Points exceeding a preset threshold are marked as false neighbors. The preset threshold is set to 10. The rationale for this threshold is that in nonlinear dynamics analysis, if the orbital stretching caused by increasing the embedding dimension results in a distance ratio exceeding an order of magnitude, it indicates that the adjacency relationships between adjacent points in the low-dimensional space are false adjacencies due to insufficient dimension. The proportion of false neighbors in all orbital vectors is statistically analyzed. When the proportion of false neighbors first drops below 5% of the total number of orbital vectors, the corresponding candidate embedding dimension is determined as the minimum embedding dimension. The search for the minimum embedding dimension starts with a candidate embedding dimension of 1 and increments by 1 each time.
[0050] Based on the optimal delay time and minimum embedding dimension, each univariate sequence is mapped to an orbital matrix in a high-dimensional space. Let the length of the univariate sequence be... The optimal delay time is The minimum embedding dimension is Univariate sequences can be constructed orbital vectors The orbital matrix of the first... Behavior No. There are orbital vectors, and the orbital matrix is a... OK The matrix of columns, the orbital matrix of the th column Line 1 The element value of the column is ,in For column indexes, The value range is 1 to Integers.
[0051] Align the two orbital matrices of the same sequence pair according to their time indices. The number of rows in each orbital matrix is determined by the length, optimal delay time, and minimum embedding dimension of each univariate sequence. The two univariate sequences in the same pair may have different lengths and reconstruction parameters, resulting in different numbers of rows in the two orbital matrices. Take the minimum number of rows in the two orbital matrices as the common row number after alignment. Extract the orbital vectors corresponding to the first common row number from the first orbital matrix, and extract the orbital vectors corresponding to the first common row number from the second orbital matrix. After extraction, the first and second orbital matrices have the same number of rows, and the first... The high-dimensional states corresponding to the rows originate from the same time index interval, thus obtaining orbital matrix pairs.
[0052] See Figure 4 The horizontal axis in the graph represents the delay time. The value ranges from 0 to 500, and the left vertical axis represents mutual information. The value ranges from 0 to 2.5; the dimension above the horizontal axis is the embedding dimension. The value ranges from 1 to 20, and the right-hand vertical axis represents the percentage of false neighbors (%), ranging from 0 to 120%. The solid green line in the figure represents mutual information. With delay time The trend of change shows an overall pattern of first rapid decline followed by slow fluctuations. The green dashed vertical line marks the position where the mutual information curve first reaches a local minimum, i.e., the optimal delay time. This value is approximately 182. The red dashed line represents the proportion of spurious neighbors as a function of the embedding dimension. The trend shows a high initial value, followed by a rapid decline to a lower level and then stabilization. The red dashed vertical line marks the embedding dimension corresponding to the first drop in the proportion of false neighbors below 5%, i.e., the minimum embedding dimension. This value is approximately 2. Based on the method described above for calculating reconstruction parameters using mutual information and spurious nearest neighbor methods, the first local minimum point of the green mutual information curve in the figure determines the optimal delay time for the univariate sequence. This is beneficial for constructing an efficient orbital matrix; the sharp drop in the proportion curve of red false neighbors and the points where the proportion first falls below 5% determine the minimum embedding dimension. This ensures sufficient embedding space dimensions and avoids false proximity problems caused by excessively low dimensions.
[0053] Example 4:
[0054] In practice, the process of calculating the cross recursion graph between each pair of orbital matrices includes setting a recursion radius threshold, marking recursive and non-recursive points, removing isolated points, and extracting the longest diagonal length. Setting the recursion radius threshold involves traversing the first and second orbital matrices in the pair and calculating the Euclidean distance between all pairs of orbital vectors in both matrices. Let the first orbital matrix contain... The first orbital vector, the second orbital matrix also contains... The orbital vector, the first orbital matrix in the first orbital matrix The orbital vectors are denoted as The second orbital matrix The orbital vectors are denoted as ,in and The value range is 1 to Calculate integers. and Euclidean distance between A total of Each Euclidean distance value. All The sum of the Euclidean distance values is divided by . This yields the mean of all pairwise Euclidean distances. Recursive radius threshold Set to the mean of all pairwise Euclidean distances. Ten percent, that is Set the recursive radius threshold. The rationale for setting the threshold to 10 percent of the mean of all pairwise Euclidean distances is that, in the recursive analysis of nonlinear time series, the recursion rate needs to be controlled within a moderate range. Using 10 percent of the mean distance as a threshold can suppress spurious recursion points caused by noise while ensuring the visibility of the recursive structure, so that the cross-recursion graph can reflect the deterministic coupling relationship between the orbital matrix pairs.
[0055] The process of marking recursive and non-recursive points is as follows: for each pair of time index states in the orbital matrix pair, the first orbital matrix is set to the position at the specified time index. orbital vectors with time indices and the second orbital matrix in the orbital vectors with time indices As a pair of time-indexed states, calculate and Euclidean distance between If Euclidean distance Less than the recursion radius threshold Then in the cross recursion graph, the first... Line 1 The corresponding positions in the column are marked as recursion points. If the Euclidean distance... Greater than or equal to the recursion radius threshold Then in the cross recursion graph, the first... Line 1 The corresponding positions in the column are marked as non-recursive points. A cross-recursive graph is a... OK The column is a two-dimensional binary matrix, where the value of a recursive point is 1 and the value of a non-recursive point is 0.
[0056] The process of removing isolated points involves performing connectivity analysis on all recursive points with a value of 1 in the cross-recursion graph. An isolated point is a recursive point whose eight neighbors contain no other recursive points. The eight neighbors are defined as the eight adjacent positions centered on a given recursive point in the cross-recursion graph: above, below, left, right, upper left, upper right, lower left, and lower right. Each recursive point in the cross-recursion graph is scanned, and it is checked whether at least one recursive point exists in its eight neighbors. If not, the recursive point is considered isolated, and the corresponding matrix element value is changed from 1 to 0. The remaining recursive points in the cross-recursion graph are all constituent points of line segments consisting of at least two consecutive recursive points. Only line segments consisting of at least two consecutive recursive points are retained in the cross-recursion graph. "Consecutive" means that the two recursive points are adjacent in column index in the same row, adjacent on the main diagonal, or adjacent in row index in the same column.
[0057] The process of extracting the longest diagonal length uses a dynamic programming algorithm to scan all line segments parallel to the main diagonal in the cross recursion graph, taking the longest consecutively occurring recursion point. The dynamic programming algorithm then processes each segment along the second diagonal of the cross recursion graph, extracting the longest segment from the second diagonal. Line 1 The element values of a column are denoted as Let an auxiliary matrix be defined. It has the same number of rows and columns as the cross recursion graph. Indicates the first Line 1 The maximum length of consecutive occurrences of recursive points among line segments parallel to the main diagonal with the endpoint listed as the endpoint. The recursive formula for the dynamic programming algorithm is:
[0058]
[0059] in, Represents the nth element in the auxiliary matrix Line 1 The element values of the column, Represents the nth element in the auxiliary matrix Line 1 The element values of the column, In the cross recursion graph, the first... Line 1 The element values of the column, This is the row index, with values ranging from 1 to... integers, For column indexes, values range from 1 to... Integer. When When the value is 0, the longest consecutive length at the corresponding position is set to 0. =1 and If it exists, increment the longest consecutive length of the previous diagonal position by 1 and assign it to the current position. =1 and When it does not exist, that is or The longest consecutive length at the current position is set to 1. After processing all positions in the cross recursive graph, the auxiliary matrix is traversed. Take the maximum value among all elements in the array and use it as the length of the longest diagonal.
[0060] Example 5:
[0061] In practice, the process of extracting coupling features includes calculating the recursion rate, calculating the longest diagonal length, and calculating the vertical recursion rate. The cross-recursion graph is a... OK A binary matrix of columns, This represents the row number of each orbital matrix in the orbital matrix pair. In the cross-recursion graph, recursive points are assigned a value of 1, and non-recursive points are assigned a value of 0. The recursion rate is calculated as the ratio of the total number of recursive points to the total number of grid points in the cross-recursion graph. Traverse all nodes in the cross-recursion graph. For each grid location, count the number of grid locations with a value of 1 to obtain the total number of recursive points. The total number of grid points is Recursion rate The calculation method is as follows .
[0062] The longest diagonal length is determined by scanning all diagonal segments parallel to the main diagonal in the cross recursion graph and selecting the longest segment with consecutive occurrences at the recursion point. For each diagonal segment parallel to the main diagonal in the cross recursion graph, the grid position value is checked one by one along the diagonal direction, and the length of the consecutive segments with consecutive values of 1 is recorded. The lengths of all consecutive segments on all diagonals are compared, and the maximum length value is determined as the longest diagonal length. .
[0063] The vertical recursion rate is calculated as the ratio of the total number of line segments formed by at least two consecutive recursive points in the vertical direction to the total number of vertical line segments. In a cross-recursion graph, scanning is performed column-by-column. The cross-recursion graph has a total of... Columns, column indexes The value range is from 1 to Integers. For the th The column is checked row by row from top to bottom, examining the values at each grid position. When a value changes from 0 to 1, it is marked as the start of a vertical line segment. The check continues downwards until a value changes from 1 to 0 or the end of the column is reached, at which point it is marked as the end of a vertical line segment. If the number of recursive points between the start and end of a vertical line segment is greater than or equal to 2, then this vertical line segment is counted in the total number of line segments formed by at least two consecutive recursive points in the vertical direction. Scan all After listing, sum up to get The total number of vertical segments is defined as the total number of vertical segments formed by adjacent grid pairs within each column of the cross recursive graph. Each column contains... Number of adjacent grid pairs, number of all vertical line segments The calculation formula is:
[0064]
[0065] in, This represents the number of rows in the orbital matrix, which is also the number of rows and columns in the cross recursion graph. Vertical recursion rate. The calculation method is as follows .
[0066] In practice, the intestinal inflammation marker concentration sequence is a comprehensive inflammatory activity sequence obtained by scaling the C-reactive protein concentration sequence and the fecal calprotectin concentration sequence to zero mean and unit variance, and then calculating the arithmetic mean point by point. The process of concatenating the coupled features from all sequence pairs into a multidimensional feature vector is performed according to the grouping order. The three coupled features of the intestinal inflammation marker concentration sequence and the heart rate variability sequence pair are arranged into the first group of features, which includes the recursion rate. Longest diagonal length and vertical recursion rate The three coupling features of the intestinal inflammatory marker concentration sequence and the body motion acceleration sequence pair are arranged into a second set of features, which includes the recurrence rate. Longest diagonal length and vertical recursion rate The three coupling features of the heart rate variability sequence and the body acceleration sequence are arranged into a third set of features, which includes the recursion rate. Longest diagonal length and vertical recursion rate By horizontally concatenating the features in the order of the first, second, and third groups, a nine-dimensional multi-dimensional feature vector is obtained, which takes the form of... .
[0067] See Figure 5 The figure shows the recurrence ratio (RR) and longest diagonal length (LD) of three sequence pairs obtained based on the coupled feature extraction method in Example 5. The statistical distribution of the recurrence rate (RR) and vertical recurrence rate (VR) is shown. The horizontal axis represents the feature dimension, with the following values in order: first group of features (intestinal inflammation and heart rate variability) RR1, L1, VR1; second group of features (intestinal inflammation and body acceleration) RR2, L2, VR2; and third group of features (heart rate variability and body acceleration) RR3, L3, VR3. The left vertical axis corresponds to the numerical range of the recurrence rate RR and vertical recurrence rate VR, which ranges from approximately 0.00 to 0.225; the right vertical axis corresponds to the length of the longest diagonal. The numerical range is approximately 0 to 250.
[0068] The three sets of features in the figure are distinguished by different colors: blue represents the coupling feature between intestinal inflammation and heart rate variability, orange represents the coupling feature between intestinal inflammation and body acceleration, and green represents the coupling feature between heart rate variability and body acceleration. Each feature is displayed in the form of a box plot. In the box plot, the boxes represent the upper and lower quartiles (Q1 and Q3), the midline of the box represents the median, the stub represents the maximum and minimum values of the data, and the scatter plot represents the distribution of all sample points.
[0069] From the comparison of recursion rates (RR), the median of the third feature group (heart rate variability and body acceleration, RR3) was the highest, at about 0.14, which was higher than that of the first group (RR1) and the second group (RR2), indicating that the proportion of coupling recursion points between heart rate variability and body acceleration was relatively high.
[0070] Longest diagonal length Among the features, all three groups showed a large range of variation. The second group (L2) and the third group (L3) both showed a large number of scatter points extending to over 200, indicating that the coupling trajectory between intestinal inflammation and heart rate variability, as well as between heart rate variability and body acceleration, has a long duration and certainty. The first group (L1) had a significantly lower median, around 25, and the data distribution was more concentrated, reflecting that the coupling trajectory between intestinal inflammation and body acceleration was relatively short.
[0071] Regarding the vertical recursion rate (VR) feature, the second group (VR2) showed the highest median, approximately 0.12, followed by the second group (VR2), while the first group (VR1) had the lowest. All three groups had a wide distribution range, indicating that the coupling relationship between heart rate variability and body acceleration is more significant in the vertical recursion structure.
[0072] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. An artificial intelligence-based method for predicting the risk of depression in IBD patients, characterized in that, Includes the following steps: Collect multidimensional physiological time-series data of IBD patients, including at least the concentration sequence of intestinal inflammatory markers, the heart rate variability sequence, and the body acceleration sequence; The multi-dimensional physiological time-series data are time-aligned and missing values are imputed to obtain a set of resampled sequences with equal intervals. For each pair of sequences in the resampled sequence set, phase space reconstruction is performed to obtain orbital matrix pairs; Calculate the cross recursion graph between each pair of orbital matrices, and extract the recursion rate, longest diagonal length, and vertical recursion rate as coupling features. The coupled features from all sequence pairs are concatenated into a multidimensional feature vector; The multidimensional feature vector is compared dimension by dimension with the pre-constructed risk threshold vector, the number of dimensions exceeding the threshold is counted, and the depression risk level is output based on the number of dimensions.
2. The method for predicting the risk of depression in IBD patients based on artificial intelligence according to claim 1, characterized in that, The collection of multidimensional physiological time-series data from IBD patients includes: The concentration sequences of intestinal inflammatory markers were obtained through the hospital information system, and the intestinal inflammatory markers included at least C-reactive protein and fecal calprotectin; Daily self-assessment symptom scores are collected through a patient's handheld device. The symptom scores include the degree of abdominal pain, frequency of bowel movements, and fatigue. Heart rate variability and body acceleration sequences are continuously collected using a wearable smart bracelet, with a sampling frequency set to once per minute.
3. The artificial intelligence-based method for predicting the risk of depression in IBD patients according to claim 1, characterized in that, The process of time alignment and missing value imputation of the multi-dimensional physiological time-series data includes: Using the high-frequency sampling time points of wearable devices as the reference time axis, the concentration sequences of intestinal inflammatory markers and the self-assessment symptom score sequences are projected onto the reference time axis through linear interpolation; For any missing time points that still exist after interpolation, a moving average of three time points before and after the interpolation is used to fill them; The padded sequences are zero-mean and unit variance scaled to obtain a standardized resampled sequence set.
4. The method for predicting the risk of depression in IBD patients based on artificial intelligence according to claim 1, characterized in that, The phase space reconstruction of each pair of sequences in the resampled sequence set includes: For each univariate sequence in each pair of sequences, the optimal delay time is calculated using the mutual information method, and the minimum embedding dimension is calculated using the spurious nearest neighbor method. Based on the optimal delay time and the minimum embedding dimension, each univariate sequence is mapped to an orbital matrix in a high-dimensional space; Align the two orbital matrices of the same sequence according to their time indices to obtain the orbital matrix pair.
5. The artificial intelligence-based method for predicting the risk of depression in IBD patients according to claim 1, characterized in that, The calculation of the cross recursive graph between each pair of orbital matrices includes: Set a recursion radius threshold, which is 10% of the mean of all pairwise Euclidean distances in the two orbital matrices; For each pair of time index states in the orbital matrix, if the Euclidean distance between them is less than the recursion radius threshold, then the corresponding position in the cross recursion graph is marked as a recursive point; otherwise, it is marked as a non-recursive point. Remove isolated points from the cross recursion graph, keeping only line segments consisting of at least two consecutive recursive points.
6. The method for predicting the risk of depression in IBD patients based on artificial intelligence according to claim 1, characterized in that, The extraction of recursion rate, longest diagonal length, and vertical recursion rate as coupling features includes: The recursion rate is calculated as the ratio of the total number of recursive points to the total number of grid points in the cross recursion graph; The longest diagonal length is determined by scanning all diagonal segments parallel to the main diagonal in the cross recursive graph and taking the longest segment length where the recursive point appears consecutively. The vertical recursion rate is calculated as the ratio of the total number of line segments formed by at least two consecutive recursive points in the vertical direction to the total number of vertical line segments.
7. The method for predicting the risk of depression in IBD patients based on artificial intelligence according to claim 1, characterized in that, The step of concatenating the coupled features from all sequence pairs into a multidimensional feature vector includes: The three coupling features of the intestinal inflammation marker concentration sequence and the heart rate variability sequence pair are arranged as the first group of features; The three coupling features of the intestinal inflammation marker concentration sequence and the body motion acceleration sequence pair are arranged as the second set of features; The three coupling features of the heart rate variability sequence and the body acceleration sequence are arranged into the third set of features; The features are horizontally concatenated into a multidimensional feature vector in the order of the first group of features, the second group of features, and the third group of features.
8. The method for predicting the risk of depression in IBD patients based on artificial intelligence according to claim 1, characterized in that, The longest diagonal length is determined by scanning all line segments parallel to the main diagonal in the cross recursive graph using a dynamic programming algorithm, and taking the longest consecutive occurrence of recursive points.
9. The artificial intelligence-based method for predicting the risk of depression in IBD patients according to claim 1, characterized in that, In the time alignment operation, the self-assessment symptom score sequence is projected onto the reference time axis using cubic spline interpolation.
10. The method for predicting the risk of depression in IBD patients based on artificial intelligence according to claim 1, characterized in that, In the missing value imputation process, segments with more than five consecutive missing time points are marked as invalid and the data for the corresponding time period is discarded.