A method for calculating and predicting the probability of occurrence of a certain complication using user characteristics

CN117672530BActive Publication Date: 2026-09-29ANDON HEALTH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311870329.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-31
Publication Date
2026-09-29
Estimated Expiration
2043-12-31

AI Technical Summary

Technical Problem

这种划分通常缺乏理论依据的支撑,造成最终结果不准确

Benefits of technology

[0045]1.在处理同类问题时,业内一般采用决策树模型对患者的病情进行预测,以判断患者是否有可能具有某种并发症。得到的结论一般都是“是”或“否”,但并不能给出量化的可能性百分比。本发明的预测方法能够对可能性进行量化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117672530B_ABST
    Figure CN117672530B_ABST
Patent Text Reader

Abstract

The application discloses a method for calculating and predicting the probability of occurrence of a certain complication by using user characteristics. It comprises: S1, data acquisition, S2, data filtering, S3, data resampling and transformation, S4, correlation calculation between data dimensions, S5, k-means clustering analysis, S6, data transformation into vectors and S7, calculation of complication occurrence probability. Specifically, the sampling data of the patient is arranged and filtered to obtain the data dimension, and the covariance calculation is carried out between the data dimension and various complications. Through the obtained results, the data dimension is clustered to obtain the relationship between the data dimension in the classified cluster and the complications. Finally, the disease probability is calculated by inputting the data of the current patient, the high-risk patients are early warned, and the health risks are avoided. The method can not only be used for predicting complications, but also can be popularized to any high-dimensional data set for predicting and evaluating the probability of occurrence of a certain key feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data and artificial intelligence technology, and in particular to a method for calculating and predicting the probability of a certain complication using user characteristics. Background Technology

[0002] Traditional treatments typically begin after symptoms appear, during which time the patient has already endured the suffering caused by the illness itself. If the risk of symptom onset could be predicted, intervention could be implemented in advance, alleviating the patient's suffering or even preventing the onset of symptoms. Therefore, the need for risk prediction is widespread in the medical field.

[0003] Currently, the most widely used prediction method in similar scenarios is decision tree-based. However, decision trees typically only handle discrete data. To use decision trees with continuous data, the range of continuous values ​​is usually divided into intervals. This division often lacks theoretical support, leading to inaccurate results. There are also methods that use clustering or other partitioning strategies for each dimension, but while these methods can improve the results, they are inaccurate when dealing with multiple correlated dimensions. Summary of the Invention

[0004] To address the problems existing in the prior art, this invention provides a highly accurate and interpretable method for calculating and predicting the probability of a certain complication using user characteristics.

[0005] Therefore, the present invention adopts the following technical solution:

[0006] A method for calculating and predicting the probability of a certain complication using user characteristics includes the following steps:

[0007] S1, Data Collection: Collect multi-dimensional data from multiple patients, including daily behavior data and various vital signs data;

[0008] The physical indicators include gender, region, age, height, weight, waist circumference, blood pressure, blood lipids, and blood glucose; the daily behavior data includes exercise time, sleep time, protein intake, carbohydrate intake, duration of participation in patient education, number of blood glucose tests, number of medications taken, and type of medication.

[0009] S2, Data Filtering: Filter the multi-dimensional data collected by S1, remove data with abnormal format and values, and store it.

[0010] S3, Data resampling and transformation:

[0011] S31. Arrange the multi-dimensional data obtained in S2 according to a uniform time interval. If there is no data for a certain dimension at the corresponding time point, select the data of this dimension that is closest to the corresponding time point and fill it into the time point. If the measurement results near the corresponding time point are more than the effective sampling range of the data of this corresponding dimension, then the value is empty. Remove the data of each dimension at the time point where the value is empty, and retain the multi-dimensional data after resampling.

[0012] S32, also treat complications as a data dimension. When sampling, if the complication occurs, take 1; otherwise, take 0.

[0013] S33, Data Transformation:

[0014] The data in the resampled multidimensional data are divided into qualitative data and quantitative data; the quantitative data directly participates in the calculation.

[0015] For simple qualitative data that are suitable for a binary distribution, use -1 or 1 as substitutes before participating in the calculation;

[0016] For textual data in qualitative data, it is converted into numerical data according to the degree of the textual description. Data that cannot be converted into numerical data is not included in the calculation. Specifically, the conversion into numerical data according to the degree of the textual description is as follows: severe is converted into 3, mild is converted into 1, and normal is converted into 0.

[0017] The quantitative data and the transformed qualitative data are denoted as data x;

[0018] S4, Calculation of correlations between data dimensions:

[0019] S41, For the data x obtained in S3, calculate its covariance matrix C′. The method for calculating the covariance matrix C′ for all the sampled data is as follows:

[0020]

[0021] Where, cov[x a x b ] = E[(x a -E[x a ])(x b -E[x b ])], a∈[1,n], b∈[1,n], E[x] represents the expectation of x.

[0022] When calculating the covariance matrix C′, items with empty values ​​are not included. When calculating the covariance matrix C′ for the aforementioned data dimension, data items with a covariance value less than 0.1 are filtered out. All items retained in the covariance matrix C′ are sorted in descending order of absolute value; items ranked higher indicate a stronger correlation between the corresponding two dimensions.

[0023] Normalize the data dimensions in the covariance matrix C′ to obtain the normalized complication covariance matrix C, and the terms C in the complication covariance matrix C. ij The calculation method is as follows:

[0024] C ij =cov[x i x j ] / |S[x i ]*S[x j ]|,

[0025] in, cov[x i x j Let ] be the covariance matrix C′, i be the index of the complication dimension in the data x, and j be the index of the other dimensions in the data x besides the complication dimension;

[0026] Let P i For the complication dimension corresponding to number i, Q j The dimension corresponding to number j;

[0027] S42, determine the data dimension Q of the non-zero term corresponding to the target complication I in the normalized complication covariance matrix C. j The set J;

[0028] S43, for any two data dimensions Q1 and Q2 in the set J, determine their results in the matrix C′. If the results are not 0, they are merged into a single dimension group d. m The subsets formed by grouping related dimensions are merged to obtain set D; if the result is 0, then the dimension that is unrelated to other dimensions and is only related to the target complication I is grouped as a separate dimension d'. m , all d' m It is also added to set D as a subset, where m is the subset number;

[0029] S5, k-meams clustering analysis:

[0030] For each subset in D obtained in S4, the k-meams clustering algorithm is used to cluster the data in the data dimensions contained in the subset. The resulting cluster set is denoted as Km. The k-meams cluster set determines different K categories for different subsets in D. For each subset, when the subset contains t data dimensions, the range of K is 2 to t.

[0031] S6, convert the data into a vector:

[0032] Using the clustering method in S5, the data for each patient in the data x is transformed into a vector (K1, K2, K3, ... K) consisting of classification sets. m ), and combine all vectors into a vector set;

[0033] S7. Calculate the probability of complications:

[0034] The probability of each classification set appearing in all vector sets described in S6 is denoted as P(K). m );

[0035] The probability of the occurrence of target complication I among multiple patients in S1 is denoted as P(I(1)), where the occurrence of target complication I is denoted as I(1);

[0036] When target complication I occurs, the probability of a certain classification set appearing is denoted as P(K). m |I(1));

[0037] All of the P(K) m P(I(1)) and P(K) m |I(1)) form a probability set;

[0038] For any new sample record, first convert it into a vector as described in S6, then use each category set in the vector to determine the corresponding probability in the probability set, and substitute it into the following formula to obtain the probability of having the target complication I:

[0039] P(I(1)|(K1,K2,K3…K m ))=P(K1|I(1))*P(K2|I(1))*P(K3|I(1))*…*P(K m |I(1))*P(I(1)) / (P(K1)*P(K2)*P(K3)*…*P(K m )).

[0040] Where: P(I(1)|(K1,K2,K3…K m ))=P(K1,K2,K3,…K m |I(1))*P(I(1)) / P(K1, K2, K3…K m), because (K1, K2, K3, ... K m They are independent of each other, therefore:

[0041] P(K1, K2, K3…K) m )=P(K1)*P(K2)*P(K3)*…*P(K m );

[0042] P(K1, K2, K3…K) m |I(1))=P(K1|I(1))*P(K2|I(1))*P(K3|I(1))*…*P(K m |I(1)).

[0043] For other target complications in the complication covariance matrix C, their occurrence probability is calculated through S5-S7.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] 1. In dealing with similar problems, the industry generally uses decision tree models to predict a patient's condition to determine whether the patient is likely to have a certain complication. The conclusions obtained are generally "yes" or "no," but they do not provide a quantifiable percentage of probability. The prediction method of this invention can quantify the probability.

[0046] 2. The prediction method of this invention uses a subset generated by the k-meams clustering algorithm that often corresponds to a common complication in reality, thus reflecting a trend of similar changes in certain dimensions of data; or some patients with common behavioral characteristics will develop the same complication, thereby making the entire model more interpretable. Existing decision tree methods do not have this property.

[0047] 3. For patients at high risk of developing the disease, the method of this invention can achieve early warning and intervention, avoiding health risks. This method can not only be used in scenarios of complication risk prediction, but also extended to datasets with any number of dimensions to predict and assess the probability of a key feature occurring. The method itself uses a large amount of data to find the probability of a certain dimension belonging to a certain value; therefore, it can be used for all scenarios that meet this data organization format. For example, in predicting shopping, it can be used to predict the probability of a customer of a certain age and gender who meets certain characteristics purchasing a certain product, or the probability of a customer purchasing another product after purchasing several other products. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of data resampling values ​​in this invention. Detailed Implementation

[0049] The method of the present invention will be described in detail below with reference to the embodiments and accompanying drawings.

[0050] The method of the present invention for calculating and predicting the probability of a certain complication using user characteristics includes the following steps:

[0051] S1, Data Acquisition:

[0052] Daily behavioral data and various vital signs data of multiple patients were collected over a period of time as multi-dimensional data. These dimensions include, but are not limited to, gender, age, height, weight, waist circumference, blood glucose measurement value, frequency, blood pressure, frequency of visits, medication, medical history, hobbies, conversation, diet, shopping, and various biochemical measurement indicators.

[0053] S2, Data Filtering:

[0054] The multi-dimensional data collected in S1 is filtered to remove data items whose format does not meet the record requirements. The records containing outliers are filtered by referring to the format and value description of each information. The processed data can be transferred to a file or stored in other databases.

[0055] S3, Data Resampling:

[0056] The data filtered by S2 is resampled. The reason for resampling is that the collection time and period for each data dimension are different; a collection might have collected some dimensions at one time, and then collected other dimensions after a period of time. Resampling ensures that each sampled record has reasonable values ​​across all dimensions.

[0057] Sampling methods such as Figure 1 As shown in the figure, the horizontal axis represents sampling time. Since the data collection times for different dimensions may not be the same, for a given sampling point, the data item for each dimension that is closest to the sampling point in time is selected. If the time of a certain dimension is beyond the valid range of the corresponding data item, then the data for that dimension is empty at that sampling point. The complication to be predicted is also considered as a dimension; during sampling, if the complication occurs, it is set to 1; otherwise, it is set to 0.

[0058] For example, to collect sampling data every three days, and to record data older than 24 hours from the sampling point as empty, the data to be collected would be as follows: blood glucose (measured before and after each meal); medication administration time recorded daily; medication dosage recorded for each course of treatment (7 days or more); presence of complications; and weight recorded daily. The first sampling point would then yield: blood glucose, medication administration time, weight, no complications, and an empty medication dosage record. After sampling, the empty records would be deleted, i.e., the medication dosage would be removed, resulting in a resampled sample.

[0059] The data includes qualitative and quantitative data. Qualitative data refers to textual data, such as region, gender, and type of medication; quantitative data refers to numerical data, such as age, height, blood sugar, and frequency of medication.

[0060] Numerical data in all quantitative data are directly used in the calculation;

[0061] For simple qualitative data that are suitable for a binary distribution, use -1 or 1 as substitutes before participating in the calculation;

[0062] For qualitative data consisting of textual descriptions, it is necessary to convert them into numerical values ​​according to the severity level described in the text. For example, if the textual descriptions are categorized as "severe," "mild," or "moderate," then they need to be converted into a numerical severity index, with severe corresponding to 3, mild to 1, and moderate to 0. Other textual data that cannot be converted into numerical values ​​are not included in the calculation.

[0063] The quantitative data and the transformed qualitative data are denoted as data x.

[0064] S4, Calculation of correlations between data dimensions:

[0065] For the data x obtained in S3, calculate its covariance matrix C′. The following is the method for calculating the covariance matrix, where C′ is represented as cov[x] a x b ]:

[0066] cov[x a x b ] = E[(x a -E[x a ])(x b -E[x b ])],

[0067] Here, 'a' and 'b' are the data dimension numbers. When calculating, ensure that items with empty values ​​are not included in the calculation, including when calculating the expected value E. When calculating the covariance of two dimensions, if either dimension in a record is empty, that record is not included in the calculation.

[0068] Normalize each term obtained in the covariance matrix C′ to obtain the normalized complication covariance matrix C. The following are the terms C in the complication covariance matrix C. ij Calculation method:

[0069] C ij =cov[x i x j ] / |S[x i ]*S[x j ]|,

[0070] Where: S[x] is half the size of the range of x, that is: and i represents the complication dimension number, and j represents the dimension numbers of various daily behavioral data and vital signs data other than complications.

[0071] Let P i For the complication dimension corresponding to number i, Q j For each dimension (j) other than the complication dimension, filter out data items in the complication covariance matrix C that are close to 0 (data less than 0.1 can be defined for filtering). Sort all remaining items by absolute value from largest to smallest. The earlier an item appears in the list, the stronger the correlation between the corresponding two dimensions.

[0072] Suppose we want to focus on target complication I, its related dimension set is J (elements Q in set J) j Satisfy C ij (Not 0). For any two dimensions Q1 and Q2 in set J, query their correlation in matrix C′. If the two dimensions are correlated (e.g., the corresponding item in C′ is not 0), then group them into a single dimension group d. m After merging the subsets formed by grouping related dimensions, we get a set of the form D = {{d1, d2, d3}, {d4, d5}, {d6, d7, d8}...}, where the dimensions in the same subset are related to each other.

[0073] Dimensions that are unrelated to other dimensions and are only related to complications are grouped as a separate dimension d' m , all d' m If a subset is also added to set D, then the set will take the form:

[0074] D={{d1, d2, d3}, {d4, d5}, {d6, d7, d8}, d'1, d'2, d'3...d' m …}}, where m is the subset number.

[0075] The purpose of dividing into multiple subsets is to ensure that the data dimensions within each subset of set D are related, while the data dimensions in different subsets are not related.

[0076] In short, the process involves first calculating the covariance matrix, then normalizing the covariance matrix, filtering out data items close to 0, sorting them from largest to smallest, selecting data dimensions related to the complication dimension to form a set, and then querying the correlation between any two data dimensions in the set in matrix C′. If the two dimensions are related, they are merged into a subset, resulting in a set D consisting of multiple subsets.

[0077] For example: to determine the set D, if C in the normalized matrix C... i1 C i2 C i3 If all values ​​are non-zero, it means that dimensions 1, 2, and 3 (j) are all related to the target complication. Then, we can examine C using the covariance matrix. 12 C 13 and C 23 The value of C. If C 12 Not equal to 0, C 13 and C 23 If both are 0, then 1 and 2 are merged into a subset d1, and 3 is another subset d'1 of set D;

[0078] S5, k-meams clustering analysis:

[0079] For each subset of D obtained in S4, the k-meams clustering algorithm (a well-known method, whose process and principle will not be discussed further) is used to cluster the data in the subset across its dimensions, and the clustering results are evaluated. The ultimate goal of clustering is to find a classification method that divides all data into K categories, denoted as K. m Under this classification method, the cohesion of each data type is high.

[0080] The k-value is determined using the silhouette coefficient. The silhouette coefficient is a well-known method for measuring the effectiveness of k-meams classification; a larger silhouette coefficient indicates better performance. The k-value measures the cohesion of a clustering result. As the value of k changes, the calculated silhouette coefficient also changes. When k changes small, the silhouette coefficient may change significantly. As k continues to change uniformly, the rate of change in the silhouette coefficient gradually decreases. The critical point at which k is found is taken as the clustering result. The value of k depends on the number of dimensions in the subsets and the correlation between them, so there is no specific range. Generally, when the number of dimensions in the subsets is t, the range of k can be considered to be 2 to t. The k-value corresponding to each dimension group is not necessarily the same; each dimension group needs to be determined individually using k-meams.

[0081] Therefore, in this method, we only need to find the k value that maximizes the contour coefficient within a certain range.

[0082] Cluster analysis is performed on several dimensions that have been identified as being associated with complications, ensuring that clustering along these dimensions is meaningful. Clustering uses several discrete values ​​to represent the distribution of all data along these dimensions. These discrete values ​​form the basis for calculating probabilities using Bayes' theorem.

[0083] S6, convert the data into a vector:

[0084] Using the clustering method in S5, each sample record is transformed into a vector (K1, K2, K3, ... K) consisting of classification sets. m ), and combine all vectors into a vector set;

[0085] For example: Suppose the sampled data is ultimately divided into two dimensional groups: d1 = {a, b, c} and d2 = {d, e}. Group d1 contains three dimensions: a, b, and c, while group d2 contains two dimensions: d and e. After S5 clustering, the data clusters along dimensions a, b, and c into two types, numbered K. a K b The two clusters along dimensions d and e form two types, numbered K respectively. d K e For a given sampled data, if its values ​​in the three dimensions a, b, and c belong to type K... a If the values ​​in dimensions d and e belong to type ke, then this sampled data is transformed into (K) a K e ).

[0086] S7, Calculate the probability of developing complications:

[0087] For example: For any sampled record V, the data can be represented in vector form as (K V1 K V2 K V3 …K Vm ).

[0088] When the sampling record is V, the probability of complication is denoted as P(I(1)|(K). V1 K V2 K V3 …K Vm )).

[0089] According to Bayes' theorem:

[0090] P(I(1)|(K V1 K V2 K V3 …K Vm ))=P(K V1 K V2 K V3 …K Vm |I(1))*P(I(1)) / P(K V1 K V2 ,

[0091] K V3 …K Vm ),

[0092] Because (K) V1 K V2 K V3 …K Vm ) is obtained from each subset of the set D, so (K V1 K V2 K V3 …K Vm )he

[0093] These are independent of each other, therefore:

[0094] P(K V1 K V2 K V3 …K Vm )=P(K v1 )*P(K V2 )*P(K V3 )*…*P(K Vm );

[0095] P(K V1 K V2 K V3 …K Vm |I(1))=P(K V1 |I(1))*P(K V2 |I(1))*P(K V3 |I(1))*…*P(K Vm |I(1)),

[0096] Where: P(K) m ) represents the probability that each category of all vectors appears in the vector set obtained in S6;

[0097] The probability of target complication I occurring in multiple patients of S1, wherein the occurrence of target complication I is denoted as I(1);

[0098] P(K m |I(1)) represents the probability of a certain category set appearing when the target complication I occurs.

[0099] P(I(1)) and P(K) m |I(1)) can all be calculated from the sampled data.

[0100] When there is enough data, P(K) can be established. m ), P(I(1)) and P(K) m The probability set of |I(1)) for newly acquired sampled data

[0101] Find the corresponding K m Find the corresponding P(K) m |I(1)) and P(Km The final result is calculated using the following method: P(I(1)|(K1, K2, K3…K) m ))=P(K1|I(1))*P(K2|I(1))*P(K3|I(1))*…*P(K m |I(1))*P(I(1)) /

[0102] (P(K1)*P(K2)*P(K3)*…*P(K m )).

[0103] For any sampled record, first convert it into vector form as described in S6, and then substitute this vector into the method described in S7.

[0104] The formula is used to calculate the probability.

[0105] Example 1

[0106] The following is the data after collection, filtering, resampling, and transformation:

[0107]

[0108] First, calculate the covariance matrix C′:

[0109]

[0110] Suppose that the target complication we are interested in is hypertension. Following the grouping method described in S4, we group the other dimensions into two subsets: {number of measurements per week} and {blood glucose measurements, whether or not diabetes is present}.

[0111] Then, k-meams clustering analysis was performed on the two subsets respectively. It can be seen that {number of measurements per week} has only one dimension, with values ​​of 4, 6, 3, and 1. Assuming that the clustering results are divided into two classes, less than 3 times are classified into class K1 and more than 3 times are classified into class K2;

[0112] For {blood glucose measurement value, presence of diabetes}, after clustering, those with diabetes whose blood glucose measurement value is higher than 7 belong to class K3, and those without diabetes whose blood glucose measurement value is lower than 7 belong to class K4. The original sampling data is then transformed into the following form:

[0113] Patient 1 K2 K3 1 Patient 2 K2 K4 1 Patient 3 K2 K3 1 Patient 4 K1 K4 -1

[0114] The data in the table shows that P(I(1)) = 75%.

[0115] Calculate each classification set K m The probability of occurrence, P(Km), is:

[0116] P(K1) = 25%;

[0117] P(K2) = 75%;

[0118] P(K3) = 50%;

[0119] P(K4) = 50%.

[0120] When the hypertension dimension is 1, the probability of each K category appearing, i.e., P(Km|I(1)), is:

[0121] P(K1|I(1))=0%;

[0122] P(K2|I(1)) = 100%;

[0123] P(K3|I(1)) = 67%;

[0124] P(K4|I(1))=33%.

[0125] Now suppose we have a patient's sample data as follows:

[0126] 5 6.3 -1

[0127] First, transform this data according to the dimensional grouping and clustering results:

[0128] K2 K4

[0129] According to the calculation formula, the probability that this patient has hypertension is:

[0130] P(I(1)|(K2,k4))=P(K2|I(1))*P(K4|I(1))*P(I(1)) / (P(K2)*P(K4)), Substituting the previously calculated probability into the result, we get 100%*33%*75% / (75%*50%)=66%.

[0131] Example 2

[0132] To verify the effectiveness of the method, multiple validation sets were selected, and both the proposed method and the decision tree approach were used simultaneously to predict whether patients had a certain complication. Since decision trees can only provide "yes" or "no" outputs and cannot quantify the result, while the output of this invention is a probability, the output of the method described in this paper was modified for comparison. It was agreed that when the output probability was higher than 50%, the output would be "yes," otherwise "no." The effectiveness of the two methods was compared by examining the proportion of correctly judged data in the validation sets.

[0133] 1 92.27% 90.16% 2 76.33% 71.26% 3 89.30% 88.31% mean 85.96% 83.24%

[0134] The table above shows a comparison of the averaging results from multiple predictions. The final results may vary depending on the selected test samples and dimensions. However, in most cases, this method achieves higher accuracy than decision trees.

[0135] Decision trees typically only handle discrete data. When comparing data using decision trees, they divide the data into intervals within the range of continuous values. However, theoretically, this division doesn't consider the actual distribution of the data. This is a limitation of decision trees. The k-meams clustering method used in this paper can improve upon this issue.

Claims

1. A method for calculating and predicting the probability of a certain complication using user characteristics, characterized in that, Includes the following steps: S1, Data Collection: Collect multi-dimensional data from multiple patients, including daily behavior data and various vital signs data; S2, Data Filtering: Filter the multi-dimensional data collected by S1, remove data with abnormal format and values, and store it. S3, Data resampling and transformation: S31. Arrange the multi-dimensional data obtained in S2 according to a uniform time interval. If there is no data for a certain dimension at the corresponding time point, select the data of this dimension that is closest to the corresponding time point and fill it into the time point. If the measurement results near the corresponding time point are more than the effective sampling range of the data of this corresponding dimension, then the value is empty. Remove the data of each dimension at the time point where the value is empty, and retain the multi-dimensional data after resampling. S32, also treat complications as a data dimension. When sampling, if the complication occurs, take 1; otherwise, take 0. S33, Data Transformation: The data in the resampled multidimensional data are divided into qualitative data and quantitative data; the quantitative data directly participates in the calculation. For simple qualitative data that are suitable for a binary distribution, use -1 or 1 as substitutes before participating in the calculation; For textual data in qualitative data, convert it into numerical data according to the degree of textual expression; data that cannot be converted into numerical data is not included in the calculation. The quantitative data and the transformed qualitative data are denoted as data x; S4, Calculation of correlations between data dimensions: S41, For the data x obtained in S3, calculate its covariance matrix C′, normalize the data dimensions in the covariance matrix C′ to obtain the normalized complication covariance matrix C, and the terms C in the complication covariance matrix C. ij The calculation method is as follows: C ij =cov[x i ,x j ] / |S[x i ]*S[x j ]|, in, cov[x i x j Let ] be the covariance matrix C′, i be the index of the complication dimension in the data x, and j be the index of the other dimensions in the data x besides the complication dimension; Let P i For the complication dimension corresponding to number i, Q j The dimension corresponding to number j; S42, determine the data dimension Q of the non-zero term corresponding to the target complication I in the normalized complication covariance matrix C. j The set J; S43, for any two data dimensions Q1 and Q2 in the set J, determine their results in the matrix C′. If the results are not 0, they are merged into a single dimension group d. m The subsets formed by grouping related dimensions are merged to obtain set D; if the result is 0, then the dimension that is unrelated to other dimensions and is only related to the target complication I is grouped as a separate dimension d'. m , all d' m It is also added to set D as a subset, where m is the subset number; S5, k-meams clustering analysis: For each subset of D obtained in S4, the k-meams clustering algorithm is used to cluster the data in the data dimension contained in the subset. The resulting clusters are denoted as K. m ; S6, convert the data into a vector: Using the clustering method in S5, the data for each patient in the data x is transformed into a vector (K1, K2, K3, ... K) consisting of classification sets. m ), and combine all vectors into a vector set; S7, Calculate the probability of complications: The probability of each classification set appearing in all vector sets described in S6 is denoted as P(K). m ); The probability of the occurrence of target complication I among multiple patients in S1 is denoted as P(I(1)), where the occurrence of target complication I is denoted as I(1); When target complication I occurs, the probability of a certain classification set appearing is denoted as P(K). m |I(1)); All of the P(K) m P(I(1)) and P(K) m |I(1)) form a probability set; For any new sample record, first convert it into a vector as described in S6, then use each category set in the vector to determine the corresponding probability in the probability set, and substitute it into the following formula to obtain the probability of having the target complication I: P(I(1)|(K1,K2,K3…K m ))=P(K1|I(1))*P(K2|I(1))*P(K3|I(1))*…*P(K m |I(1))*P(I(1)) / (P(K1)*P(K2)*P(K3)*…*P(K m ))。 2. The method according to claim 1, characterized in that, The specific conversion of the degree of severity described in S33 into numerical data is as follows: severe is converted to 3, mild is converted to 1, and normal is converted to 0.

3. The method according to claim 1, characterized in that, The method for calculating the covariance matrix C′ for all the sampled data in step S4 is as follows: Where, cov[x a x b ] = E[(x a -E[x a ])(x b -E[x b ])], a∈[1,n], b∈[1,n], E[x] represents the expectation of x.

4. The method according to claim 1, characterized in that: In S4, terms with empty values ​​are not included when calculating the covariance matrix C′.

5. The method according to claim 1, characterized in that: In S41, when calculating the covariance matrix C′ of the data dimension, data items with a value less than 0.1 in the covariance matrix are filtered out.

6. The method according to claim 5, characterized in that: For all the items retained in the covariance matrix C′, sort them from largest to smallest absolute value. The earlier the item appears in the sorted list, the greater the correlation between the corresponding two dimensions.

7. The method according to claim 1, characterized in that: In S5, the k-meams clustering set determines different K categories for different subsets of D. For each subset, when the subset contains t data dimensions, K ranges from 2 to t.

8. The method according to claim 1, characterized in that: In S7, P(I(1)|(K1, K2, K3…K m ))=P(K1,K2,K3,…K m |I(1))*P(I(1)) / P(K1, K2, K3…K m ), because (K1, K2, K3, ... K m They are independent of each other, therefore: P(K1,K2,K3…K m )=P(K1)*P(K2)*P(K3)*…*P(K m ); P(K1,K2,K3…K m |I(1))=P(K1|I(1))*P(K2|I(1))*P(K3|I(1))*…*P(K m |I(1))。 9. The method according to any one of claims 1-8, characterized in that, The vital signs data mentioned in S1 include gender, region, age, height, weight, waist circumference, blood pressure, blood lipids, and blood glucose; the daily behavior data include exercise time, sleep time, protein intake, carbohydrate intake, duration of participation in patient education, number of blood glucose tests, number of medications taken, and type of medication.

10. The method according to claim 1, characterized in that: For other target complications in the complication covariance matrix C, their occurrence probability is calculated through S5-S7.

Citation Information

Patent Citations

  • Multi-granularity breast cancer gene classification method based on dual adaptive neighborhood radius

    CN113838532A

  • Diabetes risk early warning system

    US20220301708A1