Big data-based senile disease risk prediction method and system

By combining user information and clinical data, using machine learning technology to predict the risk of geriatric diseases, the problem of insufficient prediction accuracy in the existing technology is solved, and more accurate prevention suggestions and higher prediction accuracy are achieved.

CN120199463APending Publication Date: 2025-06-24THE FIRST AFFILIATED HOSPITAL OF ZHENGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510270453.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art fails to effectively combine user symptoms and clinical data in disease prediction, resulting in insufficient prediction accuracy and inability to provide accurate prevention suggestions.

Method used

By obtaining user information, extracting user feature vectors and entering a model to obtain prediction results, derive the prediction process diagram to generate a target feature map, integrating the target feature map of cross feature parameters, inputting the clinical prediction model to obtain predicted clinical information, and adjusting the model accuracy based on this information.

Benefits of technology

It significantly improves the prediction accuracy of geriatric disease risk, provides reliable basis for clinical decision-making, and provides targeted prevention suggestions through transparent prediction processes and feature maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199463A_ABST
    Figure CN120199463A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, in particular to a big data-based senile disease risk prediction method and system. The method comprises the following steps: acquiring user information, extracting a user feature vector, inputting the user feature vector into an ith model, acquiring an ith prediction result containing a disease type and grade, exporting a prediction process chart from the ith model, and generating a target feature chart; integrating the target feature maps with the cross feature parameters and the contribution values greater than a set threshold value to obtain integrated feature maps, and inputting each integrated feature map and the unintegrated target feature maps into a clinical prediction model to obtain predicted clinical information; acquiring predicted disease information based on the predicted clinical information; and comparing with a prediction result, and adjusting the model precision. According to the method, the accuracy of senile disease risk prediction can be improved, and powerful support is provided for clinical decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and particularly to a method and system for predicting the risk of geriatric diseases based on big data. Background Art

[0002] In the prior art, a variety of disease prediction methods based on machine learning have been proposed. Similar prior arts include a Chinese patent with the publication number CN116052867A, which proposes a disease prediction method, device, equipment, and storage medium. The method includes: obtaining the current symptoms of a patient; determining at least one first historical symptom, where the first historical symptom is a historical symptom similar to the current symptom; determining the first disease probability distribution corresponding to the current symptom according to at least one first historical symptom; and predicting the disease suffered by the patient according to the first disease probability distribution, thereby improving the accuracy of disease prediction. In addition, a similar prior art is a European patent with the publication number EP3164824B1, which provides a method for predicting the risk of a disease occurring in a subject, including: a) determining the corresponding values of a plurality of disease risk factors of the subject; b) providing a database of individuals for whom the values of a plurality of disease risk factors have been determined and whether the disease has occurred or not is known; c) recording each value of the plurality of disease risk factors of the subject and the database individuals on the same disease incidence scale; d) selecting a plurality of individuals in the database that are at the lowest Euclidean distance from the subject relative to other individuals in the database, where the Euclidean distance is based on the recoded values of the plurality of disease risk factors; this method can be easily understood by clinicians and can easily incorporate new risk factors. Although the above two patent documents both solve the problem of disease prediction, they only rely on the symptoms or clinical data of users, without combining the two to more accurately predict the disease risk, and cannot give precise prevention suggestions based on the prediction basis. Summary of the Invention

[0003] The present invention provides a method for predicting the risk of geriatric diseases based on big data, and the method includes:

[0004] Obtaining user information, extracting the user feature vector corresponding to the i-th model from the user information, inputting the user feature vector into the i-th model, and obtaining the i-th prediction result, where the i-th prediction result includes the i-th disease type and the i-th disease risk value;

[0005] Deriving, by an extraction unit, a prediction process diagram corresponding to the i-th prediction result from the i-th model, and generating a target feature diagram corresponding to the i-th disease based on the prediction process diagram;

[0006] Integrate at least two of the target feature maps with cross-feature parameters and a contribution value of the cross-feature parameters greater than a set threshold to obtain an integrated feature map, input each integrated feature map and the unintegrated target feature maps into a clinical prediction model, and obtain corresponding predicted clinical information;

[0007] Based on each predicted clinical information and a disease database, a prediction disease information is obtained by an analysis unit;

[0008] Adjust the accuracy of the i-th model and the clinical prediction model based on the prediction disease information, the predicted clinical information, and the corresponding i-th prediction result.

[0009] As a preferred technical solution of the present invention, the training of the i-th model includes:

[0010] Read from a database and use the historical data corresponding to the i-th geriatric disease as the i-th target data group, extract N features corresponding to the i-th geriatric disease from the i-th target data group, and obtain multiple feature vectors corresponding to the i-th target data group based on the N features;

[0011] Select M of the feature vectors as initial centroids according to the average distribution positions of the multiple feature vectors, divide each other feature vector into the same class group as the centroid feature vector with the closest distance, obtain the average centroid of the class group, calculate the distance between the initial centroid and the average centroid, and repeat this step until the distance is less than or equal to a set distance;

[0012] Take the class group closest to the center position of the distribution area of all the feature vectors as the target class group, randomly select K of the feature vectors in the target class group as target feature vectors, sequentially select other target feature vectors from adjacent class groups, and use the target feature vectors and the other target feature vectors as training data to train the i-th model, where the i + 1 other target feature vector with the farthest distance from the i-th other target feature vector is sequentially selected from the adjacent class group of the class group where the i-th other target feature vector is located, and the first other target feature vector is the feature vector with the farthest distance from the target feature vector in the adjacent class group of the target class group.

[0013] As a preferred technical solution of the present invention, extracting the target feature map corresponding to the i-th disease based on the prediction process diagram further includes:

[0014] The prediction process diagram for obtaining the i-th prediction result based on the user information is derived from the i-th model by the extraction unit, the weight value corresponding to each feature parameter in the user information and the product of the weight value and the corresponding feature parameter value are obtained from the prediction process diagram, the feature parameters and the corresponding weight values with the product greater than a set value are extracted from the prediction process diagram, the feature parameters are used as target features, the product is used as the contribution value, and the target feature diagram is generated based on the target features and the corresponding contribution values.

[0015] As a preferred technical solution of the present invention, the acquisition of the clinical information includes:

[0016] Any two of the target feature diagrams corresponding to the i-th disease are compared, and multiple target feature diagrams with intersecting feature parameters and the contribution values of the intersecting feature parameters all greater than a set threshold are integrated to obtain an integrated feature diagram. Each integrated feature diagram and each un-integrated target feature diagram are input into the clinical prediction model to obtain the predicted clinical information, where the predicted clinical information includes predicted clinical symptom information and the risk value of the predicted clinical symptom.

[0017] As a preferred technical solution of the present invention, the acquisition of the predicted disease information includes:

[0018] The analysis unit compares the predicted clinical symptoms in the clinical information with the symptoms corresponding to each elderly disease type in the disease database, and uses the disease information of the elderly disease type closest to the predicted clinical symptoms as the predicted disease type.

[0019] As a preferred technical solution of the present invention, adjusting the accuracy of the i-th model and the clinical prediction model based on the predicted disease information, the predicted clinical information and the corresponding i-th prediction result includes:

[0020] Obtain the predicted disease type based on the predicted disease information, obtain the risk values of each predicted clinical symptom corresponding to the predicted disease type for the user based on the predicted clinical information, perform weighting, and use the weighted result as the predicted risk value of the predicted disease type. Compare the predicted disease type with the i-th disease type in the i-th prediction result to obtain a comparison result. Also obtain the difference between the predicted risk value and the i-th disease risk value. When the comparison result is that the predicted disease type is the same as the i-th disease type and the difference is less than or equal to a preset value, it is normal. Otherwise, increase the first sample data corresponding to the i-th disease type and the predicted disease type in the i-th model, and at the same time increase the second sample data corresponding to the i-th disease type and the predicted disease type in the clinical prediction model. Then, perform secondary training on the i-th model and the clinical prediction model based on the first sample data and the second sample data respectively. Among them, the second sample data includes the historical target feature map and symptom information corresponding to the i-th disease type and the predicted disease type.

[0021] As a preferred technical solution of the present invention, the training of the clinical prediction model includes:

[0022] Obtain patient sample data from the database. The patient sample data includes historical feature maps and corresponding symptom information. Among them, the historical feature maps include the contribution values of each feature parameter, and train the clinical prediction model based on the sample data.

[0023] As a preferred technical solution of the present invention, the value range of i is a positive integer greater than or equal to 1.

[0024] The present invention also provides a geriatric disease risk prediction system based on big data for implementing the above method. The system includes:

[0025] A first prediction unit for obtaining user information, extracting the user feature vector corresponding to the i-th model from the user information, inputting the user feature vector into the i-th model, and obtaining the i-th prediction result. The i-th prediction result includes the i-th disease type and the i-th disease risk value;

[0026] A generation unit for deriving a prediction process diagram corresponding to the i-th prediction result from the i-th model through an extraction unit, and generating a target feature map corresponding to the i-th disease based on the prediction process diagram;

[0027] A second prediction unit, configured to integrate at least two of the target feature maps with cross - feature parameters and whose contribution values of the cross - feature parameters are greater than a set threshold to obtain an integrated feature map, input each of the integrated feature maps and the un - integrated target feature maps into a clinical prediction model, and obtain corresponding predicted clinical information;

[0028] An analysis unit, configured to obtain predicted disease information based on each of the predicted clinical information;

[0029] An adjustment unit, configured to adjust the accuracy of the i - th model and the clinical prediction model based on the predicted disease information, the predicted clinical information, and the corresponding i - th prediction result.

[0030] The present invention also provides a computer - readable storage medium, on which instructions are stored, and when the instructions are executed by a processor, the above - mentioned method is implemented.

[0031] Effect

[0032] By combining multi - dimensional features in user information, the present invention uses big data and machine learning technologies to accurately predict the risk of geriatric diseases. Through the prediction and mutual verification of a dual - model (the i - th model and the clinical prediction model), the accuracy of the prediction is significantly improved, providing a reliable basis for clinical decision - making. At the same time, by extracting representative training data of the i - th model to improve the accuracy of the i - th model, and by exporting the prediction process diagram and generating target feature maps, the transparency of the prediction process of the i - th model is achieved. Through the logic and basis in the prediction process of the i - th model, users and medical staff can more clearly understand the prediction results and provide more targeted prevention suggestions based on the feature maps. Also, based on the cross - feature parameters, the above - mentioned target feature maps are integrated to obtain integrated feature maps, and predicted clinical information is obtained based on the integrated feature maps and the un - integrated target feature maps, and predicted disease information is obtained based on the predicted clinical information. Further, based on the predicted disease information, the predicted clinical information, and the i - th prediction result, the accuracy of the model can be dynamically adjusted. When there are deviations among the predicted disease information, the predicted clinical information, and the i - th prediction result, relevant sample data are automatically added, and the i - th model and the clinical prediction model are retrained, thereby continuously optimizing the prediction ability of the model. By providing detailed prediction processes and result explanations, prevention suggestions can be provided according to the prediction results, which helps users take timely measures to reduce the risk of disease and improve the overall health level. Through the mutual cooperation of the above - mentioned technical solutions, the risk of geriatric diseases can be improved, providing users with more accurate and reliable prediction services, and at the same time promoting the rational allocation and utilization of medical resources. Description of the Drawings

[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 Flowchart of the method for predicting the risk of geriatric diseases based on big data in the embodiments of the present invention;

[0035] Figure 2 Flowchart of the training method for the i-th model in the embodiments of the present invention;

[0036] Figure 3 Flowchart of the accuracy adjustment method for the i-th model and the clinical prediction model in the present invention;

[0037] Figure 4 Structure diagram of the system for predicting the risk of geriatric diseases based on big data in the embodiments of the present invention. Detailed implementation manners

[0038] The embodiments of the present invention provide a method and system for predicting the risk of geriatric diseases based on big data. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0039] For ease of understanding, the following describes the specific process of the embodiments of the present invention. As Figure 1 shown, the method for predicting the risk of geriatric diseases based on big data in the embodiments of the present invention includes:

[0040] Step S1: Obtain user information, extract the user feature vector corresponding to the i-th model from the user information, input the user feature vector into the i-th model, and obtain the i-th prediction result, where the i-th prediction result includes the i-th disease type and the i-th disease risk value;

[0041] Specifically, multi-dimensional user information is collected from users, including basic information such as age, gender, medical history, lifestyle habits, height, and weight, diagnosis and treatment time-series data such as examination results and treatment records at different time points, and biometric data such as blood pressure, blood sugar, and cholesterol levels. These data are processed through standardization and denoising to ensure the input quality. According to the characteristics of the i-th geriatric disease such as diabetes and hypertension, N features strongly related to the disease, such as blood sugar value and insulin resistance index, are screened from the preprocessed user information. Through feature engineering such as normalization and principal component analysis, these features are transformed into numerical vectors to form a user feature vector. The user feature vector is input into the i-th model, such as a random forest model, to output the i-th prediction result. The i-th prediction result includes the i-th disease type and the i-th disease risk value. For example, the type of geriatric disease a user may have is diabetes, and the probability of getting the disease is 85%. Through the above technical solution, the disease risk value of each geriatric disease of the patient can be obtained.

[0042] Step S2: A prediction process diagram corresponding to the i-th prediction result is derived from the i-th model through an extraction unit, and a target feature diagram corresponding to the i-th disease is generated based on the prediction process diagram;

[0043] Specifically, a visualization chart of the prediction process, such as a decision tree path or a feature importance diagram, is extracted from the i-th model to show the weight of each feature parameter on the prediction result. For example, the weight of the blood sugar value is 0.3, and the weight of age is 0.2. Features whose contribution value, that is, the product of the weight and the eigenvalue of the corresponding feature parameter, is greater than a set threshold are extracted from the prediction process diagram. These features and their contribution values are sorted by priority to generate a target feature diagram. Through the above technical solution, the key pathogenic factors and their influencing degrees can be intuitively displayed, facilitating the formulation of prevention and treatment measures, and at the same time laying a foundation for further obtaining the predicted clinical information corresponding to the i-th disease type.

[0044] Step S3: At least two of the target feature diagrams with cross feature parameters and the contribution value of the cross feature parameters greater than a set threshold are integrated to obtain an integrated feature diagram. Each integrated feature diagram and the unintegrated target feature diagram are input into a clinical prediction model, and corresponding predicted clinical information is obtained;

[0045] Specifically, by comparing the target feature diagrams of different i-th diseases, cross feature parameters will appear, such as chronic inflammation markers affecting both Alzheimer's disease and cardiovascular diseases, and their contribution values both exceed the set threshold. The target feature diagrams of the cross feature parameters are merged into an integrated feature diagram to capture the key factors common to multiple diseases. The integrated feature diagram and the unintegrated independent target feature diagrams are input into a clinical prediction model, and predicted clinical information is obtained. Through the above technical solution, the predicted clinical information of the user can be obtained, making

[0046] Step S4: The analysis unit obtains predicted disease information based on each of the predicted clinical information and the disease database;

[0047] Step S5: Adjust the accuracy of the i-th model and the clinical prediction model based on the predicted disease information, the predicted clinical information, and the corresponding i-th prediction result.

[0048] Further, the training of the i-th model, as Figure 2 shown, includes:

[0049] Read from the database and use the historical data corresponding to the i-th geriatric disease as the i-th target data group, extract N features corresponding to the i-th geriatric disease from the i-th target data group, and obtain multiple feature vectors corresponding to the i-th target data group based on the N features. Among them, the historical data includes the basic information of the patient, the diagnosis and treatment time series data, and the corresponding geriatric disease type;

[0050] Average select M of the feature vectors as the initial centroids according to the distribution positions of the multiple feature vectors, divide each other feature vector into the same class group as the centroid feature vector with the closest distance, obtain the average centroid of the class group, calculate the distance between the initial centroid and the average centroid, and repeat this step until the distance is less than or equal to the set distance;

[0051] Take the class group closest to the central position of all the feature vectors as the target class group, randomly select K of the feature vectors in the target class group as the target feature vectors, select other target feature vectors corresponding to each of the target feature vectors and farthest from the target feature vectors in other class groups, and use the target feature vectors and the other target feature vectors as training data to train the i-th model.

[0052] Specifically, historical data of geriatric patients are obtained from a database. Since there is more than one geriatric disease, such as hypertension, diabetes, Alzheimer's disease, geriatric nephropathy, etc., the historical data of geriatric patients need to be grouped and divided according to the types of geriatric diseases. Among them, each group corresponds to a disease type. The historical data includes the basic information of the patient, the chronological diagnosis and treatment data, and the corresponding geriatric disease type. The above-mentioned patient basic information includes at least the patient's gender, age, medical history, living habits, height, and weight. The above-mentioned patient chronological data includes the patient's diagnosis and treatment data collected at different time points. The above-mentioned chronological diagnosis and treatment data includes at least the examination result data (biometric data) and the treatment record data. The multiple above-mentioned historical data corresponding to the i-th geriatric disease are used as the i-th target data group, and the correlation between each category of data in the i-th target data group and the i-th geriatric disease is analyzed. Based on the above-mentioned correlation, N features are extracted from the i-th target data group. The value of N is a positive integer greater than or equal to 1. And based on the above-mentioned N features, the corresponding feature vectors are extracted from each historical data in the i-th target data group. Since the amount of data in the i-th target data group is large, when performing machine learning based on the i-th target data group, if the sample data, that is, the historical data in the i-th target data group, is not screened, it will cause a large consumption of resources in the learning process and even risk of noise amplification, thus reducing the model accuracy. Therefore, the distribution positions of each above-mentioned feature vector are evenly divided into M distribution regions. The value of M is a positive integer greater than or equal to 1. And the feature vector at the center position of each distribution region is used as the initial centroid. And each other feature vector is divided into the same group as the feature vector corresponding to the nearest centroid, that is, the above-mentioned centroid feature vector. And the average value of the feature vectors within the same group, that is, the average centroid, is calculated. The distance between the above-mentioned initial centroid and the above-mentioned average centroid is also calculated. When the above-mentioned distance is greater than or equal to the set distance, the above-mentioned average centroid is used as the above-mentioned initial centroid, and this step is repeated until the above-mentioned distance is less than the set distance, that is, the group division is completed. In order to be able to select training data with higher quality, the group closest to the center positions of all the above-mentioned feature vector distribution regions, that is, the target group, is selected. And K above-mentioned feature vectors are randomly selected from the above-mentioned target group as the above-mentioned target feature vectors. K is greater than 1. And the first other target feature vectors that are the farthest from each above-mentioned target feature vector are sequentially selected from the adjacent groups of the above-mentioned target group. And the second target feature vectors that are the farthest from the above-mentioned first target feature vector are selected from the adjacent groups of the group where the first other target feature vector is located. Similarly, the i-th other target feature vectors are obtained. And the above-mentioned target feature vectors and the above-mentioned other target feature vectors are used as the above-mentioned training data to train the i-th model. Through the above technical solution, relatively comprehensive and classical training data for the i-th model can be obtained, thereby improving the training efficiency and accuracy of the i-th model.

[0053] Further, extracting a target feature map corresponding to the i-th disease based on the prediction process map further includes:

[0054] The extraction unit derives a prediction process map for obtaining the i-th prediction result based on the user feature vector from the i-th model, obtains the weight value corresponding to each feature parameter in the user feature vector and the product of the weight value and the corresponding feature parameter value from the prediction process map, extracts the feature parameters and the corresponding weight values whose products are greater than a set value from the prediction process map, takes the feature parameters as target features, takes the products as contribution values, and generates the target feature map based on the target features and the corresponding contribution values.

[0055] Specifically, when predicting the risk of geriatric diseases of a user through the above-mentioned i-th model, the entire prediction process is opaque to the user and medical staff. The user and medical staff cannot know the entire prediction process, and the medical staff cannot give reasonable preventive suggestions based on the prediction results. Therefore, the extraction unit derives a prediction process map for obtaining the i-th prediction result based on the above-mentioned user information from the above-mentioned i-th model. The prediction process map includes the weight value corresponding to each feature parameter value in the user feature vector and the product of the feature parameter value and the corresponding weight value. The product reflects the contribution of the feature parameter value to the prediction result. Also, from the prediction process map, feature parameters with relatively large corresponding contribution values are selected as the above-mentioned target features, and the above-mentioned target features and the corresponding above-mentioned contribution values are used to generate the above-mentioned target feature map corresponding to the i-th model, that is, the user information corresponding to the i-th disease. Through the above technical solution, the user can understand the reasons for the prediction result, which is convenient for judging and giving preventive suggestions based on the user's target feature map, and also lays a foundation for obtaining corresponding disease information based on the above-mentioned target feature map.

[0056] Further, the acquisition of the clinical information includes:

[0057] Compare any two target feature maps corresponding to the i-th disease, integrate multiple target feature maps with intersecting feature parameters and the contribution values of the intersecting feature parameters all greater than a set threshold to obtain an integrated feature map, input each integrated feature map and each unintegrated target feature map into the clinical prediction model, and obtain the predicted clinical information, where the predicted clinical information includes predicted clinical symptom information and the risk value of the predicted clinical symptoms.

[0058] Specifically, since there are intersections and relatively large contribution values among the characteristic parameters of different geriatric diseases, that is, the same characteristic parameter may contribute to multiple geriatric diseases. Therefore, when obtaining the corresponding predicted clinical information from the above-mentioned target feature maps, two or more of the above-mentioned target feature maps with intersecting characteristic parameters may also overlap in clinical symptoms. For example, chronic inflammation markers are related to both Alzheimer's disease and cardiovascular diseases. Therefore, in order to improve the accuracy of clinical symptoms, any two of the above-mentioned target feature maps of the i-th disease pair are compared, at least two of the target feature maps with intersecting characteristic parameters are integrated, and the integrated images corresponding to various diseases are obtained. Each of the above-mentioned integrated images and the target feature images that do not intersect with other target feature images are respectively input into the above-mentioned clinical prediction model to obtain the corresponding predicted clinical information, that is, the above-mentioned clinical symptom information. Through the above technical solution, the disease prediction information of the user can be accurately obtained, facilitating medical staff to put forward reasonable prevention suggestions.

[0059] Further, the obtaining of the predicted disease information includes:

[0060] Analyze through the analysis unit and compare the predicted clinical symptoms in the predicted clinical information with the symptoms corresponding to each geriatric disease type in the disease database, and use the disease information of the geriatric disease type closest to the predicted clinical symptoms as the predicted disease information.

[0061] Specifically, analyze the predicted clinical symptoms in the above-mentioned predicted clinical information through the above-mentioned analysis unit, compare the predicted clinical symptoms with the symptoms corresponding to each geriatric disease type in the above-mentioned disease database, and use the disease information corresponding to the geriatric disease type closest to the above-mentioned predicted clinical symptoms as the above-mentioned predicted disease information. Among them, the above-mentioned predicted clinical symptoms correspond to at least one geriatric disease type. Through the above technical solution, the specific geriatric disease type and disease information corresponding to the above-mentioned predicted clinical symptoms can be obtained, providing more disease prediction information for patients and medical staff, and at the same time laying a foundation for further improving the accuracy of the above-mentioned i-th model and the above-mentioned clinical prediction model.

[0062] Further, adjust the accuracy of the i-th model and the clinical prediction model based on the predicted disease information, the predicted clinical information, and the corresponding i-th prediction result, as Figure 3 shown, including:

[0063] Obtain the predicted disease type based on the predicted disease information, weight the risk values of each predicted clinical symptom corresponding to the predicted disease type for the user based on the predicted clinical information, and use the weighted result as the predicted risk value of the predicted disease type. Compare the predicted disease type with the i-th disease type in the i-th predicted result to obtain a comparison result. Also obtain the difference between the predicted risk value and the i-th disease risk value. When the comparison result indicates that the predicted disease type is the same as the i-th disease type and the difference is less than or equal to a preset value, it is normal; otherwise, increase the first sample data corresponding to the i-th disease type and the predicted disease type in the i-th model, and at the same time increase the second sample data corresponding to the i-th disease type and the predicted disease type in the clinical prediction model. Then, perform secondary training on the i-th model and the clinical prediction model respectively based on the first sample data and the second sample data. Among them, the second sample data includes the historical target feature maps and symptom information corresponding to the i-th disease type and the predicted disease type.

[0064] Specifically, the above-mentioned predicted disease type is obtained from the above-mentioned predicted disease information. When the predicted clinical symptoms corresponding to the above-mentioned predicted disease information are obtained by integrating the feature maps, the predicted disease types corresponding to the above-mentioned predicted disease information may be at least two. Each of the above-mentioned predicted disease types is compared with the i-th disease type in the corresponding i-th predicted result. At the same time, the risk values of the predicted clinical symptoms corresponding to each of the above-mentioned predicted disease types are weighted. Among them, the weighting weight is set according to the typical degree of the above-mentioned predicted clinical symptoms. The greater the typical degree, the higher the weight. And the above-mentioned weighted result is used as the above-mentioned predicted risk value. The difference between the above-mentioned predicted risk value and the i-th disease risk value is calculated. Since both the above-mentioned predicted risk value and the i-th disease risk value reflect the possibility of the occurrence of the corresponding above-mentioned predicted disease type, therefore, when the comparison result is that the predicted disease type is the same as the i-th disease type in the corresponding i-th predicted result and the difference is within the above-mentioned preset value range, that is, the predicted results of the i-th model and the clinical prediction model are consistent. At this time, it is considered that the risk prediction of the i-th disease of the above-mentioned user is accurate. On the contrary, when the comparison result is inconsistent or the difference is greater than or equal to the preset value, it is considered that at least one of the i-th model and the clinical prediction model is inaccurate in predicting the i-th disease type or the above-mentioned predicted disease type. In order to further improve the accuracy of the i-th disease type and the above-mentioned predicted disease type, at the same time, the first sample data and the second sample data corresponding to the i-th disease type and the above-mentioned predicted disease type of the i-th model and the clinical prediction model are increased, and the i-th model and the clinical prediction model are retrained. Through the above technical solution, the prediction accuracy of the i-th model and the clinical prediction model can be further improved, and then the prediction accuracy of the user's risk of geriatric diseases can be improved.

[0065] Further, the training of the clinical prediction model includes:

[0066] Obtain patient sample data from the database. The patient sample data includes historical feature maps and corresponding symptom information. Among them, the historical feature maps include the contribution values of each feature parameter, and train the clinical prediction model based on the sample data.

[0067] Further, the value range of i is a positive integer greater than or equal to 1.

[0068] The present invention also provides a geriatric disease risk prediction system based on big data for implementing the above method, as Figure 4 shown, the system includes:

[0069] A first prediction unit, configured to obtain user information, extract a user feature vector corresponding to the i-th model from the user information, input the user feature vector into the i-th model, and obtain an i-th prediction result, where the i-th prediction result includes an i-th disease type and an i-th disease risk value;

[0070] A generation unit, configured to derive a prediction process diagram corresponding to the i-th prediction result from the i-th model through an extraction unit, and generate a target feature diagram corresponding to the i-th disease based on the prediction process diagram;

[0071] A second prediction unit, configured to integrate at least two of the target feature diagrams having cross feature parameters and a contribution value of the cross feature parameters greater than a set threshold to obtain an integrated feature diagram, input each of the integrated feature diagrams and the un-integrated target feature diagrams into a clinical prediction model, and obtain corresponding predicted clinical information;

[0072] An analysis unit, configured to obtain predicted disease information based on each of the predicted clinical information;

[0073] An adjustment unit, configured to adjust the accuracy of the i-th model and the clinical prediction model based on the predicted disease information, the predicted clinical information, and the corresponding i-th prediction result.

[0074] The present invention further provides a computer-readable storage medium, on which instructions are stored, and when the instructions are executed by a processor, the above method is implemented.

[0075] In summary, the present invention combines multi-dimensional features in user information, utilizes big data and machine learning technologies to accurately predict the risk of geriatric diseases, and significantly improves the prediction accuracy through the prediction and mutual verification of a dual model (the i-th model and the clinical prediction model), providing a reliable basis for clinical decision-making. At the same time, the accuracy of the i-th model is improved by extracting representative training data of the i-th model, and the transparency of the prediction process of the i-th model is achieved by exporting the prediction process diagram and generating the target feature map. Through the logic and basis in the prediction process of the i-th model, users and medical staff can more clearly understand the prediction results and provide more targeted prevention suggestions based on the feature map. The target feature map is also integrated according to the cross-feature parameters to obtain the integrated feature map, and the predicted clinical information is obtained based on the integrated feature map and the un-integrated target feature map. The predicted disease information is obtained based on the predicted clinical information. Based on the predicted disease information, predicted clinical information, and the i-th prediction result, the accuracy of the model can be dynamically adjusted. When there is a deviation between the predicted disease information, predicted clinical information, and the i-th prediction result, relevant sample data is automatically added, and the i-th model and the clinical prediction model are retrained, thus continuously optimizing the prediction ability of the model. By providing a detailed prediction process and result explanation, preventive suggestions can be provided according to the prediction results, which helps users take timely measures to reduce the risk of disease and improve the overall health level. Through the mutual cooperation of the above technical solutions, the risk of geriatric diseases can be improved, more accurate and reliable prediction services can be provided for users, and at the same time, the rational allocation and utilization of medical resources can be promoted.

[0076] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0077] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0078] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting the risk of geriatric diseases based on big data, characterized in that: The method comprises: Acquire user information, extract a user feature vector corresponding to an i-th model from the user information, input the user feature vector into the i-th model, and acquire an i-th prediction result, wherein the i-th prediction result includes an i-th disease type and an i-th disease risk value; deriving a prediction process graph corresponding to the i-th prediction result from the i-th model through an extraction unit, and generating a target feature graph corresponding to the i-th disease based on the prediction process graph; Integrate at least two of the target feature graphs that have cross-feature parameters and whose contribution values ​​are greater than a set threshold to obtain an integrated feature graph, input each of the integrated feature graphs and the unintegrated target feature graph into a clinical prediction model, and obtain corresponding predicted clinical information; Acquiring predicted disease information based on each of the predicted clinical information and the disease database by an analysis unit; The accuracy of the i-th model and the clinical prediction model is adjusted based on the predicted disease information, the predicted clinical information and the corresponding i-th prediction result.

2. The method according to claim 1, characterized in that: The training of the i-th model comprises: Reading historical data corresponding to the i-th geriatric disease from a database and using it as an i-th target data group, extracting N features corresponding to the i-th geriatric disease from the i-th target data group, and obtaining a plurality of feature vectors corresponding to the i-th target data group based on the N features, wherein the historical data includes basic information of the patient, diagnosis and treatment time series data, and the corresponding geriatric disease type; According to the distribution positions of the plurality of feature vectors, M feature vectors are selected on average as initial centroids, each other feature vector and the nearest centroid feature vector are divided into the same group, the average centroid of the group is obtained, the distance between the initial centroid and the average centroid is calculated, and this step is repeated until the distance is less than or equal to the set distance; The cluster group closest to the center position of all the feature vectors is taken as the target cluster group, and K feature vectors are randomly selected from the target cluster group as target feature vectors, and other target feature vectors corresponding to each of the target feature vectors and farthest from the target feature vector are selected from other cluster groups, and the target feature vectors and the other target feature vectors are used as training data to train the i-th model.

3. The method according to claim 1, characterized in that Extracting a target feature graph corresponding to the i-th disease based on the prediction process graph also includes: A prediction process diagram for obtaining the i-th prediction result based on the user feature vector is derived from the i-th model through the extraction unit, a weight value corresponding to each feature parameter in the user feature vector and a product of the weight value and the corresponding feature parameter value are obtained from the prediction process diagram, and the feature parameters and corresponding weight values ​​whose products are greater than a set value are extracted from the prediction process diagram, and the feature parameters are used as target features, and the product is used as a contribution value, and the target feature diagram is generated based on the target feature and the corresponding contribution value.

4. The method according to claim 1, characterized in that: The acquisition of the predicted clinical information includes: The target feature graphs corresponding to any two of the i-th diseases are compared, and multiple target feature graphs having crossed feature parameters and whose contribution values ​​of crossed feature parameters are greater than a set threshold are integrated to obtain an integrated feature graph, and each integrated feature graph and each unintegrated target feature graph are input into the clinical prediction model to obtain the predicted clinical information, wherein the predicted clinical information includes predicted clinical symptom information and predicted clinical symptom risk values.

5. The method according to claim 1, characterized in that The acquisition of the predicted disease information includes: The predicted clinical symptoms in the predicted clinical information are compared with the symptoms corresponding to each disease type in the disease database by the analysis unit, and the disease information of the disease type closest to the predicted clinical symptoms is used as the predicted disease information, wherein the predicted clinical symptoms correspond to at least one type of elderly disease.

6. The method according to claim 1, characterized in that Adjusting the accuracy of the i-th model and the clinical prediction model based on the predicted disease information, the predicted clinical information and the corresponding i-th prediction result includes: The predicted disease type is obtained based on the predicted disease information, and the risk value of each predicted clinical symptom of the user's occurrence of the predicted disease type is obtained based on the predicted clinical information, and the weighted result is used as the predicted risk value of the predicted disease type, and the predicted disease type is compared with the i-th disease type in the i-th prediction result to obtain a comparison result, and the difference between the predicted risk value and the i-th disease risk value is also obtained. When the comparison result is that the predicted disease type is consistent with the i-th disease type, and the difference is less than or equal to a preset value, it is normal. Otherwise, the first sample data corresponding to the i-th disease type and the predicted disease type in the i-th model is added, and the second sample data corresponding to the i-th disease type and the predicted disease type in the clinical prediction model is added, and the i-th model and the clinical prediction model are trained twice based on the first sample data and the second sample data, respectively, wherein the second sample data includes the historical target feature map and symptom information corresponding to the i-th disease type and the predicted disease type.

7. The method according to claim 1, characterized in that The training of the clinical prediction model comprises: Patient sample data is obtained from a database, wherein the patient sample data includes a historical feature graph and corresponding symptom information, wherein the historical feature graph includes a contribution value of each feature parameter, and the clinical prediction model is trained based on the sample data.

8. The method according to claim 1, characterized in that The i-th model and the clinical prediction model are both random forest models.

9. A system for predicting the risk of geriatric diseases based on big data, used to implement the method according to any one of claims 1 to 8, characterized in that: The system comprises: A first prediction unit is used to obtain user information, extract a user feature vector corresponding to an i-th model from the user information, input the user feature vector into the i-th model, and obtain an i-th prediction result, wherein the i-th prediction result includes an i-th disease type and an i-th disease risk value; a generating unit, configured to derive a prediction process graph corresponding to the i-th prediction result from the i-th model through the extracting unit, and generate a target feature graph corresponding to the i-th disease based on the prediction process graph; a second prediction unit, configured to integrate at least two of the target feature graphs having cross-feature parameters and whose contribution values ​​are greater than a set threshold, to obtain an integrated feature graph, input each of the integrated feature graphs and the unintegrated target feature graphs into a clinical prediction model, and obtain corresponding predicted clinical information; an analyzing unit, configured to obtain predicted disease information based on each of the predicted clinical information; An adjustment unit is used to adjust the accuracy of the i-th model and the clinical prediction model based on the predicted disease information, the predicted clinical information and the corresponding i-th prediction result.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Disease prediction method and device, equipment and storage medium

    CN116052867A

  • Method for prognosing a risk of occurrence of a disease

    EP3164824B1