A future incidence probability prediction method and device, a terminal, and a storage medium

By constructing and optimizing a disease risk prediction model and analyzing historical values ​​of user data at multiple time points, the problem of the inability to predict future disease incidence rates in existing technologies has been solved, enabling accurate prediction of future disease incidence rates and personalized service recommendations.

CN117995412BActive Publication Date: 2025-11-18INT DIGITAL ECONOMY ACAD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410404810.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-07
Publication Date
2025-11-18
Estimated Expiration
2044-04-07

AI Technical Summary

Technical Problem

Existing disease prediction models cannot predict the probability of future disease incidence, thus limiting their application scenarios.

Method used

By acquiring user data from target users, including data on influencing factors of vital signs, a disease risk prediction model is constructed and optimized. The optimized model is then used to analyze historical values ​​of user data at multiple time points to output the predicted incidence probability of the target disease at future time points.

Benefits of technology

Accurately predict the probability of a user's target disease onset at multiple future time points, provide personalized health management plans and insurance product recommendations, and improve the cost-effectiveness of risk protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117995412B_ABST
    Figure CN117995412B_ABST
Patent Text Reader

Abstract

The application discloses a future incidence probability prediction method and device, a terminal and a storage medium, and relates to the technical field of data prediction. The method comprises the following steps: obtaining user data of a target user, wherein the user data comprises influence factor data for reflecting physical sign information; obtaining an optimized disease risk prediction model; inputting historical values of the user data at a plurality of time nodes into the optimized disease risk prediction model; and outputting predicted incidence probabilities of a target disease at a plurality of future time nodes by the optimized disease risk prediction model according to the input historical values. The application can accurately predict the incidence probabilities of the target disease of the user at the plurality of future time nodes by analyzing the historical values of the user data at the plurality of time nodes by the optimized disease risk prediction model. By predicting the future incidence probabilities, the application can more accurately recommend appropriate insurance products and value-added services for the customer, provide long-term and effective protection for the customer, and improve the performance-price ratio of risk protection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data prediction, and particularly relates to a future disease incidence probability prediction method and device, a terminal and a storage medium. BACKGROUND

[0002] With the increasing maturity of Artificial Intelligence (AI) technology and theory, there currently exists a disease prediction model capable of predicting the disease risk of a user in a current state according to user data. However, the existing disease prediction model cannot predict future disease incidence probability, which limits the application scenarios of the model.

[0003] Therefore, the prior art still needs to be improved and developed. SUMMARY

[0004] The technical problem to be solved by the present application is to provide a future disease incidence probability prediction method and device, a terminal and a storage medium to solve the problem that the existing disease prediction model cannot predict future disease incidence probability.

[0005] The technical solution adopted by the present application to solve the problem is as follows:

[0006] In a first aspect, the present application provides a future disease incidence probability prediction method, which comprises:

[0007] Obtaining user data of a target user, wherein the user data comprises influence factor data reflecting physical sign information;

[0008] Obtaining an optimized disease risk prediction model, and inputting historical values of the user data at a plurality of time nodes into the optimized disease risk prediction model;

[0009] Outputting, by the optimized disease risk prediction model, predicted disease incidence probability of a target disease at a plurality of future time nodes according to the input historical values.

[0010] In an embodiment, the influence factor data comprises a plurality of types of influence factors, and each type of influence factor comprises a plurality of indexes.

[0011] In an embodiment, the user data further comprises consumption behavior data.

[0012] In an embodiment, the obtaining of the optimized disease risk prediction model comprises:

[0013] Collecting the influence factor data of a plurality of users;

[0014] inputting values of the influence factor data of several previous time nodes into an initial disease risk prediction model to obtain predicted incidence probabilities of several subsequent time nodes of the influence factor data; and obtaining real incidence probabilities of the subsequent time nodes, and calculating prediction error values of the influence factor data according to the predicted incidence probabilities and the real incidence probabilities;

[0015] optimizing the initial disease risk prediction model according to the prediction error values to obtain the optimized disease risk prediction model.

[0016] In an embodiment, the optimizing the initial disease risk prediction model according to the prediction error values comprises:

[0017] If a fitting value of the prediction error values does not meet preset requirements, the corresponding influence factor data is screened according to the prediction error values, the initial disease risk prediction model is retrained according to the screened influence factor data, until the fitting value of the prediction error values meets the preset requirements, and then the optimization is stopped.

[0018] In an embodiment, the retraining the initial disease risk prediction model according to the screened influence factor data comprises:

[0019] adjusting the screened influence factor data, wherein the adjustment manner comprises category adjustment of an influence factor in the influence factor data;

[0020] inputting the adjusted influence factor data into the initial disease risk prediction model for retraining.

[0021] In an embodiment, the method further comprises:

[0022] determining a target disease according to the predicted incidence probability output by the disease risk prediction model;

[0023] screening an associated indicator from the user data of the target user according to the target disease, wherein the user data after deleting a historical value corresponding to each associated indicator is re-input into the optimized disease risk prediction model, a difference value of the predicted incidence probabilities before and after deleting the historical value corresponding to the associated indicator is calculated, and it is determined whether the associated indicator is an associated indicator of the target disease;

[0024] generating a health management scheme of the target user according to the associated indicator of the target disease.

[0025] In an embodiment, the associated indicator of the target disease comprises a risk factor of the target disease, and the method further comprises:

[0026] For each risk factor of the target disease, the historical value of the risk factor in the user data is modified, and the optimized disease risk prediction model is caused to recalculate the predicted incidence probability of the target disease based on the user data corresponding to the modified historical value of the risk factor;

[0027] The difference between the predicted incidence probability of the target disease corresponding to the historical value of the risk factor before modification and the predicted incidence probability of the target disease corresponding to the modified historical value of the risk factor is calculated, and the risk influence of the risk factor on the target disease is determined.

[0028] In an embodiment, the influence factor in the influence factor data includes physical examination information, and an index of the physical examination information includes a physical examination item; the health management scheme includes a physical examination scheme, and the health management scheme of the target user is generated according to the associated index of the target disease, including:

[0029] The target physical examination item is determined according to the physical examination item contained in the associated index of the target disease;

[0030] The physical examination scheme of the target user is generated according to the target physical examination item.

[0031] In an embodiment, the health management scheme includes a health improvement suggestion, and the health management scheme of the target user is generated according to the associated index of the target disease, including:

[0032] The associated index of the target disease of the target user at different time nodes is aggregated to obtain a first aggregation result, and the consumption behavior data of a plurality of users associated with the target user at the same time node is aggregated to obtain a second aggregation result, wherein the associated relationship is determined based on the consumption behavior data;

[0033] The health improvement suggestion of the target user is generated according to the risk factor of the target disease, the first aggregation result and the second aggregation result.

[0034] In an embodiment, the method further includes:

[0035] The insurance package recommendation scheme corresponding to the target user is generated according to the predicted incidence probability of the target disease of the target user at each future time node.

[0036] In a second aspect, the embodiments of the present application further provide a prediction device for future incidence probability, and the device includes:

[0037] An acquisition module is configured to acquire user data of a target user, wherein the user data includes influence factor data used to reflect physical sign information;

[0038] An input module is configured to input historical values of the user data at a plurality of time nodes into the disease risk prediction model.

[0039] An output module is configured to output a predicted incidence probability of the target disease at a plurality of future time nodes according to the historical values input into the disease risk prediction model.

[0040] In a third aspect, an embodiment of the present application further provides a terminal, which comprises a memory and one or more processors; the memory stores one or more programs; the programs contain instructions for executing the method for predicting the future incidence probability; and the processors are configured to execute the programs.

[0041] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded and executed by a processor to implement the steps of the method for predicting the future incidence probability.

[0042] The embodiment of the present application can accurately predict the incidence probability of the target disease at a plurality of future time nodes by analyzing the historical values of the user data at a plurality of time nodes through the disease risk prediction model. By predicting the future incidence probability, more suitable insurance products and value-added services can be recommended to the customers, long-term effective protection can be provided for the customers, and the cost performance of the risk protection can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0044] Figure 1 FIG. 1 is a flowchart of the method for predicting the future incidence probability provided by the embodiment of the present application.

[0045] Figure 2 FIG. 2 is a module schematic diagram of the device for predicting the future incidence probability provided by the embodiment of the present application.

[0046] Figure 3 FIG. 3 is a principle block diagram of the terminal provided by the embodiment of the present application. DETAILED DESCRIPTION

[0047] The application discloses a future incidence probability prediction method and device, a terminal and a storage medium. To make the purpose, technical scheme and effects of the application clearer and more explicit, the application is further described in detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and not to limit the application.

[0048] Those skilled in the art can understand that the singular forms "a", "an" and "the" used herein include plural forms unless specifically stated otherwise. It should be further understood that the use of the term "include" in the specification of the application means that the described features, integers, steps, operations, elements, and / or components exist, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be intermediate elements. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any single unit and all combinations of the associated listed items.

[0049] Those skilled in the art can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as that generally understood by those skilled in the art to which the application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have meanings consistent with those in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such.

[0050] To solve the above-mentioned defects of the prior art, the application provides a future incidence probability prediction method, which comprises: obtaining user data of a target user, wherein the user data comprises influence factor data for reflecting physical sign information; obtaining an optimized disease risk prediction model; inputting historical values of the user data at a plurality of time nodes into the optimized disease risk prediction model; and outputting predicted incidence probabilities of a target disease at a plurality of future time nodes by the optimized disease risk prediction model according to the input historical values. The application can accurately predict the incidence probabilities of the target disease of the user at a plurality of future time nodes by analyzing and calculating the historical values of the user data at a plurality of time nodes through the disease risk prediction model obtained by pre-optimization. The problem that the disease prediction model in the prior art cannot predict future incidence probabilities is solved. Moreover, through prediction of future incidence probabilities, more suitable insurance products and value-added services can be more accurately recommended to customers, long-term effective protection can be provided for customers, and the cost performance of risk protection is improved.

[0051] For example, user data of user a is obtained, the user data including impact factor data reflecting the physical information of user a, such as physical examination information, medical treatment information, and demographic characteristics. The historical values of the user data at t1, t2, and t3 are input into the optimized disease risk prediction model, and the optimized disease risk prediction model outputs the predicted incidence probabilities of disease A of user a at future t4, t5, and t6 as 10%, 20%, and 30% respectively, where t represents a time node.

[0052] As shown in Figure 1 The method includes the following steps:

[0053] In step S100, user data of a target user is obtained, wherein the user data includes impact factor data reflecting physical information.

[0054] Specifically, the target user in the embodiment can be any user in need of future disease risk prediction. The user data is data related to the user and includes impact factor data reflecting the physical information of the user. The user data of the target user is obtained to predict the disease probability of the target user for a single disease or multiple diseases.

[0055] In an implementation manner, the impact factor data includes a plurality of types of impact factors, and each type of impact factor includes a plurality of indexes.

[0056] Specifically, the impact factor data can include a plurality of impact factors, such as physical examination information, medical treatment information, and demographic characteristics. The physical examination information can reflect the physical examination results of the user, the medical treatment information can reflect the disease condition of the user, and the demographic characteristics can reflect the population characteristics (age, gender, etc.) of the region where the user is located. Each type of impact factor can include a plurality of indexes, for example, the impact factor corresponding to the physical examination information includes a plurality of indexes, and different indexes are respectively used to reflect the information of different physical examination items. In an actual application scenario, the indexes included in different impact factors can be determined according to the type of the disease to be predicted.

[0057] In step S200, an optimized disease risk prediction model is obtained, and historical values of the user data at a plurality of time nodes are input into the optimized disease risk prediction model.

[0058] The embodiment pre-constructs and optimizes a disease risk prediction model for predicting the risk of a single disease or multiple diseases. The historical values of each type of impact factor of the user data of the target user at each time node are input into the optimized disease risk prediction model, and the optimized disease risk prediction model predicts the long-term disease risk of the target user, such as the predicted incidence probability of a certain disease at the current time, after 5 years, after 10 years, and the like.

[0059] In an implementation manner, the user data further comprises consumption behavior data.

[0060] Specifically, in order to further improve the accuracy of the disease risk prediction model, in addition to the impact factor data for reflecting the influence of the physical information of the target user, the embodiment also predicts the incidence probability of a certain disease of the target user in the future by combining the consumption behavior data of the target user. The consumption behavior data can reflect the living habits of the target user and can indirectly affect the future incidence probability of different diseases. In actual application scenarios, the disease risk prediction model that has been optimized learns the influence of different consumption behavior data on different diseases, i.e., the correlation between the consumption behavior data and the disease. Therefore, inputting the consumption behavior data and the impact factor data of the user into the disease risk prediction model that has been optimized can effectively improve the prediction accuracy of the disease risk prediction model that has been optimized for the incidence probability of each future time node.

[0061] In an implementation manner, the obtaining the disease risk prediction model that has been optimized comprises:

[0062] collecting the impact factor data of a plurality of users;

[0063] for the impact factor data of each user, inputting the values of a plurality of preceding time nodes of the impact factor data into an initial disease risk prediction model to obtain the predicted incidence probability of a plurality of subsequent time nodes of the impact factor data, and obtaining the real incidence probability of each subsequent time node, calculating the prediction error value of the impact factor data according to the predicted incidence probability and the real incidence probability;

[0064] optimizing the initial disease risk prediction model according to the prediction error value to obtain the disease risk prediction model that has been optimized.

[0065] Specifically, the embodiment defines an uncompleted disease risk prediction model as an initial disease risk prediction model, and the embodiment collects influence factor data of multiple users to optimize the initial disease risk prediction model, so as to improve the prediction accuracy of the model. Taking the influence factor data of a user as an example, the influence factor data is divided into two parts according to time nodes, the time node in front is the former time node, and the time node in back is the latter time node. It should be noted that the former time node and the latter time node are both historical time nodes. The values of each former time node in the influence factor data are input into the initial disease risk prediction model, and the initial disease risk prediction model calculates the prediction incidence probability of each latter time node. By comparing the calculated prediction incidence probability with the corresponding true incidence probability, the prediction error value of the initial disease risk prediction model on the influence factor data can be obtained. The size of the prediction error value can reflect the current prediction accuracy of the initial disease risk prediction model, and therefore the initial disease risk prediction model can be optimized in the direction of the prediction error value to improve the prediction accuracy of the model.

[0066] In an implementation manner, the optimizing the initial disease risk prediction model according to the prediction error value comprises:

[0067] If the fitting value of the prediction error value does not meet the preset requirement, the corresponding influence factor data is screened according to the prediction error value, the initial disease risk prediction model is retrained according to the screened influence factor data, and the optimization is stopped until the fitting value of the prediction error value meets the preset requirement.

[0068] Specifically, if the fitting value of the prediction error value does not meet the preset requirement, it indicates that the prediction accuracy of the initial disease risk prediction model is not enough, the influence factor data with prediction errors is screened according to the prediction error value, and the initial disease risk prediction model is retrained to improve the accuracy of the prediction result of the model and the data fitting value of the training data set. After multiple rounds of retraining of the initial disease risk prediction model, the data fitting value of the training data set is continuously improved. When the fitting value meets the preset requirement (for example, a preset convergence threshold), the optimization is stopped, and an optimized disease risk prediction model is obtained.

[0069] In an implementation manner, the retraining the initial disease risk prediction model according to the screened influence factor data comprises:

[0070] The screened influence factor data is adjusted, and the adjustment manner comprises category adjustment of an influence factor in the influence factor data.

[0071] The adjusted influence factor data is input into the initial disease risk prediction model for retraining.

[0072] Specifically, for the screened prediction error impact factor data, before retraining, the impact factor data needs to be adjusted, including but not limited to adjusting the categories of impact factors contained in the impact factor data. Then input the adjusted impact factor data into the initial disease risk prediction model for re-prediction to improve the data fitting value in the training data set. For example, the prediction error impact factor data D includes three categories of impact factors a, b and c, then delete the c category impact factor in the impact factor data D, and then retrain the initial disease risk prediction model according to the impact factor data D with only a and b categories.

[0073] For example, the impact factor data (physical examination information, medical information and demographic characteristics) of a plurality of users is obtained as training data. The impact factor data includes values of a plurality of time nodes, for example,

t1, t2, …, tm, tm+1, tm+2, …, tn

tm+1, tm+2, …, tn

[0074] In an implementation manner, the fitting value can be measured by AUC, that is, the area (Area Under Curve) surrounded by the coordinate axis of the ROC curve (Receiver Operating Characteristic curve). The preset fitting requirement can be the ROC curve range. The fitting value is compared with the ROC curve range to determine whether the tuning is completed. For example, the ROC curve range is [0.5, 1], and if the fitting value is within the ROC curve range and close to 1, it means that the tuning is completed.

[0075] Step S300: outputting, by the tuned disease risk prediction model, predicted incidence probabilities of a target disease at a plurality of future time nodes according to input historical values.

[0076] Specifically, the target disease can be one or more diseases, that is, the disease risk prediction model in this embodiment can predict the incidence probabilities of single or multiple diseases. After inputting the historical values of the user data at each time node into the tuned disease risk prediction model, the tuned disease risk prediction model performs data analysis on the input historical values and outputs the predicted incidence probabilities of the target disease at a plurality of future time nodes, for example, the predicted incidence probabilities at the present time, after 5 years, and after 10 years. In short, the tuned disease risk prediction model can predict the predicted incidence probabilities of up to N years in the future through the user data of the previous N years, where N is a natural number representing the number of years.

[0077] For example, taking the stroke risk prediction model (Frmingham Stroke Scale model) as an example, the historical values ​​of user data at multiple time points are input into the risk prediction model constructed using the Stroke Risk Rating Scale (ESRS) to obtain the stroke risk at multiple future time points.

[0078] In one implementation, the method further includes:

[0079] Based on the target disease, related indicators are selected from the user data of the target user. For each related indicator, the user data after deleting the historical values ​​corresponding to the related indicator is re-input into the optimized disease risk prediction model. The difference in the predicted incidence probability before and after deleting the historical values ​​corresponding to the related indicator is calculated to determine whether the related indicator is a related indicator of the target disease.

[0080] Based on the correlation indicators of the target disease, a health management plan is generated for the target user.

[0081] To further expand the application scenarios of the disease risk prediction model, this embodiment can also provide users with personalized health management plans based on the predicted incidence probability of the target disease output by the optimized disease risk prediction model. Specifically, if the predicted incidence probability output by the optimized disease risk prediction model reflects the risk of the target user having the target disease, then correlation indicators with high correlation to the target disease are selected from the target user's user data. For example, the correlation indicator can be a certain physical examination item or risk factor. The specific selection method is as follows: for each correlation indicator in the user data, the historical value corresponding to the correlation indicator is deleted from the user data, and the user data after deleting the historical value corresponding to the correlation indicator is used as additional test data and re-inputted into the optimized disease risk prediction model. The optimized disease risk prediction model compares the predicted incidence probabilities output by the added test data and the original user data. If the difference is large, it indicates that the presence or absence of the associated indicator has a significant impact on the predicted incidence probability of the target disease at a future time point; that is, the associated indicator has a high correlation with the incidence probability of the target disease, and is therefore considered an associated indicator for the target disease. If the difference is small, it indicates that the presence or absence of the associated indicator has a small impact on the predicted incidence probability of the target disease at a future time point; that is, the associated indicator has a low correlation with the incidence probability of the target disease, and is therefore not an associated indicator for the target disease. Developing health management plans for target users based on the selected associated indicators can effectively reduce their future incidence probability of the target disease. It is understandable that if the user data only contains influencing factor data, the associated indicator is selected from the influencing factor data; if the user data contains both influencing factor data and consumer behavior data, the associated indicator can be selected from the influencing factor data and / or the consumer behavior data. Similar to influencing factor data, consumer behavior data can also contain multiple indicators, such as health product purchase behavior, dining out behavior, etc.

[0082] For example, suppose a well-tuned disease risk prediction model outputs a predicted incidence probability indicating that a user is at risk of disease A. Taking the health checkup item B as an example, the method for selecting correlation indicators for disease A is as follows: Delete historical values ​​in the user data corresponding to health checkup item B. For example, delete the values ​​corresponding to health checkup item B [t1, t2, ..., t...]. mThe values ​​corresponding to the deleted health check item B are then removed. The user data after deleting the historical values ​​of the health check item B is used as the additional test data. This additional test data is then re-input into the optimized disease risk prediction model for prediction. The predicted incidence probability output by the optimized disease risk prediction model based on the additional test data is compared with the predicted incidence probability output based on the original user data. If the difference is greater than a preset threshold (the threshold can be determined according to actual needs), it indicates that health check item B is highly correlated with disease A. In this case, health check item B is used as a correlation indicator for disease A, and a health management plan can be developed based on health check item B. If the difference is less than or equal to the preset threshold, it indicates that health check item B is not highly correlated with disease A. In this case, health check item B is not a correlation indicator for disease A, and a health management plan does not need to be developed based on health check item B. This embodiment uses a unified indicator screening method; the screening method for risk factors can refer to the screening method for health check items described above.

[0083] In one implementation, the association indicators for the target disease include risk factors for the target disease, and the method further includes:

[0084] For each of the target diseases, the historical values ​​of the risk factors in the user data are modified, and the optimized disease risk prediction model is recalculated based on the user data corresponding to the modified historical values ​​of the risk factors;

[0085] Calculate the difference in the predicted incidence probability of the target disease before and after the historical value of the risk factor is modified, and determine the risk impact of the risk factor on the target disease.

[0086] Specifically, the selected correlation indicators can include risk factors such as hypertension and hyperlipidemia. To accurately quantify the risk impact of risk factors on the target disease, this embodiment modifies the historical values ​​of each risk factor in the user data, for example, by replacing them with standard values. Standard values ​​reflect the normal value of the risk factor and can be obtained manually. Then, the optimized disease risk prediction model recalculates the predicted incidence probability based on the user data corresponding to the modified historical values ​​of the risk factors. The recalculated predicted incidence probability is compared with the original predicted incidence probability, and the difference between the two determines the risk impact of the change in the risk factor's value on the target disease.

[0087] For example, the risk factors for disease A are three indicators: a, b, and c. Taking indicator a as an example, in the original user data, the historical t1 value of indicator a is 60, and the historical t2 value is 70. The optimized disease risk prediction model, based on the original user data, outputs a predicted incidence probability of disease A in 5 years of 80% and a predicted incidence probability in 10 years of 90%. By modifying the historical t1 value of indicator a to 20 and the historical t2 value to 30, new user data is obtained. The optimized disease risk prediction model is then recalculated based on the new user data, resulting in a predicted incidence probability of disease A in 5 years of 40% and a predicted incidence probability in 10 years of 50%. By comparing the difference in the predicted incidence probability of the target disease output by the optimized disease risk prediction model before and after modifying the historical value of indicator a (i.e., 80%-40% and 90%-50%), the impact of changes in the value of indicator a on the risk of disease A can be analyzed.

[0088] In one implementation, the impact factors in the impact factor data include physical examination information, and the indicators of the physical examination information include physical examination items; the health management plan includes a physical examination plan; generating the health management plan for the target user based on the correlation indicators of the target disease includes:

[0089] The target medical examination items are determined based on the medical examination items included in the associated indicators of the target disease;

[0090] A health check plan for the target user is generated based on the target health check items.

[0091] This embodiment, when developing a health management plan for a user, may include creating a personalized physical examination plan. Specifically, the influencing factor data contains influencing factors such as physical examination information, and each indicator included in this type of influencing factor can correspond to different physical examination items. If the selected associated indicators include physical examination items, it indicates that the physical examination item is a high-detection screening item for the target disease, and it is then used as the target physical examination item and added to the user's physical examination plan. It can be understood that regardless of whether the input data of the optimized disease risk prediction model is influencing factor data or influencing factor data and consumer behavior data, as long as the influencing factor data contains influencing factors such as physical examination information and the associated indicators include physical examination items, a physical examination plan for the target user can be generated when developing a health management plan.

[0092] In one implementation, the health management plan includes health improvement suggestions, and generating a health management plan for the target user based on the correlation indicators of the target disease includes:

[0093] Target risk factors are determined based on the risk factors included in the associated indicators of the target disease;

[0094] The association indicators of the target disease of the target user at different time points are aggregated to obtain a first aggregation result; and the consumption behavior data of several users who are related to the target user at the same time point are aggregated to obtain a second aggregation result, wherein the association is determined based on the consumption behavior data.

[0095] Based on the target risk factors of the target disease, the first aggregation result, and the second aggregation result, health improvement suggestions are generated for the target user.

[0096] Specifically, if the selected correlation indicators include risk factors, such as hypertension and hyperlipidemia, it indicates that the risk factor is a high-risk factor for the target disease, and it is then used as the target risk factor. The disease risk prediction model aggregates indicators in user data (such as standard values ​​of physical examination results) in two dimensions. One dimension aggregates the correlation indicators corresponding to the influencing factor data of a single user at different time points, that is, it summarizes and analyzes indicators such as physical examination items and risk factors that are highly related to the target disease. This dimension aggregation can be achieved through an attention mechanism to obtain the first aggregation result. The other dimension aggregates the features corresponding to the consumption behavior data of multiple related users (determined based on the user's consumption behavior data) at the same time point. That is, it determines which user group a user belongs to through consumption behavior data, and summarizes and analyzes the consumption behavior data of multiple users in that user group at the same time point. This dimension aggregation can be achieved through graph convolution to obtain the second aggregation result. This embodiment, through the target risk factor, the first aggregation result, and the second aggregation result, can provide users with health improvement suggestions from multiple perspectives such as physiological condition and lifestyle habits to reduce the probability of users developing the target disease in the future.

[0097] In one implementation, since the target risk factors may affect the prediction results, the disease risk prediction model and the model prediction results can be updated again after the target risk factors are screened out.

[0098] In one implementation, the method further includes:

[0099] Based on the predicted incidence probability of the target disease for the target user at each future time point, an insurance package recommendation plan is generated for the target user.

[0100] This embodiment can also apply the predicted incidence rate output by the disease risk prediction model to actual insurance plan recommendations, thereby expanding the model's application and refining its capabilities. Specifically, for customers, the predicted incidence rate output by the disease risk prediction model can be used to recommend suitable insurance packages, such as insurance products and value-added services, to different customers in a personalized way, thus providing customers with long-term and effective protection. For insurance companies, accurately recommending insurance packages to different customers can enable them to sell more insurance products and protection services to the same customer, increase the average transaction value and overall profit, while also reducing their own future claims risk.

[0101] Based on the above embodiments, the present invention also provides a device for predicting the probability of future disease incidence, such as... Figure 2 As shown, the device includes:

[0102] Acquisition module 01 is used to acquire user data of the target user, wherein the user data includes influencing factor data that reflects vital signs information;

[0103] Input module 02 is used to obtain the optimized disease risk prediction model and input the historical values ​​of the user data at several time points into the optimized disease risk prediction model.

[0104] Output module 03 is used to output the predicted incidence probability of the target disease at several future time points based on the input historical values ​​through the optimized disease risk prediction model.

[0105] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 3 As shown, the terminal includes a processor, memory, network interface, and display screen connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for predicting the probability of future illness. The display screen can be an LCD screen or an e-ink screen.

[0106] Those skilled in the art will understand that Figure 3 The schematic diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0107] In one implementation, the terminal's memory stores one or more programs, and these programs are configured to be executed by one or more processors, and include instructions for a method of predicting the probability of future disease occurrence.

[0108] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0109] In summary, this invention discloses a method, apparatus, terminal, and storage medium for predicting the probability of future disease onset. The method includes: acquiring user data of a target user, wherein the user data includes influencing factor data reflecting vital signs; acquiring an optimized disease risk prediction model; inputting historical values ​​of the user data at several time points into the optimized disease risk prediction model; and outputting the predicted probability of the target disease at several future time points based on the input historical values ​​using the optimized disease risk prediction model. This invention, by analyzing and calculating historical values ​​of user data at multiple time points using a pre-optimized disease risk prediction model, can accurately predict the probability of a user developing a target disease at multiple future time points. This invention can accurately predict future disease onset probabilities, solving the problem that existing disease prediction models cannot predict future disease onset probabilities. Furthermore, by predicting future disease onset probabilities, more suitable insurance products and value-added services can be recommended to customers more accurately, providing long-term effective protection and improving the cost-effectiveness of risk protection.

[0110] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A method for predicting the probability of future disease incidence, characterized in that, The method includes: Acquire user data of the target user, wherein the user data includes influencing factor data that reflects vital signs information; A tuned disease risk prediction model is obtained by inputting historical values ​​of the user data at several time points into the tuned model. The method for obtaining the tuned model includes: collecting impact factor data from several users; for each user's impact factor data, inputting the values ​​of the impact factor data at several earlier time points into an initial disease risk prediction model to obtain the predicted incidence probability of the impact factor data at several later time points; obtaining the actual incidence probability at the several later time points; and calculating the prediction error value of the impact factor data based on the predicted incidence probability and the actual incidence probability. If the fitted value of the prediction error value does not meet a preset requirement, then... Based on the prediction error value, the corresponding impact factor data is selected, and the selected impact factor data is adjusted, including adjusting the category of the impact factor in the impact factor data. The adjusted impact factor data is then input into the initial disease risk prediction model for retraining until the fitted value of the prediction error value meets the preset fitting requirements. At this point, the optimization is stopped, and the optimized disease risk prediction model is obtained. The fitted value is measured using AUC, which is the area under the ROC curve and the coordinate axis. The preset fitting requirements are the ROC curve range, which is the range of the receiver operating characteristic curve. The fitted value is compared with the ROC curve range to determine whether the optimization is complete. Based on the input historical values, the optimized disease risk prediction model outputs the predicted incidence probability of the target disease at several future time points. The method further includes: Based on the target disease, related indicators are selected from the user data of the target user. For each related indicator, the user data after deleting the historical values ​​corresponding to the related indicator is re-input into the optimized disease risk prediction model. The difference in the predicted incidence probability before and after deleting the historical values ​​corresponding to the related indicator is calculated to determine whether the related indicator is a related indicator of the target disease. Based on the correlation indicators of the target disease, a health management plan is generated for the target user; The correlation indicators of the target disease include risk factors for the target disease, and the health management plan includes health improvement suggestions. Generating a health management plan for the target user based on the correlation indicators of the target disease includes: The association indicators of the target disease of the target user at different time points are aggregated to obtain a first aggregation result; and the consumption behavior data of several users who are related to the target user at the same time point are aggregated to obtain a second aggregation result, wherein the association is determined based on the consumption behavior data. Based on the risk factors of the target disease, the first aggregation result, and the second aggregation result, health improvement suggestions are generated for the target user. The impact factors in the impact factor data include physical examination information, and the indicators of the physical examination information include physical examination items; the health management plan includes a physical examination plan, and the step of generating a health management plan for the target user based on the correlation indicators of the target disease includes: The target medical examination items are determined based on the medical examination items included in the associated indicators of the target disease; Generate a health check plan for the target user based on the target health check items; The method further includes: For each of the target diseases, the historical values ​​corresponding to the risk factors in the user data are modified to standard values, and the optimized disease risk prediction model is recalculated based on the user data corresponding to the modified historical values ​​of the risk factors to predict the incidence probability of the target disease. Calculate the difference in the predicted incidence probability of the target disease before and after the historical value of the risk factor is modified, and determine the risk impact of the risk factor on the target disease; The method further includes: Based on the predicted incidence probability of the target disease for the target user at each future time point, an insurance package recommendation plan is generated for the target user.

2. The method for predicting the probability of future disease incidence according to claim 1, characterized in that, The impact factor data includes several categories of impact factors, and each category of impact factors includes several indicators.

3. The method for predicting the probability of future disease incidence according to claim 1, characterized in that, The user data also includes consumption behavior data.

4. A device for predicting the probability of future disease onset, characterized in that, The device includes: The acquisition module is used to acquire user data of the target user, wherein the user data includes influencing factor data that reflects vital signs information; An input module is used to acquire an optimized disease risk prediction model by inputting historical values ​​of the user data at several time points into the optimized disease risk prediction model. The method for acquiring the optimized disease risk prediction model includes: collecting the influencing factor data of several users; for each user's influencing factor data, inputting the values ​​of the influencing factor data at several earlier time points into an initial disease risk prediction model to obtain the predicted incidence probability of the influencing factor data at several later time points; obtaining the actual incidence probability at the several later time points; and calculating the prediction error value of the influencing factor data based on the predicted incidence probability and the actual incidence probability. If the fitted value of the prediction error value does not meet a preset condition... If required, the corresponding impact factor data is selected based on the prediction error value, and the selected impact factor data is adjusted, including adjusting the category of the impact factor data; the adjusted impact factor data is input into the initial disease risk prediction model for retraining until the fitted value of the prediction error value meets the preset fitting requirements, then the optimization is stopped, and the optimized disease risk prediction model is obtained; wherein, the fitted value is measured by AUC, which is the area under the ROC curve and the coordinate axis; the preset fitting requirements are the ROC curve range, which is the range of the receiver operating characteristic curve; the fitted value is compared with the ROC curve range to determine whether the optimization is complete. The output module is used to output the predicted incidence probability of the target disease at several future time points based on the input historical values ​​through the optimized disease risk prediction model. The device is also used for: Based on the target disease, related indicators are selected from the user data of the target user. For each related indicator, the user data after deleting the historical values ​​corresponding to the related indicator is re-input into the optimized disease risk prediction model. The difference in the predicted incidence probability before and after deleting the historical values ​​corresponding to the related indicator is calculated to determine whether the related indicator is a related indicator of the target disease. Based on the correlation indicators of the target disease, a health management plan is generated for the target user; The correlation indicators of the target disease include risk factors for the target disease, and the health management plan includes health improvement suggestions. Generating a health management plan for the target user based on the correlation indicators of the target disease includes: The association indicators of the target disease of the target user at different time points are aggregated to obtain a first aggregation result; and the consumption behavior data of several users who are related to the target user at the same time point are aggregated to obtain a second aggregation result, wherein the association is determined based on the consumption behavior data. Based on the risk factors of the target disease, the first aggregation result, and the second aggregation result, health improvement suggestions are generated for the target user. The impact factors in the impact factor data include physical examination information, and the indicators of the physical examination information include physical examination items; the health management plan includes a physical examination plan, and the step of generating a health management plan for the target user based on the correlation indicators of the target disease includes: The target medical examination items are determined based on the medical examination items included in the associated indicators of the target disease; Generate a health check plan for the target user based on the target health check items; The device is also used for: For each of the target diseases, the historical values ​​corresponding to the risk factors in the user data are modified to standard values, and the optimized disease risk prediction model is recalculated based on the user data corresponding to the modified historical values ​​of the risk factors to predict the incidence probability of the target disease. Calculate the difference in the predicted incidence probability of the target disease before and after the historical value of the risk factor is modified, and determine the risk impact of the risk factor on the target disease; The device is also used for: Based on the predicted incidence probability of the target disease for the target user at each future time point, an insurance package recommendation plan is generated for the target user.

5. A terminal, characterized in that, The terminal includes a memory and one or more processors; the memory stores one or more programs; the programs contain instructions for executing the method for predicting the probability of future disease incidence as described in any one of claims 1-3; the processor is used to execute the programs.

6. A computer-readable storage medium storing a plurality of instructions, characterized in that, The instructions are applicable to being loaded and executed by a processor to implement the steps of the method for predicting the probability of future disease incidence as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Visualization system for prediction of gynecological neoplastic disease risks

    CN111899888A

  • Data processing system, method and device, and storage medium

    CN112102950A

  • Battery capacity prediction model training method, prediction method, device and medium

    CN116298906A

  • Prediction method of disease risk and related device

    CN117577326A