Medical risk prediction methods and devices, storage media and electronic equipment
By identifying sampling characteristics in target cities and training individual medical expense prediction models, the problem of low accuracy in medical risk prediction was solved, enabling accurate prediction of medical expenses and risks for insured individuals.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-01
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies have low accuracy in predicting medical risks, especially in accurately predicting the medical risks of people related to the business.
By determining the sampling characteristics based on the correlation between multiple candidate characteristics of the historical insured population in the target city and the city's insurance claims information, the target population is extracted. Based on the characteristic values and medical expenses of the target population, a personal medical expense prediction model is trained to predict the medical expenses and risks of the population to be insured.
This improves the accuracy of medical cost forecasting, thereby enhancing the accuracy of medical risk forecasting. It takes into account the impact of potential risks on future expenses, thus increasing the reliability of forecasts.
Smart Images

Figure CN115115408B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to a medical risk prediction method, a medical risk prediction device, a computer-readable storage medium, and an electronic device. Background Technology
[0002] Predicting medical risks can assist relevant personnel in conducting subsequent analysis and handling based on the predicted risks. For example, it can help insurance companies determine insurance premiums based on predicted medical risks.
[0003] Among related technologies, the accuracy of medical risk prediction is relatively low, especially in accurately predicting the medical risks of people related to the business.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this disclosure is to provide a method and apparatus for predicting medical risks, a computer-readable storage medium and an electronic device, thereby improving, at least to some extent, the problem of low accuracy in predicting medical costs in related technologies.
[0006] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0007] According to a first aspect of this disclosure, a medical risk prediction method is provided, comprising: determining sampling features from multiple first candidate features of a target city's historical insured population based on the correlation between these features and the city's insurance claim information; determining the distribution information of the sampling features of the historical insured population, and extracting a target population from the population with medical records in the target city based on the distribution information; training a personal medical expense prediction model based on the feature values of the target population's target features in a first historical period and the target population's personal medical expenses in a second historical period, wherein the second historical period follows the first historical period; and predicting the target medical expenses of the prospective insured population based on the personal medical expense prediction model.
[0008] In an exemplary embodiment of this disclosure, based on the foregoing scheme, the personal medical expenses include one or more of total personal medical expenses, out-of-pocket medical expenses, and out-of-pocket medical expenses; when the personal medical expenses include the total personal medical expenses, the target feature includes a first target feature, which is determined by: determining the first target feature based on a first correlation degree between a second candidate feature and the total personal medical expenses; when the personal medical expenses include the out-of-pocket medical expenses, the target feature includes a second target feature, which is determined by: determining the second target feature based on a second correlation degree between a second candidate feature and the out-of-pocket medical expenses; when the personal medical expenses include the out-of-pocket medical expenses, the target feature includes a third target feature, which is determined by: determining the third target feature based on a third correlation degree between a second candidate feature and the out-of-pocket medical expenses; wherein, the second candidate feature is determined by: acquiring historical medical data and performing dimensionality reduction on individual features and derived features in the historical medical data to obtain the second candidate feature.
[0009] In one exemplary embodiment of this disclosure, based on the foregoing scheme, the historical medical data includes various information of the medical subject, such as gender, age, past medical history, family medical history, inpatient and outpatient diagnoses, surgeries and procedures, total inpatient and outpatient medical expenses, out-of-pocket medical expenses, out-of-pocket medical expenses, and the generic name of drugs, drug dosage, drug cost, examination fees, and surgical fees in the medical expense charge details.
[0010] In an exemplary embodiment of this disclosure, based on the foregoing scheme, the individual features and derived features in the historical medical data are determined by: discretizing the continuous variables in the historical medical data to generate individual features in the historical medical data; and combining the individual features in the historical medical data to obtain the derived features.
[0011] In an exemplary embodiment of this disclosure, based on the foregoing scheme, training a personal medical expense prediction model according to the feature values of the target population's target characteristics in a first historical period and the personal medical expenses of the target population in a second historical period includes: when the personal medical expenses include the total personal medical expenses, training a total personal medical expense model based on the feature values of the first target characteristics of the target population in the first historical period and the expense range to which the total personal medical expenses of the target population belong in the second historical period; when the personal medical expenses include the out-of-pocket medical expenses, training a out-of-pocket medical expense prediction model based on the feature values of the second target characteristics of the target population in the first historical period and the expense range to which the out-of-pocket medical expenses of the target population belong in the second historical period; and when the personal medical expenses include the out-of-pocket medical expenses, training a out-of-pocket medical expense prediction model based on the feature values of the third target characteristics of the target population in the first historical period and the expense range to which the out-of-pocket medical expenses of the target population belong in the second historical period.
[0012] In an exemplary embodiment of this disclosure, based on the foregoing scheme, predicting the target medical expenses of the prospective insured population based on the personal medical expense prediction model includes: predicting the number of prospective insured individuals based on the number of historical insured individuals; determining the distribution of target characteristics of the prospective insured population based on the distribution of target characteristics of historical insured individuals; sampling from medical subjects with medical records based on the number of prospective insured individuals and the distribution of target characteristics of the prospective insured population to obtain the prospective insured population; predicting the personal target medical expenses of each prospective insured individual in the prospective insured population based on the personal medical expense prediction model; and determining the target medical expenses of the prospective insured population based on the personal target medical expenses of each prospective insured individual.
[0013] In an exemplary embodiment of this disclosure, based on the foregoing scheme, when the personal medical expense prediction model includes a decision tree-based distributed gradient boosting model, the method further includes: determining the importance level of the target feature according to the decision tree-based distributed gradient boosting model, and selecting the target feature ranked in the top N by importance level as a warning feature; determining the insured person's personal medical expenses based on the personal medical expense prediction model, and determining the insured person as a candidate for warning when the personal medical expenses meet preset conditions; and pushing prevention and control information of the disease indicated by the warning feature to the client corresponding to the candidate for warning when the warning feature of the candidate for warning exceeds a warning threshold.
[0014] According to a second aspect of this disclosure, a medical risk prediction device is provided, comprising: a sampling feature determination module configured to determine sampling features from a plurality of first candidate features based on the correlation between the historical insured population of a target city and the claim information of the city's insurance; a target population determination module configured to determine the distribution information of the sampling features of the historical insured population, and to extract a target population from the population with medical records in the target city based on the distribution information; a model training module configured to train a personal medical expense prediction model based on the feature values of the target features of the target population in a first historical period and the personal medical expenses of the target population in a second historical period, wherein the second historical period is after the first historical period; and a medical risk prediction module configured to predict the target medical expenses of the population to be insured based on the personal medical expense prediction model.
[0015] According to a third aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect of the above embodiments.
[0016] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to perform the method as described in the first aspect of the above embodiments.
[0017] As can be seen from the above technical solutions, the medical risk prediction method, medical risk prediction device, and computer-readable storage medium and electronic device for implementing the medical risk prediction method in the exemplary embodiments of this disclosure have at least the following advantages and positive effects:
[0018] In some embodiments of this disclosure, sampling characteristics are determined based on their relevance to claims information from urban medical insurance. Then, a target population is extracted from the medical group based on the distribution of these sampling characteristics from historical insured individuals. The target characteristics of this target population are used as training samples to train a medical expense prediction model. The urban insurance cost is then determined based on the medical expenses predicted by the model, thereby determining the medical risk of the insured population based on the target medical expenses. Compared to related technologies, this disclosure, through sampling characteristics, ensures that the distribution of the target population used as training samples is consistent with that of historical insured individuals, thus improving the accuracy of the model's prediction of medical expenses for the insured population and consequently enhancing the accuracy of its prediction of medical risks. Simultaneously, this disclosure trains a cost prediction model capable of predicting future medical expenses using historical data from the target population, thus considering the impact of potential risks on future medical expenses and further improving the accuracy of medical expense prediction, thereby further enhancing the accuracy of its prediction of medical risks for the insured population.
[0019] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0021] Figure 1 A flowchart illustrating a medical risk prediction method in an exemplary embodiment of this disclosure is shown.
[0022] Figure 2 A flowchart illustrating a method for determining sampling characteristics in an exemplary embodiment of this disclosure is shown.
[0023] Figure 3 A flowchart illustrating a method for determining a first target feature in an exemplary embodiment of this disclosure is shown.
[0024] Figure 4 A flowchart illustrating a method for determining city insurance costs in an exemplary embodiment of this disclosure is shown.
[0025] Figure 5 This diagram illustrates a flowchart of a method for pushing information to insured individuals in an exemplary embodiment of this disclosure.
[0026] Figure 6This diagram illustrates a framework of a medical cost prediction system according to an exemplary embodiment of the present disclosure.
[0027] Figure 7 This diagram illustrates the structure of a medical risk prediction device in an exemplary embodiment of the present disclosure.
[0028] Figure 8 A schematic diagram of the structure of an electronic device in an exemplary embodiment of this disclosure is shown. Detailed Implementation
[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0030] The terms “a,” “an,” “the,” and “the” are used in this specification to indicate the presence of one or more elements / components / etc.; the terms “including” and “having” are used to indicate an open-ended inclusion and to mean that there may be other elements / components / etc. in addition to the listed elements / components / etc.; the terms “first” and “second” are used only as markings and are not a limitation on the number of objects.
[0031] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0032] Predicting medical risks can assist relevant personnel in conducting subsequent analysis and handling based on the predicted risks. For example, it can help insurance companies determine insurance premiums based on predicted medical risks.
[0033] Among related technologies, the accuracy of medical risk prediction is relatively low, especially in accurately predicting the medical risks of people related to the business.
[0034] Urban insurance is a policy for basic medical insurance in each city's unified planning area. On the basis of basic medical insurance reimbursement, it provides secondary reimbursement for personal out-of-pocket expenses and / or personal medical expenses to reduce the personal medical burden.
[0035] Taking the analysis of medical risks for urban insured populations as an example, related technologies directly use the historical medical expenses of basic medical insurance participants to predict the medical expenses of urban insured populations, thereby determining the medical risks of urban insured populations based on the predicted medical expenses.
[0036] However, since the participants in basic medical insurance and urban insurance are not entirely the same, and there may even be significant group differences, directly using the medical expenses of basic medical insurance participants to predict the medical risks of urban insurance participants will lead to inaccurate prediction results.
[0037] In the embodiments of this disclosure, a medical risk prediction method is first provided, which at least to some extent improves the deficiencies existing in the above-mentioned prior art.
[0038] Figure 1 This diagram illustrates a flowchart of a medical risk prediction method in an exemplary embodiment of this disclosure, with reference to... Figure 1 The method includes:
[0039] Step S110: Based on the correlation between multiple first candidate features of the historical insured population in the target city and the city's insurance claims information, determine the sampling features from the multiple first candidate features;
[0040] Step S120: Determine the distribution information of the sampling characteristics of the historical insured population, and extract the target population from the population with medical records in the target city according to the distribution information.
[0041] Step S130: Based on the feature values of the target characteristics of the target population in the first historical period and the personal medical expenses of the target population in the second historical period, a personal medical expense prediction model is trained, wherein the second historical period is after the first historical period.
[0042] Step S140: Based on the personal medical expense prediction model, predict the target medical expenses of the population to be insured, so as to determine the medical risk of the population to be insured based on the target medical expenses.
[0043] exist Figure 1In the technical solution provided by the illustrated embodiment, sampling characteristics are determined based on the relevance to claims information of urban medical insurance. Then, a target population is extracted from the medical group based on the distribution of sampling characteristics of historical insured individuals. The target population's target characteristics are used as training samples to train a medical expense prediction model. The urban insurance cost is then determined based on the medical expenses predicted by the model, and the medical risk of the prospective insured population is determined based on the target medical expenses. Compared with related technologies, this disclosure, through sampling characteristics, ensures that the distribution of the target population used as training samples is consistent with that of historical insured individuals, thereby improving the accuracy of the model's prediction of medical expenses for the prospective insured population, and thus improving the accuracy of medical risk prediction. Simultaneously, this disclosure uses historical data of the target population to train a cost prediction model capable of predicting future medical expenses, thus taking into account the impact of potential risks on future expenses, further improving the accuracy of medical expense prediction, and thus further improving the accuracy of medical risk prediction.
[0044] The following are Figure 1 The specific implementation methods of each step in the illustrated embodiment are described in detail below:
[0045] In step S110, sampling features are determined from the multiple first candidate features based on the correlation between the historical insured population of the target city and the city's insurance claims information.
[0046] In one exemplary implementation, city insurance can be understood as city medical insurance. City medical insurance can supplement basic medical insurance, reimbursing out-of-pocket expenses and / or personal out-of-pocket costs covered by basic medical insurance, thereby reducing the medical burden. City medical insurance can also directly reimburse total medical expenses; this exemplary implementation does not specifically limit this. The target city can be customized according to user needs, and it can be determined based on administrative regions, such as any provincial-level region or any municipal-level region.
[0047] Historical insured individuals can be understood as those who have participated in city insurance. The primary candidate characteristics can include those who have participated in city insurance, such as gender, age, medical history, blood pressure, blood sugar, blood lipids, historical medical expenses, and historical claims expenses from city insurance.
[0048] From multiple primary candidate features, features with a high degree of relevance to urban insurance claims information can be selected as sampling features. These sampling features can then be used to obtain training samples in subsequent steps. For example, Figure 2 A flowchart illustrating a method for determining sampling characteristics in an exemplary embodiment of this disclosure is shown. (See also:) Figure 2 The method may include steps S210 to S230.
[0049] In step S210, the historical feature values of multiple first candidate features of the historical insured population in the target city and the historical claim expenses of the city insurance corresponding to the historical insured population are obtained.
[0050] In step S220, for each first candidate feature, the correlation between the first candidate feature and the city insurance claim cost is determined based on the historical feature value of the first candidate feature and the historical claim cost of the city insurance.
[0051] In step S230, the first candidate features are sorted in descending order of relevance, and the first candidate feature ranked in the top M positions is determined as the sampling feature.
[0052] For example, we can obtain the historical feature values of multiple first candidate features of the insured population in multiple different historical periods and the corresponding historical claims expenses of the city insurance. For example, we can obtain the feature values of multiple first candidate features of the insured population for each year in the past 3 years and the city insurance claims expenses of each person in the insured population in that year. Then, for each first candidate feature, we can calculate 3 correlation degrees based on the data of each year, and take a weighted average of the 3 correlation degrees to obtain the target first correlation degree of each first candidate feature. Then, we sort the multiple first candidate features in descending order according to the target correlation degree, and select the first candidate features in the top M positions as sampling features.
[0053] The degree of correlation can be determined by any method that can measure the correlation between two variables, such as the covariance coefficient, Pearson coefficient, maximum mutual information coefficient, rank correlation coefficient, etc. This exemplary implementation does not impose any special limitations on this.
[0054] Through steps S210 to S230, features with a high degree of correlation with urban insurance claims information can be identified from multiple first candidate features. These features are then used as sampling features, and training samples are determined based on the sampling features in subsequent steps. This avoids the problem of low model prediction accuracy caused by large differences between the training samples and the insured population of urban insurance.
[0055] Next, in step S120, the distribution information of the sampling characteristics of the historical insured population is determined, and the target population is extracted from the population with medical records in the target city based on the distribution information.
[0056] Once the sampling characteristics are determined, the distribution information of these characteristics among historical insured individuals can be statistically analyzed. For example, if the final determined sampling characteristics are gender, age, and pre-existing conditions (e.g., whether the insured had malignant tumors, coronary heart disease, or other critical illnesses before enrollment), the gender distribution, age distribution, and pre-existing condition distribution among historical insured individuals can be statistically analyzed. This includes the ratio of men to women, the proportion of individuals in different age groups, and the proportion of individuals who had malignant tumors, coronary heart disease, or other critical illnesses before enrollment.
[0057] In one exemplary implementation, the population with medical records can be understood as the population whose historical medical data can be obtained. The sampling characteristics of the target population and the distribution information of the sampling characteristics of the historical insured population are similar or consistent.
[0058] For example, the proportion of different characteristic values of the sampling feature in the target population is the same as the proportion of different characteristic values of the sampling feature in the historical insured population, or the difference between the two is within a certain preset threshold. For instance, if the distribution of the sampling feature of gender in the historical insured population is 70% male, then the proportion of males in the target population is 70% or close to 70%, such as any value between 65% and 75%.
[0059] Since the number of historical insured individuals and the available dimensions of their features may be limited, directly training a personal medical expense prediction model based on the characteristics of historical insured individuals may result in low prediction accuracy.
[0060] Therefore, training samples can be determined based on individuals with medical records in the target city. However, the distribution of characteristics among individuals with medical records and the distribution of characteristics among the city's insurance participants may not be entirely consistent. This discrepancy could lead to training data that fails to truly represent the attributes of the city's insurance participants, resulting in insufficient model prediction accuracy.
[0061] Based on this, this disclosure utilizes business experience and single-factor statistical methods to identify features strongly correlated with urban insurance claims information, using these features as sampling features. Then, the distribution of these sampling features among historical insured individuals is determined. Following this distribution, a target population is sampled from those with medical records in the target city. This target population forms the training sample, satisfying the requirement for a sufficient number of training samples for the training model while ensuring that the determined training sample truly represents the attributes of the urban insurance insured population, thereby improving the accuracy of predicting medical expenses for potential urban insurance insured individuals.
[0062] In one exemplary implementation, sampling can be performed using MCMC (Monte Carlo Simulation and Markov Chain) methods, such as Gibbs sampling, to obtain a sample set similar to the characteristics of the population participating in urban insurance, thereby simulating the characteristics of the population participating in urban insurance.
[0063] The Gibbs sampling method is a multivariate distribution MCMC method suitable for solving multivariate distribution problems. It selects initial values from each dimension of the multivariate distribution, samples from each dimension simultaneously, and obtains the final sample after a probability stationary process.
[0064] In this disclosure, the Gibbs sampling method can more stably and scientifically simulate the distribution characteristics of the population participating in urban insurance, providing a scientific, reasonable, and more representative sample basis for the subsequent modeling process, and improving the accuracy of model determination.
[0065] Next, continue to refer to Figure 1 In step S130, a personal medical expense prediction model is trained based on the feature values of the target characteristics of the target population in the first historical period and the personal medical expenses of the target population in the second historical period.
[0066] The second historical period follows the first historical period. For example, if the first historical period is 2019, then the second historical period could be 2020.
[0067] For example, based on the time span and data quality of the acquired historical medical data, it can be divided into observation period data and validation period data. Then, the feature values of the target features from the observation period data can be used as training data, and the personal medical expenses from the validation period data can be used as label data to generate training samples to train a personal medical expense prediction model. Here, the observation period can be understood as the first historical period mentioned above, and the validation period can be understood as the second historical period mentioned above.
[0068] In one exemplary implementation, the length of the second historical period can be customized as needed, such as 3 months, 6 months, 9 months, 12 months, etc. This allows for the creation of individual medical expense prediction models corresponding to multiple time windows, enabling the prediction of individual medical expenses for different time windows.
[0069] In one exemplary implementation, personal medical expenses include one or more of total personal medical expenses, out-of-pocket medical expenses, and out-of-pocket medical expenses. Correspondingly, when personal medical expenses include total personal medical expenses, the target feature may include a first target feature corresponding to the total personal medical expenses. When personal medical expenses include out-of-pocket expenses, the target feature may include a second target feature corresponding to the out-of-pocket medical expenses. When personal medical expenses include out-of-pocket expenses, the target feature may include one or more of a third target feature corresponding to the out-of-pocket medical expenses.
[0070] Out-of-pocket medical expenses can be understood as medical expenses incurred by an individual beyond the portion reimbursed by basic medical insurance. Expenses not covered by basic medical insurance can be understood as expenses incurred by an individual that are not covered by basic medical insurance. For example, if a patient uses certain medications that are not covered by basic medical insurance, the expenses incurred for those medications will be considered out-of-pocket medical expenses.
[0071] For example, Figure 3 This diagram illustrates a flowchart of a method for determining a first target feature in an exemplary embodiment of this disclosure. (See reference...) Figure 3 The method may include steps S310 to S330.
[0072] In step S310, historical medical data is acquired, and the individual features and derived features in the historical medical data are dimensionality reduced to obtain the second candidate features.
[0073] In one exemplary implementation, historical medical data includes the patient's gender, age, past medical history, family medical history, inpatient and outpatient diagnoses, surgeries and procedures, total inpatient and outpatient medical expenses, out-of-pocket medical expenses, out-of-pocket medical expenses, and various items in the medical expense charge details, such as generic names of drugs, drug dosages, drug costs, examination fees, and surgical fees.
[0074] For example, medical records from the past five years can be retrieved from the medical insurance system or other databases that access medical data. Sensitive personal information can be anonymized to create a dataset containing information such as the patient's name, age, gender, past medical history, family medical history, inpatient and outpatient diagnoses, surgeries and procedures, total inpatient and outpatient medical expenses, out-of-pocket medical expenses, and details of drug charges, including generic drug names, dosages, drug costs, examination fees, and surgical fees. Then, the historical medical data in this dataset can be preprocessed to obtain individual and derived features.
[0075] Preprocessing can include processes such as deduplication, missing value insertion, inaccurate data correction, and named entity name standardization to clean the data.
[0076] For example, deduplication can be understood as aggregating the medical records of a patient across different hospitals or institutions based on their unique identifier, such as ID card information, so that one medical record corresponds to one medical record. Missing value imputation can be understood as filling in missing fields in a medical record's data. For example, it can fill in the missing field values of the current medical record based on the values of the missing fields in other medical records with the highest feature similarity to the current medical record. Inaccurate data correction can be understood as modifying the values corresponding to obviously incorrect fields. For example, entering "female" for age is clearly incorrect and can be corrected to improve data quality. Named entity name standardization can be understood as unifying different names for the same thing based on natural language processing technology, such as standardizing different names for the same drug.
[0077] Of course, preprocessing may also include other data processing procedures that can improve data quality, and this exemplary embodiment does not impose any special limitations on this.
[0078] Furthermore, based on business experience and needs, single features and derived features can be generated from field names in the medical data. Single features can include attributes indicated by field names in the medical data, while derived features can include attributes obtained by combining attributes indicated by multiple field names in the medical data.
[0079] For example, individual features and derived features in historical medical data are determined by: discretizing continuous variables in the historical medical data to generate individual features in the historical medical data; and combining individual features in the historical medical data to obtain the derived features.
[0080] For example, age characteristics are correlated with personal medical expenses, and the personal medical expenses for different age groups vary greatly. Therefore, the chi-square binning method can be used to divide age into different age groups to generate age characteristics.
[0081] Similarly, derived features can be generated based on existing fields in medical data. For example, if the medical data contains monthly medical expenses, derived features such as the highest medical expenses in the past 6 months and the average medical expenses in the past 12 months can be generated by combining the monthly medical expenses.
[0082] After generating individual and derived features, dimensionality reduction can be performed on these features to obtain second candidate features. For example, methods such as PCA (principal component analysis) and LDA (linear discriminant analysis) can be used to reduce the dimensionality of individual and derived features to obtain second candidate features.
[0083] After obtaining the second candidate features, in step S320, the first correlation degree between each second candidate feature and the individual's total medical expenses is calculated.
[0084] For example, based on historical medical data, the first degree of correlation between each second candidate feature and the individual's total medical expenses can be calculated.
[0085] The degree of first correlation can also be determined by any method that can measure the correlation between two variables, such as the covariance coefficient, Pearson coefficient, maximum mutual information coefficient, rank correlation coefficient, etc. This exemplary implementation does not impose any special limitations on this.
[0086] In step S330, the second candidate features are sorted in descending order according to the first relevance, and the top K second candidate features are selected as the first target features based on the sorting result.
[0087] For example, multiple first relevance levels can be calculated using the feature values corresponding to the second candidate features from multiple different historical periods and the total personal medical expenses. Then, a weighted average of the multiple first relevance levels can be taken to obtain the target first relevance level. The second candidate features can be sorted in descending order according to the target first relevance level, and the top K second candidate features can be selected as the first target features based on the sorting results.
[0088] Of course, other filtering or packaging methods can also be used to select the first target feature from the second candidate features, and this exemplary embodiment does not impose any special limitations on this.
[0089] By using feature reduction and feature selection, the original data can be transformed into features that more accurately reflect the impact of medical costs, thereby improving the accuracy of the cost prediction model obtained through subsequent training.
[0090] When the target feature includes a second target feature, the second target feature can be determined based on the second candidate feature's second degree of correlation with the individual's out-of-pocket medical expenses.
[0091] When the target feature includes a third target feature, the third target feature can be determined based on the degree of correlation between the second candidate feature and the third personal out-of-pocket medical expenses.
[0092] The methods for determining the second and third target features are the same as those for determining the first target feature. The total personal medical expenses in steps S310 to S330 are simply replaced with the corresponding "personal out-of-pocket expenses" and "personal out-of-pocket costs". This will not be elaborated further here.
[0093] In one exemplary embodiment, the personal medical expense prediction model in this disclosure may include one or more of the following: a personal total medical expense prediction model, a personal out-of-pocket medical expense prediction model, and a personal self-paid medical expense prediction model.
[0094] Based on this, a specific implementation of step S130 may include: when the personal medical expenses include the total personal medical expenses, training a model for the total personal medical expenses based on the feature values of the first target feature of the target population in the first historical period and the cost range to which the total personal medical expenses of the target population belong in the second historical period; when the personal medical expenses include the out-of-pocket medical expenses, training a model for the out-of-pocket medical expenses based on the feature values of the second target feature of the target population in the first historical period and the cost range to which the out-of-pocket medical expenses of the target population belong in the second historical period; when the personal medical expenses include the out-of-pocket medical expenses, training a model for the out-of-pocket medical expenses based on the feature values of the third target feature of the target population in the first historical period and the cost range to which the out-of-pocket medical expenses of the target population belong in the second historical period.
[0095] For example, for city insurance, the focus should be on the overall claims situation rather than individual claims expenses. Therefore, for individual medical expense prediction models, the prediction target can be transformed into an individual medical expense range. This transforms the regression problem of medical expense prediction into a classification problem, reducing the computational complexity of the model and improving prediction efficiency.
[0096] For example, based on experience and the claims benchmarks in city insurance (such as the deductible), the total personal medical expenses, out-of-pocket medical expenses, and personal out-of-pocket medical expenses in the second historical period can be segmented into multiple expense ranges, such as: first expense range: 0-5000 yuan; second expense range: 5000-10000 yuan; third expense range: 10000-15000 yuan; fourth expense range: 15000-20000 yuan; fifth expense range: 20000-30000 yuan; sixth expense range: 30000-40000 yuan; seventh expense range: 40000-50000 yuan; and eighth expense range: above 50000 yuan.
[0097] For example, based on experience, preset cost ranges can be determined for total personal medical expenses, out-of-pocket medical expenses, and out-of-pocket medical expenses. Alternatively, preset cost ranges for different types of medical expenses can be determined separately using the chi-square binning method.
[0098] After obtaining the personal medical expenses of the target population in the second historical period, these expenses can be divided into corresponding personal medical expense ranges. Then, using the feature information of the target population's target characteristics in the first historical period as training samples, and the expense range to which the target population's personal medical expenses belonged in the second historical period as the labels of the training samples, a supervised learning training is performed on an initial model to obtain the corresponding personal medical expense prediction model.
[0099] Taking the prediction target as the range of out-of-pocket medical expenses for individuals as an example, some data from a partial sample set can be seen in Table 1.
[0100] Table 1 shows partial data from some training samples used to build a model for predicting out-of-pocket medical expenses.
[0101]
[0102] In one exemplary implementation, the personal medical expense prediction model may include LightGBM (LightGradient Boosting Machine, a distributed gradient boosting model based on decision trees) or LSTM (Long Short-Term Memory) models. LightGBM is an optimization model based on GBDT (Gradient Boosting Decision Tree). The algorithm uses regression trees as weak learners, obtaining the current residual regression tree by using the residual between each prediction result and the target value as the target for the next learning step. Each tree learns the conclusions and residuals of all previous trees, and the results of multiple decision trees are summed together as the final prediction output.
[0103] LightGBM uses a histogram algorithm with gradient one-sided sampling to pre-sort features, using the single gradient of a sample on a certain feature as the weight for training, and constructs a tree using node expansion. The algorithm implementation process is as follows: With N training samples, the first a% of the larger gradients (the gradient of the loss function with respect to the feature) are selected as training samples with large gradient values; from the remaining 1-a% of the smaller gradient values, b% are randomly selected as training samples with small gradient values; for the samples with smaller gradients, i.e., b%*N, they are amplified by (1-a) / b times when calculating the information gain. In summary, a%*N + b%*N samples are used as training samples. This construction is to maintain consistency with the overall data distribution as much as possible and ensure that samples with small gradient values are trained.
[0104] LSTM is a special type of recurrent neural network (RNN) primarily used to address the long-term dependency problem inherent in RNNs. Like RNNs, LSTM also has a chain-like structure.
[0105] The inventors of this disclosure compared the test results of several different types of machine learning models trained on the test set and found that the test results of LSTM and LightGBM both met the accuracy requirements. Therefore, in this disclosure, the personal medical expense prediction model can be either LSTM or LightGBM.
[0106] Furthermore, in this disclosure, when personal medical expenses include out-of-pocket medical expenses and / or out-of-pocket medical expenses, a personal total medical expense prediction model can also be trained, and then the accuracy and reasonableness of the prediction results of the out-of-pocket medical expense model and / or the out-of-pocket medical expense prediction model can be further verified based on the predicted personal total medical expense model.
[0107] For example, a personal total medical expense prediction model can be used to predict the total medical expenses of the target group. Then, a personal out-of-pocket expense prediction model can be used to predict the total out-of-pocket medical expenses of the target group. Based on the total out-of-pocket medical expenses and the total medical expenses, the predicted reimbursement ratio for the group's medical expenses can be determined. The predicted reimbursement ratio can be compared with the actual reimbursement ratio of basic medical insurance. If the difference is within the preset range, it indicates that the prediction results of the personal out-of-pocket medical expense prediction model are reliable. Otherwise, it indicates that the results are unreliable and the personal out-of-pocket medical expense prediction model needs to be adjusted.
[0108] Next, continue to refer to Figure 1 In step S140, the target medical expenses of the insured population are predicted based on the personal medical expense prediction model, so as to determine the medical risk of the insured population based on the target medical expenses.
[0109] In one exemplary implementation, the target medical expenses of the prospective insured population can be predicted in the following manner: predicting the number of prospective insured individuals based on the number of historical insured individuals; determining the distribution of target characteristics of the prospective insured population based on the distribution of target characteristics of historical insured individuals; sampling from medical subjects with medical records based on the number of prospective insured individuals and the distribution of target characteristics of the prospective insured population to obtain the prospective insured population; predicting the individual target medical expenses of each prospective insured individual in the prospective insured population based on the individual medical expense prediction model; and determining the target medical expenses of the prospective insured population based on the individual target medical expenses of each prospective insured individual.
[0110] For example, the number of people awaiting enrollment can be predicted based on the changing trends or curves of the number of historical enrollees over multiple historical enrollment periods; the distribution of the target characteristics of the people awaiting enrollment can be determined based on the distribution of the target characteristics of the historical enrollees. Then, based on the predicted number of people awaiting enrollment and the distribution of the target characteristics of the people awaiting enrollment, a sample is taken from the medical population to obtain the people awaiting enrollment.
[0111] After identifying the potential enrollee population, for each individual within this population, the feature value of their target characteristic is obtained. This feature value is then input into the individual medical expense prediction model to obtain the individual target medical expense for that individual. Finally, the individual target medical expenses for each individual are summed to obtain the target medical expense for the entire potential enrollee population.
[0112] Among them, the group of people who are to be insured can be understood as people who may participate in urban insurance in the future.
[0113] Once the target medical expenses for the prospective insured population are determined, their medical risk can be assessed based on these expenses. For example, if the target medical expenses exceed a first preset value, the prospective insured population is classified as having a first-level medical risk; if the target medical expenses are less than or equal to the first preset value, the prospective insured population is classified as having a second-level medical risk. The first-level risk is considered higher than the second-level risk.
[0114] Alternatively, if the target medical expenses fall within the first interval, the medical risk of the prospective insured population is determined as the first risk level; if the target medical expenses fall within the second interval, the medical risk is determined as the second risk level; and if the target medical expenses fall within the third interval, the medical risk is determined as the third risk level. In this case, the minimum value of the first interval is greater than the maximum value of the second interval, the minimum value of the second interval is greater than the maximum value of the third interval, the risk level of the first risk level is higher than that of the second risk level, and the risk level of the second risk level is higher than that of the third risk level.
[0115] In one exemplary application scenario disclosed herein, after determining the medical risks of the prospective insured population, the cost of city insurance can be determined based on these risks. For example, different city insurance costs corresponding to different medical risk levels can be pre-configured based on experience or historical claims data. After determining the medical risk level of the prospective insured population, the corresponding city insurance cost can be queried based on the determined medical risk level, thereby determining the city insurance cost based on the medical risk level. Alternatively, after determining the medical risks of the prospective insured population, the city insurance cost can be determined directly based on the determined medical risk level combined with historical experience.
[0116] In another exemplary application scenario of this disclosure, the medical risk of a prospective insured person can also be determined based on a medical expense prediction model. This prospective insured person can be understood as someone applying to participate in city insurance. Then, based on the determined medical risk of the prospective insured person, their application for insurance is reviewed. For example, the feature values of the prospective insured person's target characteristics are obtained and input into the personal medical expense prediction model to obtain the prospective insured person's predicted personal medical expenses. If the predicted personal medical expenses exceed a preset value, the prospective insured person's risk level is determined to be Level 1 risk; otherwise, it is determined to be Level 2 risk. When the prospective insured person's risk level is Level 2 risk, the application is directly approved, i.e., the application for insurance is agreed upon. When the prospective insured person's risk level is Level 1 risk, the application is transferred to a manual channel for review.
[0117] In another exemplary application scenario disclosed herein, the characteristics of historical insured individuals can be analyzed to extract common features. Then, the characteristics of candidate insured individuals are matched with these common features. If a match is successful, the candidate insured individual is designated as the target individual. The feature values of the target individual's target features are input into a personal medical expense prediction model to predict the target individual's medical expenses. If the medical expenses exceed a preset value, the individual is classified as a first risk level; otherwise, a second risk level. Then, relevant information about city insurance is pushed to target individuals at the second risk level, enabling them to determine whether to participate in city insurance based on the pushed information.
[0118] In another exemplary application scenario of this disclosure, after obtaining the target medical expenses of the population to be insured, it can assist relevant personnel of the city insurance in directly determining the city insurance expenses based on the target medical expenses of the population to be insured.
[0119] For example, Figure 4 This diagram illustrates a flowchart of a method for determining city insurance costs in an exemplary embodiment of this disclosure. (See reference...) Figure 4The method may include steps S410 to S450. Wherein:
[0120] In step S410, the personal medical expense prediction model is used to determine the expense range to which the personal medical expenses of the person to be insured belong, so as to determine the distribution of the person to be insured in each expense range.
[0121] The method for determining the population to be insured can refer to the relevant content in step S140 above, and will not be repeated here.
[0122] For example, the type of the personal medical expense prediction model in step S410 can be determined according to the claim conditions of the city insurance. For instance, if the claim conditions of the city insurance are to reimburse personal out-of-pocket medical expenses, then the personal medical expense prediction model is determined to be the personal out-of-pocket medical expense prediction model. If the claim conditions of the city insurance are to reimburse personal out-of-pocket medical expenses, then the personal medical expense prediction model is determined to be the personal out-of-pocket medical expense prediction model, and so on.
[0123] Taking the claim conditions of urban insurance as the claim for out-of-pocket medical expenses as an example, the target characteristics of the prospective insured population are input into the out-of-pocket expense prediction model to determine the out-of-pocket expense range for each prospective insured person.
[0124] After obtaining the personal out-of-pocket expense range for each person to be insured, the distribution of the out-of-pocket expense ranges for the population to be insured is statistically analyzed to determine the number of people to be insured included in each preset personal out-of-pocket expense range.
[0125] Next, in step S420, for each of the cost intervals, the individual compensation cost for the cost interval is determined based on the midpoint of the cost interval and the claim basis of the city insurance.
[0126] In one exemplary implementation, the median of the expense range can be understood as the average of the two endpoints of the expense range. For example, if the expense range is 10,000-15,000, the median is 12,500; if the expense range is 30,000-40,000, the median is 35,000. The claim basis for city insurance can include the deductible and reimbursement ratio. Taking a city insurance deductible of 10,000, a reimbursement ratio of 50%, and a claim condition of out-of-pocket medical expenses as an example, the meaning of this claim basis can be understood as follows: within the validity period of the city insurance, the portion of out-of-pocket medical expenses exceeding 100 million will be reimbursed at a ratio of 50%. For example, if the out-of-pocket medical expenses are 50,000, the corresponding personal reimbursement amount is (50,000-10,000)*50% = 20,000.
[0127] For each expense range, the individual claim expense for that range can be determined based on the difference between the median of the expense range and the claim benchmark for city insurance.
[0128] Then, in step S430, for each of the cost intervals, the first total claim cost for the cost interval is determined based on the individual claim cost and the distribution of the insured population in the cost interval.
[0129] For example, the first total claim amount for each claim range can be determined by multiplying the individual claim amount corresponding to each claim range by the number of insured persons included in that claim range.
[0130] In step S440, the second total claim cost for the insured population is determined based on the first total claim cost for each of the cost ranges.
[0131] For example, by summing the first total claim expenses corresponding to each expense range, we can determine the total claim expenses for all individuals in the prospective insured population, which is the second total claim expense.
[0132] In step S450, based on the second general claim cost and the number of people to be insured, the average claim cost of the people to be insured is determined, so as to determine the cost of the city insurance according to the average claim cost and a preset adjustment factor.
[0133] For example, by dividing the second-highest claim cost by the number of people in the uninsured population, we can obtain the average claim cost for the uninsured population. Then, based on the average claim cost for the uninsured population and the preset adjustment factor, we can determine the cost of the city insurance.
[0134] The preset adjustment factor can be determined based on the operating costs of the city insurance, and the operating costs can be determined based on historical operating experience.
[0135] Through steps S410 to S450 above, the average claim cost corresponding to different claim standards can be obtained. Then, based on the different claim costs corresponding to different claim standards, the corresponding city insurance cost can be determined, thereby improving the flexibility and efficiency of city insurance cost determination.
[0136] Furthermore, as mentioned above, this disclosure can provide a personal medical expense prediction model corresponding to different time windows. Therefore, this disclosure can also predict the claim expenses of insured individuals within different time windows during the insurance period, and then determine whether it is necessary to adjust the claim liability for diseases or drugs in real time based on the claim expenses within different time windows.
[0137] If, based on the insured population, it is predicted that the payout ratio corresponding to their claims expenses within 12 months is far lower than the target payout ratio, then certain diseases or drugs that were not originally covered by the city insurance can be added to the city insurance's payout coverage. Furthermore, the payout expenses after adding a certain disease or drug to the payout coverage can be predicted, thereby controlling the city insurance payout ratio within the expected range to avoid the city insurance payout ratio being too high or too low.
[0138] In one exemplary implementation, the range of personal medical expenses for insured individuals within the insurance period can also be predicted, and based on the predicted range of medical expenses, it can be determined whether to push prevention and control information of the target disease to the client of the insured individual.
[0139] Figure 5 This diagram illustrates a flowchart of a method for pushing information to insured individuals in an exemplary embodiment of this disclosure. (See reference...) Figure 5 The method may include steps S510 to S530.
[0140] In step S510, the importance of the target feature is determined according to the distributed gradient boosting model based on decision tree, and the target feature ranked in the top N by importance is selected as the warning feature.
[0141] In one exemplary embodiment, when the personal medical expense prediction model in this disclosure includes the aforementioned LightGBM distributed gradient boosting model based on decision trees, LightGBM can rank the input target features, and the ranking result can reflect the importance of the target features. Therefore, the target features can be determined based on the top N target features in the ranking result of LightGBM.
[0142] In step S520, the personal medical expenses of the insured person are determined based on the personal medical expense prediction model. When the personal medical expenses meet the preset conditions, the insured person is identified as an insured person to be warned.
[0143] For example, the target characteristics of insured individuals can be input into a personal medical expense prediction model to obtain the predicted range of personal medical expenses for each insured individual. When the predicted range falls within the target range, the insured individual is identified as a candidate for early warning. Insured individuals can be understood as those already enrolled in medical insurance.
[0144] The target cost range can be determined based on the deductible in the city's insurance claim criteria. For example, the target cost range can include cost ranges where the value is greater than the deductible.
[0145] Next, in step S530, when the warning characteristics of the insured person to be warned exceed the warning threshold, prevention and control information of the disease indicated by the warning characteristics is pushed to the client corresponding to the insured person to be warned.
[0146] For example, for insured individuals awaiting early warning, their early warning characteristic values can be monitored or queried. When the characteristic value exceeds the warning threshold, prevention and control information for the corresponding ailment can be pushed to the client of the insured individual awaiting early warning. Taking blood pressure as an example, after identifying insured individuals awaiting early warning, their blood pressure values can be monitored. For instance, these individuals could report their blood pressure values daily on their client. When their blood pressure value exceeds the corresponding warning threshold, management suggestions for blood pressure control can be sent to their client.
[0147] Through steps S510 to S530 described above, potential high-risk groups can be identified based on the predicted personal medical expenses of urban insurance participants and the corresponding warning characteristic values. Management recommendations can then be made for these potential high-risk groups. This can help improve the controllability of urban insurance claims costs, ensuring that this difficult-to-control target can be kept within the expected range.
[0148] Furthermore, Figure 6 This diagram illustrates the framework of a medical cost prediction system according to an exemplary embodiment of this disclosure. (Reference) Figure 6 The system framework may include a data access layer 61, a data aggregation and cleaning subsystem 62, a model sample layer 63, a machine learning model layer 64, and an application layer 65.
[0149] The data access layer 61 is used to acquire different medical data from different channels and send the different medical data from different channels to the data aggregation and cleaning subsystem 62 through a "batch channel".
[0150] The data aggregation and cleaning subsystem 62 can perform data aggregation and cleaning on the medical data sent by the data access layer 61 based on the Hadoop framework (an open-source framework for processing, storing, and analyzing massive amounts of distributed, unstructured data) to improve the data quality of the medical data. For example, the data aggregation and cleaning subsystem 62 can extract candidate features from the medical data through processes such as "data reading - data cleaning - feature extraction". Here, data cleaning can be understood as the preprocessing process mentioned above, which will not be elaborated here.
[0151] Model sample layer 63 is used to extract the target population from the dataset and cleaned medical data based on the distribution of sampling characteristics of the historical insured population, using the Gibbs sampling method, to obtain a sample set of simulated urban insurance participants. This sample set is then divided into a feature set and a label set to generate training data. The training data is further divided into a training set and a test set, facilitating the training of the personal medical expense prediction model by the machine learning model layer based on these sets. The feature set can be determined based on the characteristic data of the target population in the first historical period, and the label set can be determined based on the personal medical expenses of the target population in the second historical period.
[0152] Machine learning model layer 64 can train machine learning or deep learning models through processes such as "feature selection, model parameter tuning, model evaluation, and iterative optimization" to obtain a personal medical cost prediction model.
[0153] Application layer 65 is used to determine whether an individual is a person requiring early warning based on the personal medical expense prediction model obtained from machine learning model layer 64, thereby providing health management tips to such individuals. It can also predict the total medical expenses of the population awaiting insurance based on the personal medical expense prediction model, from a group perspective.
[0154] In this disclosure, machine learning or deep learning models can be used to comprehensively consider the impact of multiple dimensions on individual medical expenses, thereby improving the accuracy of individual medical expense prediction. Simultaneously, by simulating the distribution of urban insurance participants, the accuracy of predicting claims costs for urban insurance participants can be improved, thereby enhancing the accuracy of determining the medical risks of urban insurance participants.
[0155] In exemplary application scenarios, the cost of urban insurance can be determined based on the predicted medical expenses of the prospective insured population, thereby improving the accuracy of urban insurance cost prediction. Furthermore, during the validity period of urban insurance, the personal medical expense prediction model can be adjusted and optimized in stages based on the latest claims data, considering the impact of enrollment behavior and claims liability on personal medical expenses, and predicting the final payout ratio to achieve dynamic monitoring and early warning of the payout ratio. When the payout ratio deviates significantly from expectations, claims liability can be flexibly adjusted to keep it within the expected range as much as possible.
[0156] At the same time, it can also identify individuals requiring early warning based on their medical expenses, and conduct preventive and control management of related diseases based on the characteristics of these individuals, thereby achieving effective control and management of long-term medical expenses.
[0157] Those skilled in the art will understand that all or part of the steps of the above embodiments are implemented as a computer program executed by a CPU. When the computer program is executed by the CPU, it performs the functions defined by the method provided by the present invention. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk.
[0158] Furthermore, it should be noted that the above figures are merely illustrative representations of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0159] Figure 7 A schematic diagram of the structure of a medical risk prediction device in an exemplary embodiment of this disclosure is shown. (Reference) Figure 7 The device 700 may include a sampling feature determination module 710, a target population determination module 720, a model training module 730, and a medical risk prediction module 740. Wherein:
[0160] The sampling feature determination module 710 is configured to determine sampling features from the multiple first candidate features based on the correlation between the historical insured population of the target city and the city's insurance claims information.
[0161] The target population determination module 720 is configured to determine the distribution information of the sampling characteristics of the historical insured population, and to extract the target population from the population with medical records in the target city based on the distribution information.
[0162] The model training module 730 is configured to train a personal medical expense prediction model based on the feature values of the target characteristics of the target population in a first historical period and the personal medical expenses of the target population in a second historical period, wherein the second historical period is after the first historical period.
[0163] The medical risk prediction module 740 is configured to predict the target medical expenses of the insured population based on the personal medical expense prediction model, and determine the medical risk of the insured population based on the target medical expenses.
[0164] In some exemplary embodiments of this disclosure, based on the foregoing embodiments, the personal medical expenses include one or more of total personal medical expenses, out-of-pocket medical expenses, and out-of-pocket medical expenses; when the personal medical expenses include the total personal medical expenses, the target feature includes a first target feature, which is determined by: determining the first target feature based on a first correlation degree between a second candidate feature and the total personal medical expenses; when the personal medical expenses include the out-of-pocket medical expenses, the target feature includes a second target feature, which is determined by: determining the second target feature based on a second correlation degree between a second candidate feature and the out-of-pocket medical expenses; when the personal medical expenses include the out-of-pocket medical expenses, the target feature includes a third target feature, which is determined by: determining the third target feature based on a third correlation degree between a second candidate feature and the out-of-pocket medical expenses; wherein, the second candidate feature is determined by: acquiring historical medical data and performing dimensionality reduction on individual features and derived features in the historical medical data to obtain the second candidate feature.
[0165] In some exemplary embodiments of this disclosure, based on the foregoing embodiments, the historical medical data includes the patient's gender, age, past medical history, family medical history, inpatient and outpatient diagnoses, surgeries and procedures, total inpatient and outpatient medical expenses, out-of-pocket medical expenses, out-of-pocket medical expenses, and various items in the medical expense charge details, such as generic names of drugs, drug dosages, drug costs, examination fees, and surgical fees.
[0166] In some exemplary embodiments of this disclosure, based on the foregoing embodiments, the individual features and derived features in the historical medical data are determined by: discretizing the continuous variables in the historical medical data to generate individual features in the historical medical data; and combining the individual features in the historical medical data to obtain the derived features.
[0167] In some exemplary embodiments of this disclosure, based on the foregoing embodiments, the model training module 730 may further be specifically configured to: when the personal medical expenses include the total personal medical expenses, train a total personal medical expense model based on the feature values of the first target feature of the target population in the first historical period and the cost range to which the total personal medical expenses of the target population belong in the second historical period; when the personal medical expenses include the out-of-pocket medical expenses, train a out-of-pocket medical expense prediction model based on the feature values of the second target feature of the target population in the first historical period and the cost range to which the out-of-pocket medical expenses of the target population belong in the second historical period; when the personal medical expenses include the out-of-pocket medical expenses, train a out-of-pocket medical expense prediction model based on the feature values of the third target feature of the target population in the first historical period and the cost range to which the out-of-pocket medical expenses of the target population belong in the second historical period.
[0168] In some exemplary embodiments of this disclosure, based on the foregoing embodiments, the medical risk prediction module 740 can be specifically configured to: predict the number of people to be insured based on the number of historical insured persons; determine the distribution of target characteristics of the people to be insured based on the distribution of target characteristics of historical insured persons; sample from medical subjects with medical records based on the number of people to be insured and the distribution of target characteristics of the people to be insured to obtain the people to be insured; predict the personal target medical expenses of each person to be insured in the people to be insured based on the personal medical expense prediction model; and determine the target medical expenses of the people to be insured based on the personal target medical expenses of each person to be insured.
[0169] In some exemplary embodiments of this disclosure, based on the foregoing embodiments, the device further includes an information push module, which includes: a warning feature determination unit, configured to determine the importance of the target feature according to the decision tree-based distributed gradient boosting model, and select the target feature ranked in the top N by importance as a warning feature; a pre-warning insured person determination unit, configured to determine the insured person's personal medical expenses based on the personal medical expense prediction model, and determine the insured person as a pre-warning insured person when the personal medical expenses meet preset conditions; and a prevention and control information push unit, configured to push prevention and control information of the disease indicated by the warning feature to the client corresponding to the pre-warning insured person when the warning feature of the pre-warning insured person exceeds a warning threshold.
[0170] The specific details of each unit in the aforementioned medical risk prediction device have been described in detail in the corresponding medical risk prediction method, so they will not be repeated here.
[0171] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0172] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0173] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0174] In exemplary embodiments of this disclosure, a computer storage medium capable of implementing the above-described methods is also provided. It stores a program product capable of implementing the methods described in this specification. In some possible embodiments, various aspects of this disclosure may also be implemented as a program product including program code, which, when run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.
[0175] Embodiments of this disclosure may also include a program product for implementing the above methods. This program product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this disclosure is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0176] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0177] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0178] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0179] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0180] Furthermore, in an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.
[0181] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0182] The following reference Figure 8 To describe an electronic device 800 according to such an embodiment of the present disclosure. Figure 8 The electronic device 800 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0183] like Figure 8 As shown, the electronic device 800 is manifested in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one processing unit 810, at least one storage unit 820, a bus 830 connecting different system components (including storage unit 820 and processing unit 810), and a display unit 840.
[0184] The storage unit stores program code that can be executed by the processing unit 810, causing the processing unit 810 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 810 can perform actions such as... Figures 1 to 5 The steps shown are as follows.
[0185] Storage unit 820 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 8201 and / or cache memory 8202, and may further include a read-only memory (ROM) 8203.
[0186] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0187] Bus 830 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0188] Electronic device 800 can also communicate with one or more external devices 900 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 800, and / or with any device that enables electronic device 800 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 850. Furthermore, electronic device 800 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 860. As shown, network adapter 860 communicates with other modules of electronic device 800 via bus 830. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0189] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0190] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0191] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
Claims
1. A medical risk prediction method, characterized by, The method comprises the following steps: determining a sampling feature from a plurality of first candidate features of a historical insured population of a target city according to a correlation degree of the plurality of first candidate features and claim information of urban insurance; determining distribution information of the sampling feature of the historical insured population, and extracting a target population from a population having medical records in the target city according to the distribution information, wherein the sampling feature of the target population and the distribution information of the sampling feature of the historical insured population are the same; training a personal medical expense prediction model according to a feature value of a target feature of the target population in a first historical period and personal medical expenses of the target population in a second historical period, wherein the second historical period is after the first historical period; predicting target medical expenses of a to-be-insured population based on the personal medical expense prediction model to determine medical risks of the to-be-insured population according to the target medical expenses; wherein the determination manner of the target feature comprises: obtaining historical medical data, performing dimension reduction on single features and derived features in the historical medical data to obtain second candidate features, and determining the target feature according to a correlation degree between the second candidate features and personal medical expenses, wherein the historical medical data comprises a plurality of the following: gender, age, past medical history, family medical history, inpatient and outpatient diagnosis, surgery and operation, total medical expenses of inpatient and outpatient, personal medical expenses, personal medical expenses, drug generic name in medical expense item details, drug dosage, drug expenses, examination expenses, and surgical expenses.
2. The medical risk prediction method of claim 1, wherein, The personal medical expenses comprise one or more of the following: personal total medical expenses, personal medical expenses, and personal medical expenses; when the personal medical expenses comprise the personal total medical expenses, the target feature comprises a first target feature, and the first target feature is determined by the following manner: determining the first target feature according to a first correlation degree between the second candidate features and the personal total medical expenses; when the personal medical expenses comprise the personal medical expenses, the target feature comprises a second target feature, and the second target feature is determined by the following manner: determining the second target feature according to a second correlation degree between the second candidate features and the personal medical expenses; when the personal medical expenses comprise the personal medical expenses, the target feature comprises a third target feature, and the third target feature is determined by the following manner: determining the third target feature according to a third correlation degree between the second candidate features and the personal medical expenses.
3. The medical risk prediction method of claim 2, wherein, The single features and the derived features in the historical medical data are determined by the following manners respectively: discretizing continuous variables in the historical medical data to generate single features in the historical medical data; combining the single features in the historical medical data to obtain the derived features.
4. The medical risk prediction method of claim 2, wherein, The training of the personal medical expense prediction model according to the feature value of the target feature of the target population in the first historical period and the personal medical expenses of the target population in the second historical period comprises: When the personal medical expense comprises the total personal medical expense, the total personal medical expense model is trained according to the feature value of the first target feature of the target population in the first historical period and the expense interval to which the total personal medical expense of the target population in the second historical period belongs; When the personal medical expense comprises the personal out-of-pocket medical expense, the personal out-of-pocket medical expense prediction model is trained according to the feature value of the second target feature of the target population in the first historical period and the expense interval to which the personal out-of-pocket medical expense of the target population in the second historical period belongs; When the personal medical expense comprises the personal self-paid medical expense, the personal self-paid medical expense prediction model is trained according to the feature value of the third target feature of the target population in the first historical period and the expense interval to which the personal self-paid medical expense of the target population in the second historical period belongs.
5. The medical risk prediction method of claim 4, wherein, The method for predicting the target medical expense of the to-be-insured population based on the personal medical expense prediction model comprises: predicting the number of the to-be-insured population according to the number of the historical insured personnel; determining the distribution of the target features of the to-be-insured population based on the distribution of the target features of the historical insured personnel; sampling the to-be-insured population from the medical subjects with medical records according to the number of the to-be-insured population and the distribution of the target features of the to-be-insured population; predicting the personal target medical expense of each to-be-insured personnel in the to-be-insured population based on the personal medical expense prediction model; determining the target medical expense of the to-be-insured population according to the personal target medical expense of each to-be-insured personnel.
6. The medical risk prediction method of claim 1, wherein, When the personal medical expense prediction model comprises the distributed gradient boosting model based on the decision tree, the method further comprises: determining the importance degree of the target features according to the distributed gradient boosting model based on the decision tree, and selecting the target features with the top N importance degrees as the warning features; determining the personal medical expense of the insured personnel based on the personal medical expense prediction model, and determining the to-be-warned insured personnel when the personal medical expense meets the preset condition; pushing the prevention and control information of the disease indicated by the warning feature to the client corresponding to the to-be-warned insured personnel when the warning feature of the to-be-warned insured personnel exceeds the warning threshold.
7. A medical risk prediction apparatus characterized by comprising: comprises: a sampling feature determination module configured to determine sampling features from a plurality of first candidate features of a target city according to the relevance of the plurality of first candidate features to claim information of the city insurance; a target population determination module configured to determine distribution information of the sampling features of the historical insured population, and extract a target population from the population with medical records in the target city according to the distribution information, wherein the sampling features of the target population and the distribution information of the sampling features of the historical insured population are the same. The model training module is configured to train a personal medical expense prediction model according to the feature values of the target features of the target population in a first historical period and the personal medical expenses of the target population in a second historical period, the second historical period being after the first historical period; The medical risk prediction module is configured to predict target medical expenses of a to-be-insured population based on the personal medical expense prediction model, so as to determine the medical risk of the to-be-insured population according to the target medical expenses. The determination manner of the target features comprises: obtaining historical medical data, performing dimension reduction on single features and derived features in the historical medical data to obtain second candidate features, and determining the target features according to the correlation between the second candidate features and personal medical expenses, the historical medical data comprising multiple types of medical object features such as gender, age, past medical history, family medical history, inpatient and outpatient diagnosis, surgery and operation, total inpatient and outpatient medical expenses, personal self-payment medical expenses, personal self-financing medical expenses, drug generic name in medical expense charging item details, drug dosage, drug expenses, examination expenses, and surgery expenses.
8. A computer readable medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method of any one of claims 1 to 6.
9. An electronic device, comprising: Comprise: One or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, so that the one or more processors implement the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Medical big data analysis method and apparatus
CN107967948A
Medical expense data processing method and device
CN111815052A