Data prediction method, device, electronic device and computer-readable medium

Through the combination of pre-trained classification model and the combination of frequency and quantity distribution, the data prediction problem of incomplete factors in the existing technology is solved, and more accurate prediction results are achieved.

CN114757785BActive Publication Date: 2025-08-08BEIJING MEDICAL CROSS THE CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011591377.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-29
Publication Date
2025-08-08
Estimated Expiration
2040-12-29

AI Technical Summary

Technical Problem

The existing data prediction methods are not comprehensive enough in the medical and insurance fields, resulting in insufficient prediction results.

Method used

Through the pre-trained classification model and the pre-fitted frequency and quantity distribution, the event occurrence probability of pre-set special events and the relevant data of pre-set special items are respectively predicted. Combined with the event occurrence probability and the predicted value of the related guarantee data of the prediction object regarding the pre-set special items, the predicted value of the predicted object's related guarantee data about the pre-set special items is obtained.

Benefits of technology

It improves the accuracy of data prediction, can consider the probability of affecting events and behavioral distribution more comprehensively, and improves the accuracy of prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114757785B_ABST
    Figure CN114757785B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data prediction method, device, electronic device and computer-readable medium, and belongs to the field of data processing technology. The method includes: obtaining characteristic variables of the prediction object; inputting the characteristic variables of the prediction object into a pre-trained classification model to obtain the probability of occurrence of the prediction object regarding a preset special event; obtaining the frequency fitting distribution and quantity fitting distribution of the preset special items in the preset special event after the preset special event occurs; obtaining the predicted value of the relevant data of the preset special item based on the frequency fitting distribution and quantity fitting distribution of the preset special item; obtaining the predicted value of the relevant security data of the prediction object regarding the preset special item based on the probability of occurrence of the preset special event and the predicted value of the relevant data of the preset special item. The present disclosure can improve the accuracy of data prediction by predicting relevant data through classification models and fitting distributions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data prediction method, a data prediction device, an electronic device, and a computer-readable medium. Background Art

[0002] In medical, insurance and other related fields, it is often necessary to predict medical data or insurance data based on some existing information.

[0003] However, existing methods generally make predictions based on empirical data. However, since the factors considered are not comprehensive enough, the prediction results are often not accurate enough.

[0004] In view of this, there is an urgent need in this field for a data prediction method that can improve prediction accuracy.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0006] The purpose of the present disclosure is to provide a data prediction method, a data prediction device, an electronic device and a computer-readable medium, thereby improving the accuracy of data prediction results at least to a certain extent.

[0007] According to a first aspect of the present disclosure, a data prediction method is provided, comprising:

[0008] Get the characteristic variables of the prediction object;

[0009] Inputting the characteristic variables of the prediction object into a pre-trained classification model to obtain the probability of occurrence of the prediction object with respect to a preset special event;

[0010] Obtaining a frequency fitting distribution and a quantity fitting distribution of preset special items in the preset special event after the preset special event occurs;

[0011] Obtaining predicted values of relevant data of the preset special items based on the frequency fitting distribution and quantity fitting distribution of the preset special items;

[0012] According to the occurrence probability of the preset special event and the predicted value of the relevant data of the preset special item, the predicted value of the relevant security data of the prediction object regarding the preset special item is obtained.

[0013] In an exemplary embodiment of the present disclosure, the classification model training method includes:

[0014] Acquire training samples from a sample database, and construct a training sample set for the classification model based on the sample event types of the training samples and the characteristic variables corresponding to the training samples;

[0015] An independent variable is obtained according to the characteristic variables corresponding to the training samples in the training sample set, the sample event type is used as a dependent variable, and the classification model is trained according to the training sample set.

[0016] In an exemplary embodiment of the present disclosure, obtaining a training sample from a sample database includes:

[0017] Obtaining the variable names of the feature variables required for training the classification model;

[0018] Acquire a sample object from the sample database, and acquire a characteristic variable of the sample object according to the variable name;

[0019] The sample objects are filtered according to the preset screening conditions corresponding to the characteristic variables of the sample objects to obtain training samples.

[0020] In an exemplary embodiment of the present disclosure, after filtering the sample objects, the method further includes:

[0021] Determining a sampling classification variable from the variable name, and classifying the sample objects according to the sampling classification variable to obtain multiple sample object sets;

[0022] The sample objects in each of the sample object sets are sampled respectively to obtain the training samples.

[0023] In an exemplary embodiment of the present disclosure, the sample database includes real-world data.

[0024] In an exemplary embodiment of the present disclosure, the method for determining the frequency fitting distribution of the preset special items includes:

[0025] Acquire a fitting sample from a sample database, and obtain the frequency of the preset special item according to the fitting sample to obtain a sample frequency histogram of the preset special item;

[0026] Determining a candidate frequency fitting distribution of the preset special item according to the distribution of the sample frequency histogram;

[0027] If the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is less than or equal to a frequency difference threshold, the candidate frequency fitting distribution is determined as the frequency fitting distribution of the preset special item;

[0028] If the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is greater than the frequency difference threshold, the candidate frequency fitting distribution of the preset special item is re-determined.

[0029] In an exemplary embodiment of the present disclosure, the method for determining the quantity fitting distribution of the preset special items includes:

[0030] Acquire fitting samples from a sample database, and obtain the quantity of the preset special items within a preset time period based on the fitting samples to obtain a sample quantity histogram of the preset special items;

[0031] Determining a candidate quantity fitting distribution of the preset special item based on the distribution of the sample quantity histogram;

[0032] If the observed difference between the sample quantity histogram and the candidate quantity fitting distribution is less than or equal to a quantity difference threshold, the candidate quantity fitting distribution is determined as the preset special item quantity fitting distribution;

[0033] If the observed difference between the sample quantity histogram and the candidate quantity fitting distribution is greater than the quantity difference threshold, the candidate quantity fitting distribution of the preset special item is re-determined.

[0034] In an exemplary embodiment of the present disclosure, obtaining a predicted value of relevant data of the preset special item according to the frequency fitting distribution and the quantity fitting distribution of the preset special item includes:

[0035] Obtaining a frequency statistic of the preset special item according to a frequency fitting distribution of the preset special item, and obtaining a quantity statistic of the preset special item according to a quantity fitting distribution of the preset special item;

[0036] Single relevant data of the preset special item is obtained, and a predicted value of the relevant data of the preset special item is obtained based on the single relevant data of the preset special item, the frequency statistical value and the quantity statistical value.

[0037] In an exemplary embodiment of the present disclosure, obtaining the predicted value of the relevant guarantee data of the prediction object regarding the preset special item based on the probability of occurrence of the preset special event and the predicted value of the relevant data of the preset special item includes:

[0038] Obtaining the relevant data guarantee ratio of the prediction object regarding the preset special item;

[0039] According to the occurrence probability of the preset special event, the predicted value of the relevant data of the preset special item and the relevant data guarantee ratio, the predicted value of the relevant guarantee data of the prediction object regarding the preset special item is obtained.

[0040] According to a second aspect of the present disclosure, there is provided a data prediction device, comprising:

[0041] A feature variable acquisition module is used to obtain the feature variables of the prediction object;

[0042] An event probability prediction module is used to input the characteristic variables of the prediction object into a pre-trained classification model to obtain the event probability of the prediction object with respect to a preset special event;

[0043] A fitting distribution acquisition module, configured to acquire, after the occurrence of the preset special event, a frequency fitting distribution and a quantity fitting distribution of the preset special items in the preset special event;

[0044] A related data prediction module, configured to obtain a predicted value of the related data of the preset special item based on the frequency fitting distribution and the quantity fitting distribution of the preset special item;

[0045] The security data prediction module is used to obtain the predicted value of the security data related to the preset special item of the prediction object based on the probability of occurrence of the preset special event and the predicted value of the relevant data of the preset special item.

[0046] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned data prediction methods by executing the executable instructions.

[0047] According to a fourth aspect of the present disclosure, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements any one of the above-mentioned data prediction methods.

[0048] The exemplary embodiments of the present disclosure may have the following beneficial effects:

[0049] In the data prediction method of the example implementation method of the present disclosure, the probability of occurrence of a preset special event of the prediction object and the predicted value of the relevant data of the preset special items in the preset special event after the occurrence of the preset special event are obtained respectively through the pre-trained classification model and the pre-fitted frequency distribution and quantity distribution, thereby obtaining the predicted value of the relevant security data of the prediction object with respect to the preset special items. The data prediction method of the example implementation method of the present disclosure first predicts the probability of occurrence of an event through the classification model, and then predicts the relevant data of the preset special items generated when the event occurs through the pre-fitted frequency distribution and quantity distribution. This can more comprehensively consider the factors affecting the probability of occurrence of the event, measure the overall distribution of the behavior of the prediction object, and make the final prediction of the data more accurate.

[0050] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0052] Figure 1 A schematic flow chart showing a method for predicting data according to an exemplary embodiment of the present disclosure is provided;

[0053] Figure 2 A schematic diagram illustrating a flow chart of a method for training a classification model according to an exemplary embodiment of the present disclosure is shown;

[0054] Figure 3 A schematic diagram of a process for obtaining training samples according to an exemplary embodiment of the present disclosure is shown;

[0055] Figure 4 A schematic diagram of a process for determining distribution by KS test according to a specific embodiment of the present disclosure is shown;

[0056] Figure 5 A schematic diagram showing a process of determining a frequency fitting distribution according to an exemplary embodiment of the present disclosure is shown;

[0057] Figure 6 A schematic diagram showing a process of determining a quantity fitting distribution according to an exemplary embodiment of the present disclosure is shown;

[0058] Figure 7 A schematic diagram showing a process of determining a prediction value of relevant data according to an exemplary embodiment of the present disclosure is shown;

[0059] Figure 8 A schematic flow chart of a method for predicting data according to a specific embodiment of the present disclosure is shown;

[0060] Figure 9 A block diagram showing a data prediction apparatus according to an exemplary embodiment of the present disclosure;

[0061] Figure 10 A schematic structural diagram of a computer system suitable for implementing the electronic device according to the embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0062] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0063] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0064] In fields like healthcare and insurance, it's often necessary to predict medical or insurance data based on existing information. For example, in some implementations of specialty drug insurance, data on the medication use behavior and dosage of specialty drug users can be collected. Based on the law of large numbers, the average dosage and amount of medication used can be used to estimate the actual medication use behavior of these users. Further actuarial pricing of specialty drug insurance can then be calculated based on these averages.

[0065] However, using the law of large numbers to estimate medication collection behavior among specialty medication users may overlook the distributional characteristics of medication collection behavior. For example, when the actual medication collection distribution is skewed, estimating solely using the sample mean will be biased. Furthermore, in actual applications, samples may be missing or drop out of the study, making them less representative of the specialty medication population who consistently take their medications. Without extensive real-world clinical data to test hypotheses, it is difficult to guarantee the accuracy of the conclusions.

[0066] Based on the above problems, this exemplary embodiment first provides a data prediction method. Figure 1 As shown, the prediction method of the above data may include the following steps:

[0067] Step S110: Obtain characteristic variables of the prediction object.

[0068] Step S120: Input the characteristic variables of the prediction object into the pre-trained classification model to obtain the probability of occurrence of the prediction object with respect to the preset special event.

[0069] Step S130: Obtain frequency fitting distribution and quantity fitting distribution of preset special items in the preset special event after the preset special event occurs.

[0070] Step S140: Obtain predicted values of relevant data of the preset special items according to the frequency fitting distribution and quantity fitting distribution of the preset special items.

[0071] Step S150. According to the occurrence probability of the preset special event and the predicted value of the relevant data of the preset special item, the predicted value of the relevant security data of the prediction object regarding the preset special item is obtained.

[0072] In the data prediction method of the example implementation method of the present disclosure, the probability of occurrence of a preset special event of the prediction object and the predicted value of the relevant data of the preset special items in the preset special event after the occurrence of the preset special event are obtained respectively through the pre-trained classification model and the pre-fitted frequency distribution and quantity distribution, thereby obtaining the predicted value of the relevant security data of the prediction object with respect to the preset special items. The data prediction method of the example implementation method of the present disclosure first predicts the probability of occurrence of an event through the classification model, and then predicts the relevant data of the preset special items generated when the event occurs through the pre-fitted frequency distribution and quantity distribution. This can more comprehensively consider the factors affecting the probability of occurrence of the event, measure the overall distribution of the behavior of the prediction object, and make the final prediction of the data more accurate.

[0073] Next, combine Figures 2 to 7 The above steps of this exemplary embodiment are described in more detail.

[0074] In step S110 , characteristic variables of the prediction target are acquired.

[0075] In this example implementation, the prediction object refers to the person for whom relevant data prediction is required. For example, for a person who purchases special drug insurance, the price of the special drug insurance needs to be predicted, and the person who purchases the special drug insurance is the prediction object.

[0076] The characteristic variables of the prediction object refer to some attribute variables of the prediction object itself that are required to predict related data. Taking specialty drug insurance as an example, the characteristic variables of the prediction object generally include gender, age, past medical history, previous medication behavior and other characteristic variables, which can be used as the basic data for specialty drug insurance price prediction.

[0077] In step S120, the characteristic variables of the prediction object are input into the pre-trained classification model to obtain the event occurrence probability of the prediction object with respect to the preset special event.

[0078] In this example implementation, the classification model can determine whether a preset special event occurs in the prediction object and the probability of the preset special event occurring based on the input feature variables of the prediction object. For example, the classification model can determine the probability of a patient who purchases special drug insurance making a claim based on the input feature variables such as the patient's basic data and medical data. Figure 2 As shown, the training method of the classification model can specifically include the following steps:

[0079] Step S210: Acquire training samples from the sample database, and construct a training sample set for the classification model based on the sample event types of the training samples and the characteristic variables corresponding to the training samples.

[0080] In this example implementation, the sample event types of the training samples include two types: the occurrence of preset special events and the non-occurrence of preset special events. Taking special drug insurance as an example, the training samples are patient samples, and the sample event types include two types: the patient claims or does not claim. The characteristic variables corresponding to the training samples may include gender, age, past medical history, previous medication behavior, etc.

[0081] In this example implementation, the sample database includes real-world data. The National Center for Drug Evaluation defines "Real World Study, Real World Research, RWR" as: collecting RWD (Real World Data) related to patients in a real-world environment, and obtaining clinical evidence RWE (Real World Evidence) of the use value and potential benefits or risks of medical products through analysis. The main research types are observational studies, which can also be clinical trials. Therefore, in this example implementation, a sample set can be constructed based on the population in the real-world data. Since the real-world data comes from a real medical environment, it can reflect the actual diagnosis and treatment process and the health status of patients under real conditions.

[0082] Specifically, before establishing a training sample set, real-world data needs to be correlated, cleaned, and sampled. For example, in the case of specialty drug insurance, during model training, real-world data on medication collection for people who use a particular specialty drug is screened and sampled to obtain a representative sample set for further data analysis.

[0083] In this example implementation, Figure 3 As shown in the figure, obtaining training samples from the sample database may include the following steps:

[0084] Step S310: Obtain the variable names of the feature variables required for training the classification model.

[0085] Taking specialty drug insurance as an example, the data on specialty drug collection often comes from multiple data tables. For example, it is necessary to use the diagnosis table to frame the population that meets the special drug diagnosis requirements, obtain the doctor's prescription records through the medical order table, and obtain the in-hospital drug collection costs and dosage through the expense details table. Therefore, data association is required to determine the data fields for the required variable names. In addition to basic data such as gender, age, and previous illnesses, it can also include characteristic variables such as diagnosis results, collection date, collection amount, medical insurance type, and so on.

[0086] Step S320: Obtain a sample object from the sample database, and obtain the characteristic variables of the sample object according to the variable name.

[0087] After determining the variable names of the required characteristic variables, while obtaining sample objects from the sample database, it is also necessary to obtain the required characteristic variables of the sample objects from the sample database according to the data fields of the determined variable names.

[0088] Step S330: Filter the sample objects according to the preset screening conditions corresponding to the characteristic variables of the sample objects to obtain training samples.

[0089] After acquiring sample data, since there may be some errors or inconsistent data items in the data, it is necessary to clean and filter the data when constructing the sample set. For example, logical errors in the inspection data, including impossible birth dates, medication collection dates, inconsistent hospitalization records, and unreasonable medication payment records, can be eliminated. In addition, due to the possibility of patients abandoning treatment or transferring to other hospitals, considering the consistency of sample duration, when processing the data, filtering conditions such as medication collection time greater than the observation period can be added to ensure that the training samples are in a continuous medication collection state during the observation period.

[0090] In this example implementation, after data association and filtering of sample objects, if the sample size is sufficient, the sample objects can be sampled and then a sample set constructed. Specifically, a sampling categorical variable can be determined from the variable name, and the sample objects can be classified according to the sampling categorical variable to obtain multiple sample object sets. The sample objects in each sample object set can then be sampled to obtain training samples. For example, the sample objects can be classified according to categorical variables such as age and gender, and then sampled using stratified sampling. In addition, random sampling can also be used. The specific sampling method is not specifically limited in this example implementation.

[0091] Step S220: Obtain independent variables based on the characteristic variables corresponding to the training samples in the training sample set, take the sample event type as the dependent variable, and train the classification model based on the training sample set.

[0092] Taking specialty drug insurance as an example, after determining the training sample set, it is necessary to study and predict the insured's claim probability. In this example implementation, whether the specialty drug insurance claim conditions are met can be used as the dependent variable, and the influencing factor combination established by combining medical expert opinions and characteristic variables can be used as the independent variable, such as the presence of relevant medical history, gender, age, and previous medication collection behavior. The classification model can choose a logistic or probit model to predict the insured's claim probability.

[0093] Taking the Logistic model as an example, the specific form of Logistic regression is as follows:

[0094]

[0095]

[0096] Here, x is the independent variable, y is the dependent variable, p(y = 1|x) represents the probability of a claim, p(y = -1|x) represents the probability of not claiming, and b is a constant. Gradient descent can be used to estimate the coefficient ω, which in turn yields the probability of the insured making a claim. The estimated coefficients obtained using logistic regression can explicitly account for the contribution of each feature, providing a good fit and strong interpretability.

[0097] In step S130 , after the occurrence of the preset special event, the frequency fitting distribution and the quantity fitting distribution of the preset special items in the preset special event are obtained.

[0098] After determining the sample set, we can also study the distribution of the frequency and quantity of the sample's acquisition of pre-set special items. For example, for specialty drug insurance, we can study the sample's medication collection behavior and medication quantity data, and fit a distribution for subsequent use. The medication collection behavior can be, for example, the frequency of medication collection.

[0099] In this example implementation, the frequency fitting distribution and quantity fitting distribution of the preset special items can be obtained through the KS (Kolmogorov-Smirnov) test or the W (Shapiro-Wilk) test. By observing the distribution characteristics of the sample data, several distributions that meet the rules are selected for goodness of fit test, and finally the distribution with the highest fit is determined. Figure 4 As shown in the figure, taking the KS test as an example, the specific steps for studying the medication behavior of the sample and fitting the distribution of its medication behavior and medication dosage are as follows:

[0100] Step S410: Understand data characteristics through descriptive statistical analysis (histogram).

[0101] Step S420: Lock the distribution that can be used for fitting.

[0102] Step S430: Narrow the distribution range based on the PP graph.

[0103] Step S440: Perform KS goodness-of-fit test.

[0104] If the test fails, the process returns to step S420 to redetermine the distribution for fitting; if the test passes, the process proceeds to step S450.

[0105] Step S450: Select the distribution and calculate the base of the claim cost.

[0106] The Kolmogorov-Smirnov test can be used to test whether a distribution conforms to a theoretical distribution or to compare whether two empirical distributions differ significantly based on the cumulative distribution function. The null hypothesis H0 is: the two data distributions are consistent or the data conform to the theoretical distribution. The actual observation value D = max|f(x)-g(x)|, where f(x) is the cumulative frequency of the sample and g(x) is the cumulative probability of the theoretical distribution. When the actual observation value D>D(n,α), the H0 hypothesis is rejected; otherwise, the H0 hypothesis is accepted. Where n is the sample size and α is the confidence level, which can be values such as 0.05 or 0.005.

[0107] In this example implementation, Figure 5 As shown, the method for determining the frequency fitting distribution of the preset special items may specifically include the following steps:

[0108] Step S510: Obtain fitting samples from the sample database, and obtain the frequency of the preset special items according to the fitting samples to obtain a sample frequency histogram of the preset special items.

[0109] Fitting samples are samples obtained from a sample database to fit the frequency and quantity distribution of a predefined special item. The frequency of the predefined special item is obtained from the fitting samples, and then a sample frequency histogram of the predefined special item is generated based on the statistical frequency data. For example, you can obtain medication frequency data based on the sample's medication collection behavior and then generate a frequency histogram of the sample's medication collection behavior.

[0110] Step S520: Determine the candidate frequency fitting distribution of the preset special item according to the distribution of the sample frequency histogram.

[0111] After obtaining the sample frequency histogram, analyze the sample bias to identify multiple candidate frequency distributions that can be fitted. For example, if the sample data exhibits a long-tailed, right-skewed distribution, a gamma distribution can be considered. If the sample data exhibits a bell-shaped distribution with thick tails, a t distribution can be considered.

[0112] Step S530: If the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is less than or equal to the frequency difference threshold, the candidate frequency fitting distribution is determined as the preset frequency fitting distribution of the special item.

[0113] Subsequently, a test method such as Kolmogorov-Smirnov can be used to compare the degree of fit between the frequency distribution of the sample data and the candidate frequency fitting distribution. Specifically, it can be determined whether the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is less than or equal to a preset frequency difference threshold. If so, the current candidate frequency fitting distribution is directly determined as the preset frequency fitting distribution of the special item.

[0114] Step S540: If the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is greater than the frequency difference threshold, then re-determine the candidate frequency fitting distribution of the preset special item.

[0115] If the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is greater than the frequency difference threshold, it means that the degree of fit between the frequency distribution of the sample data and the candidate frequency fitting distribution does not meet the requirements, then return to step S520 to re-determine a candidate frequency fitting distribution and perform the fit judgment again until a candidate frequency fitting distribution that meets the requirements is found.

[0116] In this example implementation, Figure 6 As shown, the method for determining the quantity fitting distribution of preset special items may specifically include the following steps:

[0117] Step S610: Obtain fitting samples from the sample database, and obtain the number of preset special items within a preset time period based on the fitting samples to obtain a sample quantity histogram of the preset special items.

[0118] First, the number of pre-set special items is obtained based on the fitting sample, and then a sample quantity histogram of the pre-set special items is obtained based on the counted number. For example, the number of medicines taken can be obtained based on the sample's medicine-taking behavior, and then a histogram of the number of medicines taken by the sample can be generated.

[0119] Step S620: Determine the candidate quantity fitting distribution of the preset special items based on the distribution of the sample quantity histogram.

[0120] After obtaining the sample size histogram, analyze the sample bias to identify several candidate distributions that can be fitted. For example, if the sample data exhibits a long-tailed, right-skewed distribution, a gamma distribution can be considered. If the sample data exhibits a bell-shaped distribution with a thick tail, a t distribution can be considered.

[0121] Step S630: If the observed difference between the sample quantity histogram and the candidate quantity fitting distribution is less than or equal to the quantity difference threshold, the candidate quantity fitting distribution is determined as the preset special item quantity fitting distribution.

[0122] Subsequently, a Kolmogorov-Smirnov test method or other test method can be used to compare the degree of fit between the quantity distribution of the sample data and the candidate quantity fitting distribution. Specifically, it can be determined whether the observed difference between the sample quantity histogram and the candidate quantity fitting distribution is less than or equal to a preset quantity difference threshold. If so, the current candidate quantity fitting distribution is directly determined as the preset quantity fitting distribution of the special item.

[0123] Step S640: If the observed difference between the sample quantity histogram and the candidate quantity fitting distribution is greater than the quantity difference threshold, then the candidate quantity fitting distribution of the preset special item is re-determined.

[0124] If the observed difference between the sample quantity histogram and the candidate quantity fitting distribution is greater than the quantity difference threshold, it means that the degree of fit between the quantity distribution of the sample data and the candidate quantity fitting distribution does not meet the requirements, then return to step S620 to re-determine a candidate quantity fitting distribution and perform the fit judgment again until a candidate quantity fitting distribution that meets the requirements is found.

[0125] Taking specialty drug insurance as an example, applying a specific distribution to fit sample medication collection behavior and dosage before actuarial calculations can provide a more comprehensive assessment of the overall distribution, rather than simply considering the sample mean. If the frequency-fitting and quantity-fitting distributions for a specific item cannot be fitted by a single distribution, a combined distribution can be constructed to fit them. For example, pricing for specialty drug insurance can be achieved by selecting different loss distribution combinations that simultaneously describe the small, medium, and large loss components of the data.

[0126] In step S140 , a predicted value of relevant data of the preset special item is obtained according to the frequency fitting distribution and quantity fitting distribution of the preset special item.

[0127] In this example implementation, Figure 7 As shown, according to the frequency fitting distribution and quantity fitting distribution of the preset special items, the predicted value of the relevant data of the preset special items is obtained, which can specifically include the following steps:

[0128] Step S710: Obtain frequency statistics of the preset special items according to the frequency fitting distribution of the preset special items, and obtain quantity statistics of the preset special items according to the quantity fitting distribution of the preset special items.

[0129] The frequency statistics of the preset special items may be, for example, the average or median of the frequency obtained by frequency fitting distribution, and the quantity statistics of the preset special items may be, for example, the average or median of the quantity obtained by quantity fitting distribution.

[0130] Step S720: Acquire single relevant data of the preset special item, and obtain a predicted value of the relevant data of the preset special item based on the single relevant data, frequency statistics, and quantity statistics of the preset special item.

[0131] Taking specialty drug insurance as an example, the single relevant data refers to the unit price of the drug. Based on the unit price of the drug as well as the frequency statistics and quantity statistics, the predicted value of the drug expenses incurred under the conditions of the claim can be calculated.

[0132] In step S150, based on the occurrence probability of the preset special event and the predicted value of the relevant data of the preset special item, the predicted value of the relevant security data of the prediction object regarding the preset special item is obtained.

[0133] In this example implementation, the relevant data assurance ratio of the prediction object regarding the preset special item can be obtained first, and then the predicted value of the relevant assurance data of the prediction object regarding the preset special item can be obtained based on the probability of occurrence of the preset special event, the predicted value of the relevant data of the preset special item and the relevant data assurance ratio.

[0134] Taking specialty drug insurance as an example, by estimating the probability of a claim through a classification model and then fitting the distribution to obtain the predicted value of the drug costs incurred under the conditions of the claim, we can obtain the predicted value of the relevant protection data for the preset special items of the prediction object, that is, the specialty drug insurance price. For example, the premium for short-term reimbursement medical insurance is:

[0135]

[0136] Where p is the premium, q is the probability of the covered event occurring, k is the average claim cost within the coverage area, e is the expense surcharge rate, and t is the safety surcharge. For determining the rate of drug expense reimbursement in specialty drug insurance, the insured event rate q is the probability of a claim, and the average claim cost k is the average annual individual medical expense cost associated with specialty drugs. This is calculated by deducting the amount paid by the basic medical insurance pooling fund from the individual's annual medical expenses associated with specialty drugs and multiplying the result by the reimbursement ratio agreed upon in the commercial insurance contract, i.e., the relevant data coverage ratio.

[0137] like Figure 8 The figure shows a complete flow chart of a data prediction method in a specific embodiment of the present disclosure, which is applied to the pricing of specialty drug insurance and is an example of the above steps in this example embodiment. The specific steps of the flow chart are as follows:

[0138] Step S810: Sampling and constructing a sample set based on real-world data.

[0139] Construct a sample set and sample the population using specialty drugs in real-world data.

[0140] Step S820: Data cleaning and data association.

[0141] Step S830: Fitting the sample drug dosage and drug taking behavior distribution.

[0142] The medication behavior of the study samples was studied, and the distribution of medication taking behavior and medication amount was fitted.

[0143] Step S840: Actuarial calculation of special drug insurance prices.

[0144] The price of specialty drug insurance is actuarially calculated based on the fitted distribution of medication behavior and medication quantity.

[0145] It should be noted that although the steps of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all steps must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0146] Furthermore, the present disclosure also provides a data prediction device. Figure 9 As shown, the data prediction device may include a feature variable acquisition module 910, an event probability prediction module 920, a fitting distribution acquisition module 930, a related data prediction module 940, and a guarantee data prediction module 950. Among them:

[0147] The feature variable acquisition module 910 can be used to obtain the feature variables of the prediction object;

[0148] The event probability prediction module 920 can be used to input the characteristic variables of the prediction object into the pre-trained classification model to obtain the event probability of the prediction object regarding the preset special event;

[0149] The fitting distribution acquisition module 930 may be used to acquire the frequency fitting distribution and quantity fitting distribution of the preset special items in the preset special event after the preset special event occurs;

[0150] The relevant data prediction module 940 may be used to obtain a predicted value of relevant data of a preset special item based on the frequency fitting distribution and quantity fitting distribution of the preset special item;

[0151] The security data prediction module 950 can be used to obtain the predicted value of the security data related to the preset special item of the prediction object based on the probability of occurrence of the preset special event and the predicted value of the related data of the preset special item.

[0152] In some exemplary embodiments of the present disclosure, a data prediction device provided by the present disclosure may further include a classification model training module, which may include a training sample set construction unit and a classification model training unit.

[0153] The training sample set construction unit can be used to obtain training samples from the sample database and construct a training sample set for the classification model based on the sample event type of the training sample and the characteristic variables corresponding to the training sample;

[0154] The classification model training unit can be used to obtain independent variables based on the characteristic variables corresponding to the training samples in the training sample set, use the sample event type as the dependent variable, and train the classification model based on the training sample set.

[0155] In some exemplary embodiments of the present disclosure, the training sample set construction unit may include a variable name acquisition unit, a feature variable acquisition unit, and a sample object filtering unit.

[0156] The variable name acquisition unit can be used to obtain the variable names of the feature variables required for training the classification model;

[0157] The characteristic variable acquisition unit can be used to acquire a sample object from a sample database and acquire the characteristic variable of the sample object according to the variable name;

[0158] The sample object filtering unit can be used to filter the sample objects according to the preset screening conditions corresponding to each characteristic variable of the sample objects to obtain training samples.

[0159] In some exemplary embodiments of the present disclosure, the training sample set construction unit may further include a sample object classification unit and a sample object sampling unit.

[0160] The sample object classification unit can be used to determine a sampling classification variable from the variable name, and classify the sample objects according to the sampling classification variable to obtain multiple sample object sets;

[0161] The sample object sampling unit can be used to sample the sample objects in each sample object set respectively to obtain training samples.

[0162] In some exemplary embodiments of the present disclosure, a data prediction device provided by the present disclosure may further include a frequency fitting distribution determination module, which may include a frequency histogram determination unit, a candidate frequency fitting distribution determination unit, a frequency fitting distribution determination unit, and a candidate frequency fitting distribution update unit. Among them:

[0163] The frequency histogram determination unit may be used to obtain a fitting sample from a sample database, and obtain the frequency of a preset special item based on the fitting sample to obtain a sample frequency histogram of the preset special item;

[0164] The candidate frequency fitting distribution determining unit may be used to determine the candidate frequency fitting distribution of a preset special item according to the distribution of the sample frequency histogram;

[0165] The frequency fitting distribution determining unit may be configured to determine the candidate frequency fitting distribution as the preset frequency fitting distribution of the special item if the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is less than or equal to a frequency difference threshold;

[0166] The candidate frequency fitting distribution updating unit may be configured to redetermine the candidate frequency fitting distribution of the preset special item if the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is greater than a frequency difference threshold.

[0167] In some exemplary embodiments of the present disclosure, a data prediction device provided by the present disclosure may further include a quantity fitting distribution determination module, which may include a quantity histogram determination unit, a candidate quantity fitting distribution determination unit, a quantity fitting distribution determination unit, and a candidate quantity fitting distribution update unit. Among them:

[0168] The quantity histogram determination unit may be used to obtain fitting samples from a sample database, and obtain the quantity of preset special items within a preset time period based on the fitting samples, to obtain a sample quantity histogram of the preset special items;

[0169] The candidate quantity fitting distribution determining unit may be used to determine the candidate quantity fitting distribution of the preset special item according to the distribution of the sample quantity histogram;

[0170] The quantity fitting distribution determining unit may be configured to determine the candidate quantity fitting distribution as the preset special item quantity fitting distribution if an observed difference between the sample quantity histogram and the candidate quantity fitting distribution is less than or equal to a quantity difference threshold;

[0171] The candidate quantity fitting distribution updating unit may be configured to re-determine the candidate quantity fitting distribution of the preset special item if the observed difference between the sample quantity histogram and the candidate quantity fitting distribution is greater than a quantity difference threshold.

[0172] In some exemplary embodiments of the present disclosure, the relevant data prediction module 940 may include a statistical value determination unit and a prediction value determination unit.

[0173] The statistical value determination unit may be configured to obtain a frequency statistical value of the preset special item according to a frequency fitting distribution of the preset special item, and obtain a quantity statistical value of the preset special item according to a quantity fitting distribution of the preset special item;

[0174] The prediction value determination unit can be used to obtain single relevant data of the preset special item, and obtain the prediction value of the relevant data of the preset special item based on the single relevant data, frequency statistics and quantity statistics of the preset special item.

[0175] In some exemplary embodiments of the present disclosure, the guarantee data prediction module 950 may include a guarantee ratio acquisition unit and a guarantee data prediction unit.

[0176] The guarantee ratio acquisition unit can be used to obtain the guarantee ratio of relevant data of the prediction object regarding the preset special item;

[0177] The security data prediction unit can be used to obtain the predicted value of the relevant security data of the prediction object regarding the preset special item based on the probability of occurrence of the preset special event, the predicted value of the relevant data of the preset special item and the relevant data security ratio.

[0178] The specific details of each module / unit in the above data prediction device have been described in detail in the corresponding method embodiment part and will not be repeated here.

[0179] Figure 10 A schematic structural diagram of a computer system suitable for implementing an electronic device according to an embodiment of the present invention is shown.

[0180] It should be noted that Figure 10 The computer system 1000 of the electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0181] like Figure 10 As shown, computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for system operation are also stored in RAM 1003. CPU 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0182] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.

[0183] In particular, according to an embodiment of the present invention, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009 and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the various functions defined in the system of the present application are performed.

[0184] It should be noted that the computer-readable medium described in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0185] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0186] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the methods described in the following embodiments.

[0187] It should be noted that although several modules of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided into multiple modules to be embodied.

[0188] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0189] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A data prediction method, characterized in that: include: Get the characteristic variables of the prediction object; Inputting the characteristic variables of the prediction object into a pre-trained classification model to obtain the probability of occurrence of the prediction object with respect to a preset special event; Obtaining a frequency fitting distribution and a quantity fitting distribution of a preset special item in the preset special event after the preset special event occurs; wherein, if the frequency fitting distribution and the quantity fitting distribution of the preset special item cannot be fitted by a single distribution, then fitting them by selecting different loss distribution combinations and constructing a combined distribution; Obtaining predicted values of relevant data of the preset special items based on the frequency fitting distribution and quantity fitting distribution of the preset special items; According to the occurrence probability of the preset special event and the predicted value of the relevant data of the preset special item, the predicted value of the relevant security data of the prediction object regarding the preset special item is obtained.

2. The data prediction method according to claim 1, characterized in that: The training method of the classification model includes: Acquire training samples from a sample database, and construct a training sample set for the classification model based on the sample event types of the training samples and the characteristic variables corresponding to the training samples; An independent variable is obtained according to the characteristic variables corresponding to the training samples in the training sample set, the sample event type is used as a dependent variable, and the classification model is trained according to the training sample set.

3. The data prediction method according to claim 2, characterized in that: The step of obtaining a training sample from a sample database includes: Obtaining the variable names of the feature variables required for training the classification model; Acquire a sample object from the sample database, and acquire a characteristic variable of the sample object according to the variable name; The sample objects are filtered according to the preset screening conditions corresponding to the characteristic variables of the sample objects to obtain training samples.

4. The data prediction method according to claim 3, characterized in that: After filtering the sample objects, the method further includes: Determining a sampling classification variable from the variable name, and classifying the sample objects according to the sampling classification variable to obtain multiple sample object sets; The sample objects in each of the sample object sets are sampled respectively to obtain the training samples.

5. The data prediction method according to claim 4, characterized in that: The sample database includes real-world data.

6. The data prediction method according to claim 1, characterized in that: The method for determining the frequency fitting distribution of the preset special items includes: Acquire a fitting sample from a sample database, and obtain the frequency of the preset special item according to the fitting sample to obtain a sample frequency histogram of the preset special item; Determining a candidate frequency fitting distribution of the preset special item according to the distribution of the sample frequency histogram; If the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is less than or equal to a frequency difference threshold, the candidate frequency fitting distribution is determined as the frequency fitting distribution of the preset special item; If the observed difference between the sample frequency histogram and the candidate frequency fitting distribution is greater than the frequency difference threshold, the candidate frequency fitting distribution of the preset special item is re-determined.

7. The data prediction method according to claim 1, characterized in that: The method for determining the quantity fitting distribution of the preset special items includes: Acquire fitting samples from a sample database, and obtain the quantity of the preset special items within a preset time period based on the fitting samples to obtain a sample quantity histogram of the preset special items; Determining a candidate quantity fitting distribution of the preset special item based on the distribution of the sample quantity histogram; If the observed difference between the sample quantity histogram and the candidate quantity fitting distribution is less than or equal to a quantity difference threshold, the candidate quantity fitting distribution is determined as the preset special item quantity fitting distribution; If the observed difference between the sample quantity histogram and the candidate quantity fitting distribution is greater than the quantity difference threshold, the candidate quantity fitting distribution of the preset special item is re-determined.

8. The data prediction method according to claim 1, characterized in that: The step of obtaining a predicted value of relevant data of the preset special item according to the frequency fitting distribution and the quantity fitting distribution of the preset special item includes: Obtaining a frequency statistic of the preset special item according to a frequency fitting distribution of the preset special item, and obtaining a quantity statistic of the preset special item according to a quantity fitting distribution of the preset special item; Single relevant data of the preset special item is obtained, and a predicted value of the relevant data of the preset special item is obtained based on the single relevant data of the preset special item, the frequency statistical value and the quantity statistical value.

9. The data prediction method according to claim 1, characterized in that: The step of obtaining the predicted value of the relevant security data of the prediction object regarding the preset special item based on the probability of occurrence of the preset special event and the predicted value of the relevant data of the preset special item includes: Obtaining the data guarantee ratio of the prediction object regarding the preset special item; According to the occurrence probability of the preset special event, the predicted value of the relevant data of the preset special item and the relevant data guarantee ratio, the predicted value of the relevant guarantee data of the prediction object regarding the preset special item is obtained.

10. A data prediction device, characterized in that: include: A feature variable acquisition module is used to obtain the feature variables of the prediction object; An event probability prediction module is used to input the characteristic variables of the prediction object into a pre-trained classification model to obtain the event probability of the prediction object with respect to a preset special event; A fitting distribution acquisition module is used to obtain, after the occurrence of the preset special event, a frequency fitting distribution and a quantity fitting distribution of a preset special item in the preset special event; wherein, if the frequency fitting distribution and the quantity fitting distribution of the preset special item cannot be fitted by a single distribution, then a combination distribution is constructed by selecting different loss distribution combinations to fit the distribution; A related data prediction module, configured to obtain a predicted value of the related data of the preset special item based on the frequency fitting distribution and the quantity fitting distribution of the preset special item; The security data prediction module is used to obtain the predicted value of the security data related to the preset special item of the prediction object based on the probability of occurrence of the preset special event and the predicted value of the relevant data of the preset special item.

11. An electronic device, characterized in that: include: processor; as well as A memory for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the data prediction method according to any one of claims 1 to 9.

12. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the data prediction method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Insurance pricing method, device and equipment based on big data and readable storage medium

    CN109636638A

  • Disease prediction system

    CN110957034A