Granularity loss ratio prediction method and related equipment for group insurance
By acquiring data at the granularity of corporate insurance types, cleaning dirty data, and using a binary classification model group to determine the target sub-range of the claims ratio within the preset range, the problems of complex algorithm dependence and lack of confidentiality in the prediction of group insurance claims ratios are solved, and efficient, accurate, and confidential prediction results are achieved.
Patent Information
- Application Number
- CN202210221137.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-03-08
AI Technical Summary
Existing technologies for predicting claims ratios for group insurance products rely on complex algorithms and high-intensity manual judgment, and lack data cleaning, resulting in high prediction difficulty and insufficient confidentiality of results.
Data is obtained at the granularity of corporate insurance types. After dirty data cleaning, a pre-trained binary classification model group is used to determine the target sub-interval of the claims ratio in the preset distribution range, and the sub-interval label is output as the prediction result, which simplifies the algorithm design and achieves confidentiality.
It reduces the difficulty of claim ratio prediction, improves the accuracy and efficiency of prediction, and at the same time ensures the confidentiality of claim ratio prediction results, making the results more intuitive.
Smart Images

Figure CN114841817B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of insurance business technology, and in particular to a method for predicting granular claims ratios for group insurance products and related equipment. Background Art
[0002] Group insurance (group insurance) refers to insurance that provides protection for multiple insured persons under a single policy. Group insurance is underwritten by a collective entity, with the insurance company and the collective entity as the two parties, and the contract is signed using a single insurance policy. Typically, the collective entity is the policyholder, and the employees within the entity are the insured. Insurance types are broadly categorized into social insurance and commercial insurance. Social insurance includes pension insurance, medical insurance, unemployment insurance, work-related injury insurance, and maternity insurance. Commercial insurance is divided into property insurance and life insurance. Property insurance is further divided into three categories: property damage insurance, liability insurance, and credit guarantee insurance. Granularity refers to the degree of statistical coarseness of data within a given dimension. In the computer field, granularity refers to the minimum incremental value of system memory expansion. The higher the level of granularity, the smaller the granularity; conversely, the lower the level of granularity, the larger the granularity. Salespeople and reviewers subjectively consider whether to accept insurance based on information such as claims records and the loss ratio over the past three years. This subjective approach to making decisions based on historical information places high demands on salespeople and reviewers. Summary of the Invention
[0003] In view of this, the purpose of this application is to propose a group insurance granularity loss ratio prediction method and related equipment to solve the above problems.
[0004] Based on the above objectives, the first aspect of the present application provides a method for predicting the granularity loss ratio of group insurance products, comprising:
[0005] Obtain current portfolio feature data at the granularity of corporate insurance types;
[0006] Cleaning the current combined feature data to obtain valid current data;
[0007] Determining a target sub-interval of the predicted loss ratio within a preset distribution interval based on the valid current data and the pre-trained binary classification model group;
[0008] The subinterval label corresponding to the target subinterval is output as a prediction result.
[0009] A second aspect of the present application provides a group insurance type granularity loss ratio prediction device, comprising:
[0010] The feature data selection module is configured to: obtain current combination feature data based on the granularity of enterprise insurance types;
[0011] a cleaning module configured to: clean dirty data from the current combined feature data to obtain valid current data;
[0012] A prediction module is configured to: determine a target sub-range of the predicted loss ratio within a preset distribution range based on the valid current data and a pre-trained binary classification model group;
[0013] The output module is configured to output the sub-interval label corresponding to the target sub-interval as a prediction result.
[0014] The third aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method provided in the first aspect of the present application is implemented.
[0015] A fourth aspect of the present application provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute the method provided by the first aspect of the present application.
[0016] From the above, it can be seen that the group insurance type granularity claims ratio prediction method and related equipment provided by this application, first, obtain the current combination feature data with the enterprise insurance type as the granularity, and use the enterprise insurance type as the granularity to indicate that the smallest unit of data is the insurance type, and there is no need to obtain overly detailed data, which provides convenience for subsequent data use. Secondly, the current combination feature data is cleaned of dirty data to obtain valid current data. Cleaning the dirty data can improve the accuracy and efficiency of the claims ratio prediction. Then, based on the valid current data and the pre-trained binary classification model group, the target sub-interval of the predicted claims ratio in the preset distribution interval is determined, and the binary classification model group is used to predict the claims ratio, which frees the staff from the high-intensity labor brought by manual prediction, and there is no need to use complex algorithms to confirm the specific value of the claims ratio. It is only necessary to judge the position of the predicted claims ratio in the preset distribution interval, which reduces the difficulty of the claims ratio prediction. Finally, the sub-interval label corresponding to the target sub-interval is output as the prediction result. The predicted value of the payout ratio and the position of the target sub-interval in the distribution interval are internal data and do not need to be announced. Therefore, the sub-interval label corresponding to the target sub-interval is selected as the prediction result for output, which achieves a certain degree of confidentiality and makes the prediction result more intuitive. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 This is a flow chart of the method for predicting the granularity loss ratio of group insurance products according to an embodiment of the present application;
[0019] Figure 2 A flowchart of the preparation process of an embodiment of the present application;
[0020] Figure 3 This is a flow chart of dividing distribution intervals according to an embodiment of the present application;
[0021] Figure 4 This is a flowchart of the binary classification model group training in an embodiment of the present application;
[0022] Figure 5 This is a flowchart of the loss ratio prediction of an embodiment of the present application;
[0023] Figure 6 This is a flow chart of determining the position of a payout ratio in a distribution interval according to an embodiment of the present application;
[0024] Figure 7 A flowchart of the result output of an embodiment of the present application;
[0025] Figure 8 This is a flowchart of obtaining current combined feature data according to an embodiment of the present application;
[0026] Figure 9 This is a structural diagram of a device for predicting granular claims ratios for group insurance according to an embodiment of the present application;
[0027] Figure 10 This is a structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to make the objectives, technical solutions and advantages of this application more clear, this application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0029] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0030] In the related art, the prediction of loss ratio is mainly based on the ARMA model (auto regressive moving average model). The ARMA model is one of the high-resolution spectral analysis methods of the model parameter method. This method is a typical method for studying the rational spectrum of stationary random processes and is applicable to a large class of practical problems. The ARMA model method has more accurate spectral estimation and better spectral resolution performance, but its parameter estimation is relatively cumbersome, and this method uses the loss ratio as a measurement indicator to predict the future development of the insurance industry. What is predicted is the change pattern of the loss ratio over time, that is, the prediction is made with time as the granularity. In the related art, there is also a prediction method that uses complex algorithms to process and predict valid current data. The purpose is to make the processed predicted loss ratio fall directly into the corresponding target sub-interval or obtain a specific loss ratio value through a single calculation. The calculation method is complex, the algorithm design is difficult, and the requirements for input data are high.
[0031] The group insurance granularity loss ratio prediction method provided in the embodiments of the present application utilizes a binary classification model group to predict the loss ratio. This eliminates the need for complex algorithms to confirm the specific value of the loss ratio. Instead, it simply requires a step-by-step determination of the position of the predicted loss ratio within a preset distribution interval, reducing the difficulty of loss ratio prediction. Finally, the subinterval label corresponding to the target subinterval is output as the prediction result. The specific value of the loss ratio and the position of the target subinterval within the distribution interval are internal data and do not need to be published. Therefore, the subinterval label corresponding to the target subinterval is selected as the prediction result for output, achieving a certain degree of confidentiality and making the prediction result more intuitive.
[0032] In some embodiments, as Figure 1 As shown, a method for predicting the granular loss ratio of group insurance products includes:
[0033] Step 100: Obtain current combination feature data at the granularity of enterprise insurance types.
[0034] In this step, we obtain the one-year insurance policies and related insurance information that have been effective within a certain time period at the enterprise insurance type granularity, and examine the compensation amount and number of people paid at the insurance type granularity during the effective period of the policy. We also examine the insurance and compensation situation of the enterprise in the past year, and collect the insurance type compensation information at the city and occupation granularity in the past year. The data of the above dimensions are extracted and combined as the current combined feature data. The final current combined feature data specifically includes the following information: (1) enterprise information; (2) group information; (3) group insurance type information; (4) policy and insurance type claim information.
[0035] Step 200: Clean the dirty data of the current combined feature data to obtain valid current data.
[0036] In this step, data cleaning is the process of re-examining and verifying the data in order to remove duplicate information, correct existing errors, and provide data consistency.
[0037] Among them, because the data in the data warehouse is a collection of data for a certain topic, these data are extracted from multiple business systems and contain historical data. It is inevitable that some data are wrong data and some data are conflicting. These wrong or conflicting data are obviously unwanted and are called "dirty data". According to certain rules, "dirty data" is "cleaned". This is dirty data cleaning. Data that does not meet the requirements, namely "dirty data", mainly consists of three categories: incomplete data, wrong data, and duplicate data. Data cleaning should comply with the following principles: (1) Completeness: whether there are null values in a single data item and whether the statistical fields are complete. (2) Comprehensiveness: for a set of values, the maximum value, minimum value, average value and other data definition values can be compared to determine whether the set of value data is comprehensive. (3) Legality: whether the type, content and size of the value meet the pre-set threshold. For example: if the insured person is over 200 years old, this data is illegal. (4) Uniqueness: whether the data is recorded repeatedly. For example: the data of an insured person is recorded repeatedly multiple times.
[0038] Step 300: Based on the valid current data and the pre-trained binary classification model group, determine the target sub-interval of the predicted loss ratio in the preset distribution interval.
[0039] In this step, valid current data is input into the pre-trained binary classification model group, and the binary classification model group is used to predict the claims ratio. This frees the staff from the high-intensity labor brought by manual prediction, and there is no need to use complex algorithms to confirm the specific value of the claims ratio. It is only necessary to determine the position of the predicted claims ratio in the preset distribution range, which reduces the difficulty of claims ratio prediction.
[0040] Step 400: Output the sub-interval label corresponding to the target sub-interval as a prediction result.
[0041] In this step, the sub-interval label corresponding to the target sub-interval is output as the prediction result because the predicted value of the payout ratio and the position of the target sub-interval in the distribution interval are internal data and do not need to be announced. Therefore, the sub-interval label corresponding to the target sub-interval is selected as the prediction result for output, which achieves a certain degree of confidentiality and makes the prediction result more intuitive.
[0042] In some embodiments, as Figure 2 As shown in the figure, the pre-training of the binary classification model group specifically includes:
[0043] Step 010: Obtain historical combination feature data and historical loss ratios corresponding to the historical combination feature data at the granularity of enterprise insurance types.
[0044] In this step, the historical combination feature data includes: policyholder information, plan information and historical claims information, among which the policyholder information includes the insurance time, unit nature, industry category, occupation category, number of members, number of employees, number of insured persons, age distribution of insured persons, etc.; plan information includes plan type and contract form, major business category, group, insurance type, insurance amount, premium, discount rate, commission rate, sales area, whether there is a retroactive designated effective date, etc.; historical claims information includes main insurance type, effective date, expiration date, policy beginning, policy end, policy claims status, etc.
[0045] Step 020: Determine the distribution range of the payout ratio based on the historical payout ratio.
[0046] In this step, the historical claims ratio is counted and calculated, and the historical claims ratio is divided into several groups according to the profit and loss relationship after the payment. Each group corresponds to a sub-interval, and all sub-intervals corresponding to all groups together constitute the distribution interval.
[0047] Step 030: Train a binary classification model group based on historical combination feature data and historical claims ratio.
[0048] In this step, based on the historical combination feature data and historical claims ratio, xgboost (extreme gradient boosting algorithm) is continuously used to train multiple binary classification models, and multiple binary classification models constitute the binary classification model group.
[0049] Among them, if the training effect is not ideal, it can be replaced with other machine learning classification algorithms such as LR (logistic regression algorithm) and SVM (support vector machine algorithm) to supplement the training of the binary classification model until the training results meet the user's needs.
[0050] In some embodiments, as Figure 3 As shown, step 020: determining the distribution range of the loss ratio based on the historical loss ratio, specifically including:
[0051] Step 021: Determine the first boundary payout ratio, the second boundary payout ratio and the third boundary payout ratio based on the historical payout ratio.
[0052] In this step, based on the statistics of historical payout ratios, the first boundary payout ratio, the second boundary payout ratio and the third boundary payout ratio are determined as the boundaries for dividing different sub-intervals.
[0053] Optionally, 70% is selected as the first boundary payout ratio, 40% is selected as the second boundary payout ratio, and 90% is selected as the third boundary payout ratio.
[0054] Step 022: Based on the first boundary payout ratio, the payout ratio is divided into the first half interval and the second half interval.
[0055] In this step, the first half of the interval is divided into those with a payout ratio less than or equal to the first threshold, and the second half is divided into those with a payout ratio greater than the first threshold. The first half indicates a profit when the payout ratio is less than or equal to the first threshold, while the second half indicates a loss when the payout ratio is greater than the first threshold. The distribution range is first divided into two larger intervals, each representing an opposing meaning, to qualitatively characterize the payout ratio as either profit or loss.
[0056] Step 023: Based on the second boundary payout ratio, the first half interval is divided into a first sub-interval and a second sub-interval.
[0057] In this step, once profitability has been determined, the first half of the interval is further divided based on the second boundary claims ratio according to the magnitude of profitability. The first sub-interval is defined as the interval where the claims ratio is less than or equal to the second boundary claims ratio, while the second sub-interval is defined as the interval where the claims ratio is greater than the second boundary claims ratio. The first sub-interval indicates greater profitability when the claims ratio is greater than zero and less than or equal to the second boundary claims ratio, while the second sub-interval indicates less profitability when the claims ratio is greater than the second boundary claims ratio but less than the first boundary claims ratio. Dividing the first half of the profitable interval into two sub-intervals with different profit margins allows for qualitative analysis of the profitability of the claims ratio to determine the merits of the group insurance policy.
[0058] Step 024: Based on the third boundary payout ratio, the second half interval is divided into a third sub-interval and a fourth sub-interval.
[0059] In this step, once the loss has been determined, the second half of the interval is further divided based on the magnitude of the loss, based on the third boundary claims ratio. The third sub-interval is defined for losses less than or equal to the third boundary claims ratio, while the second sub-interval is defined for losses greater than the third boundary claims ratio. The third sub-interval indicates that losses are smaller when the claims ratio is greater than the first boundary claims ratio but less than or equal to the third boundary claims ratio, while the fourth sub-interval indicates that losses are greater when the claims ratio is greater than the third boundary claims ratio. Dividing the second half of the loss interval into two sub-intervals with different loss magnitudes allows for qualitative characterization of the loss situation based on the claims ratio to determine the merits of the group insurance policy.
[0060] Step 025: Combine the first sub-interval, the second sub-interval, the third sub-interval, and the fourth sub-interval to obtain a distribution interval.
[0061] In this step, the first sub-interval, the second sub-interval, the third sub-interval and the fourth sub-interval are combined to obtain distribution intervals from zero to the second boundary payout ratio, the second boundary payout ratio to the first boundary payout ratio, the first boundary payout ratio to the third boundary payout ratio, and greater than the third boundary payout ratio.
[0062] Optionally, the distribution intervals are (0-40%, 40%-70%, 70%-90%, and greater than 90%), where (0-40%) is the first sub-interval with greater profits, (40%-70%) is the second sub-interval with smaller profits, (70%-90%) is the third sub-interval with smaller losses, and (greater than 90%) is the fourth sub-interval with greater losses. If the user requires further segmentation, any sub-interval can be divided, but this requires training more and more accurate binary classification model groups.
[0063] In some embodiments, as Figure 4 As shown, step 030: training a binary classification model group based on historical combination feature data and historical loss ratios, specifically including:
[0064] Step 031: Clean the historical combination feature data to obtain valid historical data.
[0065] Data cleaning is the process of re-examining and verifying data to remove duplicate information, correct existing errors, and provide data consistency. Since the current combined feature data needs to be cleaned of dirty data in step 200 to obtain valid current data, and this valid current data is used as the input for the binary classification model group, when training the binary classification model group, the historical combined feature data also needs to be cleaned, and then the cleaned valid historical data is used to train the binary classification model group.
[0066] Step 032: Split the valid historical data into a training data set and a validation data set.
[0067] In this step, in order to verify the training results of the binary classification model group, a control group needs to be set up, so the valid historical data is divided into a training data set and a validation data set. The training data set is used to train the binary classification model group, and the validation data set is used to verify the training results of the binary classification model group.
[0068] Among them, the effective historical data set is subjected to feature engineering processing such as sampling, embedding, normalization, and feature combination, and the effective historical data is cross-validated and divided into a training data set and a validation data set.
[0069] Step 033: Based on the training data set, train the binary classification model group through the machine learning algorithm.
[0070] Optionally, the xgboost algorithm (extreme gradient boosting algorithm) is first tried to train the binary classification model group. If it is verified to meet the requirements of the salesperson and the approver, there is no need to proceed to the next step of training. If it does not meet the requirements of the salesperson and the approver after verification, you can choose classification algorithms such as random forest algorithm, svm algorithm (support vector machine algorithm) or LR algorithm (logistic regression algorithm) to train the multi-classification model group, and then adjust the model parameters and optimize the prediction algorithm.
[0071] Step 034: Use the trained binary classification model group to predict the loss ratio of the validation data set and determine the target sub-interval that the loss ratio falls into.
[0072] In this step, the binary classification model group has been preliminarily trained and takes the verification data set as input. The binary classification model group will output the target sub-interval in which the claims ratio corresponding to the verification data set falls. The target sub-interval is one of the first sub-interval, the second sub-interval, the third sub-interval and the fourth sub-interval. Since the verification data set is obtained from the historical combination feature data, the true value claims ratio of the verification data set is the historical claims ratio.
[0073] Step 035: In response to determining that the historical loss ratio is within the target sub-interval, it is determined that the binary classification model group meets the business requirements and the training is completed.
[0074] In this step, if the historical loss ratio is within the target sub-interval, that is, the loss ratio predicted by the binary classification model group and the historical loss ratio are within the same sub-interval, it means that the loss ratio predicted by the binary classification model group meets the requirements and the binary classification model group training is completed. If the historical loss ratio is not within the target sub-interval, that is, the loss ratio predicted by the binary classification model group and the historical loss ratio are not within the same sub-interval, it means that the loss ratio predicted by the binary classification model group does not meet the requirements and it is necessary to further train the binary classification model group using the algorithm in step 033 until the binary classification model group meets the requirements of the salesperson and the approver, that is, the loss ratio predicted by the binary classification model group and the historical loss ratio are within the same target sub-interval.
[0075] In some embodiments, as Figure 5 and Figure 6 As shown, step 300: based on the valid current data and the pre-trained binary classification model group, determining the target sub-interval of the predicted loss ratio within the preset distribution interval, specifically includes:
[0076] Step 310: Divide the binary classification model group into a first binary classification model, a second binary classification model, and a third binary classification model.
[0077] In this step, corresponding to step 021, the first two-classification model corresponds to the first boundary payout ratio, which is used to determine whether the predicted payout ratio is in the first half of the interval or the second half of the interval; the second two-classification model corresponds to the second boundary payout ratio, which is used to determine whether the predicted payout ratio is in the first sub-interval or the second sub-interval; the third two-classification model corresponds to the third boundary payout ratio, which is used to determine whether the predicted payout ratio is in the third sub-interval or the fourth sub-interval.
[0078] Step 320: Based on the valid current data and the first boundary payout ratio, determine the position of the predicted payout ratio in the distribution range through the first binary classification model.
[0079] In this step, the first binary classification model is used to process the input valid current data, and the predicted payout ratio obtained after processing is compared with the first boundary payout ratio. If the predicted payout ratio is greater than the first boundary payout ratio, it means that the predicted payout ratio is in the second half of the interval; if the predicted payout ratio is less than the first boundary payout ratio, it means that the predicted payout ratio is in the first half of the interval. Since it is only necessary to judge the size relationship between the predicted payout ratio and the first boundary payout ratio, there is no need to determine the specific value of the predicted payout ratio, so the first binary classification model is selected for judgment. Compared with using complex algorithms to process the valid current data to obtain the specific value of the payout ratio, or using a four-classification model or even a multi-classification model to process the current valid data, the final target sub-interval where the predicted payout ratio is located is directly obtained in one step, and multiple binary classification model groups are used for step-by-step judgment, which reduces the difficulty of predicting the predicted payout ratio and simplifies the design of the algorithm for the judgment process. The relatively simple algorithm can reduce the error probability of each binary classification model in the binary classification model group.
[0080] Step 330: In response to determining that the predicted odds ratio is in the first half of the interval, based on the valid current data and the second boundary odds ratio, in the first sub-interval and the second sub-interval, a second binary classification model is used to determine the target sub-interval into which the predicted odds ratio falls.
[0081] In this step, after determining that the predicted payout ratio falls within the first half of the interval, the second binary classification model is used to process the input valid current data. The resulting predicted payout ratio is then compared with the second boundary payout ratio. If the predicted payout ratio is greater than the second boundary payout ratio, it indicates that the predicted payout ratio is within the second sub-interval; if the predicted payout ratio is less than the second boundary payout ratio, it indicates that the predicted payout ratio is within the first sub-interval. Since only the relationship between the predicted payout ratio and the first boundary payout ratio needs to be determined, and there is no need to determine the specific value of the predicted payout ratio, the second binary classification model is selected for judgment, which reduces the difficulty of predicting the predicted payout ratio and simplifies the design of the algorithm for the judgment process. The relatively simple algorithm can reduce the error probability of the second binary classification model.
[0082] Step 340: In response to determining that the predicted odds ratio is in the second half of the interval, based on the valid current data and the third boundary odds ratio, in the third sub-interval and the fourth sub-interval, the third binary classification model is used to determine the target sub-interval into which the predicted odds ratio falls.
[0083] In this step, after determining that the predicted payout ratio falls within the second half of the interval, the third binary classification model is used to process the input valid current data. The resulting predicted payout ratio is then compared with the third boundary payout ratio. If the predicted payout ratio is greater than the third boundary payout ratio, it indicates that the predicted payout ratio is within the fourth sub-interval; if the predicted payout ratio is less than the third boundary payout ratio, it indicates that the predicted payout ratio is within the third sub-interval. Since only the relationship between the predicted payout ratio and the third boundary payout ratio needs to be determined, and there is no need to determine the specific value of the predicted payout ratio, the third binary classification model is selected for judgment. This reduces the difficulty of predicting the predicted payout ratio and simplifies the design of the algorithm for the judgment process. The relatively simple algorithm can reduce the error probability of the third binary classification model.
[0084] In some embodiments, as Figure 7 As shown, step 400: outputting the subinterval label corresponding to the target subinterval as a prediction result, specifically includes:
[0085] Step 410: In response to determining that the target subinterval is the first subinterval, output a first label corresponding to the first subinterval as a prediction result.
[0086] In this step, optionally, if the target subinterval is determined to be the first subinterval, the first label "large profit" corresponding to the first subinterval is selected and output as the prediction result.
[0087] Step 420: In response to determining that the target subinterval is the second subinterval, output a second label corresponding to the second subinterval as a prediction result.
[0088] In this step, optionally, if it is determined that the target subinterval is the second subinterval, the second label "small profit" corresponding to the second subinterval is selected and output as the prediction result.
[0089] Step 430: In response to determining that the target subinterval is the third subinterval, output a third label corresponding to the third subinterval as a prediction result.
[0090] In this step, optionally, if the target subinterval is determined to be the third subinterval, the third label "small loss" corresponding to the third subinterval is selected and output as the prediction result.
[0091] Step 440: In response to determining that the target subinterval is the fourth subinterval, output a fourth label corresponding to the fourth subinterval as a prediction result.
[0092] In this step, optionally, if the target subinterval is determined to be the th subinterval, the fourth label "large loss" corresponding to the fourth subinterval is selected and output as the prediction result.
[0093] In some embodiments, as Figure 8 As shown, step 100: obtaining current combination feature data based on the enterprise insurance type granularity, specifically including:
[0094] Step 110: Taking the enterprise insurance type as the granularity, the one-year insurance policies and related insurance information that have taken effect within the preset time interval are obtained as the first feature information.
[0095] Step 120: The compensation amount and number of people paid at the insurance type granularity during the effective period of the insurance policy are used as the second feature information.
[0096] Step 130: The insurance and compensation status of the enterprise in the past year is used as the third feature information.
[0097] Step 140: The city where the enterprise is located and the compensation situation of the enterprise's insured professions in the past year are used as the fourth feature information.
[0098] Step 150: Combine the first feature information, the second feature information, the third feature information, and the fourth feature information to obtain current combined feature data.
[0099] The information obtained specifically includes: policyholder information and plan information, where policyholder information includes insurance time, unit nature, industry category, occupation category, number of members, number of employees, number of insured persons, age distribution of policyholders, etc.; plan information includes plan type, contract form, business category, insurance type, group, insurance amount, premium, discount rate, commission rate, sales area, whether there is a retroactive designated effective date, etc. The data of the above dimensions are extracted and combined as the current combined feature data, and the final current combined feature data specifically includes the following information: (1) enterprise information; (2) group information; (3) group insurance type information; (4) policy and insurance type claim information.
[0100] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.
[0101] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0102] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a group insurance type granularity claim ratio prediction device.
[0103] refer to Figure 9 The group insurance type granularity loss ratio prediction device includes:
[0104] The feature data selection module 1 is configured to: obtain the current combination feature data based on the granularity of the enterprise insurance type;
[0105] Cleaning module 2 is configured to: clean dirty data of current combined feature data to obtain valid current data;
[0106] Prediction module 3 is configured to: determine a target sub-interval of the predicted loss ratio within a preset distribution interval based on valid current data and a pre-trained binary classification model group;
[0107] The output module 4 is configured to output the sub-interval label corresponding to the target sub-interval as a prediction result.
[0108] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0109] The device of the above embodiment is used to implement the corresponding group insurance type granularity loss ratio prediction method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0110] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the group insurance type granularity claim ratio prediction method described in any of the above embodiments is implemented.
[0111] Figure 1010 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0112] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0113] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0114] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0115] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0116] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0117] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0118] The electronic device of the above embodiment is used to implement the corresponding group insurance type granularity loss ratio prediction method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0119] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the group insurance type granularity claim ratio prediction method as described in any of the above embodiments.
[0120] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0121] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the group insurance type granularity loss ratio prediction method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0122] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. Within the scope of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0123] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0124] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.
[0125] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this application.
Claims
1. A method for predicting the granularity loss ratio of group insurance, characterized by: include: Obtain current portfolio feature data at the granularity of corporate insurance types; Cleaning the current combined feature data to obtain valid current data; Determining a target sub-interval of the predicted loss ratio within a preset distribution interval based on the valid current data and the pre-trained binary classification model group; The binary classification model group includes a first binary classification model, a second binary classification model, and a third binary classification model, as well as a first boundary loss ratio corresponding to the first binary classification model, a second boundary loss ratio corresponding to the second binary classification model, and a third boundary loss ratio corresponding to the third binary classification model; determining a target subrange of the predicted loss ratio within a preset distribution range based on the valid current data and the pre-trained binary classification model group includes: Based on the valid current data and the first boundary loss ratio, determining the position of the predicted loss ratio within the distribution interval using the first binary classification model; wherein the position of the distribution interval includes the first half interval or the second half interval, the first half interval includes the first subinterval and the second subinterval, and the second half subinterval includes the third subinterval and the fourth subinterval; In response to determining that the predicted odds ratio is located in the first half of the interval, based on the valid current data and the second boundary odds ratio, using a second binary classification model to determine the target sub-interval in which the predicted odds ratio falls within the first sub-interval and the second sub-interval; In response to determining that the predicted odds ratio is in the second half of the interval, based on the valid current data and the third boundary odds ratio, using a third binary classification model to determine the target sub-interval that the predicted odds ratio falls into within the third sub-interval and the fourth sub-interval; The subinterval label corresponding to the target subinterval is output as a prediction result.
2. The method according to claim 1, characterized in that The pre-training of the binary classification model group specifically includes: Obtaining historical combination characteristic data and historical loss ratios corresponding to the historical combination characteristic data based on the granularity of the enterprise insurance type; Determining the distribution range of the payout ratio based on the historical payout ratio; The binary classification model group is trained based on the historical combination feature data and the historical claims ratio.
3. The method according to claim 2, characterized in that The distribution range of the payout ratio determined based on the historical payout ratio specifically includes: Determining a first boundary payout ratio, a second boundary payout ratio, and a third boundary payout ratio based on the historical payout ratio; Based on the first boundary payout ratio, the payout ratio is divided into a first half interval and a second half interval; Based on the second boundary payout ratio, the first half interval is divided into a first sub-interval and a second sub-interval; Based on the third boundary payout ratio, the second half interval is divided into a third sub-interval and a fourth sub-interval; The first sub-interval, the second sub-interval, the third sub-interval and the fourth sub-interval are combined to obtain the distribution interval.
4. The method according to claim 3, characterized in that Training the binary classification model group based on the historical combination feature data and the historical loss ratio specifically includes: Performing data cleaning on the historical combination feature data to obtain valid historical data; Splitting the valid historical data into a training data set and a validation data set; Based on the training data set, training the binary classification model group by a machine learning algorithm; Using the trained binary classification model group to predict the loss ratio of the validation data set, and determining the target sub-interval that the loss ratio falls into; In response to determining that the historical claims ratio is within the target sub-interval, it is determined that the binary classification model group meets the business requirements and the training is completed.
5. The method according to claim 4, characterized in that Outputting the subinterval label corresponding to the target subinterval as a prediction result specifically includes: In response to determining that the target subinterval is the first subinterval, outputting a first label corresponding to the first subinterval as a prediction result; In response to determining that the target subinterval is the second subinterval, outputting a second label corresponding to the second subinterval as a prediction result; In response to determining that the target subinterval is the third subinterval, outputting a third label corresponding to the third subinterval as a prediction result; In response to determining that the target subinterval is the fourth subinterval, a fourth label corresponding to the fourth subinterval is output as a prediction result.
6. The method according to claim 4, characterized in that The acquisition of current combination feature data based on the enterprise insurance type granularity specifically includes: Taking the enterprise insurance type as the granularity, the one-year insurance policies and related insurance information that have been effective within the preset time interval are obtained as the first feature information; The amount of compensation paid and the number of people receiving compensation at the insurance type granularity during the effective period of the insurance policy are used as the second feature information; The insurance coverage and claims payment status of the enterprise in the past year are used as the third feature information; The city where the enterprise is located and the claims paid by the enterprise's insured occupations in the past year are used as the fourth characteristic information; The first feature information, the second feature information, the third feature information, and the fourth feature information are combined to obtain the current combined feature data.
7. A device for predicting the granularity loss ratio of group insurance, characterized in that: include: The feature data selection module is configured to: obtain current combination feature data based on the granularity of enterprise insurance types; a cleaning module configured to: clean dirty data from the current combined feature data to obtain valid current data; A prediction module is configured to: determine a target sub-range of the predicted loss ratio within a preset distribution range based on the valid current data and a pre-trained binary classification model group; The binary classification model group includes a first binary classification model, a second binary classification model, and a third binary classification model, as well as a first boundary loss ratio corresponding to the first binary classification model, a second boundary loss ratio corresponding to the second binary classification model, and a third boundary loss ratio corresponding to the third binary classification model; determining a target subrange of the predicted loss ratio within a preset distribution range based on the valid current data and the pre-trained binary classification model group includes: Based on the valid current data and the first boundary loss ratio, determining the position of the predicted loss ratio within the distribution interval using the first binary classification model; wherein the position of the distribution interval includes the first half interval or the second half interval, the first half interval includes the first subinterval and the second subinterval, and the second half subinterval includes the third subinterval and the fourth subinterval; In response to determining that the predicted odds ratio is located in the first half of the interval, based on the valid current data and the second boundary odds ratio, using a second binary classification model to determine the target sub-interval in which the predicted odds ratio falls within the first sub-interval and the second sub-interval; In response to determining that the predicted odds ratio is in the second half of the interval, based on the valid current data and the third boundary odds ratio, using a third binary classification model to determine the target sub-interval that the predicted odds ratio falls into within the third sub-interval and the fourth sub-interval; The output module is configured to output the sub-interval label corresponding to the target sub-interval as a prediction result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 6 when executing the program. 9 . A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to claim 1 .
Citation Information
Patent Citations
Group insurance service prediction method and device
CN112330476A
Traffic classification method and traffic management equipment
CN112953851A