A prediction method and device for insurance business
Through machine learning methods, the time-consuming and labor-consuming problem of the P&C insights in the existing technology is solved, efficient and accurate prediction of P&C insights, and the work efficiency of P&C business is improved.
Patent Information
- Application Number
- CN202210396949.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-04-15
AI Technical Summary
In the existing technology, Jiabao insight relies on the cooperation of professional data scientists and business personnel, which makes the modeling and analysis process time-consuming and labor-intensive, and the model output results need to be processed and sorted before they can be applied to business scenarios, which has a high threshold.
Data modeling is carried out through machine learning methods, and the target annotation tool is used to generate a set of additional and guaranteed insight samples, train an additional and guaranteed insight models, and optimize the model through confusion matrix and model evaluation methods to achieve accurate prediction analysis of insured customers.
The process of complex modeling and analysis is processed and intelligent, allowing business personnel to quickly realize full-process modeling with one click, improve the efficiency and accuracy of the forecast for the insurance business, and can analyze from multiple angles of insurance products and customer characteristics to provide accurate prediction results.
Smart Images

Figure CN114861989B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data analysis, and in particular to a prediction method and device for insurance premium services. Background Art
[0002] Currently, more and more companies are beginning to attach importance to customer management, and use technical means to conduct insight analysis on customer data throughout their life cycle, hoping to discover valuable features to support subsequent business expansion. Insurance companies' customer insights are mainly concentrated in key business scenarios such as customer acquisition, promotion, underwriting, additional insurance, and claims. Additional insurance refers to promoting historical customers to insure again, which is a customer insight scenario that sales staff focus on. Through a comprehensive analysis of historical customer information, sales staff can help grasp the important characteristics of additional insurance customers, such as region, gender, age, insured object, insurance timing, etc., to provide customers with corresponding business activities and improve the success rate of additional insurance.
[0003] Currently, the insurance insight for customers relies on the cooperation of professional data scientists and business personnel. First, the insurance insight problem is converted into a data modeling and analysis problem, and then various advanced model algorithms are used to extract valuable information from large amounts of data. Even senior experts need to go through a series of complex operation processes such as data processing, feature extraction, model selection, parameter tuning, model evaluation, etc. when conducting modeling and analysis. They also need to use experience to make repeated iterative adjustments. The whole process is time-consuming and labor-intensive. Although various tool platforms have been launched to support modeling and analysis work, there is still a high threshold and technical personnel are required to use them. The model output results also need to be processed and sorted by experts before they can be applied to business scenarios. Summary of the invention
[0004] In view of this, the embodiment of the present application provides a prediction method for additional insurance business, which can realize predictive analysis of insurance big data through data modeling in machine learning, so as to accurately understand the future trend of customers' additional insurance and improve the work efficiency of business personnel.
[0005] In a first aspect, an embodiment of the present application provides a method for predicting an insurance service, including:
[0006] Use the target labeling tool to label the insurance data and obtain the insurance insight sample set;
[0007] The Jiabao insight model is trained according to the training set of the Jiabao insight sample set to obtain a candidate Jiabao insight model;
[0008] According to the confusion matrix and model evaluation method, the candidate Jiabao Insight model is evaluated using the test set of the Jiabao Insight sample set to obtain the model evaluation result;
[0009] According to the model evaluation results, determine the insurance insight model that meets the detection accuracy;
[0010] Based on the insurance add-on insight model, predictive analysis is performed on all insured customers to obtain the predicted results of customers adding insurance.
[0011] In combination with the first aspect, the embodiment of the present application provides a first possible implementation of the first aspect, wherein before obtaining the insurance insight sample set, the method further includes:
[0012] Determine the additional insurance target based on the characteristic parameters of the insured customer and the characteristic parameters of the insurance time point;
[0013] Obtain the insurance data of the insured target, which includes: customer gender, age, marital status, occupation, income, historical premiums, historical insurance types, number of insurance purchases, insurance amount, order of insurance products, and product attention data;
[0014] Label the target data in the insurance data to obtain the training set and test set of the insurance insight sample set.
[0015] In combination with the first possible implementation of the first aspect, the embodiment of the present application provides a second possible implementation of the first aspect, wherein, before the insurance insight model is trained according to the training set of the insurance insight sample set to obtain the candidate insurance insight model, the method further includes:
[0016] The binary code of the training set is converted into a format, and the data samples of the training set are obtained after the conversion;
[0017] A feature screening strategy is used to screen the features of each converted data sample to obtain the target features;
[0018] Input the feature vector of the target feature into the added insurance logic model, and use multiple algorithms to perform cyclic training on the added insurance logic model to obtain a candidate added insurance logic model, among which the multiple algorithms include: logistic regression algorithm, greedy algorithm, decision tree C4.5 algorithm, classification regression tree algorithm, XGBoost tree class algorithm, LightGBM optimization algorithm and neural network algorithm;
[0019] According to the model evaluation method, the candidate insurance logic model is evaluated using the test set of the insurance insight sample set to obtain the model evaluation result;
[0020] According to the evaluation results, determine the selected insurance logic model that meets the requirements;
[0021] The Bayesian optimization algorithm is used to tune the hyperparameters of the selected insurance logic model to obtain the insurance insight model with the optimal parameters.
[0022] In combination with the first possible implementation or the second possible implementation of the first aspect, the embodiment of the present application provides a third possible implementation of the first aspect, wherein, according to the model evaluation result, determining the insurance insight model that meets the detection accuracy includes:
[0023] If the detection accuracy of the candidate insight model for insurance coverage reaches the threshold, it is determined as the insight model for insurance coverage;
[0024] If the detection accuracy of the candidate insurance insight model does not reach the threshold, the candidate insurance insight model is retrained.
[0025] In combination with the first possible implementation manner or the second possible implementation manner of the first aspect, the embodiment of the present application provides a fourth possible implementation manner of the first aspect, wherein a prediction analysis is performed on all insured customers according to the insurance adding insight model to obtain a prediction result of the customer adding insurance, including:
[0026] According to the insurance adding insight model, a predictive analysis is conducted on all insured customers to obtain the prediction results of customers adding insurance, among which the prediction results of customers adding insurance are: customer insurance adding model evaluation results, customer insurance adding rating, impact value of insurance adding results and customer classification insurance adding probability value.
[0027] In combination with the first possible implementation manner or the second possible implementation manner of the first aspect, the embodiment of the present application provides a fifth possible implementation manner of the first aspect, wherein a prediction analysis is performed on all insured customers according to the insurance adding insight model to obtain a prediction result of the customer adding insurance, specifically including:
[0028] If the evaluation index of the insurance insight model exceeds the first preset threshold, the customer insurance model evaluation result is excellent;
[0029] If the evaluation index of the insurance insight model is equal to the second preset threshold, the evaluation result of the customer insurance model is good;
[0030] If the evaluation index of the insurance insight model is equal to the third preset threshold, the customer insurance model evaluation result is optional;
[0031] If the evaluation index of the insurance insight model is less than the fourth preset threshold, the customer insurance model evaluation result is unavailable.
[0032] In combination with the first possible implementation manner or the second possible implementation manner of the first aspect, the embodiment of the present application provides a sixth possible implementation manner of the first aspect, wherein a prediction analysis is performed on all insured customers according to the insurance adding insight model to obtain a prediction result of the customer adding insurance, specifically including:
[0033] Divide the proportion of the number of customers into multiple levels;
[0034] All insured customers are rated according to each level to obtain the probability of each customer's insurance rating. The multiple levels of customer numbers are 10%, 20%, 30% and 40% respectively.
[0035] In combination with the first possible implementation manner or the second possible implementation manner of the first aspect, the embodiment of the present application provides a seventh possible implementation manner of the first aspect, wherein a prediction analysis is performed on all insured customers according to the insurance adding insight model to obtain a prediction result of the customer adding insurance, specifically including:
[0036] Calculate the predicted value of the insured customer based on the target characteristics of the insurance insight model;
[0037] According to the predicted value of the insured customer, the contribution value of each target feature to the predicted value is calculated, and the contribution value is used as the impact value of the target feature on the insurance result.
[0038] In combination with the first possible implementation manner or the second possible implementation manner of the first aspect, the embodiment of the present application provides an eighth possible implementation manner of the first aspect, wherein a prediction analysis is performed on all insured customers according to the insurance adding insight model to obtain a prediction result of the customer adding insurance, specifically including:
[0039] According to different types of target characteristics, calculate the probability value of adding insurance for each type of customer classification;
[0040] Based on the continuous target characteristics, the customer is divided into designated categories and the probability of adding insurance for each type of customer is calculated separately. The continuous variables include: age and income.
[0041] In a second aspect, the embodiment of the present application further provides a prediction device for insurance service, the device comprising:
[0042] Determine the sample module, which is used to target the insurance data according to the target labeling tool to obtain the insurance insight sample set;
[0043] A model training module is used to train the insurance insight model according to the training set of the insurance insight sample set to obtain a candidate insurance insight model;
[0044] A model evaluation module is used to evaluate the candidate insurance insight model using the test set of the insurance insight sample set according to the confusion matrix and the model evaluation method to obtain a model evaluation result;
[0045] The model determination module is used to determine the insurance insight model that meets the detection accuracy based on the model evaluation results;
[0046] The predictive analysis model is used to conduct predictive analysis on all insured customers based on the insurance addition insight model to obtain the predicted results of customers adding insurance.
[0047] A prediction method for additional insurance business provided in an embodiment of the present application is compared with the prior art that relies on data experts in the insurance industry to perform additional insurance analysis on insured customers. The method targets the insurance data according to the target labeling tool to obtain an additional insurance insight sample set; trains the additional insurance insight model according to the training set of the additional insurance insight sample set to obtain a candidate additional insurance insight model; uses the test set of the additional insurance insight sample set to perform model evaluation on the candidate additional insurance insight model according to the confusion matrix and the model evaluation method to obtain a model evaluation result; determines the additional insurance insight model that meets the detection accuracy according to the model evaluation result; performs prediction analysis on all insured customers according to the additional insurance insight model to obtain a prediction result of the customer's additional insurance. Specifically, by labeling targets on historical insurance data, we obtain an insurance insight sample set, conduct model training and model evaluation on the insurance insight model, and obtain the insurance insight model based on the evaluation results. We can use machine learning methods to perform data modeling to achieve predictive analysis of insurance big data, and use descriptive statistics to interpret the model evaluation data. We can analyze the impact of each target feature in the training sample set on the insurance insight model, thereby accurately gaining insight into the future trend of customers' insurance purchases and improving the work efficiency of business personnel. At the same time, we can analyze customers in multiple types from two aspects: insurance products and customer characteristics, and finally obtain an accurate prediction result.
[0048] Furthermore, the prediction method for additional insurance business provided in the embodiment of the present application has the beneficial effects of the dependent claims: based on machine learning technology, historical insurance data is preprocessed, feature screened, model selected, parameter tuned and model evaluated, and the complex modeling and analysis process is streamlined and intelligent, so that business personnel can quickly realize full-process modeling with one click, and can be applied to a variety of insurance business scenarios to conduct a comprehensive analysis of historical customer information.
[0049] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0051] Figure 1 A flow chart of a prediction method for an insurance service provided in an embodiment of the present application is shown.
[0052] Figure 2A schematic diagram of a process for obtaining an insurance insight sample set in a prediction method for insurance business provided in an embodiment of the present application is shown.
[0053] Figure 3 A schematic diagram of a process of training an insurance logic model in a prediction method for insurance business provided in an embodiment of the present application is shown.
[0054] Figure 4 A schematic diagram of a process for calculating a customer insurance policy model evaluation result in a method for predicting insurance policy services provided in an embodiment of the present application is shown.
[0055] Figure 5 A schematic diagram of a process for predicting a customer's insurance rating in a method for predicting insurance business provided in an embodiment of the present application is shown.
[0056] Figure 6 A schematic diagram of a process for predicting an impact value of an insurance policy result in a prediction method for an insurance policy provided in an embodiment of the present application is shown.
[0057] Figure 7 A schematic diagram of a process for predicting the probability value of adding insurance by customer classification in a prediction method for adding insurance business provided in an embodiment of the present application is shown.
[0058] Figure 8 A schematic diagram of the structure of a prediction device for insurance service provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0059] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0060] Taking into account that currently industry data experts rely on experience to conduct big data analysis on insured customers, based on this, an embodiment of the present application provides a prediction method for additional insurance business, which is described below through an embodiment.
[0061] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0062] Figure 1 FIG. 1 shows a flow chart of a prediction method for an insurance service provided in an embodiment of the present application; Figure 1 As shown, the specific steps include:
[0063] Step S100, target labeling is performed on the insurance data using a target labeling tool to obtain an insurance insight sample set.
[0064] When step S100 is implemented specifically, before the target labeling of the insurance data, it also includes: obtaining the required insurance data summarized according to historical experience through data docking, wherein the data docking method includes: database sharing data method, interface transmission data method, local data upload method and C / S interaction method, etc.; the above-mentioned insurance data includes: customer gender, age, marital status, occupation, income, historical premiums, historical insurance types, number of insurances, insurance amount, order of insured products and product attention data, and the insurance data is targeted and labeled according to the target labeling tool to obtain an insurance insight sample set, and the insurance insight sample set is divided into a training set and a test set. The above-mentioned insurance addition is to increase the protection on the basis of the original policy protection. There are generally two ways to add insurance, one is to insure a new policy based on the current age, occupation, physical condition and other conditions, and the other is to add insurance through policy upgrade.
[0065] Step S200: training the insurance insight model according to the training set of the insurance insight sample set to obtain a candidate insurance insight model.
[0066] In the specific implementation of step S200, the above-mentioned insurance insight model can be a decision tree model or a neural network model. Before model training is performed on the insurance insight model, it is also necessary to perform data preprocessing on the data samples in the obtained insurance insight sample set, randomly sort the data samples using a feature screening strategy, and calculate the score of each feature in the data sample in turn. If the score exceeds the set threshold, the target feature is obtained. The above-mentioned set threshold is the benchmark value divided by the value of the third quartile in the training set; the target feature is input into the insurance logic model for cyclic training to obtain a candidate insurance logic model; according to the model evaluation method, the candidate insurance logic model is evaluated using the test set; according to the model evaluation result, the selected insurance logic model is obtained, and then the hyperparameters of the selected insurance logic model are tuned. After limited cyclic training, the insurance insight model with the optimal parameters is obtained; two-thirds of the data samples are selected from the insurance insight sample set as the training set, and the selected training set is input into the insurance insight model for model training until the model reaches the number of iterations to obtain the candidate insurance insight model.
[0067] Step S300, based on the confusion matrix and the model evaluation method, the candidate insurance insight model is evaluated using the test set of the insurance insight sample set to obtain a model evaluation result.
[0068] In the specific implementation of step S300, one-third of the data samples are selected from the insurance insight sample set as the test set. When the candidate insurance insight model is evaluated using the data samples of the test set, the influence of the data samples on the prediction results output by the candidate insurance insight model is analyzed through multiple data analysis algorithms to obtain the prediction results of the candidate insurance insight model for the test sample set, and the model evaluation data of the candidate insurance insight model is calculated based on the prediction results. Then, the model evaluation data is statistically analyzed based on the data samples to obtain the model evaluation results. The above data samples also include at least one of the information of the collection time and the location information of the collection test sample set. The model accuracy is measured according to the confusion matrix, and each row in the matrix represents The prediction results of the test sample set, each column represents the real information of the test sample set, and the cell data in the matrix represents the number of samples of different types. For example, taking the historical insurance type data sample as an example, the cell data in the third row and first column of the mixed matrix is 10. The row where the cell is located indicates that the prediction result is major disease insurance, and the column where the cell is located indicates that the real information is financial insurance. Then, the candidate insurance insight model predicts that the number of major disease insurance samples as financial insurance samples is 10; then, the prediction results of the test sample set are compared with the real information of the test sample set, and according to the type of data samples, the model evaluation data of the candidate insurance insight model is calculated, and statistical feature analysis is performed based on the model evaluation data to obtain the model evaluation results.
[0069] Step S400, determining an insurance insight model that meets the detection accuracy based on the model evaluation result;
[0070] During the specific implementation of step S400, after obtaining the evaluation results of the candidate insurance insight models, it is determined whether the evaluation results of the candidate insurance insight models are greater than a preset threshold. If so, the insurance insight model is selected from the candidate insurance insight models greater than the preset threshold.
[0071] Step S500, performing a prediction analysis on all insured customers according to the insurance adding insight model to obtain the prediction results of the customers adding insurance.
[0072] During the specific implementation of step S500, the gender, age, marital status, occupation, income, historical premiums, historical insurance types, number of insurance purchases, insurance amount, order of insured products and product attention data of each insured customer are predicted and analyzed according to the insurance insight model to obtain the customer insurance model evaluation results, customer insurance rating, insurance result impact value and customer classification insurance probability value.
[0073] In one possible implementation, Figure 2 A schematic diagram of a process of obtaining an insurance insight sample set in a prediction method for insurance business provided in an embodiment of the present application is shown; before executing the above step S100, it also includes:
[0074] Step S10, determining the insurance target according to the characteristic parameters of the insured customer and the characteristic parameters of the insurance time point.
[0075] Step S20, obtaining the insurance data of the insurance target, the insurance data includes: customer gender, age, marital status, occupation, income, historical premiums, historical insurance types, number of insurances, insurance amount, order of insurance products and product attention data.
[0076] Step S30, labeling the target data in the insurance data to obtain the training set and the test set of the insurance insight sample set.
[0077] During the specific implementation of steps S10, S20, and S30, a customer insurance target is set in a computer device, wherein the insurance target is divided into product insurance items and customer insurance items. The product insurance items are used to gain insights into the characteristic parameters of the insured customers of a certain insurance product; the customer insurance items are used to gain insights into the customer's preference characteristics such as the insured object, insured product, and insurance time point; insurance data is summarized based on the insurance target and history, such as the customer's gender, age, marital status, occupation, income, historical premiums, historical insurance types, number of insurances, insurance amount, order of insured products, and product attention data; relevant insurance data is obtained from the database, client, and server through a database sharing data method, an interface transmission data method, a local data upload method, and a C / S interaction method, and the annotation information of the insurance data is targeted and annotated according to a target annotation tool to obtain an insurance insight sample set, and the insurance insight sample set is divided into a training set and a test set according to a custom data volume ratio or a set time point.
[0078] In one possible implementation, Figure 3 A schematic diagram of a process of training a model for an insurance logic model in a prediction method for an insurance service provided in an embodiment of the present application is shown; before executing the above step S200 to train the insurance insight model according to the training set of the insurance insight sample set to obtain a candidate insurance insight model, the method further includes:
[0079] Step S40, converting the format of the binary code of the training set to obtain data samples of the training set.
[0080] Step S50, using a feature screening strategy to perform feature screening on each converted data sample to obtain target features.
[0081] Step S60, input the feature vector of the target feature into the added protection logic model, and use multiple algorithms to perform cyclic training on the added protection logic model to obtain a candidate added protection logic model, where the multiple algorithms include: logistic regression algorithm, decision tree C4.5 algorithm, classification regression tree algorithm, XGBoost tree class algorithm, LightGBM optimization algorithm and neural network algorithm.
[0082] Step S70, according to the model evaluation method, use the test set of the insurance insight sample set to perform model evaluation on the candidate insurance logic model to obtain a model evaluation result.
[0083] Step S80: Determine the selected insurance logic model that meets the requirements based on the evaluation results.
[0084] Step S90, using a Bayesian optimization algorithm to tune the hyperparameters of the selected insurance logic model to obtain an insurance insight model with optimal parameters.
[0085] In the specific implementation of steps S40, S50, S60, S70, S80, and S90, the data samples in the obtained insurance insight sample set are verified, the missing situation of each feature data sample is analyzed, all null values or single-value features are filtered out, and multiple features are analyzed and processed at the same time, and the features with high correlation are reduced in dimension to remove redundant data to obtain the processed insurance insight sample set, and then the binary code of the training set is converted to obtain a data sample in the form of a string, and then a feature screening strategy is adopted to record each feature g1, g2, ..., g n The importance of base1, base2, ..., base n ; Then randomly sort the data samples of the training set. After repeating the above operation multiple times, the importance of multiple features after different random sorting is obtained. For example, the importance set based on feature gi is nset i ={nset i1 ,nset i2 ,…,new m}; Calculate the score of each feature, score = base i / percent 0.75 (nset i ), the score is the benchmark value divided by the third quartile value in the training set; and the score is screened, and the feature with the highest score is taken as the final feature.
[0086] After screening the features, the feature vector of the target feature is input into the insurance logic model, and a variety of algorithms are used to perform cyclic training on the insurance logic model. The algorithms include: logistic regression algorithm, decision tree C4.5 algorithm, classification regression tree algorithm, XGBoost tree class algorithm, LightGBM optimization algorithm and neural network algorithm to obtain a candidate insurance logic model. According to the model evaluation method, the candidate insurance logic model is evaluated using the test set to obtain model evaluation data. The model evaluation data is between 0.5 and 1. The larger the model evaluation data, the higher the model accuracy. The model evaluation results are obtained based on the model evaluation data, and the selected insurance logic model is obtained based on the model evaluation results.
[0087] After selecting the selected insurance logic model, the Bayesian optimization algorithm is used to automatically select the optimal hyperparameters of the selected insurance logic model. After finite cycle training, the insurance insight model with the optimal parameters is obtained; the specific calculation method is as follows: Assume X = x1, x2, ..., x n , represents the search space of hyperparameters, each X represents a set of hyperparameter combinations, f represents the hyperparameter optimization of the selected hyperparameter optimization model, and the hyperparameter optimization insight model is obtained. The data set D = {(x1, y1), (x2, y2), …, (x n ,y n )}; That is, if y i =f(x i ) represents a set of hyperparameters. X obtains the corresponding result y by selecting the guaranteed logistic model f. Since each time the parameters are selected, f(x i ), each calculation consumes a lot of parameters, so it is necessary to set a fixed number of parameter selections to T times, each t = 1, ..., T. In each cycle calculation, a set of hyperparameters x is input i , we get y i =f(x i ), select x in t cycles t, So that the optimal result can be obtained under a limited number of cycles;
[0088] For example: How to choose which x to observe in each iteration loop? t , x t is achieved by optimizing another function acquisition function (α t ) to choose, that is
[0089]
[0090] Select Gaussian Process-Upper Confidence Bound (GP-UCB) as the acquisition function:
[0091] Right now,
[0092] Among them, μ t-1 (x) and σ t-1 (x) is the data set D = {(x1,y1),(x2,y2),…,(x t-1 ,y t-1 )}, the β line represents the default setting of a constant, and the optimal hyperparameter x is finally obtained through multiple cycles. t ; Get the optimal parameters of the insurance insight model.
[0093] In a feasible implementation scheme, in the above step S400, determining the insurance insight model that meets the detection accuracy according to the model evaluation result includes:
[0094] Step 4001: If the detection accuracy of the candidate insight model for added protection reaches a threshold, it is determined as the insight model for added protection.
[0095] Step 4002: if the detection accuracy of the candidate insurance insight model does not reach the threshold, the candidate insurance insight model is retrained.
[0096] During the specific implementation of steps 4001 and 4002, it is determined whether the evaluation score of the candidate insight model is greater than a preset threshold based on the model evaluation results. If it is greater than the threshold, it indicates that the candidate insight model meets the detection accuracy. Then, the insight model is added from the candidate insight models greater than the preset threshold. If the evaluation score of the candidate insight model is less than the preset threshold, the model is discarded or retrained.
[0097] In one possible implementation, Figure 4 A schematic diagram of a process for predicting the evaluation result of a customer insurance policy adding model in a prediction method for insurance policy adding business provided in an embodiment of the present application is shown; in the above step S500, a prediction analysis is performed on all insured customers according to the insurance policy adding insight model to obtain a prediction result of the customer insurance adding, including:
[0098] Step S5001: If the evaluation index of the insurance insight model exceeds the first preset threshold, the evaluation result of the customer insurance model is excellent.
[0099] Step S5002: If the evaluation index of the insurance insight model is equal to the second preset threshold, the evaluation result of the customer insurance model is good.
[0100] Step S5003, if the evaluation index of the insurance insight model is equal to the third preset threshold, the customer insurance model evaluation result is optional.
[0101] Step S5004: If the evaluation index of the insurance insight model is less than the fourth preset threshold, the customer insurance model evaluation result is unavailable.
[0102] During the specific implementation of steps S5001, S5002, S5003, and S5004, the customer insurance prediction results of the insurance insight model are output according to each feature, the selected algorithm, and the hyperparameters, and the evaluation results of the customer insurance model are described in segments, wherein if the evaluation index of the insurance insight model exceeds the first preset threshold, the customer insurance model evaluation result is excellent; if the evaluation index of the insurance insight model is equal to the second preset threshold, the customer insurance model evaluation result is good; if the evaluation index of the insurance insight model is equal to the third preset threshold, the customer insurance model evaluation result is optional; if the evaluation index of the insurance insight model is less than the fourth preset threshold, the customer insurance model evaluation result is unavailable, wherein the preset thresholds are 0.8 for excellent, 0.6-0.8 for good, 0.5-0.6 for unavailable, and less than 0.5 for unavailable.
[0103] In one possible implementation, Figure 5 The flowchart of predicting a customer's insurance rating in a prediction method for insurance business provided in an embodiment of the present application is shown; in the above step S500, a prediction analysis is performed on all insured customers according to the insurance insight model to obtain a prediction result of the customer's insurance, including:
[0104] Step S5005, dividing the proportion of the number of customers into multiple levels.
[0105] Step S5006, all insured customers are evaluated according to each level to obtain the probability of each customer's insurance rating, and the multiple levels of customer numbers are 10%, 20%, 30% and 40% respectively.
[0106] During the specific implementation of steps S5005 and S5006, a prediction analysis is performed on all customers based on the insurance adding insight model to obtain a predicted probability value for each customer to add insurance, and the predicted probability values for adding insurance are sorted from large to small. Then, the customers are divided into 4 levels according to different quantities, such as 10%, 20%, 30% and 40%, and the customers are evaluated according to different levels. The higher the level, the greater the probability of the customer adding insurance.
[0107] In one possible implementation, Figure 6 A schematic diagram of a process for calculating the impact value of an insurance addition result in a prediction method for an insurance addition service provided in an embodiment of the present application is shown; in the above step S500, a prediction analysis is performed on all insured customers according to the insurance addition insight model to obtain a prediction result of the customer's insurance addition, including:
[0108] Step S5007, calculating the predicted value of the insured customer based on the target features of the insurance insight model.
[0109] Step S5008, based on the predicted value of the insured customer, calculate the contribution value of each target feature to the predicted value, and the contribution value serves as the impact value of the target feature on the insurance result.
[0110] In the specific implementation of steps S5007 and S5008, according to the characteristics of the insurance insight model, the target feature is selected, and the contribution value of the corresponding feature in each data sample to the final predicted value is calculated according to the target feature. The average value of the different contribution values of the feature in all feature sequences is taken to obtain the influence probability value of the feature; for example: a single feature A is used to construct an insurance insight model alone to generate a prediction result Predict(A); feature B is added to the insurance insight model to generate a prediction result Predict(A, B), then the contribution value of the prediction value of B is Predict(A, B)-Predict(A); a global feature sequence is generated for each combination in multiple feature sequences, and the predicted value of one of the features is calculated. According to the predicted value of the insured customer, the contribution value of each target feature to the predicted value is calculated and the average value is taken to obtain the influence probability value of the feature.
[0111] In one possible implementation, Figure 7 The flowchart of predicting the probability value of adding insurance by customer classification in a prediction method of adding insurance business provided by an embodiment of the present application is shown; in the above step S500, a prediction analysis is performed on all insured customers according to the insurance adding insight model to obtain the prediction result of the customer adding insurance, including:
[0112] Step S5009, calculate the probability value of adding insurance for each type of customer classification according to different types of target features.
[0113] Step S5010, based on continuous target features, divide according to designated categories and calculate the probability value of customer insurance addition for each category respectively. The continuous target features include: age and income.
[0114] During the specific implementation of steps S5009 and S5010, the probability values of adding insurance for each category of customers are calculated respectively according to gender, occupation, and marital status. Based on continuous target features such as age and income, they are divided into specified types, and the probability values of adding insurance for each category of customers are calculated respectively, and a descriptive statistical analysis chart of feature importance is generated.
[0115] Figure 8 FIG. 1 shows a schematic diagram of a prediction device structure for an insurance service provided in an embodiment of the present application. Figure 8 As shown, the above device comprises:
[0116] The sample determination module 6001 is used to target the insurance data according to the target labeling tool to obtain the insurance insight sample set;
[0117] The model training module 6002 is used to perform model training on the insurance insight model according to the training set of the insurance insight sample set to obtain a candidate insurance insight model;
[0118] The model evaluation module 6003 is used to evaluate the candidate insurance insight model using the test set of the insurance insight sample set according to the confusion matrix and the model evaluation method to obtain the model evaluation result;
[0119] The determination model module 6004 is used to determine the added security insight model that meets the detection accuracy according to the model evaluation results;
[0120] The prediction analysis model 6005 is used to perform prediction analysis on all insured customers based on the insurance adding insight model to obtain the prediction results of the customers' insurance adding.
[0121] In specific implementation, before the target labeling of the insurance data, it also includes obtaining the required insurance data summarized according to historical experience through data docking, wherein the data docking methods include: database sharing data method, interface transmission data method, local data upload method and C / S interaction method, etc.; the above insurance data includes: customer gender, age, marital status, occupation, income, historical premiums, historical insurance types, number of insurances, insurance amount, order of insurance products and product attention data, target labeling of the insurance data according to the target labeling tool, and obtaining the insurance insight sample set, which is divided into a training set and a test set. The above insurance addition is to increase the protection on the basis of the original insurance policy protection. There are generally two ways to add insurance, one is to insure a new policy based on the current age, occupation, physical condition and other conditions, and the other is to add insurance through policy upgrade;
[0122] Perform data preprocessing on the data samples in the obtained insurance insight sample set, randomly sort the data samples using a feature screening strategy, calculate the score of each feature in the data sample in turn, and if the score exceeds the set threshold, obtain the target feature, which is the benchmark value divided by the third quartile value in the training set; input the target feature into the insurance logic model for cyclic training to obtain a candidate insurance logic model, and use the test set to evaluate the candidate insurance logic model according to the model evaluation method. According to the model evaluation results, obtain the selected insurance logic model, and then optimize the parameters of the selected insurance logic model according to the hyperparameters. After limited cyclic training, obtain the insurance insight model with the optimal parameters; select two-thirds of the data samples from the insurance insight sample set as the training set, and input the selected training set into the insurance insight model for model training until the model reaches the number of iterations to obtain the candidate insurance insight model;
[0123] One third of the data samples are selected from the insurance insight sample set as the test set. When the data samples of the test set are used to evaluate the candidate insurance insight model, the influence of the data samples on the prediction results output by the candidate insurance insight model is analyzed through multiple data analysis algorithms to obtain the prediction results of the candidate insurance insight model for the test sample set. The model evaluation data of the candidate insurance insight model is calculated based on the prediction results. Then, the model evaluation data is statistically analyzed based on the data samples to obtain the model evaluation results.
[0124] After obtaining the evaluation results of the candidate insurance insight models, it is determined whether the evaluation results of the candidate insurance insight models are greater than a preset threshold. If the evaluation results are greater than the threshold, the insurance insight model is selected from the candidate insurance insight models greater than the preset threshold;
[0125] Based on the insurance adding insight model, we conduct predictive analysis on the gender, age, marital status, occupation, income, historical premiums, historical insurance types, number of insurance purchases, insurance amount, order of insured products and product attention data of each insured customer, and obtain the customer insurance adding model evaluation results, customer insurance adding rating, impact value of insurance adding results and customer classification insurance adding probability value.
[0126] Based on the above analysis, it can be seen that compared with the related technology that relies on the cooperation of professional data scientists and business personnel to complete the customer's insurance insight analysis, the embodiment of the present application provides a predictive analysis of insurance big data through data modeling through machine learning methods, which can analyze the impact of the corresponding features of each target feature in the training sample set on the insurance insight model, thereby accurately gaining insight into the customer's future insurance trends and improving the work efficiency of business personnel; at the same time, it can analyze customers in multiple types from two aspects: insurance products and customer characteristics, and finally obtain an accurate prediction result.
[0127] The prediction device for the insurance service provided in the embodiment of the present application can be specific hardware on the device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in the embodiment of the present application are the same as those of the aforementioned method embodiment. For the sake of brief description, the parts not mentioned in the device embodiment can refer to the corresponding contents in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here.
[0128] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0129] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0130] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0131] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.
[0132] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.
[0133] Finally, it should be noted that the above embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the aforementioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the aforementioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A method for predicting an insurance service, characterized in that: include: Use the target labeling tool to label the insurance data and obtain the insurance insight sample set; Performing model training on the insurance insight model according to the training set of the insurance insight sample set to obtain a candidate insurance insight model; Before the model training of the insurance insight model, the method also includes: performing data preprocessing on the data samples in the obtained insurance insight sample set, randomly sorting the data samples using a feature screening strategy, and calculating the score of each feature in the data sample in turn. If the score exceeds the set threshold, the target feature is obtained, and the set threshold is the benchmark value divided by the value of the third quartile in the training set; inputting the target feature into the insurance logic model for cyclic training to obtain a candidate insurance logic model, and using the test set to evaluate the candidate insurance logic model according to the model evaluation method, and obtaining a selected insurance logic model according to the model evaluation result, and then tuning the hyperparameters of the selected insurance logic model, and obtaining the insurance insight model with the optimal parameters after limited cyclic training; selecting two-thirds of the data samples from the insurance insight sample set as the training set, and inputting the selected training set into the insurance insight model for model training until the model reaches the number of iterations to obtain the candidate insurance insight model; According to the confusion matrix and the model evaluation method, the candidate insurance insight model is evaluated using the test set of the insurance insight sample set to obtain a model evaluation result; Determine an additional insight model that meets the detection accuracy based on the model evaluation results; According to the insurance adding insight model, a prediction analysis is performed on all insured customers to obtain the prediction results of the customers adding insurance.
2. The method for predicting the insured service according to claim 1, characterized in that: Before getting the Jiabao Insight sample set, it also includes: Determine the additional insurance target based on the characteristic parameters of the insured customer and the characteristic parameters of the insurance time point; Obtain the insurance data of the insured target, which includes: customer gender, age, marital status, occupation, income, historical premiums, historical insurance types, number of insurance purchases, insurance amount, order of insurance products, and product attention data; Label the target data in the insurance data to obtain the training set and test set of the insurance insight sample set.
3. The method for predicting the insured service according to claim 1, characterized in that: The model training of the insurance insight model is performed according to the training set of the insurance insight sample set. Before obtaining the candidate insurance insight model, the following steps are also included: The binary code of the training set is converted into a format, and the data samples of the training set are obtained after the conversion; A feature screening strategy is used to screen the features of each converted data sample to obtain the target features; Input the feature vector of the target feature into the added insurance logic model, and use multiple algorithms to perform cyclic training on the added insurance logic model to obtain a candidate added insurance logic model, among which the multiple algorithms include: logistic regression algorithm, greedy algorithm, decision tree C4.5 algorithm, classification regression tree algorithm, XGBoost tree class algorithm, LightGBM optimization algorithm and neural network algorithm; According to the model evaluation method, the candidate insurance logic model is evaluated using the test set of the insurance insight sample set to obtain the model evaluation result; According to the evaluation results, determine the selected insurance logic model that meets the requirements; The Bayesian optimization algorithm is used to tune the hyperparameters of the selected insurance logic model to obtain the insurance insight model with the optimal parameters.
4. The method for predicting the insured service according to claim 1, characterized in that: Based on the model evaluation results, determine the insurance insight model that meets the detection accuracy, including: If the detection accuracy of the candidate added-protection insight model reaches the threshold, it is determined as the added-protection insight model; If the detection accuracy of the candidate insurance insight model does not reach the threshold, the candidate insurance insight model is retrained.
5. The method for predicting the insured service according to claim 1, characterized in that: Based on the insurance insight model, we conduct a predictive analysis on all insured customers to obtain the predicted results of customers adding insurance, including: According to the insurance adding insight model, a predictive analysis is conducted on all insured customers to obtain the prediction results of customers adding insurance, among which the prediction results of customers adding insurance are: customer insurance adding model evaluation results, customer insurance adding rating, impact value of insurance adding results and customer classification insurance adding probability value.
6. The method for predicting the insured service according to claim 5, characterized in that: According to the insurance insight model, all insured customers are predicted and analyzed to obtain the prediction results of customers adding insurance, including: If the evaluation index of the insurance insight model exceeds the first preset threshold, the customer insurance model evaluation result is excellent; If the evaluation index of the insurance insight model is equal to the second preset threshold, the evaluation result of the customer insurance model is good; If the evaluation index of the insurance insight model is equal to the third preset threshold, the customer insurance model evaluation result is optional; If the evaluation index of the insurance insight model is less than the fourth preset threshold, the customer insurance model evaluation result is unavailable.
7. The method for predicting the insured service according to claim 5, characterized in that: According to the insurance insight model, all insured customers are predicted and analyzed to obtain the prediction results of customers adding insurance, including: Divide the proportion of the number of customers into multiple levels; All insured customers are rated according to each level to obtain the probability of each customer's insurance rating. The multiple levels of customer numbers are 10%, 20%, 30% and 40% respectively.
8. The method for predicting the insured service according to claim 5, characterized in that: According to the insurance insight model, all insured customers are predicted and analyzed to obtain the prediction results of customers adding insurance, including: Calculate the predicted value of the insured customer based on the target characteristics of the insurance insight model; According to the predicted value of the insured customer, the contribution value of each target feature to the predicted value is calculated, and the contribution value is used as the impact value of the target feature on the insurance result.
9. The method for predicting the insured service according to claim 5, characterized in that: According to the insurance insight model, all insured customers are predicted and analyzed to obtain the prediction results of customers adding insurance, including: According to different types of target characteristics, calculate the probability value of adding insurance for each type of customer classification; Based on the continuous target characteristics, the customer is divided into designated categories and the probability of adding insurance for each type of customer is calculated separately. The continuous variables include: age and income.
10. A prediction device for insurance service, characterized in that: The device includes: Determine the sample module, which is used to target the insurance data according to the target labeling tool to obtain the insurance insight sample set; A model training module is used to perform model training on the insurance insight model according to the training set of the insurance insight sample set to obtain a candidate insurance insight model; the model training module is specifically used to: before the model training of the insurance insight model, perform data preprocessing on the data samples in the obtained insurance insight sample set, randomly sort the data samples using a feature screening strategy, and calculate the score of each feature in the data sample in turn. If the score exceeds the set threshold, the target feature is obtained, and the set threshold is the benchmark value divided by the value of the third quartile in the training set; the target feature is input into the insurance logic model for cyclic training to obtain a candidate insurance logic model, and according to the model evaluation method, the candidate insurance logic model is evaluated using the test set, and the selected insurance logic model is obtained according to the model evaluation result, and then the hyperparameters of the selected insurance logic model are tuned, and after limited cyclic training, the insurance insight model with the optimal parameters is obtained; two-thirds of the data samples are selected from the insurance insight sample set as the training set, and the selected training set is input into the insurance insight model for model training until the model reaches the number of iterations to obtain the candidate insurance insight model; A model evaluation module is used to evaluate the candidate insurance insight model using the test set of the insurance insight sample set according to the confusion matrix and the model evaluation method to obtain a model evaluation result; The model determination module is used to determine the insurance insight model that meets the detection accuracy based on the model evaluation results; The predictive analysis model is used to conduct predictive analysis on all insured customers based on the insurance addition insight model to obtain the predicted results of customers adding insurance.
Citation Information
Patent Citations
Insurance renewal prediction method, device, computer equipment and storage medium
CN108830734A
Vehicle insurance renewal prediction method and system
CN109978257A
User intention prediction method and device, computer equipment and storage medium
CN110389970A
Automatic model updating method, device and system and electronic equipment
CN113011596A
Enterprise customer insurance renewal prediction method and device, medium and electronic equipment
CN113822724A