Policy Processing Method, Device, Medium and Electronic Device
By analyzing historical policy data and training prediction models, insurance companies can predict the policy's refund risks, solving the problem that customers' refunds are difficult to predict, reducing losses and retaining customer resources.
Patent Information
- Application Number
- CN201911164268.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-11-25
AI Technical Summary
After the policy takes effect, it is difficult for insurance companies to predict whether customers will sue the insurance, resulting in losses and waste of resources.
By obtaining and analyzing historical policy data, extracting feature data and training gradient enhancement classifiers, predicting the probability of surrender of the current policy.
Effectively predict whether there is a risk of surrendering the policy, help insurance companies take measures to retain customers, reduce losses and retain customer resources.
Smart Images

Figure CN112837168B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a policy processing method, device, medium and electronic device. Background Art
[0002] With the rapid development of Internet technology, customers can purchase insurance online, which leads to an increasing number of insurance products of insurance companies and thus a large number of policies are generated. At present, after a policy takes effect, customers often terminate the insurance contract for various reasons, that is, initiate surrender actively. There are many policies surrendered by insurance companies every year. Surrender by customers will bring certain losses to both customers and insurance companies. The occurrence of surrender has great unpredictability and it is generally difficult to intervene in advance to retain customers. Therefore, how to predict whether a customer will surrender based on policy data has become a technical problem faced by insurance companies currently. In view of this technical problem, the present invention proposes a policy processing method.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a policy processing method, device, medium and electronic device, so as to at least avoid customer surrender to a certain extent, reduce the losses of insurance companies, and thus retain customer resources for insurance companies.
[0005] Other features and advantages of the present invention will become apparent through the following detailed description, or be learned in part through the practice of the present invention.
[0006] According to the first aspect of the embodiments of the present invention, a policy processing method is provided, including: obtaining historical policy data, where the historical policy data includes first historical policy data and second historical policy data, the first historical policy data includes historical policy data containing surrender records, and the second historical policy data includes historical policy data without surrender records; extracting first feature data from the first historical policy data and second feature data from the second historical policy data, where the first feature data includes the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, and the second feature data includes the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data; training a gradient boosting classifier based on the attributes of the first historical policy data, the attribute values of the attributes of the first historical policy data, the attributes of the second historical policy data, and the attribute values of the attributes of the second historical policy data to obtain the preset model; obtaining current policy data; extracting feature data from the current policy data, where the feature data includes the attributes of the current policy data and the attribute values of the attributes; determining the surrender probability and / or non-surrender probability of the current policy data according to the attributes of the current policy data and the attribute values of the attributes through the preset model.
[0007] In some embodiments of the present invention, the prediction model includes a decision tree, and each branch of the decision tree includes partial attributes of the historical policy data, attribute values of partial attributes of the historical policy data, and surrender probability and / or non-surrender probability.
[0008] In some embodiments of the present invention, training the gradient boosting classifier based on the attributes of the first historical policy data, the attribute values of the attributes of the first historical policy data, the attributes of the second historical policy data, and the attribute values of the attributes of the second historical policy data includes: calculating the information gain of each attribute in the historical policy data according to the attributes of the first historical policy data, the attribute values of the attributes of the first historical policy data, the attributes of the second historical policy data, and the attribute values of the attributes of the second historical policy data through the gradient boosting decision tree algorithm; determining the distribution of each attribute in the decision tree according to the information gain of each attribute in the historical policy data; training the gradient boosting classifier based on the distribution of each attribute in the decision tree and the marked surrender records in the historical policy data.
[0009] In some embodiments of the present invention, the formula of the gradient boosting decision tree algorithm is:
[0010]
[0011] Wherein, S is a set of historical policy data, A is an attribute in the historical policy data, Entropy(S) is the information entropy of S, Value(A) is a set of attribute values of A, v is an attribute value in the set of attribute values of A, and S v is a set of historical data when the attribute value of A in S is v, and Entropy(S v ) is the information entropy of S v .
[0012] In some embodiments of the present invention, the method further includes: obtaining test data, where the test data includes historical policy data with surrender records and historical policy data without surrender records; obtaining the surrender probability of the test data through the preset model; sorting the surrender probabilities of the test data, and using the test data with surrender probabilities greater than or equal to a preset threshold as the data to be surrendered according to the sorting result; determining the surrender coverage rate of the test data according to the actual surrender quantity of the test data and the actual surrender quantity of the data to be surrendered; determining the surrender accuracy rate of the test data according to the actual surrender quantity of the data to be surrendered and the quantity of the data to be surrendered; and evaluating the preset model according to the surrender coverage rate and the surrender accuracy rate of the test data.
[0013] In some embodiments of the present invention, determining the surrender probability and / or non-surrender probability of the current policy data through the preset model according to the attribute of the current policy data and the attribute value of the attribute includes: determining the branch of the decision tree in the preset model where the attribute of the current policy data and the attribute value of the attribute are located; and determining the surrender probability and / or non-surrender probability of the current policy data according to the information of the branch.
[0014] In some embodiments of the present invention, determining the surrender probability and / or non-surrender probability of the current policy data according to the information of the branch includes: determining the surrender probability and / or non-surrender probability of the current policy data according to the attributes, attribute values, and surrender probability and / or non-surrender probability included in the branch.
[0015] According to the second aspect of the embodiments of the present invention, a policy processing device is provided, including: a first acquisition module, configured to acquire historical policy data, where the historical policy data includes first historical policy data and second historical policy data, the first historical policy data includes historical policy data containing surrender records, and the second historical policy data includes historical policy data without surrender records; a first extraction module, configured to extract first feature data from the first historical policy data and extract second feature data from the second historical policy data, where the first feature data includes the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, and the second feature data includes the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data; a training module, configured to train a gradient boosting classifier based on the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data and the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data to obtain the preset model; a second acquisition module, configured to acquire policy data; a second extraction module, configured to extract feature data from the policy data, where the feature data includes the attributes of the policy data and the attribute values of the attributes; a first determination module, configured to determine the surrender probability and / or non-surrender probability of the policy data according to the attributes of the policy data and the attribute values of the attributes through the preset model.
[0016] In some embodiments of the present invention, the prediction model includes a decision tree, and each branch of the decision tree includes partial attributes of the historical policy data, attribute values of partial attributes of the historical policy data, and a surrender probability and / or non-surrender probability.
[0017] In some embodiments of the present invention, the above training module includes: a calculation module, configured to calculate the information gain of each attribute in the historical policy data through the gradient boosting decision tree algorithm according to the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data and the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data; a second determination module, configured to determine the distribution of each attribute in the decision tree according to the information gain of each attribute in the historical policy data; a training sub-module, configured to train the gradient boosting classifier based on the distribution of each attribute in the decision tree and the surrender records marked in the historical policy data.
[0018] In some embodiments of the present invention, the formula of the gradient boosting decision tree algorithm is:
[0019]
[0020] Wherein, S is a set of historical policy data, A is an attribute in the historical policy data, Entropy(S) is the information entropy of S, Value(A) is a set of attribute values of A, v is an attribute value in the set of attribute values of A, and S v is a set of historical data when the attribute value of A in S is v, and Entropy(S v ) is the information entropy of S v .
[0021] In some embodiments of the present invention, the device further includes: a third acquisition module configured to acquire test data, where the test data includes historical policy data with surrender records and historical policy data without surrender records; a fourth acquisition module configured to acquire the surrender probability of the test data through the preset model; a sorting module configured to sort the surrender probabilities of the test data and use the test data with surrender probabilities greater than or equal to a preset threshold as the data to be surrendered according to the sorting result; a third determination module configured to determine the surrender coverage rate of the test data according to the actual surrender quantity of the test data and the actual surrender quantity of the data to be surrendered; a fourth determination module configured to determine the surrender accuracy rate of the test data according to the actual surrender quantity of the data to be surrendered and the quantity of the data to be surrendered; and an evaluation module configured to evaluate the preset model according to the surrender coverage rate of the test data and the surrender accuracy rate of the test data.
[0022] In some embodiments of the present invention, the above-mentioned first determination module includes: a determination branch module configured to determine the attributes of the current policy data and the attribute values of the attributes in the branches of the decision tree according to the decision tree in the preset model; a determination probability module configured to determine the surrender probability and / or non-surrender probability of the current policy data according to the information of the current branch.
[0023] In some embodiments of the present invention, the above-mentioned determination probability module is configured to: determine the surrender probability and / or non-surrender probability of the current policy data according to the attributes, attribute values, and surrender probability and / or non-surrender probability included in the branch.
[0024] According to a third aspect of the embodiments of the present invention, there is provided an electronic device, including: one or more processors; a storage device configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the one or more processors to implement the policy processing method as described in the first aspect of the above embodiments.
[0025] According to a fourth aspect of the embodiments of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, it implements the policy processing method as described in the first aspect of the above embodiments.
[0026] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0027] In the technical solutions provided by some embodiments of the present invention, policy data is obtained, and feature data is extracted from the policy data. The feature data includes the attributes of the policy data and the attribute values of the attributes. Then, a preset model determines the surrender probability and / or non-surrender probability of the policy data according to the attributes of the policy data and the attribute values of the attributes. In this way, the risk of surrender of the policy data can be effectively predicted, thereby avoiding customer surrender and reducing the losses of the insurance company, and thus retaining customer resources for the insurance company.
[0028] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0030] Figure 1 A schematic diagram showing an exemplary system architecture to which the policy processing method or policy processing device according to the embodiments of the present invention can be applied;
[0031] Figure 2 A flowchart schematically showing the policy processing method according to an embodiment of the present invention;
[0032] Figure 3 A flowchart schematically showing the policy processing method according to another embodiment of the present invention;
[0033] Figure 4 A schematic diagram showing a decision tree in a preset model according to an embodiment of the present invention;
[0034] Figure 5 A flowchart schematically showing the policy processing method according to another embodiment of the present invention;
[0035] Figure 6 A flowchart schematically showing the policy processing method according to another embodiment of the present invention;
[0036] Figure 7 A block diagram schematically showing the policy processing device according to an embodiment of the present invention;
[0037] Figure 8 Schematically shows a block diagram of an insurance policy processing device according to another embodiment of the present invention;
[0038] Figure 9 Schematically shows a block diagram of an insurance policy processing device according to another embodiment of the present invention;
[0039] Figure 10 Schematically shows a block diagram of an insurance policy processing device according to another embodiment of the present invention;
[0040] Figure 11 Shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present invention. Detailed implementation manners
[0041] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this invention will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0042] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present invention. However, those skilled in the art will realize that the technical solutions of the present invention can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present invention.
[0043] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0044] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0045] Figure 1 Shows a schematic diagram of an exemplary system architecture to which the insurance policy processing method or insurance policy processing device of the embodiments of the present invention can be applied.
[0046] As Figure 1As shown, the system architecture 100 may include one or more of the terminal devices 101, 102, 103, the network 104, and the server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0047] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0048] are merely illustrative. According to implementation requirements, there can be any number of terminal devices, networks, and servers. For example, the server 105 can be a server cluster composed of multiple servers, etc.
[0049] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 can be various electronic devices with a display screen, including but not limited to smartphones, tablets, portable computers, and desktop computers, etc.
[0050] In some embodiments, the policy processing method provided by the embodiments of the present invention is generally executed by the server 105. Correspondingly, the policy processing device is generally set in the server 105. In other embodiments, some terminals may have functions similar to those of the server to execute this method. Therefore, the policy processing method provided by the embodiments of the present invention is not limited to being executed on the server side.
[0051] Figure 2 Schematically shows a flowchart of the policy processing method according to an embodiment of the present invention.
[0052] As Figure 2 shown, the policy processing method may include steps S210 to S260.
[0053] In step S210, historical policy data is obtained. The historical policy data includes first historical policy data and second historical policy data. The first historical policy data includes historical policy data containing surrender records, and the second historical policy data includes historical policy data without surrender records.
[0054] In step S220, first feature data is extracted from the first historical policy data, and second feature data is extracted from the second historical policy. The first feature data includes the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data. The second feature data includes the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data.
[0055] In step S230, a gradient boosting classifier is trained based on the attributes of the first historical policy data, the attribute values of the attributes of the first historical policy data, the attributes of the second historical policy data, and the attribute values of the attributes of the second historical policy data to obtain the preset model.
[0056] In step S240, current policy data is obtained.
[0057] In step S250, feature data is extracted from the current policy data. The feature data includes the attributes of the current policy data and the attribute values of the attributes.
[0058] In step S260, the surrender probability and / or non-surrender probability of the current policy data is determined by the preset model according to the attributes of the current policy data and the attribute values of the attributes.
[0059] This method can train a gradient boosting classifier based on historical policy data with surrender records and historical policy data without surrender records to obtain the preset model. The preset model trained in this way is more accurate in determining the surrender probability when judging whether there is a surrender risk in the current policy data. For example, current policy data is obtained, and feature data is extracted from the current policy data. The feature data includes the attributes of the current policy data and the attribute values of the attributes. Then, the surrender probability and / or non-surrender probability of the current policy data is determined by the preset model according to the attributes of the current policy data and the attribute values of the attributes. In this way, the risk of surrender in the policy data can be effectively predicted, thereby avoiding customer surrender and reducing the losses of the insurance company, and thus retaining customer resources for the insurance company.
[0060] In an embodiment of the present invention, the above historical policy data includes first historical policy data and second historical policy data. Among them, the first historical policy data includes historical policy data containing surrender records. For example, a certain customer submitted a surrender application for an insurance product insured on September 12, 2018, and successfully completed the surrender process. The policy data of this customer can be regarded as the first historical policy data. Another example, a certain customer insured an insurance product on July 22, 2019, and successfully completed the insurance process, and no surrender application from this customer has been received so far. The policy data of this customer can be regarded as the second historical policy data.
[0061] In an embodiment of the present invention, the above historical policy data containing surrender records may include, but is not limited to, the situation of claim settlement records, the types of insurance products, and the ages of the policyholders or insured persons. Among them, the situation of claim settlement records may include having claim settlement records and having no claim settlement records. The types of insurance products may include insurance products that are easy to surrender and insurance products that are not easy to surrender. The ages of the policyholders or insured persons may include ages that are easy to surrender and ages that are not easy to surrender.
[0062] In an embodiment of the present invention, feature data is extracted from the above historical policy data containing surrender records. The feature data includes the attributes of the policy data (i.e., the attributes of the first historical policy data) and the attribute values of the attributes of the policy data (i.e., the attribute values of the attributes of the first historical policy data). Among them, the attributes of the policy data may include, but are not limited to, claim settlement records, the ages of the policyholders or insured persons, and the types of insurance products. When the attribute of the policy data is a claim settlement record, the attribute value of this attribute is having a claim settlement record and having no claim settlement record. When the attribute of the policy data is the age of the policyholder or insured person, the attribute value of this attribute is an age that is easy to surrender and an age that is not easy to surrender. When the attribute of the policy data is the type of insurance product, the attribute value of this attribute is an insurance product that is easy to surrender and an insurance product that is not easy to surrender.
[0063] In an embodiment of the present invention, the above historical policy data without surrender records may include, but is not limited to, the situation of claim settlement records, the types of insurance products, and the ages of the policyholders or insured persons. Among them, the situation of claim settlement records may include having claim settlement records and having no claim settlement records. The types of insurance products may include insurance products that are easy to surrender and insurance products that are not easy to surrender. The ages of the policyholders or insured persons may include ages that are easy to surrender and ages that are not easy to surrender.
[0064] In an embodiment of the present invention, feature data is extracted from the above-mentioned historical policy data without surrender records. The feature data includes the attributes of the policy data (i.e., the attributes of the second historical policy data) and the attribute values of the attributes of the policy data (i.e., the attribute values of the attributes of the second historical policy data). Among them, the attributes of the policy data may include, but are not limited to, claim records, the age of the policyholder or the insured, and the type of insurance product. When the attribute of the policy data is a claim record, the attribute value is "with claim record" or "without claim record". When the attribute of the policy data is the age of the policyholder or the insured, the attribute value is "easy-to-surrender age" or "not easy-to-surrender age". When the attribute of the policy data is the type of insurance product, the attribute value is "easy-to-surrender insurance product" or "not easy-to-surrender insurance product".
[0065] In an embodiment of the present invention, the above prediction model includes a decision tree. Each branch of the decision tree includes some attributes of the historical policy data, the attribute values of some attributes of the historical policy data, and the surrender probability and / or non-surrender probability. Specifically, reference can be made to Figure 4 . The surrender probability in the decision tree 200 can be determined based on the surrender records and branch information in the historical data.
[0066] In an embodiment of the present invention, the above current policy data may include, but are not limited to, the situation of claim records, the type of insurance product, and the age of the policyholder or the insured. Among them, the situation of claim records may include "with claim record" or "without claim record". The type of insurance product may include "easy-to-surrender insurance product" or "not easy-to-surrender insurance product". The age of the policyholder or the insured may include "easy-to-surrender age" or "not easy-to-surrender age".
[0067] In an embodiment of the present invention, feature data is extracted from the above current policy data. The feature data includes the attributes of the current policy data and the attribute values of the attributes of the current policy data. Among them, the attributes of the current policy data may include, but are not limited to, claim records, the age of the policyholder or the insured, and the type of insurance product. When the attribute of the current policy data is a claim record, the attribute value is "with claim record" or "without claim record". When the attribute of the current policy data is the age of the policyholder or the insured, the attribute value is "easy-to-surrender age" or "not easy-to-surrender age". When the attribute of the current policy data is the type of insurance product, the attribute value is "easy-to-surrender insurance product" or "not easy-to-surrender insurance product".
[0068] In an embodiment of the present invention, the above preset model can be trained based on historical policy data. After the training is completed, the preset model includes a decision tree for determining whether there is a surrender risk in the current policy data, such as Figure 4The decision tree 200 shown includes a total of four branches from left to right, namely claim record - yes - product type - easily surrenderable product - surrender probability 50%, claim record - yes - product type - not easily surrenderable product - surrender probability 20%, claim record - no - age - easily surrenderable age - surrender probability 80%, and claim record - no - age - not easily surrenderable age - surrender probability 30%. Among them, claim record, product type, and age in the decision tree 200 are attributes of historical policy data, and yes, no, easily surrenderable product, not easily surrenderable product, easily surrenderable age, and not easily surrenderable age in the decision tree 200 are attribute values of the attributes of historical policy data. The surrender probability of 50% means that if a customer's policy data contains a claim record and the insured product is an easily surrenderable product, the surrender probability of this customer is 20%. The surrender probability of 20% means that if a customer's policy data contains a claim record and the insured product is a not easily surrenderable product, the surrender probability of this customer is 20%. The surrender probability of 80% means that if a customer's policy data contains no claim record and the age of the policyholder or the insured is an easily surrenderable age, the surrender probability of this customer is 80%. The surrender probability of 30% means that if a customer's policy data contains no claim record and the age of the policyholder or the insured is a not easily surrenderable age, the surrender probability of this customer is 30%.
[0069] It should be noted that yes in the decision tree 200 means having a claim record, no means having no claim record, and product type means the type of insurance product.
[0070] Based on the foregoing solution, the surrender probability or non - surrender probability of the current policy data in step S260 can be determined through the above - mentioned decision tree 200. Among them, the non - surrender probability is 1 - the surrender probability in the decision tree 200. The non - surrender probability is not shown in the decision tree 200, and in this example, the non - surrender probability can be regarded as the probability that the customer continues to renew the policy. Based on this surrender probability, it can be predicted whether there is a surrender risk for this policy data.
[0071] It should be noted that the above - mentioned decision tree 200 only exemplarily shows some attributes and attribute values in the insurance industry. That is to say, as long as the factors that cause customers to surrender the policy in other insurance industries can be shown in the decision tree 200, and the specific attributes and attribute values shown in the decision tree 200 are determined according to the historical policy data for training the preset model.
[0072] Figure 3 The flowchart of the policy processing method according to another embodiment of the present invention is schematically shown.
[0073] As Figure 3 shown, the above - mentioned step S250 may specifically include step S310 and step S320.
[0074] In step S310, determine the attributes of the current policy data and the branches of the attribute values of the attributes in the decision tree according to the decision tree in the preset model.
[0075] In step S320, determine the surrender probability and / or non-surrender probability of the current policy data according to the information of the branch.
[0076] The method can determine the attributes of the current policy data and the branches of the attribute values of the attributes in the above-mentioned preset model according to the decision tree in the preset model, and then determine the surrender probability and / or non-surrender probability of the current policy data according to the information of the branch. In this way, the determined surrender probability and / or non-surrender probability are more accurate.
[0077] In an embodiment of the present invention, determining the surrender probability and / or non-surrender probability of the current policy data according to the information of the above branch includes: determining the surrender probability and / or non-surrender probability of the current policy data according to the attributes, attribute values, and surrender probability and / or non-surrender probability included in the branch.
[0078] Reference Figure 4 , the decision tree 200 includes a total of four branches from left to right. The information of each branch is claim record - yes - product type - easily surrenderable product - surrender probability 50%, claim record - yes - product type - not easily surrenderable product - surrender probability 20%, claim record - no - age - easily surrenderable age - surrender probability 80%, claim record - no - age - not easily surrenderable age - surrender probability 30%.
[0079] In this example, if the attributes of the obtained current policy data include claim record, type of insurance product, and age of the policyholder or insured, the decision tree 200 can determine that these attributes are distributed on the four branches of the decision tree 200. Further, if the attribute value of the claim record is yes and the type of insurance product is an easily surrenderable insurance product, the decision tree 200 can determine that the current policy data is distributed on the first branch from the left of the decision tree 200, that is, according to the surrender probability of this branch, the surrender probability of the current policy data can be determined to be 50% and / or the non-surrender probability is also 50% (not shown in the figure).
[0080] For another example, if the attribute value of the claim record is no and the age of the policyholder or insured is an easily surrenderable age, the decision tree 200 can determine that the current policy data is distributed on the third branch from the left of the decision tree 200, that is, according to the surrender probability of this branch, the surrender probability of the current policy data can be determined to be 80% and / or the non-surrender probability is also 20% (not shown in the figure).
[0081] In an embodiment of the present invention, the surrender risk can be divided into three levels according to the surrender probability. For example, a surrender probability of 1% - 40% is defined as a low surrender risk. A surrender probability of 40% - 70% is defined as a medium surrender risk. A surrender probability of 70% - 100% is defined as a high surrender risk. When it is determined through a preset model that the policy data is a high surrender risk, relevant personnel of the insurance company need to contact the customer of the policy data to retain the customer in a timely manner.
[0082] Figure 5 Schematically shows a flowchart of a policy processing method according to another embodiment of the present invention.
[0083] As Figure 5 shown, the above step 230 may include steps S510 to S530.
[0084] In step S510, the information gain of each attribute in the historical policy data is calculated by the gradient boosting decision tree algorithm according to the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, as well as the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data.
[0085] In step S520, the distribution of each attribute in the decision tree is determined according to the information gain of each attribute in the historical policy data.
[0086] In step S530, the gradient boosting classifier is trained based on the distribution of each attribute in the decision tree and the surrender records marked in the historical policy data.
[0087] This method can determine the distribution of each attribute in the decision tree according to the information gain of each attribute in the historical policy data. In this way, attributes with high discriminative power can be set at the root node of the decision tree, that is, in this way, the main factors affecting customer surrender can be easily found. Then, based on the distribution of each attribute in the decision tree and the surrender records marked in the historical policy data, the gradient boosting classifier is trained. In this way, when training the gradient boosting classifier to determine whether there is a surrender risk in the current policy data in the future, the determined surrender probability is more accurate.
[0088] In an embodiment of the present invention, the formula of the above gradient boosting decision tree algorithm is:
[0089]
[0090] Where S is the set of historical policy data, A is the attribute in the historical policy data, Entropy(S) is the information entropy of S, Value(A) is the set of attribute values of A, v is an attribute value in the set of attribute values of A, Sv is the set of historical data when the attribute value of A in S is v, and Entropy(S v ) is the information entropy of S v .
[0091] In one embodiment of the present invention, the information entropy of each attribute in the historical policy data can be determined through the above formula. The position of each attribute in the decision tree can be determined according to the information entropy of each attribute. For example, the historical policy data contains three attributes, namely claim record, product type, and age. Calculate the information entropy of the attributes of claim record, product type, and age respectively according to the above formula. The calculation result shows that the information entropy of the claim record is the highest (that is, the claim record is the most discriminative attribute), and the information entropies of the product type and age are similar. The attribute distribution situation that can be determined based on this calculation result is referred to Figure 4 , take the claim record as the root node, and take the product type and age as the child nodes of the claim record.
[0092] Figure 6 Schematically shows a flowchart of a policy processing method according to another embodiment of the present invention.
[0093] As Figure 6 shown, after the training of the gradient boosting classifier is completed, the above method further includes steps S610 to S660.
[0094] In step S610, obtain test data, where the test data includes historical policy data with surrender records and historical policy data without surrender records.
[0095] In step S620, obtain the surrender probability of the test data through the preset model.
[0096] In step S630, sort the surrender probabilities of the test data, and use the test data with surrender probabilities greater than or equal to the preset threshold as the data to be surrendered according to the sorting result.
[0097] In step S640, determine the surrender coverage rate of the test data according to the actual surrender quantity of the test data and the actual surrender quantity of the data to be surrendered.
[0098] In step S650, determine the surrender accuracy rate of the test data according to the actual surrender quantity of the data to be surrendered and the quantity of the data to be surrendered.
[0099] In step S660, evaluate the preset model according to the surrender coverage rate and the surrender accuracy rate of the test data.
[0100] This method can evaluate the above-mentioned preset model according to the surrender coverage rate of the test data and the surrender accuracy rate of the test data, which can further improve the accuracy of determining the surrender probability by the preset model subsequently.
[0101] In an embodiment of the present invention, the test data may include historical policy data of multiple customers. For example, the test data is the historical policy data of 1000 customers, among which, 300 customers have surrendered their policies, and 700 customers have not surrendered their policies. In this example, by processing the historical policy data of 1000 customers through the above-mentioned preset model, 1000 surrender probabilities can be obtained, and they are sorted in ascending order. The surrender probabilities greater than or equal to the preset threshold are regarded as customers who will surrender their policies. For example, the preset threshold is 70%, and it can be adjusted according to the actual situation specifically.
[0102] Based on the foregoing example, according to the sorting result, it is determined that 200 surrender probabilities are greater than or equal to the above-mentioned preset threshold, and the historical policy data of these 200 customers is used as the data of customers to be surrendered. Among them, 180 customers in the historical policy data of these 200 customers have actually surrendered their policies. According to the actual surrender quantity (300) of the test data and the actual surrender quantity (180) of the data of customers to be surrendered, the surrender coverage rate of the test data is determined to be 180 / 300 = 60%. According to the actual surrender quantity (180) of the data of customers to be surrendered and the quantity of customers to be surrendered (200) of the data of customers to be surrendered, the surrender accuracy rate of the test data is determined to be 180 / 200 = 90%. In this case, the preset model is evaluated according to the surrender coverage rate of the test data and the surrender accuracy rate of the test data. Generally, the greater the surrender coverage rate of the test data and the surrender accuracy rate of the test data, the more stable the evaluation result of the preset model. In this embodiment, the surrender coverage rate is relatively low, and the historical policy data can be continuously increased to retrain the gradient boosting classifier again, or the preset threshold can also be adjusted, and then the preset model is re-evaluated based on the test data.
[0103] Figure 7 A block diagram of a policy processing device according to an embodiment of the present invention is schematically shown.
[0104] As Figure 7 shown, the policy processing device 700 includes a first acquisition module 701, a first extraction module 702, a training module 703, a second acquisition module 704, a second extraction module 705, and a first determination module 706.
[0105] Specifically, the first acquisition module 701 is configured to acquire historical policy data, where the historical policy data includes first historical policy data and second historical policy data. The first historical policy data includes historical policy data with surrender records, and the second historical policy data includes historical policy data without surrender records.
[0106] The first extraction module 702 is configured to extract first feature data from the first historical policy data and second feature data from the second historical policy. The first feature data includes the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data. The second feature data includes the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data
[0107] The training module 703 trains a gradient boosting classifier based on the attributes of the first historical policy data, the attribute values of the attributes of the first historical policy data, the attributes of the second historical policy data, and the attribute values of the attributes of the second historical policy data to obtain the preset model
[0108] The second acquisition module 704 is configured to acquire the current policy data.
[0109] The second extraction module 705 is configured to extract feature data from the current policy data. The feature data includes the attributes of the policy data and the attribute values of the attributes.
[0110] The first determination module 706 is configured to determine the surrender probability and / or non-surrender probability of the current policy data according to the attributes of the current policy data and the attribute values of the attributes through the preset model.
[0111] The policy processing device 700 can train a gradient boosting classifier based on historical policy data with surrender records and historical policy data without surrender records to obtain the preset model. The surrender probability determined by the preset model trained in this way is more accurate when judging whether there is a surrender risk in the current policy data. For example, acquire the current policy data, extract feature data from the current policy data. The feature data includes the attributes of the current policy data and the attribute values of the attributes, and then determine the surrender probability and / or non-surrender probability of the current policy data according to the attributes of the current policy data and the attribute values of the attributes through the preset model. In this way, it can effectively predict whether there is a risk of surrender in the policy data, thereby avoiding customer surrender, reducing the losses of the insurance company, and retaining customer resources for the insurance company.
[0112] According to an embodiment of the present invention, the policy processing device 700 can be used to implement Figure 2 The policy processing method described in the embodiment.
[0113] Figure 8 Schematically shows a block diagram of a policy processing device according to another embodiment of the present invention.
[0114] Such as Figure 8As shown, the above-mentioned first determination module 706 includes a determination branch module 706-1 and a determination probability module 706-2.
[0115] Specifically, the determination branch module 706-1 is configured to determine the attributes of the current insurance policy data and the branches of the attribute values of the attributes in the decision tree according to the decision tree in the preset model.
[0116] The determination probability module 706-2 is configured to determine the surrender probability and / or non-surrender probability of the current insurance policy data according to the information of the branch.
[0117] The above-mentioned first determination module 706 can determine the attributes of the current insurance policy data and the branches of the attribute values of the attributes in the above-mentioned decision tree according to the decision tree in the above-mentioned preset model, and then determine the surrender probability and / or non-surrender probability of the current insurance policy data according to the information of the branch. In this way, the determined surrender probability and / or non-surrender probability are more accurate.
[0118] According to an embodiment of the present invention, the above-mentioned first determination module 706 can be used to implement Figure 3 the insurance policy processing method described in the embodiment.
[0119] Figure 9 Schematically shows a block diagram of an insurance policy processing device according to another embodiment of the present invention.
[0120] As Figure 9 shown, the above-mentioned training module 703 includes a calculation module 703-1, a second determination module 703-2, and a training sub-module 703-3.
[0121] Specifically, the calculation module 703-1 is configured to calculate the information gain of each attribute in the historical insurance policy data by using the gradient boosting decision tree algorithm according to the attributes of the first historical insurance policy data and the attribute values of the attributes of the first historical insurance policy data, and the attributes of the second historical insurance policy data and the attribute values of the attributes of the second historical insurance policy data.
[0122] The second determination module 703-2 is configured to determine the distribution of each attribute in the decision tree according to the information gain of each attribute in the historical insurance policy data.
[0123] The training sub-module 703-3 trains the gradient boosting classifier based on the distribution of each attribute in the decision tree and the surrender records marked in the historical insurance policy data.
[0124] The training module 703 can determine the distribution of each attribute in the decision tree according to the information gain of each attribute in the historical policy data. In this way, attributes with high discriminative power can be set at the root node of the decision tree, that is, in this way, the main factors affecting customers' surrendering of policies can be easily found. Then, based on the distribution of each attribute in the decision tree and the surrender records marked in the historical policy data, the gradient boosting classifier is trained. In this way, when the gradient boosting classifier determines whether there is a surrender risk in the current policy data in the subsequent judgment, the determined surrender probability is more accurate.
[0125] According to an embodiment of the present invention, the training module 703 can be used to implement Figure 5 the policy processing method described in the embodiment.
[0126] Figure 10 The block diagram of a policy processing device according to another embodiment of the present invention is schematically shown.
[0127] As Figure 10 shown, the above-mentioned policy processing device 700 includes a third acquisition module 707, a fourth acquisition module 708, a sorting module 709, a third determination module 710, a fourth determination module 711, and an evaluation module 712.
[0128] Specifically, the third acquisition module 707 is configured to acquire test data, where the test data includes historical policy data with surrender records and historical policy data without surrender records.
[0129] The fourth acquisition module 708 is configured to obtain the surrender probability of the test data through the preset model.
[0130] The sorting module 709 is configured to sort the surrender probabilities of the test data, and use the test data with surrender probabilities greater than or equal to a preset threshold as the data to be surrendered according to the sorting result.
[0131] The third determination module 710 is configured to determine the surrender coverage rate of the test data according to the actual surrender quantity of the test data and the actual surrender quantity of the data to be surrendered.
[0132] The fourth determination module 711 is configured to determine the surrender accuracy rate of the test data according to the actual surrender quantity of the data to be surrendered and the quantity of the data to be surrendered.
[0133] The evaluation module 712 is configured to evaluate the preset model according to the surrender coverage rate of the test data and the surrender accuracy rate of the test data.
[0134] The policy processing device 700 can evaluate the above preset model according to the surrender coverage rate of the test data and the surrender accuracy rate of the test data, which can further improve the accuracy of determining the surrender probability through the preset model subsequently.
[0135] According to an embodiment of the present invention, the policy processing device 700 can be used to implement Figure 6 the policy processing method described in the embodiment.
[0136] Since each module of the policy processing device in the exemplary embodiment of the present invention can be used to implement the steps of the exemplary embodiment of the above-described 2- Figure 6 described policy processing method, for details not disclosed in the device embodiment of the present invention, please refer to the embodiment of the above-described policy processing method of the present invention.
[0137] It can be understood that the first acquisition module 701, the first extraction module 702, the training module 703, the calculation module 703-1, the second determination module 703-2, the training sub-module 703-3, the second acquisition module 704, the second extraction module 705, the first determination module 706, the determination branch module 706-1, the determination probability module 706-2, the third acquisition module 707, the fourth acquisition module 708, the sorting module 709, the third determination module 710, the fourth determination module 711, and the evaluation module 712 can be implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the first acquisition module 701, the first extraction module 702, the training module 703, the calculation module 703-1, the second determination module 703-2, the training sub-module 703-3, the second acquisition module 704, the second extraction module 705, the first determination module 706, the determination branch module 706-1, the determination probability module 706-2, the third acquisition module 707, the fourth acquisition module 708, the sorting module 709, the third determination module 710, the fourth determination module 711, and the evaluation module 712 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented in any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or implemented in an appropriate combination of software, hardware, and firmware. Or, at least one of the first acquisition module 701, the first extraction module 702, the training module 703, the calculation module 703-1, the second determination module 703-2, the training sub-module 703-3, the second acquisition module 704, the second extraction module 705, the first determination module 706, the determination branch module 706-1, the determination probability module 706-2, the third acquisition module 707, the fourth acquisition module 708, the sorting module 709, the third determination module 710, the fourth determination module 711, and the evaluation module 712 can be at least partially implemented as a computer program module, and when the program is run on a computer, it can execute the functions of the corresponding module.
[0138] Reference is made below to Figure 11 , which shows a schematic structural diagram of a computer system 800 of an electronic device suitable for implementing the embodiments of the present invention. Figure 11 The computer system 800 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0139] As Figure 11As shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage section 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for system operation are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0140] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0141] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, the above functions defined in the system of the present application are executed.
[0142] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0144] The units involved in the embodiments of the present invention can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.
[0145] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device is enabled to implement the policy processing method as described in the above embodiments.
[0146] For example, the electronic device may implement as Figure 2 shown in: In step S210, historical policy data is obtained, and the historical policy data includes first historical policy data and second historical policy data. The first historical policy data includes historical policy data containing surrender records, and the second historical policy data includes historical policy data without surrender records. In step S220, first feature data is extracted from the first historical policy data, and second feature data is extracted from the second historical policy. The first feature data includes the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, and the second feature data includes the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data. In step S230, a gradient boosting classifier is trained based on the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data and the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data to obtain the preset model. In step S240, current policy data is obtained. In step S250, feature data is extracted from the current policy data, and the feature data includes the attributes of the current policy data and the attribute values of the attributes. In step S240, the surrender probability and / or non-surrender probability of the current policy data is determined by the preset model according to the attributes of the current policy data and the attribute values of the attributes.
[0147] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.
[0148] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present invention.
[0149] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present invention. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.
[0150] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A policy processing method, characterized in that, it includes: Obtain historical policy data, where the historical policy data includes first historical policy data and second historical policy data. The first historical policy data includes historical policy data containing surrender records, and the second historical policy data includes historical policy data without surrender records; Extract first feature data from the first historical policy data, and extract second feature data from the second historical policy. The first feature data includes the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, and the second feature data includes the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data; Train a gradient boosting classifier based on the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, and the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data to obtain a preset model; Obtain current policy data; Extract feature data from the current policy data, and the feature data includes the attributes of the current policy data and the attribute values of the attributes; Determine the surrender probability and / or non-surrender probability of the current policy data through the preset model according to the attributes of the current policy data and the attribute values of the attributes; Among them, training the gradient boosting classifier based on the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, and the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data includes: calculating the information gain of each attribute in the historical policy data through the gradient boosting decision tree algorithm according to the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, and the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data; determining the distribution of each attribute in the decision tree according to the information gain of each attribute in the historical policy data; training the gradient boosting classifier based on the distribution of each attribute in the decision tree and the surrender records marked in the historical policy data.
2. The method according to claim 1, characterized in that, The preset model contains a decision tree, and each branch of the decision tree contains partial attributes of the historical policy data, attribute values of partial attributes of the historical policy data, and surrender probability and / or non-surrender probability.
3. The method according to claim 1, characterized in that, The formula of the gradient boosting decision tree algorithm is: ; Among them, S is the set of historical policy data, A is the attribute in the historical policy data, Entropy(S) is the information entropy of S, Value(A) is the set of attribute values of A, v is an attribute value in the set of attribute values of A, and S v is the set of historical data when the attribute value of A in S is v, Entropy(S v ) is the information entropy of S v .
4. The method according to claim 1, characterized in that, The method further includes: Obtain test data, where the test data includes historical policy data containing surrender records and historical policy data without surrender records; Obtain the surrender probability of the test data through the preset model; Sort the surrender probabilities of the test data, and use the test data with surrender probabilities greater than or equal to a preset threshold as the data to be surrendered according to the sorting result; Determine the surrender coverage rate of the test data based on the actual surrender quantity of the test data and the actual surrender quantity of the data to be surrendered; Determine the surrender accuracy rate of the test data based on the actual surrender quantity of the data to be surrendered and the quantity of the data to be surrendered; and Evaluate the preset model based on the surrender coverage rate and the surrender accuracy rate of the test data.
5. The method according to claim 1, wherein, Determining the surrender probability and / or non-surrender probability of the current policy data by a preset model according to the attributes of the current policy data and the attribute values of the attributes includes: Determine the attributes of the current policy data and the attribute values of the attributes in the branches of the decision tree in the preset model; Determine the surrender probability and / or non-surrender probability of the current policy data according to the information of the branches.
6. The method according to claim 5, wherein, Determining the surrender probability and / or non-surrender probability of the current policy data according to the information of the branches includes: Determine the surrender probability and / or non-surrender probability of the current policy data according to the attributes, attribute values, and surrender probability and / or non-surrender probability included in the branches.
7. A policy processing device, wherein, comprising: A first acquisition module for acquiring historical policy data, the historical policy data including first historical policy data and second historical policy data, the first historical policy data including historical policy data with surrender records, and the second historical policy data including historical policy data without surrender records; A first extraction module for extracting first feature data from the first historical policy data and second feature data from the second historical policy, the first feature data including the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, and the second feature data including the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data; A training module for training a gradient boosting classifier based on the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data and the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data to obtain a preset model; A second acquisition module for acquiring current policy data; A second extraction module for extracting feature data from the current policy data, the feature data including the attributes of the current policy data and the attribute values of the attributes; A first determination module for determining the surrender probability and / or non-surrender probability of the current policy data by a preset model according to the attributes of the current policy data and the attribute values of the attributes; Among them, the above training module includes: a calculation module, a second determination module, and a training sub-module. The calculation module is used to calculate the information gain of each attribute in the historical policy data through the gradient boosting decision tree algorithm according to the attributes of the first historical policy data and the attribute values of the attributes of the first historical policy data, as well as the attributes of the second historical policy data and the attribute values of the attributes of the second historical policy data. The second determination module is used to determine the distribution of each attribute in the decision tree according to the information gain of each attribute in the historical policy data. The training sub-module is used to train the gradient boosting classifier based on the distribution of each attribute in the decision tree and the surrender records marked in the historical policy data.
8. An electronic device, comprising: one or more processors; and a storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein, the program when executed by a processor implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for client service
CN106875225A
Collaborative backup method, system, computer device and storage medium for product data
CN109272410A