Underwriting risk assessment method and system

By applying machine learning models in insurance underwriting, collecting insurance data, identifying risk factors and encoding, the problems of slow speed, strong subjectivity and difficulty in adapting to complex actual conditions in traditional underwriting process are solved, and the accurate assessment and interpretability of underwriting risks are achieved.

CN120219088AInactive Publication Date: 2025-06-27AIA LIFE INSURANCE CO LTD

Patent Information

Application Number
CN202510288160.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the traditional insurance underwriting process, there are problems such as slow review speed, strong subjectivity, high cost and difficulty in adapting to complex actual situations, especially when dealing with complex ICD disease codes, the risk cannot be accurately evaluated.

Method used

The machine learning model is adopted to collect insurance data, identify the risk factors affecting underwriting, encode discrete factors, and use multiple sub-models to score in different risk scenarios, and finally obtain the total risk score through weighted calculations.

Benefits of technology

It realizes an accurate assessment of underwriting risks, is interpretable, can explain decision-making basis to users, improves customer satisfaction and trust, and adapts to complex discrete factors such as ICD disease codes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219088A_ABST
    Figure CN120219088A_ABST
Patent Text Reader

Abstract

The invention relates to the field of risk management and control of insurance underwriting business, and discloses an underwriting risk assessment method and system, and the method comprises the steps: collecting insurance data of a target customer, and recognizing risk factors affecting underwriting in the insurance data; the risk factors comprise discrete factors; performing feature conversion on the discrete factors according to a preset factor splitting election or factor splitting expansion method, and encoding the discrete factors into a numerical form; selecting key features suitable for different risk scenes from the risk factors, and inputting the key features into a plurality of trained sub-models to obtain scores of the target customer in different risk scenes and interpretable results corresponding to the scores; the plurality of sub-models comprise a BMI intelligent identification sub-model, a claim settlement history risk control sub-model, a disease risk level classification sub-model and a delayed risk re-mining sub-model; and performing weighted calculation on the scores obtained by the sub-models, and finally obtaining the total risk score of the target customer in the current business.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of risk management and control of insurance underwriting business, and specifically, to an underwriting risk assessment method and system using a machine learning model. Background Art

[0002] Underwriting, also known as risk selection or risk assessment, is the process by which insurance companies review, screen, and classify insurance products based on different risk levels to determine whether to provide insurance and the conditions for providing insurance. In the traditional insurance underwriting process, insurance companies rely on personal judgment or rule engines to evaluate customers' insurance applications. Although manual underwriting has certain flexibility and inclusiveness when dealing with complex situations, it also has disadvantages such as slow review speed, strong subjectivity, and high cost. Although the underwriting system based on fixed rules can help identify risks to a certain extent, it also has obvious limitations: existing underwriting rules are often static and difficult to adapt to complex and changing actual conditions. For example, when faced with complex ICD disease codes, it is impossible to accurately assess risks.

[0003] Therefore, designing an underwriting risk assessment method that can use machine learning models to make it flexible enough to handle complex discrete factors and have interpretability to explain its decision-making basis to users has become a trend in the development of underwriting models. Summary of the invention

[0004] In order to solve the above technical problems, this application provides an underwriting risk assessment method and system, which customizes and builds machine learning models for different underwriting scenarios. The model can not only obtain accurate risk assessment results, but also clearly understand the logic and basis behind the risk assessment to assist in making underwriting decisions. Specifically, the technical solution of this application is as follows:

[0005] In a first aspect, the present application discloses an underwriting risk assessment method, comprising the following steps:

[0006] Collecting insurance data of target customers and identifying risk factors in the insurance data that affect underwriting; the risk factors include discrete factors;

[0007] Encoding the discrete factors according to a preset factor splitting selection or factor splitting expansion method;

[0008] Select key features suitable for different risk scenarios from the risk factors and input them into the trained multiple sub-models to obtain the scores of target customers in different risk scenarios and the interpretable results corresponding to the scores; the risk scenarios include: physical abnormality scenarios, claims history scenarios, disease history scenarios, and delay and rejection history scenarios; the multiple sub-models include: BMI intelligent identification sub-model, claims history risk control sub-model, disease risk level classification sub-model, and delay and rejection risk re-mining sub-model;

[0009] Perform weighted calculation on the scores obtained by each of the sub-models, and finally obtain the total risk score of the target customer in the current business.

[0010] In some embodiments, in the step of selecting key features suitable for different risk scenarios from the risk factors and inputting them into multiple trained sub-models to obtain the scores of the target customer in different risk scenarios and the interpretable results corresponding to the scores, the following steps are included:

[0011] Use a preset rule to determine whether the current business of the target customer hits the risk scenario;

[0012] If one or more of the risk scenarios are hit, input the key features corresponding to the risk scenario into the machine learning model corresponding to the risk scenario; the machine learning model is an XGBoost model;

[0013] Output the prediction result scores of the key features on each leaf node of the machine learning model and the prediction process trajectory; generate the interpretable results of all input features in the machine learning model based on the prediction process trajectory;

[0014] Sum the scores of the key features on each of the leaf nodes, and then calculate through an activation function to obtain the score of the current business of the target customer in this risk scenario.

[0015] In some embodiments, in the step of performing weighted calculation on the scores obtained by each of the sub-models and finally obtaining the risk score of the target customer in the current business, the following steps are included:

[0016] Perform normalization processing on the score results of each sub-model;

[0017] According to the importance of each risk scenario and the performance of historical data, assign corresponding weights to the scores of each sub-model;

[0018] Multiply the normalized scores by their respective weights, and then perform weighted summation to obtain the total risk score of the current business of the target customer.

[0019] In some embodiments, in the step of encoding the discrete factors according to a preset factor splitting election or factor splitting expansion method, the following steps are included:

[0020] Before processing the insurance data of the target customer, obtain a large number of discrete factor samples;

[0021] Perform factor splitting and election for each of the discrete factors in the discrete factor sample: Screen for high-risk factors among them and sort them. Perform feature transformation on the high-risk factors according to the sorted list to obtain a feature transformation list. When processing the insurance application data of the target customer, substitute the target discrete factor into the feature transformation list; obtain the transformed feature and then encode it into a numerical form;

[0022] Alternatively, perform factor splitting and expansion for each of the discrete factors in the discrete factor sample: Extract unique values for each of the discrete factors and create dummy variables; when processing the insurance application data of the target customer, look up the dummy variable corresponding to the target discrete factor;

[0023] The discrete factors include disease classification codes (ICD).

[0024] In some embodiments, the performing factor splitting and election for each of the discrete factors in the discrete factor sample; specifically includes the following steps:

[0025] Split the values of the discrete factor sample into a list;

[0026] Set retention parameters, including the maximum retention number and the minimum proportion; to allocate the model training set, test set, and validation set;

[0027] Count the number of each of the discrete factors after splitting, and calculate the proportion of each of the discrete factors in the positive sample or negative sample;

[0028] Perform filtering and retention processing on each of the discrete factors to screen out high-risk factors: Filter out the discrete factors with a proportion less than the minimum proportion; perform a first sorting on the remaining discrete factors according to the total quantity size, and retain a number of the discrete factors with the top order in the list according to the first sorting as the high-risk factors, and the retained quantity is the maximum retention number;

[0029] Perform a second sorting on the retained high-risk factors according to the proportion in the positive sample, and perform feature transformation on the original values of the high-risk factors according to the list of the second sorting.

[0030] In other embodiments, the performing factor splitting and expansion for each of the discrete factors in the discrete factor sample; specifically includes the following steps:

[0031] Split the values of the discrete factor sample into a list;

[0032] Identify the unique feature values of each of the discrete factors; for each of the discrete factors to obtain the unique feature values, generate a corresponding dummy variable.

[0033] In some embodiments, the underwriting risk assessment method further includes: constructing and training corresponding sub-models for different risk scenarios; specifically including the following steps:

[0034] Collect customer information from multiple channels, preprocess the customer information, and extract the general features and scenario features that affect risk assessment therein as the training set;

[0035] Divide the risk scenarios applicable to the sub-models, and independently establish machine learning models for each of the risk scenarios;

[0036] The machine learning model is a tree structure. Each input feature is used as a leaf node of the machine learning model, and the XGBoost algorithm is used to train the machine learning model to understand the degree of dependence of the machine learning model on different input features during the decision-making process;

[0037] Add each of the trained machine learning models to the underwriting risk assessment system in the form of sub-models.

[0038] Optionally, the underwriting risk assessment method further includes the following steps:

[0039] Validate and test each of the trained machine learning models using the validation set and the test set respectively; based on the comprehensive results of validation and testing, optimize the model to avoid overfitting and underfitting of the model;

[0040] Obtain the analysis and feedback of relevant technical personnel on special cases to dynamically update the selected input features and model parameters, and further optimize the model output results.

[0041] In some embodiments, the underwriting risk assessment method further includes the following steps: calculate the distribution difference between the model training data set and the actual application data set for the input features of each sub-model; for each sub-model, calculate its stability, performance, discrimination degree for positive and negative samples, and sorting index to monitor the stability of the model and the changing trend of the model effect.

[0042] In a second aspect, the present application also discloses an underwriting risk assessment system for implementing the underwriting risk assessment method described in any one of the above embodiments, specifically including:

[0043] A data collection module for collecting the insurance application data of target customers;

[0044] A risk identification module for identifying the risk factors affecting underwriting in the insurance application data; the risk factors include discrete factors;

[0045] The risk identification module is also used to encode the discrete factors according to a preset factor splitting and election or factor splitting and expansion method;

[0046] The model processing module is used to select key features suitable for different risk scenarios from the risk factors and input them into multiple trained sub-models to obtain the scores of the target customer in different risk scenarios and the interpretable results corresponding to the scores; the risk scenarios include: physical abnormality scenario, claim history scenario, disease history scenario, and rejection history scenario; the multiple sub-models include: BMI intelligent identification sub-model, claim history risk control sub-model, disease risk level classification sub-model, and rejection risk re-mining sub-model;

[0047] The risk aggregation module is used to perform weighted calculation on the scores obtained by each sub-model to finally obtain the total risk score of the target customer in the current business.

[0048] Compared with the prior art, the present application has at least one of the following beneficial effects:

[0049] 1. The present application realizes the extraction of underwriting ideas by means of machine learning. For different underwriting scenarios, including physical abnormality scenario, claim history scenario, disease history scenario, and rejection history scenario, machine learning models are customized and built. The models can obtain accurate risk assessment results. The present application selects case types that can be modeled, extracts audit professional experience and ideas. Learn the underwriting ideas of historical samples, find the common features of the samples, and predict the risk level of customers. In cases where manual review determines that they can pass, the model should be able to make a decision result that is highly consistent with it with high probability.

[0050] 2. The model of the present application is interpretable. The conclusion of each underwriting case can be quantified and explained. Continuously optimize the training to ensure the overall accuracy and stability of the model. Users can not only obtain accurate risk assessment results, but also clearly understand the logic and basis behind the risk assessment. This improvement in transparency greatly improves customer satisfaction and trust, making it easier for customers to accept underwriting conclusions and reducing disputes caused by opaque decision-making processes. By introducing a highly interpretable machine learning model, underwriters can quickly make more informed choices based on the detailed explanations provided by the model, reducing the time cost of manual review.

[0051] 3. In this application, risk factors are selected according to the characteristics of the model and continuously improved, and the risk factors are used for algorithm modeling to output underwriting prediction conclusions. Especially for discrete factors including disease classification codes (ICD), etc., this application proposes model algorithm prediction, factor splitting election, and factor splitting expansion. This intelligent factor processing method ensures that the model can make full use of the information in complex discrete features, significantly improving the model's ability to process complex discrete features such as ICD disease codes, thus bringing a significant improvement in the modeling effect. Specifically, it shows higher prediction accuracy, better stability, and stronger business interpretability.

[0052] 4. The model of this application supports dynamic optimization. This application not only customizes and constructs multiple machine learning sub-models for different underwriting scenarios, but also pays special attention to the flexibility and adaptability of the system. By dynamically adjusting the model weights and thresholds, and regularly reviewing and updating the model parameters, this application can adapt to the changing market environment and customer needs, ensuring the continuous effectiveness and stability of the model. After the model is put into production and application, it still continuously feeds back model training based on the changes in intelligent questionnaires, underwriting rules, reinsurance assessment criteria, and manual review scales, and dynamically optimizes the model output results. At the same time, the model can be continuously updated over time to adapt to new data trends, maintaining a high level of interpretability and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The above features, technical features, advantages and their implementation manners of this application will be further described below in a clear and understandable manner in combination with the drawings in the preferred embodiments.

[0054] Figure 1 It is a flowchart of the steps of an embodiment of a method for underwriting risk assessment of this application;

[0055] Figure 2 It is a schematic diagram of the first tree structure of the xgboost algorithm in an embodiment of this application;

[0056] Figure 3 It is a schematic diagram of the second tree structure of the xgboost algorithm in an embodiment of this application;

[0057] Figure 4 It is a block diagram of the structure of an embodiment of an underwriting risk assessment system of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] In the following description, specific details such as specific system architectures and technologies are presented for purposes of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.

[0059] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0060] To simplify the drawings, only the parts related to the invention are schematically shown in each figure, and they do not represent the actual structure of the product. Additionally, to simplify the drawings for easier understanding, for components with the same structure or function in some figures, only one of them is schematically illustrated, or only one of them is labeled. In this document, "one" not only means "only this one" but also can mean "more than one" situation.

[0061] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will describe the specific embodiments of the present application with reference to the accompanying drawings. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, and other embodiments can be obtained.

[0062] Insurance underwriting refers to the process in which an insurance company, based on a comprehensive understanding and verification of the information of the insured subject matter, evaluates and classifies insurable risks, and then decides whether to underwrite and under what conditions. Underwriting is the core of the insurance company's underwriting link. Through underwriting, risks that are not insurable can be prevented from being brought in, and unqualified insurance subjects can be excluded. The main purpose of underwriting is to identify the degree of risk of the insurance subject matter, classify the insurance subject matter accordingly, underwrite and set rates according to different standards, so as to ensure the quality of the underwritten business. The quality of underwriting work is directly related to whether the insurance contract can be smoothly performed, the underwriting profit and loss of the insurance company, and its financial stability. Therefore, strictly standardizing underwriting work is the key to reducing the loss ratio and increasing the profit of the insurance company, and it is also an important indicator to measure the level of the insurance company's operation and management.

[0063] There are mainly the following two existing insurance underwriting methods. The first is manual underwriting, which has always been the traditional underwriting method. However, with the development of technology and the trend of intelligentization, some disadvantages of manual underwriting have gradually emerged. For example, manual underwriting requires professional underwriters to review the information and situations of customers one by one, and the entire review process is relatively cumbersome and time-consuming. When dealing with a large number of insurance application materials manually, the speed is limited and highly dependent on the professional knowledge and experience of underwriters. This not only prolongs the underwriting cycle and affects the customer experience, but also there may be subjective biases in the interpretation and judgment of the same information by different underwriters, especially in some cases with fuzzy boundaries, such as the judgment of early symptoms of certain diseases or critical indicators, which may also lead to inconsistent underwriting results.

[0064] The second is the automatic underwriting method relying on a rule engine. Although this system based on fixed rules can help identify risks to a certain extent, it also has obvious limitations: the rules are often static and difficult to adapt to complex and changing actual situations. For example, newly emerging disease types or special occupational risks may not be included in the preset rules, resulting in the system being unable to accurately assess risks. In addition, the system cannot automatically learn and improve from historical underwriting cases. Even when encountering a large number of similar insurance application situations and underwriting result feedback, the system cannot automatically optimize its underwriting rules and algorithms, and long-term operation may lead to more and more misjudgments.

[0065] With the development of big data and artificial intelligence, some technical personnel have begun to introduce machine learning models in order to overcome the above problems. However, ordinary machine learning models often lack transparency and are difficult to explain their decision-making basis to users, which is a major challenge for the insurance industry that requires a high degree of trust and compliance with regulations. The model may make judgment errors due to the deviation of some historical data, thus affecting users' confidence. In addition, these models usually do not have enough flexibility to handle complex discrete factors, especially enumeration value splicing features such as ICD disease codes (International Classification of Diseases). Existing coding methods, such as One-Hot Encoding, label-encoder, WOE (Weight of Evidence) transformation, etc., cannot fully explore the internal information associations of such features, thus limiting the expressiveness of the model.

[0066] Therefore, from the perspective of business operations, relevant technical personnel urgently expect to build an intelligent underwriting risk assessment method and system that is both efficient and highly interpretable, enabling it to have sufficient flexibility to handle complex discrete factors and having the characteristic of interpretability to explain its decision-making basis to users. This system can not only simulate the gestures of manual decision-making, that is, in cases where manual review determines that it can pass, it has a high probability of making a decision result consistent with it, but also give a clear reason for why such a conclusion is made.

[0067] By combining advanced machine learning techniques and novel data preprocessing means, this application is committed to providing a more objective, accurate and efficient insurance underwriting risk assessment framework, with particular emphasis on the interpretability and adaptability of the model, thereby enhancing customer trust and supporting regulatory compliance. At the same time, this application also emphasizes the flexibility of the system to ensure that it can cope with the changing market environment and customer needs, providing scientific and reliable decision-making support for the risk management of insurance companies.

[0068] Referring to the appended Figure 1 description, an embodiment of an underwriting risk assessment method provided by this application includes the following steps:

[0069] S100, collect the insurance application data of the target customer. Specifically, the source of the insurance application data is mainly the insurance application form filled out by the applicant. The insurance application form is the first-hand information for underwriting and the most original insurance record. The insurer can obtain information from the filled-in items of the insurance application form to select risks. Secondly, for the situation of the insured object and the insured person that cannot be reflected on the insurance application form, it can also be further understood from the salesperson or the applicant. Finally, in addition to reviewing the insurance application form and directly understanding the situation from the salesperson and the applicant, the insurer can also conduct an actual survey of the risk situation faced by the insured object and the insured person, which is called underwriting survey.

[0070] S200. Identify the risk factors in the insurance application data that affect underwriting. The risk factors include discrete factors. Specifically, combining business experience and data-driven approaches, identify the key features that affect risk. In the specific business scenario of underwriting, for each discrete factor, especially the factors derived from ICD disease codes that are strongly related to the underwriting scenario, discrete factors usually need to be encoded so that the model can understand and process them. Common encoding methods include label-encoder and One-Hot Encoding. After using these methods, the feature dimension increases significantly, which leads to an increase in the complexity of the model. The increased complexity not only increases the computational cost and storage cost, but may also lead to an extension of the model training time, thus affecting the performance of the model. In a high-dimensional space, the distance between data points becomes unclear, which makes it difficult for the model to learn effective feature representations and results in feature loss. Therefore, how to select and innovatively design appropriate encoding methods is a difficult point in feature processing.

[0071] S300. Encode the discrete factors according to a preset factor splitting election or factor splitting expansion method.

[0072] Specifically, factor splitting election aims to screen out the ICD codes with the highest risk through splitting, setting retention parameters, statistical calculation, filtering, sorting, and feature transformation and encoding. Specifically referring to an implementation manner of this embodiment, in S310, before processing the insurance application data of the target customer, obtain a large number of discrete factor samples. Perform factor splitting election on each discrete factor in the discrete factor samples: screen out the high-risk factors and sort them, and perform feature transformation on the high-risk factors according to the sorted list to obtain a feature transformation list. When processing the insurance application data of the target customer, substitute the target discrete factor into the feature transformation list. Obtain the transformed feature, and then encode it into a numerical form.

[0073] Factor splitting expansion aims to extract unique values and create dummy variables, and equally process each ICD code to enhance the expressiveness and discrimination of the data. Specifically referring to another implementation manner of this embodiment, in S320, before processing the insurance application data of the target customer, obtain a large number of discrete factor samples. Perform factor splitting expansion on each discrete factor in the discrete factor samples: extract unique values for each discrete factor and create dummy variables. When processing the insurance application data of the target customer, search for the dummy variable corresponding to the target discrete factor.

[0074] S400. Select key features suitable for different risk scenarios from the risk factors and input them into multiple trained sub-models to obtain the scores of the target customer in different risk scenarios and the interpretable results corresponding to the scores.

[0075] Preferably, the risk scenarios include: physical abnormality scenarios, claim history scenarios, disease history scenarios, and rejection history scenarios. The multiple sub-models include: a BMI intelligent recognition sub-model, a claim history risk control sub-model, a disease risk level classification sub-model, and a rejection risk re-mining sub-model.

[0076] Specifically, the key features included in the model generally include general features and scenario features. The general features are general features developed based on the entire underwriting scenario, such as demographic features like age, gender, and residential city. The scenario features are different relevant features developed based on each specific underwriting scenario, including scenarios such as abnormal body mass index (BMI), claim history, disease history, and rejection history.

[0077] S500, calculate the weighted scores obtained by each of the sub-models, and finally obtain the total risk score of the target customer in the current business. Specifically, after the sub-model scores in each risk scenario (excessive body mass index (BMI), claim history, disease history, rejection history) are completed, we comprehensively process the score results of each sub-model to generate the overall risk score of the customer.

[0078] In some embodiments, for the scenario of abnormal body mass index (BMI), the model input features include: age, whether the height and weight of minors (<16 years old) are normal, insurance type, benefit type, body mass index (BMI) value, premium, insurance amount, occupation, branch company, and the business management office of the insurance agent.

[0079] The model is trained using a historical data set containing the above features, and the target Y variable is set to the insurance products with excessive body mass index (BMI) in history but judged to be of low risk manually. Train the machine learning model to identify the relationship between excessive body mass index (BMI) and underwriting risk and generate a risk score.

[0080] In some embodiments, for the claim history scenario, the model input features include: age, whether the height and weight of minors (<16 years old) are normal, the number of days hospitalized in the last 2 years, insurance amount, premium, the number of high-risk disease codes violated in previous claims, the number of times the main diagnosis type was chronic in the last 5 years, occupation, total number of days of hospitalization, total invoice amount, the city where the last claim hospital is located, the number of months from the last discharge date to the purchase of the insurance policy, the number of accidents in the last 5 years, cumulative risk insurance amount of accident insurance, medical insurance amount enjoyed, accident codes in history, the number of times the main diagnosis type was a major disease in the last 5 years, cumulative risk insurance amount of driving and riding accidents, endowment insurance amount enjoyed, cumulative risk insurance amount of critical illness insurance, number of surgeries, surgery codes, number of critical illness insurance policies enjoyed, and accident and medical insurance amount enjoyed, etc.

[0081] Model training uses the historical dataset of the above-mentioned historical claim-related features of customers, sets the target Y variable as the insured products that have a claim history but are judged to be of low risk manually. Train a machine learning model to identify the relationship between the claim history scenario and underwriting risk, and generate a risk score.

[0082] In some other embodiments, for the disease history scenario. Model input features: type of benefit in the insurance questionnaire, insured amount, medical insurance amount enjoyed, premium, shortest medical history time from this insurance application, disease code corresponding to the maximum remaining time of the control period, whether there is a history of declination, number of high-risk disease codes violated by previous diseases, age, longest control period of the disease, cumulative risk insured amount of critical illness insurance, number of diseases corresponding to the maximum remaining time of the control period, whether there is a history of deferral, number of months from the most recent insurance application to this application month, life insurance amount enjoyed, accident and medical insurance amount enjoyed, cumulative risk insured amount of life insurance (physical examination), body mass index (BMI), number of disease insurance policies enjoyed, cumulative risk insured amount of critical illness insurance, occupation, marital status, accident insurance amount enjoyed, whether there is an exclusion history, critical illness insurance amount enjoyed, annual income, diseases in the past two years, cumulative risk insured amount of accident insurance, number of life insurance policies enjoyed, etc.

[0083] Model training: Use the historical dataset of the above-mentioned disease history-related features of customers, set the target Y variable as the insured products that have a disease history but are judged to be of low risk manually. Train a machine learning model to identify the relationship between the disease history scenario and underwriting risk, and generate a risk score.

[0084] In some other embodiments, for the history of deferral and declination scenario, since the historical samples of the deferral and declination scenario are relatively few and the features are clear, the rule-based method is used to judge the relationship between the history of deferral and declination and underwriting risk. The model judgment rules include: 1. The current policy is not a long-term care insurance. 2. The current policy has no other risk scenarios and only has a history of deferral and declination. 3. The current policy has a standard body underwriting record with full underwriting after the most recent deferral and declination.

[0085] Another embodiment of the underwriting risk assessment method of the present application, on the basis of an embodiment of the above method, step S400 shown: Select key features suitable for different risk scenarios from the risk factors and input them into multiple trained sub-models to obtain the scores of the target customer in different risk scenarios and the interpretable results corresponding to the scores. Specifically, it includes the following sub-steps:

[0086] S410, use a preset rule to judge whether the current business of the target customer hits the risk scenario.

[0087] S420, if one or more of the risk scenarios are hit, input the key features corresponding to the risk scenario into the machine learning model corresponding to the risk scenario. The machine learning model is an XGBoost model.

[0088] S430 outputs the prediction result scores of the key features at each leaf node of the machine learning model and the prediction process trajectory. An interpretability result of all the input features in the machine learning model is generated based on the prediction process trajectory.

[0089] S440 sums the scores of the key features at each of the leaf nodes, and then calculates the score of the target customer's current business in this risk scenario through an activation function.

[0090] The intelligent underwriting system designed in this application covers multiple underwriting scenarios, independently models each underwriting scenario, and adds the trained models to the intelligent underwriting system in the form of sub-models.

[0091] Specifically, rules are used to determine whether the target customer's current business hits different rules. If it hits, it will lead to different paths. For example, the rule for the physical health index (BMI) risk scenario is that the policy customer has an overweight or underweight physical health index (BMI). The scenario with a claim history is that the policy customer has a claim history, and so on. If a policy customer hits different rules, the policy customer will go through all the paths and get a comprehensive score.

[0092] This application realizes the interpretability of the machine learning model at the sample level based on the XGBoost algorithm. In XGBoost, a gradient-based optimization algorithm is used to train the model, and the prediction ability of the model is gradually improved through an iterative method. The xgboost algorithm trains the model through additive training. Generally, the xgboost algorithm has n trees (for example, n = 200). When training the t-th tree, it will fit the residuals of the previous t - 1 trees. Based on the previous t - 1 trees, the loss function (including two parts: the true and predicted loss values and the regularization term) is minimized. Specifically, first, a model is initialized, then the gradients of the current model and the second-order derivatives of the objective function are calculated, and this information is used to generate a new decision tree model. Next, the best tree node splitting position is selected by minimizing the objective function, and the regularization term is applied to prevent overfitting. Finally, the newly generated decision tree model is weighted and fused with the previous model, and the prediction result of the model is updated. The purpose of this application is to accurately show the action direction and influence degree of each feature on the score of each sample, that is, to realize the interpretability of the model at the sample level. This will help users intuitively understand the model connotation, quickly identify factors that are contrary to the business logic, and then inject business insights into the optimization and iteration of the model, and finally achieve an efficient interaction and fit between the model and users.

[0093] Refer to the attached Figure 2 as shown in Figure 2Shows the structure diagram of the [x]th tree of the XGBoost algorithm. Optionally, for each sample, it can be divided into a certain leaf node at the bottom according to this tree structure diagram. This leaf node will have a score. In addition, the XGBoost algorithm has n trees, and after training is completed, n such tree structure diagrams will be obtained. Each tree is used to correct the current predicted value in order to minimize the loss function during the training process. The output of each tree (i.e., its score) will be added to the previous prediction result, thereby gradually approaching the true label distribution. For a certain sample, the final predicted score of the model for it is the sum of the scores of the n nodes where the sample is divided by these n tree structure diagrams, and it is converted into a probability value in the range of [0,1] through an activation function. This probability value represents the possibility that the sample is predicted as the positive class.

[0094] For example, referring to Figure 2 as shown, for a certain sample, after data processing, the value of the "sum_inv_amt" feature is 10, and the value of the "icdcod_der" feature is 2. Then this sample will be divided into the leftmost leaf node at the bottom according to the Figure 2 tree structure diagram, and the score of its leaf node = -0.0873684213. For the XGBoost algorithm, the final predicted score of a certain sample is calculated by summing the scores of the leaf nodes where the sample falls in all trees and then passing through the sigmoid activation function. For this sample, for a certain tree, according to its tree structure diagram, after successive splitting of the features, it finally falls into a certain leaf node. The score of this leaf node can be evenly distributed to each feature that splits this sample. From the Figure 2 tree structure diagram, it can be seen that the scores of the leaf nodes are positive and negative. In this way, for a certain sample, the sum of the scores of all the leaf nodes it belongs to may also be positive or negative. In one implementation manner of this embodiment, the activation function used is the sigmoid function. The sigmoid function is a monotonically increasing function and its value range is between 0 and 1 (satisfying the probability property).

[0095] Refer to the attached Figure 3 as shown, Figure 3 shows the second tree structure diagram of the XGBoost algorithm. Taking Figure 3For example, assume that for a certain sample, after data processing, the value of the "hosp_insp_fee_amt" feature is 700 and the value of the "ins_acc_age" feature is 20. Then, this sample is assigned to the rightmost node on the second tree structure diagram, and the leaf node score = 0.0354459323. On this tree, when this sample reaches the leaf node, it is split by the two features "hosp_insp_fee_amt" and "ins_acc_age". Then, it is considered that the influences of these two features are the same, and the leaf node score 0.0354459323 is evenly distributed, that is, the scores of these two features are both 0.0354459323 / 2 = 0.01772296615.

[0096] For a certain sample, the feature scores of each tree can be calculated. By summing up the feature scores of all trees separately for each feature, the influence score of this sample with respect to each feature can be obtained. The higher the influence score of a single sample for a single feature, the higher the final prediction score of the model for this sample, and the more inclined the model is to judge it as a positive sample. On the contrary, the lower the influence score of a single sample for a single feature, the lower the final prediction score of the model for this sample, and the more inclined the model is to judge it as a negative sample.

[0097] According to the above method, we can obtain the influence scores of each feature on the prediction score at the sample level, as well as the original values of the features, and provide them to the user together. According to the absolute value of the influence score, arrange them in descending order and provide detailed information. For example: Sample 1 (feature influence score): {"Factor A": -1.702587, "Factor B": 0.549438,...}; Sample 1 (feature original value): {"Factor A": 25, "Factor B": 37,...}; Sample 2 (feature influence score): {"Factor B": 1.340655, "Factor A": 1.195621,...}; Sample 2 (feature original value): {"Factor B": 155, "Factor A": 65,...}.

[0098] The user combines the influence score and the original value of the feature, and from a business perspective, judges its reasonable interpretability. Taking Factor A as an example: In Sample 1, Factor A has the greatest influence on the sample prediction score. When its value is 25, the influence score on the sample prediction score is -1.702587, and the influence direction is negative. In Sample 2, the influence of Factor A on the sample prediction score ranks second. When the value of Factor A is 65, the influence score on the sample prediction score is 1.195621, and the influence direction is positive. However, the influence degree is less than that of Factor B. Relevant technical personnel need to comprehensively evaluate the above information to ensure the reliability of the model output decision.

[0099] Based on this, the conclusion of each underwriting case can be quantified and explained in the machine learning model. Users can not only obtain accurate risk assessment results but also clearly understand the logic and basis behind the risk assessment. This improvement in transparency greatly enhances customer satisfaction and trust, making it easier for customers to accept the underwriting conclusion and reducing disputes caused by an opaque decision-making process.

[0100] Another embodiment of the underwriting risk assessment method of the present application, based on an embodiment of the above method, after the sub-model scoring for each risk scenario (exceeding the body health index (BMI), having a claim history, a disease history, a history of extension and rejection) is completed, we comprehensively process the scoring results of each sub-model to generate an overall risk score for the customer. The step S500: Perform a weighted calculation on the scores obtained by each sub-model, and finally obtain the risk score of the target customer in the current business. Specifically, it includes the following sub-steps:

[0101] S510, perform a normalization process on the scoring result of each sub-model.

[0102] Specifically, first perform a normalization process on the scoring result of each sub-model to ensure that the scores of each sub-model are in the same dimension, facilitating subsequent comprehensive calculations.

[0103] S520, assign corresponding weights to the scores of each sub-model according to the importance of each risk scenario and the performance of historical data.

[0104] Specifically, assign corresponding weights to the scores of each sub-model according to the importance of each risk scenario and the performance of historical data. The setting of the weights combines expert knowledge and actual business needs to ensure the reliability and practicality of the model.

[0105] S530, multiply the normalized scores by their respective weights, and then perform a weighted sum to obtain the total risk score of the target customer in the current business.

[0106] The specific weight values are set by relevant technical personnel or extracted by an automated tool according to historical underwriting characteristics. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

[0107] Optionally, in some other embodiments of this embodiment, it further includes the step: S540: Set different risk level thresholds according to the overall risk score. For example, customers with a score lower than a certain threshold are regarded as low-risk, and customers with a score higher than another threshold are regarded as high-risk.

[0108] In some other embodiments of this embodiment, a risk assessment method for underwriting in this application further includes: dynamically adjusting each sub-model. Specifically, it includes: regularly reviewing and updating the weights and thresholds of the model to adapt to the changing market environment and customer needs, and ensuring the continuous effectiveness of the model.

[0109] Optionally, not only are multiple machine learning sub-models customized and constructed for different underwriting scenarios, but also special attention is paid to the flexibility and adaptability of the system. By dynamically adjusting the model weights and thresholds, and regularly reviewing and updating the model parameters, this application can adapt to the changing market environment and customer needs, ensuring the continuous effectiveness and stability of the model. After the model is put into production and application, it still continuously feeds back model training based on the changes in intelligent questionnaires, underwriting rules, reinsurance assessment criteria, and manual review scales, and dynamically optimizes the model output results. At the same time, the model can be continuously updated over time to adapt to new data trends, maintaining a high level of interpretability and reliability. Through this series of processes, we can more accurately and comprehensively evaluate the risk status of customers, providing scientific and reliable decision-making support for insurance companies.

[0110] To solve the technical problem that in this specific business scenario of underwriting, the current processing methods for discrete factors cannot fully exploit the internal information associations of such features, thus limiting the expressiveness of the model and resulting in the loss of key information. Another embodiment of a risk assessment method for underwriting in this application details the processing of discrete factors during the model usage and training processes.

[0111] The discrete factor in this embodiment is illustrated by taking the disease classification code (ICD) as an example. The processing methods for other types of discrete factors should be the same as those described in this embodiment. Specifically, the ICD disease code, that is, the International Classification of Diseases (ICD), is an internationally unified disease classification method formulated by the World Health Organization (WHO). It classifies diseases according to characteristics such as the cause, pathology, clinical manifestations, and anatomical location of the disease, and represents them using code methods.

[0112] The steps of this application encode the discrete factor according to a preset factor splitting election or factor splitting expansion method, which can be specifically divided into two processing methods. Relevant technical personnel can choose to perform factor splitting election processing or factor splitting expansion processing on the discrete factor based on different processing requirements.

[0113] Specifically, in one implementation of this embodiment, in step S310, before processing the insurance application data of the target customer, a large number of discrete factor samples are obtained. For each discrete factor in the discrete factor samples, factor splitting and selection are performed: high-risk factors are screened and sorted, and feature transformation is performed on the high-risk factors according to the sorted list to obtain a feature transformation list. When processing the insurance application data of the target customer, the target discrete factor is brought into the feature transformation list. The transformed feature is obtained and then encoded into a numerical form. Specifically, it includes the following sub-steps:

[0114] S311, obtain a large number of discrete factor samples, and split the values of the discrete factors into lists.

[0115] Specifically, assume there are multiple samples, and the values of the discrete factor "previous ICD disease code" are respectively: "A001|B101|A102", "B101|C101|D103", "D102|C101", "A102"... After splitting according to the delimiter "|", they are respectively obtained as ["A001", "B101", "A102"], ["B101", "C101", "D103"], ["D102", "C101"], ["A102"]...

[0116] S312, set retention parameters, including the maximum retention number and the minimum proportion, to allocate the model training set, test set, and validation set.

[0117] Specifically, set the maximum retention number Top N and the minimum proportion parameter min ratio, which are equivalent to the parameters in the modeling process and can be debugged multiple times to achieve the balance of the effects on the model training set, test set, and OOT (Out-of-Time dataset, that is, the dataset used to verify the stability and adaptability of the model) data.

[0118] S313, count the quantities of each discrete factor after splitting, and calculate the proportions of each discrete factor in the positive samples or negative samples.

[0119] Specifically, count the quantities of each ICD disease code after splitting, and calculate various quantities and proportions. For example: For the ICD code [A001], the quantity in the negative samples is 10; the quantity in the positive samples is 20; the proportion in the positive samples is 0.666%; the total proportion is 0.03%. For the ICD code [A102], the quantity in the negative samples is 20; the quantity in the positive samples is 10; the proportion in the positive samples is 0.333%; the total proportion is 0.03%.

[0120] S314. Filter and retain each of the discrete factors to screen out high-risk factors: Filter out the discrete factors with a proportion less than the minimum proportion. Sort the remaining discrete factors in descending order according to the total quantity, and retain a number of the discrete factors with the top order in the list obtained from the first sorting as the high-risk factors, and the number of retained factors is the maximum retention number.

[0121] Specifically, first, according to the total quantity proportion after each splitting, if the value is less than the min ratio, then discard the ICD code. Here, the min ratio can also be set to 0, indicating that no ICD code is discarded. For the remaining ICD codes, sort them in descending order according to their total quantity (negative quantity + positive quantity), and retain at most the top N.

[0122] S315. Sort the retained high-risk factors in descending order according to their proportion in the positive samples, and convert the original values of the high-risk factors into numerical forms through feature encoding according to the list obtained from the second sorting.

[0123] Specifically, sort the top N retained ICD codes in descending order according to the proportion in the positive samples: After the above steps of processing the "previous ICD disease codes", the final sorted top N ICD list is obtained, for example: ["A001", "D103", "A102", "B101", ……, "C101"].

[0124] Based on the obtained descending list, perform feature transformation on its original values. For example, for a certain sample, the original value of this factor is "A001|B101|A102". First, split it according to the delimiter "|", and traverse each element in the top N ICD list ["A001", "D103", "A102", "B101", ……, "C101"] generated in Step 5 in turn until the corresponding ICD code is matched, then use this ICD code as the finally transformed ICD code. If the corresponding ICD code cannot be matched, fill it with "-99999". In this example, the finally transformed ICD code is "A001". After the above feature transformation, the enumeration quantity of this factor is reduced to the top N. On this basis, encode the factor into a numerical value through conventional encoding methods such as one-hot encoding, label encoding, WOE (Weight of Evidence) transformation, etc., and provide it to the model for learning.

[0125] In another implementation manner of this embodiment, in step S320: Before processing the insurance application data of the target customer, a large number of discrete factor samples are obtained. For each discrete factor in the discrete factor samples, factor splitting and expansion are performed: For each discrete factor, unique values are extracted and dummy variables are created. When processing the insurance application data of the target customer, the dummy variables corresponding to the target discrete factors are searched. The specific steps are as follows:

[0126] S321, extract unique values. Consistent with step S311, the factor values need to be split into a list. Then, all unique characteristic values of the ICD codes are identified from it.

[0127] S322, create dummy variables, which can more clearly describe the unique properties of each ICD and will not introduce unnecessary order relationships, making the data more expressive and distinguishable in the model. For each unique ICD code, such as "A001", a corresponding dummy variable is_A001 is generated and filled with 1 (yes) or 0 (no), indicating whether "A001" exists in the original concatenated content respectively. Suppose there are 1000 ICD codes generated by Step1 in total, and for a sample with the original ICD code of "A001|A002", the corresponding dummy variables is_A001 and is_A002 have values of 1, and the values of the remaining 998 dummy variables are 0.

[0128] The method of factor splitting and selection aims to screen out the ICD codes with the highest risks and can efficiently extract risks. The method of factor splitting and expansion aims to equally process each ICD code. The two methods complement each other. Using the processed discrete factors as model input features for machine learning training can ensure that the model can fully utilize the information in complex discrete features, significantly improve the model's ability to process complex discrete features such as ICD disease codes, and thus bring a significant improvement in the modeling effect. Specifically, it is manifested as higher prediction accuracy, better stability, and stronger business interpretability.

[0129] In another embodiment of a risk assessment method for underwriting insurance in this application, based on an embodiment of any of the above methods, before using each sub-model, the following steps are further included: For different risk scenarios, corresponding sub-models are constructed and trained. The general process of constructing each sub-model specifically includes the following steps:

[0130] S1, collect customer information from multiple channels, preprocess the customer information, and extract the general features and scenario features that affect risk assessment as the training set.

[0131] Specifically, customer information is obtained from multiple sources, including but not limited to demographic information, medical records, historical insurance application data, historical claim data, historical disease information, etc. Based on the principles of business logic and data-driven, we carefully select and process the key features that affect risk assessment to ensure that the data input into the model is both representative and easy to interpret.

[0132] Optionally, it also includes cleaning the originally collected data, removing outliers, filling in missing values, etc. The collected data is sorted, and the data set is divided into a training set, a validation set, a test set, etc. In an implementation manner of this embodiment, step S1 further includes: performing factor splitting election or factor splitting expansion processing on the discrete factor. The specific technical details are the same as those in the above embodiment, and will not be repeated in this embodiment.

[0133] S2, divide the risk scenarios applicable to the sub-models, and independently establish machine learning models for each of the risk scenarios.

[0134] Specifically, this embodiment constructs at least four different machine learning sub-models. Each model is customized for a specific underwriting scenario, including the physical abnormality scenario, the claim history scenario, the disease history scenario, and the rejection history scenario, and the model can obtain accurate risk assessment results.

[0135] S3, the machine learning model is a tree structure. Each input feature is used as a leaf node of the machine learning model, and the XGBoost algorithm is used to train the machine learning model to understand the dependence degree of the machine learning model on different input features during the decision-making process.

[0136] Specifically, this embodiment uses the XGBoost algorithm for model training to increase the model interpretability, finds a suitable hyperparameter combination through grid search, and selects the model based on K-fold cross-validation. Considering the performance of the comprehensive test set and the validation set (OOT, Out-of-Time data set), overfitting or underfitting is avoided, and the model performance is optimized.

[0137] S4, add each of the trained machine learning models to the underwriting risk assessment system in the form of sub-models.

[0138] Specifically, after the model training, testing, and validation are completed, the trained model is applied to the actual underwriting process. When the model is used, the business system can call the model interface service to obtain the model prediction result and get the model underwriting conclusion.

[0139] In another implementation of the above embodiments, after model training, the following steps are further included: S5. Use the validation set and the test set to validate and test each of the trained machine learning models respectively. Based on the comprehensive validation and test results, optimize the model to avoid overfitting and underfitting of the model.

[0140] S6. Obtain the analysis and feedback of relevant technical personnel on special cases, so as to dynamically update the selected input features and model parameters, and further optimize the model output results.

[0141] Specifically, based on the analysis and feedback of underwriting experts on false negative and false positive cases, optimize the feature selection and modeling sample selection again. By setting appropriate parameter thresholds, the underwriting scale of the model output can be determined. The threshold takes into account STP and risk control, neither over-labeling risks nor causing risk omissions, and continuously adjusts and optimizes. Output the underwriting conclusion according to the model prediction and threshold setting. In some other implementations, the accuracy of the model results can also be ensured by setting the system parallel period and increasing manual inspections.

[0142] In another embodiment of the method of the present application, based on the above embodiments, a method for underwriting risk assessment of the present application further includes: S7. Calculate the distribution difference between the model training data set and the actual application data set for the input features of each of the sub-models. For each of the sub-models, calculate its stability, performance, discrimination degree for positive and negative samples, and sorting index to monitor the stability of the model and the change trend of the model effect.

[0143] Specifically, calculate the feature PSI (Population Stability Index) for the input features of each sub-model to monitor the feature stability. And for each sub-model, calculate indicators such as PSI (Population Stability Index), AUC (Area Under the Curve), KS (Kolmogorov-Smirnov), and sorting to monitor the stability of the model and the change trend of the model effect. For the dynamic optimization of the model.

[0144] Based on the same technical concept, the present application also discloses an underwriting risk assessment system, which can be used to implement any of the above underwriting risk assessment methods. Specifically, an embodiment of an underwriting risk assessment system of the present application is as shown in the Figure 4 specification appendix

[0145] A data collection module, configured to collect the insurance application data of target customers.

[0146] A risk identification module, configured to identify the risk factors affecting underwriting in the insurance application data. The risk factors include discrete factors.

[0147] The risk identification module is further configured to encode the discrete factors according to a preset factor splitting and election or factor splitting and expansion method.

[0148] Specifically, two encoding methods, namely factor splitting and election and factor splitting and expansion, are innovatively proposed, which are specifically used to optimize the preprocessing of complex discrete features such as ICD disease codes, significantly improving the effectiveness and stability of the model. In addition, the system is adjusted in combination with expert knowledge to ensure the reliability and practicality of the model, while maintaining a high level of transparency and interpretability.

[0149] The model processing module is configured to select key features suitable for different risk scenarios from the risk factors and input them into multiple trained sub-models to obtain the scores of the target customer in different risk scenarios and the corresponding interpretable results. The multiple sub-models include: a BMI intelligent identification sub-model, a claim history risk control sub-model, a disease risk level classification sub-model, and a deferred rejection risk re-mining sub-model.

[0150] Specifically, each sub-model is for a different underwriting scenario. In the scenario where the body mass index (BMI) is abnormal, the BMI intelligent identification sub-model is used to mine relevant features, including but not limited to factors such as height, weight, and whether the standard developed based on the underwriting manual is within the normal range, and a machine learning model specifically applicable to this scenario is developed based on these features. In the scenario with a claim history, analyze the reasons, amounts, frequencies, time intervals of past claims, and the relationship between claims and the subsequent health conditions and insurance application behaviors of the policyholders, and build a model through these features to accurately judge the risk level of policyholders with a claim history when applying for insurance again. In the scenario with a disease history, carefully sort out the types, severity, treatment processes, recovery conditions of various diseases, and their mutual influences, and at the same time consider the interaction between the disease history and factors such as the age, gender, and family medical history of the policyholder, and extract representative and predictive features to build a model, so as to accurately evaluate the impact of the disease history on underwriting risks. In the scenario with a deferred rejection history, following the principle of "following precedent", for customers with a deferred rejection history, if there is a record of underwriting a profitable policy with high or the same review requirements recently, after evaluating and confirming that there are no new risk points, it is automatically passed.

[0151] The risk summarization module is configured to perform weighted calculation on the scores obtained by each sub-model to finally obtain the total risk score of the target customer in the current business.

[0152] A method and a system for underwriting risk assessment in this application have the same technical concept, and the technical details of the embodiments of the two can be mutually applicable. To avoid repetition, they will not be elaborated here.

[0153] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each program module is used as an example. In actual applications, the above functions can be allocated to different program modules as needed, that is, the internal structure of the device can be divided into different program units or modules to complete all or part of the functions described above. Each program module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into a processing unit. The above integrated units can be implemented in the form of hardware or in the form of software program units. In addition, the specific names of each program module are only for the convenience of mutual distinction and do not limit the protection scope of the present application.

[0154] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

Claims

1. A method for underwriting risk assessment, characterized in that: The steps include: Collecting insurance data of target customers; identifying risk factors in the insurance data that affect underwriting; the risk factors include discrete factors; Encoding the discrete factors according to a preset factor splitting selection or factor splitting expansion method; Select key features suitable for different risk scenarios from the risk factors and input them into multiple trained sub-models to obtain scores of target customers under different risk scenarios and interpretable results corresponding to the scores; The multiple sub-models include: BMI intelligent identification sub-model, claim history risk control sub-model, disease risk level classification sub-model, and delay and rejection risk re-mining sub-model; The scores obtained by each of the sub-models are weighted and calculated to finally obtain the total risk score of the target customer in the current business.

2. The underwriting risk assessment method according to claim 1, characterized in that: The step of selecting key features suitable for different risk scenarios from the risk factors and inputting them into the trained multiple sub-models to obtain the scores of target customers in different risk scenarios and the interpretable results corresponding to the scores comprises the following steps: Using preset rules to determine whether the current business of the target customer hits the risk scenario; If one or more of the risk scenarios are hit, the key features corresponding to the risk scenarios are input into the machine learning model corresponding to the risk scenarios; The machine learning model is an XGBoost model; Outputting the prediction result score of the machine learning model for the key feature at each leaf node and the prediction process trajectory; generating interpretable results of all input features in the machine learning model based on the prediction process trajectory; The scores of the key features on each leaf node are summed up, and then the score of the target customer's current business in the risk scenario is obtained through activation function calculation.

3. The underwriting risk assessment method according to claim 1, characterized in that: The scores obtained by each of the sub-models are weighted and calculated to finally obtain the risk score of the target customer in the current business; The steps include: Standardizing the scoring results of each sub-model; According to the importance of each risk scenario and the performance of historical data, a corresponding weight is assigned to the score of each sub-model; The standardized scores are multiplied by their respective weights, and then weighted summed to obtain the total risk score of the target customer's current business.

4. The underwriting risk assessment method according to claim 1, characterized in that: The step of encoding the discrete factors according to a preset factor splitting selection or factor splitting expansion method comprises the following steps: Before processing the insurance data of the target customers, a large number of discrete factor samples are obtained; Performing factor splitting and selection on each discrete factor in the discrete factor sample: screening and sorting the high-risk factors, performing feature conversion on the high-risk factors according to the sorted list, and obtaining a feature conversion list; When processing the insurance data of the target customer, the target discrete factor is brought into the feature conversion list; Get the converted features and encode them into numerical form; Or, factor splitting and expansion are performed on each discrete factor in the discrete factor sample: a unique value is extracted for each discrete factor and a dummy variable is created; when processing the insurance data of the target customer, the dummy variable corresponding to the target discrete factor is searched; The discrete factors include disease classification codes.

5. The underwriting risk assessment method according to claim 4, characterized in that: The step of performing factor splitting and selection on each discrete factor in the discrete factor sample specifically comprises the following steps: Splitting the values ​​of the discrete factor sample into lists; Set retention parameters, including the maximum number of retentions and the minimum percentage; to allocate model training sets, test sets, and validation sets; Counting the number of each discrete factor after splitting, and calculating the proportion of each discrete factor in the positive sample or the negative sample; Filter and retain each of the discrete factors to screen out high-risk factors: filter the discrete factors whose proportion is less than the minimum proportion; sort the remaining discrete factors for the first time according to the total number, and retain a number of the discrete factors in the first order as the high-risk factors according to the first sorted list, and the number of retained factors is the maximum retained number; The retained high-risk factors are sorted for the second time according to their proportions in the positive samples, and the original values ​​of the high-risk factors are subjected to feature conversion according to the second sorted list.

6. The underwriting risk assessment method according to claim 4, characterized in that: The step of factoring each discrete factor in the discrete factor sample comprises the following steps: Splitting the values ​​of the discrete factor sample into lists; Identify the unique eigenvalue of each discrete factor; and generate a corresponding dummy variable for each unique eigenvalue obtained for the discrete factor.

7. An underwriting risk assessment method according to any one of claims 1 to 6, characterized in that: Before the model is used, the method further includes: constructing and training corresponding sub-models for different risk scenarios; specifically, the steps include: Collect customer information from multiple channels, pre-process the customer information, and extract common features and scenario features that affect risk assessment as training sets; Divide the risk scenarios applicable to the sub-models, and independently establish a machine learning model for each of the risk scenarios; The machine learning model is a tree structure, each input feature is used as a leaf node of the machine learning model, and the machine learning model is trained using the XGBoost algorithm to understand the degree of dependence of the machine learning model on different input features in the decision-making process; The trained machine learning models are added to the underwriting risk assessment system in the form of sub-models.

8. The underwriting risk assessment method according to claim 7, characterized in that: The following steps are also included: Using the validation set and the test set to validate and test the trained machine learning models respectively; optimizing the model by combining the validation and test structures to avoid overfitting and underfitting of the model; Obtain analysis and feedback from relevant technical personnel on special cases in order to dynamically update the selected input features and model parameters and further optimize the model output results.

9. The underwriting risk assessment method according to claim 8, characterized in that: The method also includes the following steps: for the input features of each sub-model, calculating the distribution difference between the model training data set and the actual application data set; for each sub-model, calculating its stability, performance, discrimination between positive and negative samples, and ranking index, so as to monitor the stability of the model and the changing trend of the model effect.

10. An underwriting risk assessment system, characterized in that: The method for implementing the underwriting risk assessment method according to any one of claims 1 to 9 above specifically comprises: Data collection module, used to collect insurance data of target customers; A risk identification module, used to identify risk factors affecting underwriting in the insurance application data; the risk factors include discrete factors; The risk identification module is further used to encode the discrete factors according to a preset factor splitting selection or factor splitting expansion method; A model processing module is used to select key features suitable for different risk scenarios from the risk factors and input them into multiple trained sub-models to obtain the scores of target customers in different risk scenarios and the interpretable results corresponding to the scores; the multiple sub-models include: BMI intelligent identification sub-model, claim history risk control sub-model, disease risk level classification sub-model, and delay and rejection risk re-mining sub-model; The risk summary module is used to perform weighted calculation on the scores obtained by each sub-model to finally obtain the total risk score of the target customer in the current business.

Citation Information

Patent Citations

  • Credit investigation scoring method and system integrating multiple machine learning models

    CN110956273A

  • Claim settlement method and system based on machine learning

    CN114648413A

  • Default rate prediction method and system based on mixed effect logistic regression

    CN116258574A

  • Risk assessment method and device, model training method and device, medium and equipment

    CN116797343A

  • Risk portrait model construction method and device, electronic equipment and storage medium

    CN116861243A

Cited By

  • Squeezing risk assessment device and method

    CN120809241A