Claim settlement recommend and survey point location recommendation system based on random forest algorithm

Through the claims and survey point recommendation system based on the random forest algorithm, the problems of insufficient precision and lack of flexibility in the prior art are solved, and more efficient and accurate claims and survey point recommendations are achieved.

CN120067436APending Publication Date: 2025-05-30PICC HEALTH INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411984980.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing automatic claims adjustment system relies on manual experience or preset rules, resulting in insufficient precision and lack of flexibility, and ineffective reduction or cost control.

Method used

The claim preparatory and survey point recommendation system based on the random forest algorithm is adopted, and through data collection, model construction and point determination modules, an automatic adjustment model is built and survey points are recommended, and the claims operators make decisions.

Benefits of technology

It improves the accuracy and flexibility of claims and prescriptions, reduces human errors, improves the effectiveness of the survey points, and ensures the accuracy and efficiency of the prescription process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067436A_ABST
    Figure CN120067436A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of insurance claim settlement investigation, in particular to a claim settlement calling and investigation point location recommendation system based on a random forest algorithm. According to the method, claim settlement information and investigation information are extracted from a claim settlement module database and an investigation module database to form a sample data set, an automatic calling model based on a random forest algorithm is established according to optimal split features, whether a claim settlement case needs to be called is judged according to the automatic calling model, and feature evaluation is performed on the claim settlement case needing to be called. And determining a survey point location. According to the invention, the characteristic data of the claim settlement case is substituted into the automatic calling model, whether the claim settlement case is called is judged, the claim settlement worker is assisted to make a decision for initiating investigation, the situation that the calling is not accurate and effective due to insufficient experience of the claim settlement worker or system problems is avoided, and the investigation point location is recommended, so that the investigation efficiency is improved. And investigation personnel are assisted to select related investigation points for investigation, so that the effectiveness of the investigation points is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of insurance claim investigation, and specifically, to a claim transfer and investigation point recommendation system based on a random forest algorithm. Background Art

[0002] Claim automatic transfer usually refers to the use of computer technology and artificial intelligence in the insurance claim process to quickly process and settle claim cases. This process evaluates losses, reviews materials, determines damages, and settles accounts in an automated manner, thus greatly shortening the claim time, reducing human errors, and improving the efficiency and accuracy of claims.

[0003] Currently, after a claim case is accepted, claim handlers judge whether to transfer based on experience, or the system automatically transfers according to preset transfer rules. Manual judgment of whether to transfer depends on the work experience of claim handlers, which may lead to missed transfers or over-transfers, and is not conducive to loss reduction or cost control. Using preset rules for transfer has limitations in covering scenarios and lacks flexibility. In order to be able to use the random forest algorithm to recommend whether to transfer, assist claim handlers in making decisions on initiating investigations, avoid inaccurate and ineffective transfers caused by insufficient experience of claim handlers or system problems, and at the same time, intelligently recommend investigation points to assist investigators in selecting relevant investigation points for investigation and improving the effectiveness of investigation points. Therefore, we propose a claim transfer and investigation point recommendation system based on a random forest algorithm. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems that manual judgment of whether to transfer depends on the work experience of claim handlers, which may lead to missed transfers or over-transfers, and is not conducive to loss reduction or cost control. Using preset rules for transfer has limitations in covering scenarios and lacks flexibility. In order to be able to use the random forest algorithm to recommend whether to transfer, assist claim handlers in making decisions on initiating investigations, avoid inaccurate and ineffective transfers caused by insufficient experience of claim handlers or system problems, and at the same time, intelligently recommend investigation points to assist investigators in selecting relevant investigation points for investigation and improving the effectiveness of investigation points.

[0005] To achieve the above object, the present invention provides a claim transfer and investigation point recommendation system based on a random forest algorithm, including a data acquisition module, a model construction module, and a point determination module;

[0006] The data acquisition module extracts claim information and investigation information from the databases of the claim module and the investigation module, numerically processes the data, and forms a sample data set;

[0007] The model construction module randomly selects a sample set from the sample data set to construct a decision tree. At each split point of the decision tree, a random feature subset is selected, and the best split feature is found from it to establish an automatic adjustment model based on the random forest algorithm. The data set is divided into a training set, a test set, and a validation set to train the automatic adjustment model;

[0008] The location determination module determines whether a claim case needs to be adjusted according to the automatic adjustment model. For claim cases that need to be adjusted, feature evaluation is performed to determine the investigation location.

[0009] As a further improvement of this technical solution, the model construction module establishes an automatic adjustment model based on the random forest algorithm, and its method steps are as follows:

[0010] S2.1.1. Select samples: Use the bootstrap sampling method to perform random sampling with replacement on the sample data set;

[0011] S2.1.2. Construct a decision tree: Select the best split feature from the feature subset, split the node according to this feature, form the branches of the tree, and each branch forms a new node. Repeat this process;

[0012] S2.1.3. Train the automatic adjustment model: Use the training set data, measure the difference between the predicted probability and the true label with the cross-entropy loss function, and train the automatic adjustment model;

[0013] S2.1.4. Evaluate and optimize the automatic adjustment model: Use the accuracy of the model prediction as the evaluation index, and use the validation set data to evaluate and optimize the automatic adjustment model.

[0014] As a further improvement of this technical solution, in the S2.1.2 for constructing the decision tree, the best split feature is determined by calculating the information gain of each feature. The calculation formula for the information gain is:

[0015]

[0016] where Gain(D,A) is the information gain, Entropy(D) is the entropy of the data set D, Entropy(D i ) is the entropy of the data subset D i , v is the number of different values, D is the data set, D i is the data subset, |D| is the number of samples in the data set D, |D i | is the number of samples in the data subset D i , and A is the feature.

[0017] As a further improvement of this technical solution, in the S2.1.3 for training the automatic adjustment model, the grid search combined with cross-validation method is used to adjust the parameters of the automatic adjustment model.

[0018] As a further improvement of this technical solution, in S2.1.3, the automatic adjustment model is trained with the minimum cross-entropy loss function, and the gradient descent algorithm is used to further optimize the parameters in the model. The formula of the cross-entropy loss function is as follows:

[0019]

[0020] where L is the cross-entropy loss function, n is the claim settlement sample, p i is the probability that needs to be adjusted, y i is the true label, and i is the sample sorting number.

[0021] As a further improvement of this technical solution, the positioning point module determines the investigation points, and the method steps are as follows:

[0022] S3.1.1. Calculate the Gini impurity of the original data set;

[0023] S3.1.2. Divide the data set according to the features and calculate the Gini impurity of the sub-data sets;

[0024] S3.1.3. Calculate the Gini index of the features.

[0025] As a further improvement of this technical solution, in S3.1.1, the Gini impurity of the original data set is calculated, and the formula for calculating the Gini impurity is:

[0026]

[0027] where Gini(D) is the Gini impurity, K is the number of categories, D is the data set, k is the category ordinal number, and n k is the number of samples in the k-th category.

[0028] As a further improvement of this technical solution, in S3.1.3, the Gini index of the features is calculated, and the formula for calculating the Gini index is:

[0029]

[0030] where Gini(D,A) is the Gini index, Gini(D i ) is the Gini impurity of the subset D i , v is the number of different values, D is the data set, D i is the data subset, |D| is the number of samples in the data set D, |D i | is the number of samples in the data subset D i , and A is the feature.

[0031] As a further improvement of the technical solution, the point determination module determines the investigation points, sorts the importance of features according to the Gini index value of the features, determines the most important feature points as the extraction and investigation points, and recommends the investigation points.

[0032] As a further improvement of the technical solution, when the data acquisition module numerically processes the data, it uses the normalization method to map the feature values into the numerical interval.

[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0034] 1. For the claims extraction and investigation point recommendation system based on the random forest algorithm, the data acquisition module extracts claims information and investigation information from the databases of the claims module and the investigation module, numerically processes the data, and forms a sample data set. The model construction module randomly selects a sample set from the sample data set to construct a decision tree. At each split point of the decision tree, it randomly selects a feature subset and finds the best split feature from it to establish an automatic extraction model based on the random forest algorithm. Substitute the feature data of the claims case into this automatic extraction model to judge whether the claims case needs to be extracted, assisting the claims operators to make a decision on initiating an investigation, and avoiding inaccurate and ineffective extraction caused by insufficient experience of the claims operators or system problems.

[0035] 2. When the automatic extraction model in the point determination module determines that the claims case needs to be extracted, it calculates the Gini index of each feature, sorts the importance of the features based on the Gini index value of the features, determines the most important feature points as the extraction and investigation points, and recommends the investigation points, assisting the investigation personnel to select relevant investigation points for investigation, improving the effectiveness of the investigation points. When the claims case does not need to be extracted, it returns to the claims module, and the case continues the claims process. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is the overall process schematic diagram of the present invention;

[0037] Figure 2 is the method process schematic diagram of S2 of the present invention;

[0038] Figure 3 is the method process schematic diagram of S3 of the present invention.

[0039] The meanings of the various labels in the figure are as follows:

[0040] 100, data acquisition module; 200, model construction module; 300, point determination module. DETAILED DESCRIPTION OF THE INVENTION

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0042] Currently, manually judging whether to initiate an investigation depends on the work experience of claims settlement operators, which may lead to missed or excessive initiation, and is not conducive to loss reduction or cost control. Using preset rules to cover scenarios for initiating an investigation has limitations and lacks flexibility. In order to be able to use the random forest algorithm to recommend whether to initiate an investigation, assist claims settlement operators in making decisions on initiating investigations, and avoid inaccurate and ineffective initiation due to insufficient experience of claims settlement operators or system problems. At the same time, intelligently recommend investigation points, assist investigators in selecting relevant investigation points for investigation, and improve the effectiveness of investigation points.

[0043] Therefore, the present invention proposes to collect claim information and investigation information from the databases of the claims settlement module and the investigation module through a data collection module to form a sample data set. The model construction module establishes an automatic investigation initiation model based on the random forest algorithm according to the best splitting feature, and judges whether a claim case needs to be investigated according to the automatic investigation initiation model. The location determination module conducts feature evaluation on the claim cases that need to be investigated to determine the investigation points.

[0044] Specifically as follows:

[0045] Please refer to Figure 1 As shown in the figure, the present invention provides a claims investigation initiation and investigation point recommendation system based on the random forest algorithm, including a data collection module 100, a model construction module 200, and a location determination module 300;

[0046] The data collection module 100 extracts claim information and investigation information from the databases of the claims settlement module and the investigation module, numerically processes the data, and forms a sample data set;

[0047] The model construction module 200 randomly selects a sample set from the sample data set to construct a decision tree. At each splitting point of the decision tree, it randomly selects a feature subset and finds the best splitting feature from it to establish an automatic investigation initiation model based on the random forest algorithm, and divides the data set into a training set, a test set, and a validation set to train the automatic investigation initiation model;

[0048] The location determination module 300 judges whether a claim case needs to be investigated according to the automatic investigation initiation model, conducts feature evaluation on the claim cases that need to be investigated, and determines the investigation points.

[0049] During the construction of a decision tree, the traditional method is to consider all available features at each node split to select the best splitting feature. However, this approach may lead to overfitting of the decision tree to the training data because it will use every detailed feature in the data as much as possible to build a complex tree. The method of randomly selecting a subset of features is to, at each split point, only look for the best splitting feature from a random subset of all features;

[0050] For example, suppose there are 10 features. At a certain node split, randomly select 3 of them to form a subset, and then determine the best splitting feature only among these 3 features. The advantage of doing this is that different subtrees will have differences during the construction process due to different subsets of features considered, thus increasing the diversity of the model.

[0051] As Figure 2 shown, among them, the model construction module 200 builds an automatic adjustment model based on the random forest algorithm, and its method steps are as follows:

[0052] S2.1.1, Select samples: Use the bootstrap sampling method to randomly sample the sample dataset with replacement;

[0053] S2.1.2, Build a decision tree: Select the best splitting feature from the subset of features, split the node according to this feature to form branches of the tree, and each branch forms a new node. Repeat this process;

[0054] S2.1.3, Train the automatic adjustment model: Use the training set data, measure the difference between the predicted probability and the true label with the cross-entropy loss function, and train the automatic adjustment model;

[0055] S2.1.4, Evaluate and optimize the automatic adjustment model: Use the accuracy of the model prediction as the evaluation index, and use the validation set data to evaluate and optimize the automatic adjustment model;

[0056] The training dataset comes from the historical stock data of the claims settlement module and the investigation module. Extract the claims settlement information and investigation information from the databases of the claims settlement module and the investigation module, and numerically process the data to form a sample dataset. Usually, a sample contains multiple feature values and a label value (whether to adjust), and each feature describes different aspects of the claims settlement;

[0057] Perform feature engineering on the data to extract and construct valuable features. For example, the claims settlement frequency and average claims settlement amount of the policyholder can be calculated, and the accident type can be encoded, etc. At the same time, standardize or normalize some continuous features to improve the performance of the model;

[0058] Based on the bootstrap sampling method, samples are randomly drawn with replacement from an original training set containing n samples. Each time, one sample is drawn and then put back into the original training set before the next draw, which means this sample may still be selected in the next sampling. After sampling n times, a bootstrap set composed of n samples of the same size as the original training set is finally obtained. Due to random sampling, each bootstrap set is different from the original data set and other sampling sets, and it is possible to freely create an inexhaustible and distinct set of bootstrap sets. Training base classifiers with these bootstrap sets will also be different, improving the generalization performance of the model.

[0059] In order to more accurately determine the best splitting feature, in S2.1.2, a decision tree is constructed, and the best splitting feature is determined by calculating the information gain of each feature. The calculation formula for information gain is:

[0060]

[0061] Among them, Gain(D,A) is the information gain, Entropy(D) is the entropy of the data set D, Entropy(D i ) is the entropy of the data subset D i , v is the number of different values, D is the data set, D i is the data subset, |D| is the number of samples in the data set D, |D i | is the number of samples in the data subset D i , and A is the feature.

[0062] The information gain measures the degree of reduction in the uncertainty of the sample set before and after feature splitting. The larger the information gain, the better the classification effect of using this feature for splitting on the samples;

[0063] Suppose we are building an automatic insurance claim adjustment model to determine whether an insurance claim case requires further investigation (adjustment). We have a claim data set with the following features:

[0064] Claim amount (numerical): Divided into three intervals: low, medium, and high. For example, 0 - 1000 yuan is low, 1000 - 10000 yuan is medium, and above 10000 yuan is high;

[0065] Accident type (categorical): Including traffic accidents, disease medical treatment, and property damage;

[0066] Number of claims (numerical): The number of claims made by the customer in the past period;

[0067] Insurance product type (categorical): Such as life insurance, car insurance, property insurance;

[0068] Target variable: Whether adjustment is required (Yes / No);

[0069] First, let's calculate the information entropy of the original dataset. Assume there are 100 claim cases in the dataset, among which 60 do not require adjustment and 40 require adjustment. According to the information entropy formula:

[0070]

[0071] where H(D) is the information entropy, K is the number of categories, D is the dataset, k is the category ordinal number, and n k is the number of samples in the k-th category, and N is the total number of samples.

[0072] Calculated according to the formula:

[0073] Calculate the information gain of each feature. Among them, for the claim amount feature, the claim amount is divided into 3 intervals. There are 30 cases in the low-amount interval, among which 25 do not require adjustment and 5 require adjustment. There are 40 cases in the medium-amount interval, among which 20 do not require adjustment and 20 require adjustment. There are 30 cases in the high-amount interval, among which 15 do not require adjustment and 15 require adjustment. Calculate the information entropy of each amount interval according to the information entropy formula, and substitute the corresponding values into the calculation formula of information gain, as follows:

[0074]

[0075] Similarly, calculate the information gain of the accident type feature, claim frequency feature, and insurance product type feature according to the above steps. By comparing the information gain of each feature, the feature with the largest information gain value is the best splitting feature. Then continue to calculate the information gain within each subset to find the best splitting feature of the next layer, and so on to construct a decision tree. In the random forest algorithm, each decision tree will repeat this process to construct the tree structure, and finally form an integrated model for the judgment of automatic insurance claim adjustment.

[0076] In order to better adjust the model parameters, among them, in S2.1.3, train the automatic adjustment model, and use the method of grid search combined with cross-validation to adjust the parameters of the automatic adjustment model;

[0077] More trees can improve the performance of the model, but at the same time increase the computational time and memory consumption. Too few trees may not be able to take advantage of the random forest, because the generalization ability of the model may not be strong enough to capture the complex relationships in the data. The methods of tuning parameters involved in selecting the appropriate number of trees are: grid search (by trying different numbers of trees within a predefined range and combining cross-validation to select the parameter combination with the best performance), random search (randomly selecting a set of parameters in the parameter space for trial and also combining cross-validation to evaluate the performance);

[0078] First, it is necessary to determine the random forest model parameters to be adjusted and their possible value ranges. For example, for the random forest classifier RandomForestClassifier, common parameters include:

[0079] n_estimators (the number of decision trees): It can be tried to take values at intervals of 50 from 50 to 500, such as [50, 100, 150, 200, 250, 300, 350, 400, 450, 500];

[0080] max_depth (the maximum depth of the tree): The value range can be from 3 to 15, such as [3, 5, 7, 9, 11, 13, 15];

[0081] max_features (the maximum number of features considered): It can be an integer, a floating point number, or a specific string (such as'sqrt', indicating the square root of the total number of features). If it is an integer, [2, 3, 4, 5] can be tried; if it is a floating point number, [0.2, 0.3, 0.4, 0.5] can be considered;

[0082] min_samples_split (the minimum number of samples required for further splitting of internal nodes): It can start from 2 and take values at intervals of 2, such as [2, 4, 6, 8, 10];

[0083] min_samples_leaf (the minimum number of samples in a leaf node): Similar to min_samples_split, the value range can be [1, 2, 3, 4, 5];

[0084] Select the number of folds for cross-validation. Usually, 5-fold cross-validation (cv = 5) or 10-fold cross-validation (cv = 10) is used. For example, 5-fold cross-validation will divide the dataset into 5 subsets. Each time, 4 subsets are used as the training set and 1 subset is used as the validation set, and such 5 training and validation processes will be carried out;

[0085] Using Python tools, the GridSearchCV class is used to perform grid search and cross-validation. The initialized model, parameter grid, and cross-validation strategy are passed into the constructor of GridSearchCV. Here, estimator is the model whose parameters are to be adjusted, param_grid is the parameter grid, cv is the number of cross-validation folds, and scoring is the evaluation metric (accuracy is used here). The grid_search.fit method will perform grid search and cross-validation on the training set, traverse all parameter combinations in the parameter grid, and for each combination, perform 5-fold cross-validation and calculate the evaluation metric.

[0086] After the grid search is completed, the best parameter combination can be obtained through grid_search.best_params_. For example, the returned result may be {'n_estimators': 150,'max_depth': 7,'min_samples_split': 4,'min_samples_leaf': 2,'max_features':'sqrt'}. This means that under the given parameter range and evaluation metric, this parameter combination is the optimal one. The average score (accuracy here) of the best parameter combination in cross-validation can be obtained through grid_search.best_score_. The model with the best parameters can be obtained using grid_search.best_estimator_, and predictions can be made on the test set to calculate the evaluation metrics (such as accuracy, recall, F1-score, etc.) to evaluate the generalization ability of the model.

[0087] In order to better optimize the parameters in the model, in S2.1.3, the automatic tuning model is trained with the minimum cross-entropy loss function as the objective, and the gradient descent algorithm is used to further optimize the parameters in the model. The formula for the cross-entropy loss function is:

[0088]

[0089] where L is the cross-entropy loss function, n is the claim sample, p i is the probability to be tuned, y i is the true label, and i is the sample sorting number.

[0090] For example, there are 3 claim samples, and the model prediction probabilities are p 1 = 0.8, p 2 = 0.3, p 3 = 0.6;

[0091] The corresponding true labels are y 1 = 1, y 2 = 0, y3 = 1;

[0092] Substituting the corresponding values into the formula of the cross - entropy loss function gives:

[0093]

[0094] When constructing an insurance claim adjustment model using the cross - entropy loss function (assuming a binary classification problem, i.e., adjust or not adjust), the output of the model is usually a probability value, representing the probability that a claim needs to be adjusted. For example, if the model output is 0.7, it means the model believes there is a 70% probability that this claim case needs to be adjusted;

[0095] Generally speaking, the reduction of cross - entropy loss is usually accompanied by the improvement of these evaluation metrics, but the relationship between them is not completely linear. For example, in the case of class imbalance in the data, even if the cross - entropy loss decreases, the accuracy may seem high due to the dominance of the majority class, but the recall rate may be low. At this time, multiple evaluation metrics need to be considered comprehensively to evaluate the performance of the model.

[0096] As Figure 3 shown, among them, the determination point module 300 determines the investigation points, and its method steps are as follows:

[0097] S3.1.1. Calculate the Gini impurity of the original data set;

[0098] S3.1.2. Divide the data set by features and calculate the Gini impurity of the sub - data sets;

[0099] S3.1.3. Calculate the Gini index of the features;

[0100] In the Python tool, directly obtain the importance scores of the features. For example, in the scikit - learn library of Python, after training, the random forest model has a feature_importances_ attribute, which represents the relative importance of each feature for model prediction. These importance scores are calculated based on the contribution of the features to reducing impurity (such as Gini impurity) during the construction of the decision tree;

[0101] Sort all the features according to the importance scores. The higher the score of a feature, the more critical it is for claim adjustment prediction. For example, if the importance score of the "claim amount" feature is very high, it indicates that it plays an important role in judging whether adjustment is needed;

[0102] Focus on the information sources related to important features as the investigation points. If the "accident location" is an important feature, then the surrounding environment, traffic conditions, locations of relevant monitoring devices, etc. around the accident site can be used as investigation points. For highly important customer-related features, such as customer occupation, income level, etc., the customer's workplace, relevant financial institutions, etc. can be considered as investigation points.

[0103] In order to better calculate the Gini impurity, where S3.1.1 calculates the Gini impurity of the original data set, and the formula for calculating the Gini impurity is:

[0104]

[0105] Where Gini(D) is the Gini impurity, K is the number of classes, D is the data set, k is the class ordinal number, n k is the number of samples in the k-th class, and N is the total number of samples.

[0106] For example, there are 40 samples in class 0, 60 samples in class 1, and the total number of samples is 100;

[0107] Then the probability of class 0 is 0.4, and the probability of class 1 is 0.6;

[0108] Substituting them into the formula for Gini impurity, we get: Gini(D) = 1 - (0.4 2 + 0.6 2 ) = 0.48;

[0109] For each feature A, count the number of its different values. Assume that feature A has v different values, and divide the data set D into v subsets according to the values of feature A, and calculate the Gini impurity of each subset respectively.

[0110] In order to better calculate the Gini index, where S3.1.3 calculates the Gini index of the feature, and the formula for calculating the Gini index is:

[0111]

[0112] Where Gini(D,A) is the Gini index, Gini(D i ) is the Gini impurity of the subset D i , v is the number of different values, D is the data set, D i is the data subset, |D| is the number of samples in the data set D, |D i | is the number of samples in the data subset D i , and A is the feature.

[0113] If two subsets are obtained after partitioning according to a certain feature, subset 1 has 30 samples with a Gini impurity of 0.4, subset 2 has 70 samples with a Gini impurity of 0.5. The Gini impurity of the original dataset is 0.48 and the total number of samples is 100;

[0114] According to the formula of the Gini index:

[0115] Similarly, calculate the Gini index of each feature according to the above steps.

[0116] In order to better recommend investigation points, among them, S3 determines the investigation points. Taking the Gini index value of the feature as an index, sort the importance of the features, determine the feature points with the highest importance as the points to be investigated, and recommend the investigation points;

[0117] In this automatic investigation point determination model, the feature importance is usually measured by calculating the average impurity reduction of each feature in all decision trees. That is, it is calculated based on the feature importance results of multiple decision trees, called the mean decrease in impurity (MDI). The main calculation process of MDI is to take an average of the feature importance values of multiple decision trees. The decision tree evaluates the importance of features based on the decrease in Gini purity. For the features that divide each node in the decision tree, calculate the Gini index of each feature according to the above steps. Taking the Gini index value of the feature as an index, sort the importance of the features, determine the feature points with the highest importance as the points to be investigated. The automatic investigation point determination model returns the information of whether to investigate and the recommended investigation points to the claims settlement module. If it is recommended to initiate an investigation, it is transferred to the investigation module. After the investigation module finishes processing, it returns to the claims settlement module. At the same time, the investigation data is sent to the model to help the model iterate and update, making the model's prediction ability more and more accurate.

[0118] In order to better perform numerical processing on data, among them, when the data acquisition module 100 performs numerical processing on data, it uses the normalization method to map the feature values into a numerical interval;

[0119] Min-max normalization is one of the most common normalization methods. Its basic principle is to map the feature values of the original data to a specified interval through linear transformation, usually the interval [0,1], which is convenient for subsequent calculation of features.

[0120] In summary, the working principle of this solution is as follows:

[0121] The claims adjustment and investigation location recommendation system based on the random forest algorithm. The data collection module 100 extracts claims information and investigation information from the databases of the claims module and the investigation module, numerically processes the data, and forms a sample data set. The model construction module 200 randomly selects a sample set from the sample data set to construct a decision tree. At each splitting point of the decision tree, a random subset of features is selected, and the best splitting feature is found from it to establish an automatic adjustment model based on the random forest algorithm. The characteristic data of the claims case is substituted into the automatic adjustment model to judge whether the claims case needs to be adjusted, assisting the claims operators to make a decision on initiating an investigation, avoiding inaccurate and ineffective adjustment caused by insufficient experience of the claims operators or system problems. When the automatic adjustment model determines that the claims case needs to be adjusted, the location determination module 300 calculates the Gini index of each feature, ranks the importance of the features based on the Gini index values of the features, determines the most important feature location as the adjustment point, and recommends the investigation location, assisting the investigation personnel to select relevant investigation locations for investigation and improving the effectiveness of the investigation locations. When the claims case does not need to be adjusted, it returns to the claims module, and the case continues the claims process.

[0122] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A claim adjustment and investigation point recommendation system based on random forest algorithm, characterized by: It comprises a data collection module (100), a model building module (200) and a point determination module (300); The data collection module (100) extracts claim settlement information and investigation information from the claim settlement module and investigation module databases, performs numerical processing on the data, and forms a sample data set; The model building module (200) randomly selects a sample set from the sample data set to build a decision tree, randomly selects a feature subset at each split point of the decision tree, and finds the best split feature to establish an automatic training model based on a random forest algorithm, and divides the data set into a training set, a test set and a validation set to train the automatic training model; The determination point module (300) determines whether a claim case needs to be investigated based on the automatic investigation model, performs feature evaluation on the claim case that needs to be investigated, and determines the investigation point.

2. The claim adjustment and investigation point recommendation system based on random forest algorithm according to claim 1 is characterized by: The model building module (200) establishes an automatic adjustment model based on a random forest algorithm, and the method steps are as follows: S2.1.

1. Sample selection: Use the bootstrap sampling method to perform random sampling with replacement on the sample data set; S2.1.2, build a decision tree: select the best split feature from the feature subset, split the node according to the feature, form branches of the tree, each branch forms a new node, and repeat this process; S2.1.

3. Training the automatic training model: Use the training set data and the cross entropy loss function to measure the difference between the predicted probability and the true label to train the automatic training model. S2.1.

4. Evaluate and optimize the automatic tuning model: Use the accuracy of model prediction as the evaluation indicator and use the validation set data to evaluate and optimize the automatic tuning model.

3. The claim adjustment and investigation point recommendation system based on random forest algorithm according to claim 2 is characterized by: S2.1.2 constructs a decision tree and determines the best split feature by calculating the information gain of each feature. The calculation formula of information gain is: Where Gain(D,A) is the information gain, Entropy(D) is the entropy of the data set D, and Entropy(D i ) is the data subset D i The entropy of v is the number of different values, D is the data set, and D i is a subset of data, |D| is the number of samples in the dataset D, | D i | is the data subset D i The number of samples, A is the feature.

4. The claim adjustment and investigation point recommendation system based on random forest algorithm according to claim 2 is characterized by: The S2.1.3 trains the automatic tuning model and adjusts the parameters of the automatic tuning model using a grid search combined with a cross-validation method.

5. The claim adjustment and investigation point recommendation system based on random forest algorithm according to claim 4 is characterized by: S2.1.3 trains the automatic adjustment model with the minimum cross entropy loss function as the goal, and uses the gradient descent algorithm to further optimize the parameters in the model. The formula of the cross entropy loss function is: Among them, L is the cross entropy loss function, n is the claim sample, p i is the probability that needs to be raised, y i is the true label, and i is the number of sample rankings.

6. The claim adjustment and investigation point recommendation system based on random forest algorithm according to claim 1 is characterized by: The point determination module (300) determines the survey point, and the method steps are as follows: S3.1.

1. Calculate the Gini impurity of the original data set; S3.1.

2. Divide the dataset by features and calculate the Gini impurity of the sub-datasets; S3.1.

3. Calculate the Gini index of the feature.

7. The claim adjustment and investigation point recommendation system based on random forest algorithm according to claim 6 is characterized by: S3.1.1 calculates the Gini impurity of the original data set. The formula for calculating the Gini impurity is: Among them, Gini(D) is the Gini impurity, K is the number of categories, D is the data set, k is the category ordinal, and n k is the number of samples in the kth category, and N is the total number of samples.

8. The claim adjustment and investigation point recommendation system based on random forest algorithm according to claim 6 is characterized by: S3.1.3 calculates the Gini index of the feature. The formula for calculating the Gini index is: Among them, Gini(D,A) is the Gini index, Gini(D i ) is a subset D i Gini impurity, v is the number of different values, D is the data set, D i is a subset of data, | D is the number of samples in data set D, | D i | is the data subset D i The number of samples, A is the feature.

9. The claim adjustment and investigation point recommendation system based on random forest algorithm according to claim 6 is characterized by: The point determination module (300) determines the survey points, uses the Gini index value of the feature as an indicator, sorts the importance of the features, determines the most important feature points as the recommended points, and recommends the survey points.

10. The claim adjustment and investigation point recommendation system based on random forest algorithm according to claim 1 is characterized by: When the data acquisition module (100) performs numerical processing on the data, a normalization method is used to map the characteristic value into a numerical range.