Rule and model fusion intelligent decision-making method and system based on multi-source insurance data
Patent Information
- Application Number
- CN202610880078.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]但是还存在如下不足,由上述的陈述可知,该现有方案仅依托理赔请求、同类型历史理赔案件单一数据源开展风险评估,数据维度单一、风险研判依据片面;未搭建刚性业务规则与智能模型并行双判别体系,无法完成拒赔、直赔、待复核三类案件划分以及双维度风险量化评分;且规则判定结果与模型评分缺乏统一风险层级标准,难以量化二者结果分歧,缺失层级差计算与偏差补偿融合机制,大幅降低理赔争议案件判定精度
Smart Images

Figure CN122736780A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of insurance risk control technology, specifically to an intelligent decision-making method and system based on the fusion of rules and models from multi-source insurance data. Background Technology
[0002] Auto insurance claims are a core part of insurance business operations. Leveraging data and intelligent technologies to automatically identify claims risks and make intelligent decisions is key to improving claims efficiency, preventing insurance fraud, and reducing manual operating costs. Currently, most mainstream intelligent claims solutions in the industry use a single model or rule for risk assessment, while some solutions combine risk models with historical data to complete risk evaluation.
[0003] In the prior art, CN110610431A discloses a big data-based intelligent claims settlement method, which includes: obtaining information on pending claims based on received claims request information; identifying the claims type of the pending claims; obtaining information on historical claims consistent with the claims type from a historical claims database; determining the historical risk coefficient of the historical claims consistent with the claims type and the risk coefficient of the pending claims using a risk coefficient model; calculating the risk score of the pending claims and predicting the risk level of the pending claims based on the historical risk coefficient, the risk coefficient of the pending claims, and a preset algorithm; and determining preset rules in the claims processing database that match the pending claims based on the risk level, and sending the claims information to the claims terminal so that the claims terminal can process the claims based on the claims information. This method can improve the accuracy of risk prediction results and the quality of claims services.
[0004] However, the following shortcomings still exist. As can be seen from the above statements, the existing solution relies solely on a single data source—claims requests and historical claims of the same type—to conduct risk assessments. This results in a single data dimension and a one-sided basis for risk judgment. Furthermore, it lacks a dual-judgment system that combines rigid business rules with intelligent models, making it impossible to classify cases into three categories: rejected claims, direct claims, and cases awaiting review, as well as to quantify the risk in both dimensions. Moreover, the rule-based judgment results and model scores lack a unified risk level standard, making it difficult to quantify the discrepancy between the two results. The absence of a mechanism for calculating level differences and compensating for deviations significantly reduces the accuracy of judgments in disputed claims cases.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide an intelligent decision-making method and system based on the fusion of rules and models of multi-source insurance data, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A rule-model fusion intelligent decision-making method based on multi-source insurance data includes the following steps: S1. Collect multi-source heterogeneous insurance data for auto insurance, including policy underwriting data, historical claims data, vehicle condition and traffic violation data, historical repair data of repair shops, and spatiotemporal and meteorological data of accidents. Preprocess the various types of raw data to construct a normalized multi-source feature dataset. S2. Retrieve the normalized multi-source feature dataset of claims pending review, import it into a rigid business rule base built based on the auto insurance industry's claims compliance standards and historical risk control experience, classify the cases into rejected claims, direct claims, and cases pending review according to the rules, and output the rule risk coefficients corresponding to the cases pending review. S3. Input the normalized multi-source feature dataset corresponding to the case to be reviewed into the trained claims risk identification model, and output the risk scores of five dimensions: abnormal amount, accident frequency, vehicle condition violation, abnormal repair, and time, space and weather. Count the number of dimensions falling into the high-risk range, and select the corresponding risk score as the model risk score in combination with the preset risk score matching rules. S4. Pre-set unified multi-level risk intervals for rule risk coefficients and model risk scores, determine the risk intervals to which the rule risk coefficients and model risk scores belong respectively, calculate the interval level difference between the two types of scores, match the deviation compensation coefficient based on the interval level difference, and determine the comprehensive risk level of the case to be reviewed based on the interval level of the rule risk coefficient, the interval level of the model risk score, the interval level difference and the deviation compensation coefficient, so as to determine the final claim risk level and output the corresponding intelligent claim decision.
[0008] Furthermore, the policy underwriting data includes the insured amount and the agreed single claim limit; the historical claims data includes the cumulative number of claims within one year and the time interval between two adjacent claims; the vehicle condition violation data includes the number of violations within 6 months and the length of time the vehicle's annual inspection is overdue; the repair shop's historical repair data includes the amount of a single repair and the number of times the repair was repeated within 3 months; the time and weather data of the accident include the time period of the accident and the rainfall at the time of the accident. The normalized multi-source feature dataset includes the insured amount, agreed single claim limit, cumulative number of accidents within one year, time interval between two adjacent accidents, number of traffic violations within 6 months, vehicle annual inspection overdue duration, single repair amount, number of repeated repairs within 3 months, time of accident, and rainfall at the time of the accident.
[0009] Furthermore, the normalized multi-source feature dataset of claims pending review is imported into a rigid business rule base built based on auto insurance industry claims compliance standards and historical risk control experience. Cases are then categorized into three types according to these rules, with the specific logic as follows: The rigid business rule base includes two types of fixed judgment rules: multiple denial judgment rules and multiple direct payment judgment rules. The judgment logic of each judgment rule cannot be manually and dynamically adjusted. Compare the feature parameters of the normalized multi-source feature dataset of claims pending review with the rules in the rigid business rule base: When the characteristic parameters of a claim case pending review meet any one of the denial judgment rules in the database, the case is classified as a denial case. When all characteristic parameters of a claim case pending review meet the direct payment determination rules, the case will be classified as a direct payment case. When the characteristic parameters of a claim pending review do not trigger the arbitrary denial rule, nor do they fully meet the direct payment rule, the case is classified as a case pending review.
[0010] Furthermore, the rule risk coefficient corresponding to the case pending review is output, with the specific logic as follows: The cases pending review did not trigger the arbitrary denial of claim rules, and only some characteristic parameters did not meet the direct claim rules. Based on the abnormal feature parameters in the case pending review that do not meet the direct claim determination rules, the number of abnormal feature parameters in the case is counted. For each abnormal feature parameter, the absolute value of the difference between the actual value of the parameter and the determination standard value of the direct claim determination rule is calculated to obtain the deviation range of the corresponding abnormal feature parameter. The rule risk coefficient corresponding to the case pending review is obtained by weighting the number of abnormal feature parameters and the deviation range of each abnormal feature parameter.
[0011] Furthermore, the number of dimensions falling into the high-risk range is counted, and the corresponding risk score is selected based on the preset risk score matching rules as the model's risk score. The specific logic is as follows: Each of the five risk categories—abnormal amount, frequency of accidents, vehicle condition violations, abnormal maintenance, and time, space and weather—is assessed to determine whether its risk score falls into the preset high-risk range, and the number of dimensions falling into the high-risk range is counted. The preset risk score matching rules are as follows: The high-risk zone is pre-divided into three zones based on the number of dimensions. When the number of dimensions in the high-risk zone is 0 to 1, it is classified as a low-risk zone; when the number of dimensions in the high-risk zone is 2 to 3, it is classified as a medium-risk zone; and when the number of dimensions in the high-risk zone is 4 to 5, it is classified as a high-risk zone. Each zone corresponds to a preset risk score. Based on the statistically obtained number of dimensions and risk score matching rules of the high-risk intervals, the corresponding level is matched, and the risk score corresponding to that level is selected as the model risk score.
[0012] Furthermore, a unified multi-level risk range is pre-defined for the rule risk coefficient and the model risk score, and the risk range to which the rule risk coefficient and the model risk score belong are determined respectively. The specific logic is as follows: Three risk levels are pre-defined: Level 1, Level 2, and Level 3. The corresponding levels are denoted as Level 1, Level 2, and Level 3, respectively. This set of levels applies to both rule-based risk coefficients and model risk scores. The rule risk coefficient and model risk score of the reviewed case are compared with the interval thresholds of each level of risk interval. The risk interval to which the rule risk coefficient and model risk score belong are determined in turn, and the corresponding interval level is recorded.
[0013] Furthermore, the interval level difference between the two types of scores is calculated, and the deviation compensation coefficient is matched based on the interval level difference. The specific logic is as follows: Calculate the absolute value of the difference between the rule risk coefficient and the model risk score corresponding to the respective interval levels to obtain the interval level difference; Establish a one-to-one correspondence between interval level differences and deviation compensation coefficients in advance: When the interval level difference is 0, the first deviation compensation coefficient is matched. ; When the interval level difference is 1, the second deviation compensation coefficient is matched. ; When the interval level difference is 2, the third deviation compensation coefficient is matched. ; in, ; Based on the correspondence between the interval level difference of the rule risk coefficient and the model risk score and the above, the corresponding deviation compensation coefficient is retrieved.
[0014] Furthermore, the comprehensive risk level of the case to be reviewed is determined based on the interval level of the rule risk coefficient, the interval level of the model risk score, the interval level difference, and the deviation compensation coefficient. The specific logic is as follows:
[0015] in, The overall risk level of the cases pending review. For the interval level of the rule risk coefficient, For the risk scoring interval levels of the model, This is the deviation compensation coefficient. .
[0016] Furthermore, the final claim risk level is determined, and a corresponding intelligent claim decision is output. The specific logic is as follows: The calculated comprehensive risk level is compared with the first and second preset thresholds to classify claims risk into three categories: low risk, medium risk, and high risk, as detailed below: When the overall risk level is less than the first preset threshold, it is determined to be a low-risk level; When the overall risk level is between the first and second preset thresholds, it is determined to be a medium risk level; When the overall risk level is greater than the second preset threshold, it is determined to be a high-risk level; Wherein, the first preset threshold is less than the second preset threshold; Based on the determined claim risk level, a matching intelligent claim decision is output.
[0017] To achieve the above objectives, the present invention also provides the following technical solution: A rule-model fusion intelligent decision-making system based on multi-source insurance data, the system being used to execute any of the aforementioned rule-model fusion intelligent decision-making methods based on multi-source insurance data, comprising: The data collection module is used to collect multi-source heterogeneous insurance data for auto insurance, including policy underwriting data, historical claims data, vehicle condition and traffic violation data, historical repair data from repair shops, and spatiotemporal and meteorological data of accidents. It preprocesses various types of raw data to construct a normalized multi-source feature dataset. The rule discrimination module is used to retrieve the normalized multi-source feature dataset of claims pending review, import it into a rigid business rule library built based on the auto insurance industry's claims compliance standards and historical risk control experience, and classify the cases into rejected claims, direct claims, and cases pending review according to the rules, while outputting the rule risk coefficient corresponding to the cases pending review. The model scoring module is used to input the normalized multi-source feature dataset corresponding to the case to be reviewed into the trained claims risk identification model, and output risk scores for five dimensions: abnormal amount, accident frequency, vehicle condition violations, abnormal repair, and spatiotemporal and meteorological. It counts the number of dimensions falling into the high-risk range, and selects the corresponding risk score as the model risk score by combining the preset risk score matching rules. The integrated decision-making module is used to pre-set unified multi-level risk intervals for rule risk coefficients and model risk scores. It determines the risk intervals to which the rule risk coefficients and model risk scores belong, calculates the interval level difference between the two types of scores, matches the deviation compensation coefficient based on the interval level difference, and determines the comprehensive risk level of the case to be reviewed based on the interval level of the rule risk coefficient, the interval level of the model risk score, the interval level difference, and the deviation compensation coefficient. This determines the final claim risk level and outputs the corresponding intelligent claim decision.
[0018] Compared with the prior art, the beneficial effects of the present invention are: This invention overcomes the shortcomings of existing technologies, such as single data sources and one-sided risk assessment criteria, by aggregating five types of multi-source heterogeneous data on auto insurance: policy underwriting, historical claims, vehicle condition and violations, repairs, and weather conditions at the time of the accident, and constructing a normalized feature dataset. It establishes a rigid business rule base to automatically classify cases into three categories: rejected claims, direct claims, and cases awaiting review, and outputs rule-based risk coefficients. Furthermore, it utilizes a claims risk identification model to output risk scores across five dimensions and aggregates these scores to obtain a model risk score. This constructs a parallel dual-discrimination system of rules and models, achieving dual-system risk quantification and scoring, thus solving the problem of existing technologies lacking a parallel discriminant system. This addresses the issues of inability to finely classify cases and the lack of two-way scoring in existing technologies. By setting unified multi-level risk intervals for rule-based risk coefficients and model-based risk scores, calculating the level difference between the two types of scoring intervals and matching corresponding deviation compensation coefficients, and integrating the two levels and compensation coefficients to solve for the comprehensive risk level and output claims decisions, the discrepancies between the two types of risk control results are eliminated. This addresses the shortcomings of existing technologies, such as the lack of a unified risk level, the inability to quantify result deviations, and the lack of a deviation compensation fusion mechanism. It effectively optimizes the identification effect of claims borderline cases and disputed cases where rule and model judgment results contradict each other, and improves the reliability of intelligent decision-making in auto insurance claims. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the overall method flow of the present invention; Figure 2 This is a block diagram of the module composition of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0021] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0022] Example: Please see Figure 1 The present invention provides a technical solution: A rule-model fusion intelligent decision-making method based on multi-source insurance data includes the following steps: S1. Collect multi-source heterogeneous auto insurance data, including policy underwriting data, historical claims data, vehicle condition and violation data, historical repair data from repair shops, and spatiotemporal and meteorological data related to accidents. Preprocess all types of raw data to construct a normalized multi-source feature dataset. The collection process involves unified aggregation, integration, and alignment across platforms, systems, and ledgers.
[0023] Based on the above embodiments, the policy underwriting data includes the insured amount and the agreed single claim limit; the historical claims data includes the cumulative number of accidents within one year and the time interval between two adjacent accidents; the vehicle condition violation data includes the number of violations within 6 months and the length of time the vehicle's annual inspection is overdue; the repair shop's historical repair data includes the amount of a single repair and the number of times the repair was repeated within 3 months; the accident time and weather data includes the time period of the accident and the rainfall at the time of the accident. Based on the above embodiments, the specific methods for obtaining each characteristic parameter are as follows: Insured amount and agreed single-claim limit: The original policy filing ledger is retrieved from the vehicle insurance underwriting business backend database, and the underwriting assessment parameters entered in the policy signing are directly extracted; Cumulative number of claims within one year and time interval between two adjacent claims: Historical claim archives bound to the vehicle involved are retrieved from the insurance company's claims business database, and the time-series claim parameters are calculated statistically; Number of traffic violations within 6 months and vehicle annual inspection overdue duration: Vehicle compliance supervision ledger data is retrieved synchronously by connecting to the traffic management vehicle supervision open interface and the motor vehicle annual inspection filing system; Single repair amount and number of repeated repairs within 3 months: Historical repair and damage assessment work order archives are extracted by connecting to the cooperative repair shops and the auto repair damage assessment platform database; Accident time and rainfall at the time of the accident: The time of the accident is automatically determined by capturing the timestamp from the claims reporting system when the case is reported; The regional meteorological monitoring database is linked to match the accident location and time, and corresponding real-time rainfall meteorological data is retrieved synchronously.
[0024] This includes preprocessing various types of raw data, such as data cleaning, data alignment, missing value imputation, outlier removal, and unit normalization. Specifically, the process involves: cleaning the data by removing duplicate claims and invalid null values; aligning heterogeneous data from different formats and collection ports according to unique case numbers; filling missing values for numerical features such as insured amount, number of claims, and repair costs with the average values of the same vehicle model and claim scenario; removing abnormal data such as overdue annual inspection time and number of traffic violations exceeding reasonable business thresholds; and finally, normalizing all feature parameters to eliminate differences in data magnitude and generate a normalized multi-source feature dataset.
[0025] The normalized multi-source feature dataset includes the insured amount, agreed single claim limit, cumulative number of accidents within one year, time interval between two adjacent accidents, number of traffic violations within 6 months, vehicle annual inspection overdue duration, single repair amount, number of repeated repairs within 3 months, time of accident, and rainfall at the time of the accident.
[0026] S2. Retrieve the normalized multi-source feature dataset of claims pending review, import it into a rigid business rule base built based on the auto insurance industry's claims compliance standards and historical risk control experience, classify the cases into rejected claims, direct claims, and cases pending review according to the rules, and output the rule risk coefficients corresponding to the cases pending review. The specific construction process of the rigid business rule library is as follows: Combining the auto insurance regulatory compliance provisions, the unified claims review standards of the auto insurance industry, and the insurance company's more than ten years of historical claims risk control and case closure experience, as well as the sample data of fraud cases, the rule library is built by binding the normalized feature parameters of this invention. First, the legally mandated refusal compliance clauses and the bottom line standards for direct claims access in auto insurance are sorted out to extract hard compliance constraints. Second, typical cases of historical fraud, malicious claims, compliant direct claims, and manual review are summarized to extract high consensus and undisputed risk control judgment thresholds. Then, fixed refusal judgment rule groups and direct claims judgment rule groups are generated, and all rule thresholds and judgment logic are solidified and encapsulated. Finally, the rules are locked so that they cannot be manually modified or temporarily adjusted, forming a closed rigid business rule library. Among them, the refusal rules are built around excessive compensation, overdue annual inspection, short-term intensive accidents, and repeated malicious repairs, while the direct claims rules are built around low violations, low-frequency accidents, compliant insurance amounts, routine repairs, and reasonable accident environments.
[0027] Based on the above embodiments, the normalized multi-source feature dataset of claims pending review is imported into a rigid business rule base constructed according to the auto insurance industry's claims compliance standards and historical risk control experience. Cases are then categorized into three types based on these rules, with the specific logic as follows: The rigid business rule base includes two types of fixed judgment rules: multiple rules for denying claims and multiple rules for direct claims. Compare the feature parameters of the normalized multi-source feature dataset of claims pending review with the rules in the rigid business rule base: When the characteristic parameters of a claim pending review meet any of the rejection criteria in the database, it means that the case has touched the red line of claim compliance, has a legally deficient qualification for rejection, or has a significant risk of malicious insurance fraud, and the case is classified as a rejected case. When all characteristic parameters of a claim case pending review meet the direct payment determination rules, it means that the case is compliant with all-dimensional risk control indicators, has no abnormal claims risks, and belongs to a low-risk, high-quality claim case. The case will be classified as a direct payment case. When the characteristic parameters of a pending claim case neither trigger the arbitrary denial judgment rule nor fully meet the direct payment judgment rule, it means that the case has no fatal compliance defects, but there are abnormalities in some characteristic parameters, and the risk level is at a critical state. The case is classified as a pending review case. Among them, the cases pending review have no violations of claim denial regulations, only the direct payment indicators are not up to standard, and the risks all come from the characteristic parameters that do not meet the direct payment determination rules. Therefore, the characteristic parameters of these cases are selected to calculate the rule risk coefficient.
[0028] Based on the above embodiments, the rule risk coefficient corresponding to the case to be reviewed is output, and the specific logic is as follows: The cases pending review did not trigger the arbitrary denial of claim rules, and only some characteristic parameters did not meet the direct claim rules. Based on the abnormal characteristic parameters in the pending review case that do not meet the direct claim determination rules, the number of abnormal characteristic parameters in the case is counted. For each abnormal characteristic parameter, the absolute value of the difference between the actual value of the parameter and the determination standard value of the direct claim determination rule is calculated to obtain the deviation range of the corresponding abnormal characteristic parameter. Combining the number of abnormal characteristic parameters and the deviation range of each abnormal characteristic parameter, a weighted calculation is performed to obtain the rule risk coefficient corresponding to the pending review case. The formula used is as follows:
[0029] in, The rule risk coefficient is used to assess the risk control risk of the case in the rule dimension by combining two indicators: the number of abnormal feature parameters and the deviation of each abnormal feature parameter. The larger the rule risk coefficient, the greater the risk control risk in the rule dimension of the case. In the formula, The number of abnormal characteristic parameters in cases pending review that do not meet the direct compensation determination rules. For the first The deviation of the abnormal feature parameters, For the first The actual values of the abnormal feature parameters. For the first The judgment standard value of the direct compensation determination rule corresponding to the abnormal feature parameter. For the index of the abnormal feature parameters, and .
[0030] The above formula couples two core risk control indicators: the number of abnormal feature parameters and the deviation of a single feature parameter. It is highly positively correlated with the rule risk coefficient and perfectly matches the risk quantification logic of the case to be reviewed. Among them, the number of abnormal features represents the overall abnormal coverage of the case. The more abnormal parameters there are, the more compliance defects the case has, the higher the basic risk control risk, and the higher the rule risk coefficient is. The deviation of a single feature... It represents the degree to which a single feature deviates from the standard value of the direct compensation determination rule. The more obvious the deviation of the parameter from the standard value of the rule, the more prominent the single risk control risk, which amplifies the overall rule risk of the case. Therefore, the more significant the deviation of abnormal feature parameters from the standard value of the direct compensation determination rule, the more prominent the individual risk control risks, which will amplify the overall rule risk of the case. Since the number of abnormal feature parameters in the case to be reviewed is not fixed, the total deviation of all abnormal features can be collected by summation, which measures the severity of the case's overall parameter exceedance. At the same time, all non-compliant abnormal features contribute to the risk control risk, and the summation can objectively summarize the overall rule violation degree of the case. This weighted calculation formula can accurately quantify the synergistic impact of two types of indicators, namely the number of abnormal features and the magnitude of parameter deviation, on risk control risk, realize the standardized and quantitative solution of rule risk coefficients, objectively characterize the rule dimension risk level of the case to be reviewed, avoid the problems of strong subjectivity and inconsistent standards in manual scoring, and also provide accurate data support for subsequent rule risk level classification, model scoring fusion and compensation calculation.
[0031] In the formula, For the weighting coefficients of the number of abnormal feature parameters and the weighting coefficients of the deviation magnitude of abnormal feature parameters; The number of abnormal features only represents the coverage of abnormal indicators, belonging to the macro-level basic risk, while the magnitude of parameter deviation represents the severity of the indicators exceeding the limits, belonging to the core substantive risk control risks. The weight of substantive risk should be higher than that of quantity risk, conforming to the calibration pattern of historical claims risk control samples. In this case, let .
[0032] As one implementation method, The weight range is , The weight range is The specific value is set by those skilled in the art based on the actual situation, and is not limited here.
[0033] S3. Input the normalized multi-source feature dataset corresponding to the case to be reviewed into the trained claims risk identification model, and output the risk scores of five dimensions: abnormal amount, accident frequency, vehicle condition violation, abnormal repair, and time, space and weather. Count the number of dimensions falling into the high-risk range, and select the corresponding risk score as the model risk score in combination with the preset risk score matching rules. Based on the above embodiments, the claims risk identification model is constructed using a deep learning network based on a multilayer perceptron. The deep neural network of the multilayer perceptron includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer. The first hidden layer, the second hidden layer, and the third hidden layer each have at least two neurons and all use ReLU (Rectified Linear Unit) as the activation function. Multiple sets of normalized multi-source feature datasets of historical auto insurance claims are collected, and corresponding five-dimensional manually labeled real risk scores are used to construct a model training sample set. The samples are divided into a training set, a validation set, and a test set in a 7:2:1 ratio. The training set is used for iterative learning of the model network parameters; the validation set is used to adjust the network hyperparameters during training, suppress model overfitting, and improve the fitting effect of risk control features; the test set is used to evaluate the generalization discrimination ability and risk scoring accuracy of the claims risk identification model after training. The deep learning network structure of the multilayer perceptron is as follows: Input layer: receives all feature parameters from the normalized multi-source feature dataset of the case to be reviewed; First hidden layer: has 128 neurons, using ReLU as the activation function; Second hidden layer: has 64 neurons, also using the ReLU activation function; Third hidden layer: has 32 neurons, using the ReLU activation function; Output layer: has 5 neurons, corresponding to the risk scores of five dimensions: abnormal output amount, accident frequency, vehicle condition violations, abnormal maintenance, and spatiotemporal and meteorological factors; The training process of the deep learning model for claims risk identification is as follows: The model input is a multi-source feature parameter normalized from multiple sets of historical auto insurance claims. The real risk scores of the five dimensions, which are manually verified and labeled according to the corresponding cases, are used as real supervision labels. The mean absolute error loss function is selected, and the difference between the five-dimensional risk scores predicted by the model and the corresponding real supervision labels is calculated to obtain the overall loss value of the model. The deep learning model network parameters are iteratively updated through the backpropagation algorithm. When the model loss value is within the preset convergence interval and the loss does not decrease significantly for 20 consecutive training rounds, the claims risk identification model is determined to have converged and training is stopped, thus obtaining the trained claims risk identification model.
[0034] Based on the above embodiments, the number of dimensions falling into the high-risk range is counted, and the corresponding risk score is selected as the model risk score by combining the preset risk score matching rules. The specific logic is as follows: Each of the five risk categories—abnormal amount, frequency of accidents, vehicle condition violations, abnormal maintenance, and time, space and weather—is assessed to determine whether its risk score falls into the preset high-risk range, and the number of dimensions falling into the high-risk range is counted. The preset risk score matching rules are as follows: Based on the number of dimensions in the high-risk range, the case is pre-divided into three ranges. When the number of dimensions in the high-risk range is 0-1, only a single dimension exhibits minor risk anomalies, while the other four core risk control dimensions are compliant and meet standards. The overall probability of fraudulent claims, collusion, and irregular claims is extremely low, and the overall risk is minimal; therefore, it is classified as low-risk. When the number of dimensions in the high-risk range is 2-3, approximately half of the risk control modules in the case simultaneously show abnormal risks, indicating multiple related risk control vulnerabilities and suspected deliberate manipulation of claim information or localized irregular claims. The risk is at a critical level, matching the characteristics of cases awaiting review, and is therefore classified as medium risk. When the number of dimensions in the high-risk range is 4 to 5, more than half or even all of the risk control dimensions simultaneously trigger high-risk judgments. The multi-dimensional risks are cross-coupled, which belongs to typical high-risk claims cases such as group insurance fraud, false claims, and repair shops colluding to abscond with insurance. The probability of illegal claims is extremely high, so it is classified as high risk. Each level corresponds to a preset risk score. This segmentation method conforms to the superposition law of risk control in auto insurance, and the number of abnormal dimensions is positively correlated with the overall risk control risk of the case model. Based on the statistically obtained number of dimensions and risk score matching rules of the high-risk intervals, the corresponding level is matched, and the risk score corresponding to that level is selected as the model risk score.
[0035] S4. Pre-set unified multi-level risk intervals for rule risk coefficients and model risk scores, determine the risk intervals to which the rule risk coefficients and model risk scores belong respectively, calculate the interval level difference between the two types of scores, match the deviation compensation coefficient based on the interval level difference, and determine the comprehensive risk level of the case to be reviewed based on the interval level of the rule risk coefficient, the interval level of the model risk score, the interval level difference and the deviation compensation coefficient, so as to determine the final claim risk level and output the corresponding intelligent claim decision.
[0036] Based on the above embodiments, a unified multi-level risk interval is pre-set for the rule risk coefficient and the model risk score, and the risk interval to which the rule risk coefficient and the model risk score belong are determined respectively. The specific logic is as follows: Three risk levels are pre-defined: Level 1, Level 2, and Level 3. The corresponding levels are denoted as Level 1, Level 2, and Level 3, respectively. This set of levels applies to both rule-based risk coefficients and model risk scores. The rule risk coefficient and model risk score of the reviewed case are compared with the interval thresholds of each level of risk interval. The risk interval to which the rule risk coefficient and model risk score belong are determined in turn, and the corresponding interval level is recorded.
[0037] Interval Threshold Setting Instructions: Both the rule risk coefficient and the model risk score are normalized to the 0-1 interval. Combining historical risk control samples of auto insurance, model risk levels, and rule risk intensity, two levels of critical thresholds are evenly optimized and defined. The two indicators share the same set of thresholds to divide the three-level risk intervals, avoiding overlapping and gaps in classification, and adapting to subsequent level difference and compensation coefficient calculations. The critical thresholds are set at 0.35 and 0.70; corresponding to: Level 1 risk interval: 0-0.35, Level 2 risk interval: 0.35-0.70, Level 3 risk interval: 0.70-1.0.
[0038] Based on the above, it should be noted that: The purpose of setting up a unified multi-level risk range is to eliminate the differences in the dimensions and evaluation benchmarks of two heterogeneous risk control indicators, namely rule-based risk coefficients and model-based risk scores, and to solve the pain point of the existing technology's fragmented dual risk control results standards and inability to be benchmarked and compared. At the same time, it provides a standardized calculation benchmark for the calculation of hierarchical differences and the matching of deviation compensation coefficients, quantifying the discrepancies in the judgments of the two systems. It adapts to the risk control needs of claims borderline cases and disputed cases, and completes two-way deviation compensation based on hierarchical differences, reducing the error of a single judgment system. Moreover, the interval classification is highly compatible with the rule-based risk coefficients and model-based risk score levels, which fits the business logic of auto insurance risk control and ensures more stable and unified output of comprehensive risk classification and claims decision-making.
[0039] Based on the above embodiments, the interval level difference between the two types of scores is calculated, and the deviation compensation coefficient is matched based on the interval level difference. The specific logic is as follows: Calculate the absolute value of the difference between the rule risk coefficient and the model risk score corresponding to the respective interval levels to obtain the interval level difference; The interval hierarchy has only three levels: 1, 2, and 3. The level difference has only three possible results: 0, 1, and 2. A one-to-one correspondence between the interval level difference and the deviation compensation coefficient is established in advance. When the interval level difference is 0, the two risk control decisions are completely consistent, with no discrepancy error, and the first deviation compensation coefficient is matched. ; When the interval level difference is 1, the dual-system has a slight judgment bias, and the second bias compensation coefficient is matched. ; When the interval level difference is 2, the risk ratings of the two systems are completely contradictory, and the case discrepancies reach extreme values. A third deviation compensation coefficient is then used to match this. ; Based on the correspondence between the interval level difference of the rule risk coefficient and the model risk score and the above, the corresponding deviation compensation coefficient is retrieved.
[0040] Based on the above, it should be noted that: The larger the hierarchical difference value, the greater the discrepancy between the risk assessment conclusions of rigid rule-based risk control and claims risk identification model-based risk control, and the more significant the deviation in the judgment of cases pending review. Therefore, a greater degree of deviation correction compensation is required. Thus, setting the first, second, and third deviation compensation coefficients satisfies... Gradient increasing relationship; combining the massive historical review cases of auto insurance, iterative calibration of fraud samples, and fusion debugging of risk control models in this invention, a set of optimal compliance values is given: First deviation compensation coefficient Second deviation compensation coefficient Third deviation compensation coefficient .
[0041] Based on the above embodiments, the comprehensive risk level of the case to be reviewed is determined based on the interval level of the rule risk coefficient, the interval level of the model risk score, the interval level difference, and the deviation compensation coefficient. The specific logic is as follows:
[0042] in, The overall risk level of the cases pending review. For the interval level of the rule risk coefficient, For the risk scoring interval levels of the model, This is the deviation compensation coefficient. .
[0043] Based on the above, it should be noted that: Formula Antecedent To obtain the average benchmark values of the two risk control levels, namely rigid rules and claims risk identification models, the average benchmark values of the two risk control levels are calculated to achieve basic integration of the rule level and the model level. This takes into account the prior risk control experience of business rules and the risk control results driven by model data, and avoids the one-sidedness of a single judgment system. The interval level difference obtained in the previous solution accurately characterizes the degree of divergence in the risk rating of the two systems, and the embedded formula realizes the quantitative coupling of divergence measurement; the coupling deviation compensation coefficient Complete the bias weighted correction, combined with Gradient coefficient characteristics: the larger the level difference, the higher the deviation compensation coefficient. The larger the value, the stronger the correction for the bias in the dual-system judgment, which aligns with the correction needs of cases awaiting review. Therefore, the above function is used to calculate the comprehensive risk level of cases awaiting review.
[0044] Based on the above embodiments, the final claim risk level is determined, and the corresponding intelligent claim decision is output. The specific logic is as follows: The calculated comprehensive risk level is compared with the first and second preset thresholds to classify claims risk into three categories: low risk, medium risk, and high risk, as detailed below: When the overall risk level is less than the first preset threshold, it is determined to be a low-risk level; When the overall risk level is between the first and second preset thresholds, it is determined to be a medium risk level; When the overall risk level is greater than the second preset threshold, it is determined to be a high-risk level; Wherein, the first preset threshold is less than the second preset threshold; Based on the determined claim risk level, a matching intelligent claim decision is output. The matching logic is as follows: Low-risk levels correspond to intelligent claims decision-making: the risk of cases is extremely low in both dimensions, there are no malicious claims or abnormal qualifications, the system skips manual review, automatically completes the claims approval and payment transfer, realizes automatic payment, improves claims efficiency and saves manual review costs.
[0045] Medium-risk level corresponds to intelligent claims decision-making: cases pending review have minor abnormalities, slight discrepancies in the dual-system judgment, and no evidence of insurance fraud; automatic claims are suspended, and only online special reviews are conducted. After verification, the final manual review and claims payment are made, balancing review efficiency and risk control security.
[0046] High-risk levels correspond to intelligent claims decision-making: Cases awaiting review have multiple abnormally coupled features, with extremely high risks of malicious insurance fraud and joint insurance fraud; the automatic claims channel is closed, and in-depth manual review is issued, with multiple parties involved in tracing the source and collecting evidence; if violations are found, the claims application is rejected, and special claims approval is initiated only after no abnormalities are found.
[0047] Specifically, by combining the original Level 1, Level 2, and Level 3 basic risk level benchmarks, the vehicle insurance risk control business classification standards, and the historical review case sample calibration, the first preset threshold and the second preset threshold are set with the midpoint of the basic level as the dividing benchmark. The three-level risk classification is completed based on the dual benchmark thresholds, and is aligned one by one with the original risk levels.
[0048] Please see Figure 2 The present invention also provides a technical solution: A rule-model fusion intelligent decision-making system based on multi-source insurance data, the system being used to execute any of the aforementioned rule-model fusion intelligent decision-making methods based on multi-source insurance data, comprising: The data collection module is used to collect multi-source heterogeneous insurance data for auto insurance, including policy underwriting data, historical claims data, vehicle condition and traffic violation data, historical repair data from repair shops, and spatiotemporal and meteorological data of accidents. It preprocesses various types of raw data to construct a normalized multi-source feature dataset. The rule discrimination module is used to retrieve the normalized multi-source feature dataset of claims pending review, import it into a rigid business rule library built based on the auto insurance industry's claims compliance standards and historical risk control experience, and classify the cases into rejected claims, direct claims, and cases pending review according to the rules, while outputting the rule risk coefficient corresponding to the cases pending review. The model scoring module is used to input the normalized multi-source feature dataset corresponding to the case to be reviewed into the trained claims risk identification model, and output risk scores for five dimensions: abnormal amount, accident frequency, vehicle condition violations, abnormal repair, and spatiotemporal and meteorological. It counts the number of dimensions falling into the high-risk range, and selects the corresponding risk score as the model risk score by combining the preset risk score matching rules. The integrated decision-making module is used to pre-set unified multi-level risk intervals for rule risk coefficients and model risk scores. It determines the risk intervals to which the rule risk coefficients and model risk scores belong, calculates the interval level difference between the two types of scores, matches the deviation compensation coefficient based on the interval level difference, and determines the comprehensive risk level of the case to be reviewed based on the interval level of the rule risk coefficient, the interval level of the model risk score, the interval level difference, and the deviation compensation coefficient. This determines the final claim risk level and outputs the corresponding intelligent claim decision.
[0049] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0050] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by software, electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0051] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0052] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A rule and model fusion intelligent decision-making method based on multi-source insurance data, characterized in that, The specific steps include: S1. Collect multi-source heterogeneous insurance data for auto insurance, including policy underwriting data, historical claims data, vehicle condition and traffic violation data, historical repair data of repair shops, and spatiotemporal and meteorological data of accidents. Preprocess the various types of raw data to construct a normalized multi-source feature dataset. S2. Retrieve the normalized multi-source feature dataset of claims pending review, import it into a rigid business rule base built based on the auto insurance industry's claims compliance standards and historical risk control experience, classify the cases into rejected claims, direct claims, and cases pending review according to the rules, and output the rule risk coefficients corresponding to the cases pending review. S3. Input the normalized multi-source feature dataset corresponding to the case to be reviewed into the trained claims risk identification model, and output the risk scores of five dimensions: abnormal amount, accident frequency, vehicle condition violation, abnormal repair, and time, space and weather. Count the number of dimensions falling into the high-risk range, and select the corresponding risk score as the model risk score in combination with the preset risk score matching rules. S4. Pre-set unified multi-level risk intervals for rule risk coefficients and model risk scores, determine the risk intervals to which the rule risk coefficients and model risk scores belong respectively, calculate the interval level difference between the two types of scores, match the deviation compensation coefficient based on the interval level difference, and determine the comprehensive risk level of the case to be reviewed based on the interval level of the rule risk coefficient, the interval level of the model risk score, the interval level difference and the deviation compensation coefficient, so as to determine the final claim risk level and output the corresponding intelligent claim decision. 2.The method of claim 1, wherein, The policy underwriting data includes the insured amount and the agreed single claim limit; the historical claims data includes the cumulative number of claims within one year and the time interval between two adjacent claims; the vehicle condition violation data includes the number of violations within 6 months and the length of time the vehicle's annual inspection is overdue; the repair shop's historical repair data includes the amount of a single repair and the number of times the repair was repeated within 3 months; the time and weather data of the accident include the time period of the accident and the rainfall at the time of the accident. The normalized multi-source feature dataset includes the insured amount, agreed single claim limit, cumulative number of accidents within one year, time interval between two adjacent accidents, number of traffic violations within 6 months, vehicle annual inspection overdue duration, single repair amount, number of repeated repairs within 3 months, time of accident, and rainfall at the time of the accident. 3.The method of claim 2, wherein, The normalized multi-source feature dataset of claims pending review is imported into a rigid business rule base built based on auto insurance industry claims compliance standards and historical risk control experience. Cases are then categorized into three types according to these rules, with the specific logic as follows: The rigid business rule base includes two types of fixed judgment rules: multiple denial judgment rules and multiple direct payment judgment rules. The judgment logic of each judgment rule cannot be manually and dynamically adjusted. Compare the feature parameters of the normalized multi-source feature dataset of claims pending review with the rules in the rigid business rule base: When the characteristic parameters of a claim case pending review meet any one of the denial judgment rules in the database, the case is classified as a denial case. When all characteristic parameters of a claim case pending review meet the direct payment determination rules, the case will be classified as a direct payment case. When the characteristic parameters of a claim pending review do not trigger the arbitrary denial rule, nor do they fully meet the direct payment rule, the case is classified as a case pending review.
4. The intelligent decision-making method based on rule and model fusion of multi-source insurance data according to claim 3, characterized in that, Output the rule risk coefficient corresponding to the case pending review. The specific logic is as follows: The cases pending review did not trigger the arbitrary denial of claim rules, and only some characteristic parameters did not meet the direct claim rules. Based on the abnormal feature parameters in the case pending review that do not meet the direct claim determination rules, the number of abnormal feature parameters in the case is counted. For each abnormal feature parameter, the absolute value of the difference between the actual value of the parameter and the determination standard value of the direct claim determination rule is calculated to obtain the deviation range of the corresponding abnormal feature parameter. The rule risk coefficient corresponding to the case pending review is obtained by weighting the number of abnormal feature parameters and the deviation range of each abnormal feature parameter.
5. The intelligent decision-making method based on rule and model fusion of multi-source insurance data according to claim 1, characterized in that, The number of dimensions falling into the high-risk range is counted, and the corresponding risk score is selected based on the preset risk score matching rules as the model's risk score. The specific logic is as follows: Each of the five risk categories—abnormal amount, frequency of accidents, vehicle condition violations, abnormal maintenance, and time, space and weather—is assessed to determine whether its risk score falls into the preset high-risk range, and the number of dimensions falling into the high-risk range is counted. The preset risk score matching rules are as follows: The high-risk zone is pre-divided into three zones based on the number of dimensions. When the number of dimensions in the high-risk zone is 0 to 1, it is classified as a low-risk zone; when the number of dimensions in the high-risk zone is 2 to 3, it is classified as a medium-risk zone; and when the number of dimensions in the high-risk zone is 4 to 5, it is classified as a high-risk zone. Each zone corresponds to a preset risk score. Based on the statistically obtained number of dimensions and risk score matching rules of the high-risk intervals, the corresponding level is matched, and the risk score corresponding to that level is selected as the model risk score.
6. The intelligent decision-making method based on rule and model fusion of multi-source insurance data according to claim 1, characterized in that, A unified multi-level risk range is pre-defined for the rule risk coefficient and the model risk score. The risk range to which the rule risk coefficient and the model risk score belong is then determined separately. The specific logic is as follows: Three risk levels are pre-defined: Level 1, Level 2, and Level 3. The corresponding levels are denoted as Level 1, Level 2, and Level 3, respectively. This set of levels applies to both rule-based risk coefficients and model risk scores. The rule risk coefficient and model risk score of the reviewed case are compared with the interval thresholds of each level of risk interval. The risk interval to which the rule risk coefficient and model risk score belong are determined in turn, and the corresponding interval level is recorded.
7. The intelligent decision-making method based on rule and model fusion of multi-source insurance data according to claim 6, characterized in that, The interval level difference between the two types of scores is calculated, and the deviation compensation coefficient is matched based on the interval level difference. The specific logic is as follows: Calculate the absolute value of the difference between the rule risk coefficient and the model risk score corresponding to the respective interval levels to obtain the interval level difference; Establish a one-to-one correspondence between interval level differences and deviation compensation coefficients in advance: When the interval level difference is 0, the first deviation compensation coefficient is matched. ; When the interval level difference is 1, the second deviation compensation coefficient is matched. ; When the interval level difference is 2, the third deviation compensation coefficient is matched. ; in, ; Based on the correspondence between the interval level difference of the rule risk coefficient and the model risk score and the above, the corresponding deviation compensation coefficient is retrieved.
8. The intelligent decision-making method based on rule and model fusion of multi-source insurance data according to claim 7, characterized in that, The comprehensive risk level of a case to be reviewed is determined based on the interval levels of the rule-based risk coefficient, the interval levels of the model risk score, the interval level difference, and the deviation compensation coefficient. The specific logic is as follows: in, The overall risk level of the cases pending review. For the interval level of the rule risk coefficient, For the risk scoring interval levels of the model, This is the deviation compensation coefficient. .
9. The intelligent decision-making method based on rule and model fusion of multi-source insurance data according to claim 8, characterized in that, The final claim risk level is determined, and the corresponding intelligent claim decision is output. The specific logic is as follows: The calculated comprehensive risk level is compared with the first and second preset thresholds to classify claims risk into three categories: low risk, medium risk, and high risk, as detailed below: When the overall risk level is less than the first preset threshold, it is determined to be a low-risk level; When the overall risk level is between the first and second preset thresholds, it is determined to be a medium risk level; When the overall risk level is greater than the second preset threshold, it is determined to be a high-risk level; Wherein, the first preset threshold is less than the second preset threshold; Based on the determined claim risk level, a matching intelligent claim decision is output.
10. A rule-model fusion intelligent decision-making system based on multi-source insurance data, the system being used to execute the rule-model fusion intelligent decision-making method based on multi-source insurance data as described in any one of claims 1-9, characterized in that, include: The data collection module is used to collect multi-source heterogeneous insurance data for auto insurance, including policy underwriting data, historical claims data, vehicle condition and traffic violation data, historical repair data from repair shops, and spatiotemporal and meteorological data of accidents. It preprocesses various types of raw data to construct a normalized multi-source feature dataset. The rule discrimination module is used to retrieve the normalized multi-source feature dataset of claims pending review, import it into a rigid business rule library built based on the auto insurance industry's claims compliance standards and historical risk control experience, and classify the cases into rejected claims, direct claims, and cases pending review according to the rules, while outputting the rule risk coefficient corresponding to the cases pending review. The model scoring module is used to input the normalized multi-source feature dataset corresponding to the case to be reviewed into the trained claims risk identification model, and output risk scores for five dimensions: abnormal amount, accident frequency, vehicle condition violations, abnormal repair, and spatiotemporal and meteorological. It counts the number of dimensions falling into the high-risk range, and selects the corresponding risk score as the model risk score by combining the preset risk score matching rules. The integrated decision-making module is used to pre-set unified multi-level risk intervals for rule risk coefficients and model risk scores. It determines the risk intervals to which the rule risk coefficients and model risk scores belong, calculates the interval level difference between the two types of scores, matches the deviation compensation coefficient based on the interval level difference, and determines the comprehensive risk level of the case to be reviewed based on the interval level of the rule risk coefficient, the interval level of the model risk score, the interval level difference, and the deviation compensation coefficient. This determines the final claim risk level and outputs the corresponding intelligent claim decision.
Citation Information
Patent Citations
Intelligent claim settlement method and system based on big data
CN110610431A