Insurance claim settlement case identification model generation method and device, and case identification method and device
By generating an insurance claim case identification model and using binary tree technology to identify the case types of cases to be filed, the problem of low efficiency of traditional manual screening is solved, and rapid and accurate case identification and screening is achieved.
Patent Information
- Application Number
- CN202510279683.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-10
AI Technical Summary
The increase in the number of insurance claims cases has led to cost control pressure. Traditional manual screening and verification are inefficient and prone to missing false cases, which cannot meet the current management requirements.
By generating an insurance claim case identification model, using binary tree technology, target claims cases are extracted from claims and outlier claims cases, and a binary tree is generated based on their characteristic values and compensation amounts, and added to the identification model until a preset number of binary trees is included. The characteristic values of the case to be filed are input into the model, their outliers are calculated by the path length, and the case type is determined.
It realizes the rapid and accurate identification of the types of cases to be filed, reduces the cost and time of manual verification, and improves the efficiency and accuracy of case screening.
Smart Images

Figure CN120123946A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method for generating an insurance claim case recognition model, a case recognition method, and a device therefor. Background Art
[0002] Currently, with the development of the business, the number of insurance claim cases is increasing day by day, and insurance companies are facing increasing pressure on cost control. To solve the claim cost problem, on the one hand, it is necessary to quickly realize the online transformation of claims settlement, and on the other hand, it is necessary to compress the "moisture" in claims settlement as much as possible. The former not only meets the service needs of customers but also reduces the labor cost of claims settlement. However, this method inevitably brings a larger space for false claim operations. The traditional method of manually screening and verifying false or exaggerated claim cases one by one is time-consuming and laborious, and spot checks are likely to result in "fish slipping through the net". Neither the labor cost nor the accuracy can meet the current management requirements. Summary of the Invention
[0003] The present application aims to at least partly solve one of the technical problems in the related art.
[0004] To this end, the first object of the present application is to propose a method for generating an insurance claim case recognition model.
[0005] The second object of the present application is to propose an insurance claim case recognition method.
[0006] The third object of the present application is to propose a device for generating an insurance claim case recognition model.
[0007] The fourth object of the present application is to propose an insurance claim case recognition device.
[0008] The fifth object of the present application is to propose an electronic device.
[0009] The sixth object of the present application is to propose a computer-readable storage medium.
[0010] The seventh object of the present application is to propose a computer program product.
[0011] To achieve the above object, an embodiment of the first aspect of the present application proposes a method for generating an insurance claim case recognition model, including:
[0012] Obtain a training data set, where the training data set includes the feature values corresponding to each settled claim case under multiple target features, and the feature values corresponding to each outlier claim case under multiple target features, and among the multiple target features, the claim amount is included;
[0013] Extract a first number of target claim cases from the settled claim cases and the outlier claim cases;
[0014] Generate a binary tree based on the feature values and claim amounts corresponding to each target claim case under multiple target features;
[0015] Add the binary tree to the insurance claim case recognition model, and return to execute extracting a preset number of target claim cases from the settled claim cases and the outlier claim cases until the insurance claim case recognition model contains a second number of binary trees.
[0016] To achieve the above object, an embodiment of the second aspect of the present application proposes an insurance claim case recognition method, including:
[0017] Obtain the feature values corresponding to the claim case to be settled under multiple target features;
[0018] Input the feature values corresponding to the claim case to be settled under the target features into each binary tree in the insurance claim case recognition model to obtain the target node to which the claim case to be settled belongs in each binary tree, where the insurance claim case recognition model is generated according to the method described in the embodiment of the first aspect;
[0019] Determine the path length between the target node and the root node in each binary tree,
[0020] Determine the target outlier score corresponding to the claim case to be settled according to the average value of the second number of the path lengths;
[0021] Determine the case type corresponding to the claim case to be settled according to the target outlier score.
[0022] To achieve the above object, an embodiment of the third aspect of the present application proposes a generating device for an insurance claim case recognition model, including:
[0023] An obtaining module, configured to obtain a training data set, where the training data set includes the feature values corresponding to each settled claim case under multiple target features, and the feature values corresponding to each outlier claim case under multiple target features, and among the multiple target features, a claim amount is included;
[0024] An extracting module, configured to extract a first number of target claim cases from the settled claim cases and the outlier claim cases;
[0025] A first generating module, configured to generate a binary tree based on the feature values and claim amounts corresponding to each target claim case under multiple target features;
[0026] A second generation module, configured to add the binary tree to an insurance claim case recognition model, and return to execute extracting a preset number of target claim cases from the settled claim cases and the outlier claim cases until the insurance claim case recognition model contains a second number of binary trees.
[0027] To achieve the above object, an embodiment of the fourth aspect of the present application provides an insurance claim case recognition device, including:
[0028] An acquisition module, configured to acquire the feature values corresponding to a claim case to be settled under multiple target features;
[0029] A processing module, configured to input the feature values corresponding to the claim case to be settled under the target features into each binary tree in the insurance claim case recognition model, and obtain the target nodes to which the claim case to be settled belongs in each of the binary trees, where the insurance claim case recognition model is generated according to the method described in any one of claims 1-3;
[0030] A first determination module, configured to determine the path length between the target node and the root node in each of the binary trees,
[0031] A second determination module, configured to determine the target outlier score corresponding to the claim case to be settled according to the average value of the second number of the path lengths;
[0032] A third determination module, configured to determine the case type corresponding to the claim case to be settled according to the target outlier score.
[0033] To achieve the above object, an embodiment of the fifth aspect of the present application provides an electronic device, including:
[0034] A processor, and a memory communicatively connected to the processor;
[0035] The memory stores computer-executable instructions;
[0036] The processor executes the computer-executable instructions stored in the memory to implement the method described in the embodiment of the first aspect, or implement the method described in the embodiment of the second aspect.
[0037] To achieve the above object, an embodiment of the sixth aspect of the present application provides a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the embodiment of the first aspect, or implement the method described in the embodiment of the second aspect.
[0038] To achieve the above object, an embodiment of the seventh aspect of the present application provides a computer program product and a computer program. When the computer program is executed by a processor, it implements the method of the embodiment of the first aspect or the method described in the embodiment of the second aspect.
[0039] In the embodiments of the present disclosure, the eigenvalue corresponding to a claim case under multiple target features may be obtained first. Then, the eigenvalue corresponding to the claim case under the target features is input into each binary tree in the insurance claim case recognition model to obtain the target node to which the claim case belongs in each binary tree. The path length between the target node and the root node in each binary tree is determined. According to the average value of the second number of the path lengths, the target outlier score corresponding to the claim case is determined. According to the target outlier score, the case type corresponding to the claim case is determined. Thus, based on the insurance claim case recognition model, the target outlier score of each claim case can be determined, and based on the target outlier score, the case category corresponding to the claim case can be accurately determined, so that in the insurance claim process, each claim case can be detected, and abnormal cases can be accurately, quickly, and timely identified.
[0040] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0042] Figure 1 is a schematic flowchart of a method for generating an insurance claim case recognition model provided by an embodiment of the present application;
[0043] Figure 2 is a schematic flowchart of a method for recognizing an insurance claim case provided by an embodiment of the present application; and
[0044] Figure 3 is a schematic structural diagram of a device for generating an insurance claim case recognition model provided by an embodiment of the present application;
[0045] Figure 4 is a schematic structural diagram of a device for recognizing an insurance claim case provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.
[0047] Among them, a method and apparatus for generating an insurance claim case recognition model according to an embodiment of the present application will be described below with reference to the accompanying drawings.
[0048] Figure 1 It is a schematic flowchart of a method for generating an insurance claim case recognition model provided by an embodiment of the present application.
[0049] As Figure 1 shown, the method for generating the insurance claim case recognition model includes the following steps:
[0050] Step 101, obtain a training data set, where the training data set includes the feature values corresponding to each settled claim case under multiple target features, and the feature values corresponding to each outlier claim case under multiple target features, where the multiple target features include the claim amount.
[0051] In some embodiments, the settled claim cases in the training set may be vehicle insurance claim cases.
[0052] In some embodiments, obtain an initial data set, where the initial data set includes the feature values corresponding to each settled claim case under multiple initial features, where the initial features include the claim amount, and use the principal component analysis method to determine the target features from the other initial features except the claim amount.
[0053] In some embodiments, settled claim cases may be extracted from the business database of an insurance branch company, and all features related to the claim behavior in the settled claim cases are determined as the initial features.
[0054] In some embodiments, new initial features may also be constructed according to some initial features. For example, according to the policy signing date and the accident date, the number of days from signing to the accident is constructed.
[0055] In some embodiments, if the feature values corresponding to some initial features are numerical, such as vehicle age, total premium, total claim, etc., they can be directly used. If the feature values corresponding to some initial features are non-numerical, such as category, the waste resin type can be converted into a one-hot code. For example, for vehicle types, there are 27 types, corresponding to generating 27 segments, and for a certain type, a certain field takes 1 and the rest of the fields take 0.
[0056] It should be noted that since there are many initial features corresponding to each settled claim case, it may contain a lot of redundant information. And too many initial features are also not conducive to subsequent processing. Therefore, the present disclosure may need to adopt the principal component analysis method to analyze the initial features corresponding to each settled claim case to obtain the target features.
[0057] In some embodiments, a third number of settled claim cases are extracted from multiple settled claim cases as the claim cases to be processed; the eigenvalue corresponding to each claim case to be processed under the claim amount is increased to obtain the eigenvalue corresponding to each claim case to be processed under the claim amount; the eigenvalue corresponding to each claim case to be processed under multiple target features is determined as the eigenvalue corresponding to each outlier claim case under multiple target features.
[0058] In some embodiments, the eigenvalue corresponding to the claim case to be processed under the claim amount can be randomly increased by 10% to 80%. The present disclosure does not limit this.
[0059] In some embodiments, according to the eigenvalue corresponding to each claim case to be processed under each target feature, the value range corresponding to each target feature is determined, and based on the value range corresponding to each target feature, a value is randomly selected as the eigenvalue corresponding to the outlier claim case under the target feature.
[0060] Step 102, extract a first number of target claim cases from the settled claim cases and outlier claim cases.
[0061] Among them, the first number can be 100, 200, etc. The present disclosure does not limit this.
[0062] Step 103, generate a binary tree according to the eigenvalue corresponding to each target claim case under multiple target features and the claim amount.
[0063] In some embodiments, (1) the first number of target claim cases are placed in the root node of the tree as the current node; (2) a target feature is randomly specified, and a cutting point is randomly generated within the range of the target claim cases of the current node; (3) the target claim cases with the eigenvalue less than the cutting point selected by the current node are divided into one group, and the samples greater than or equal to the cutting point are divided into another group. The two groups of target claim cases are respectively used as the two sub-nodes of the current node; (4) the two sub-nodes are respectively regarded as the current node, and the processes (2) to (4) are repeated until there is only one sample in the sub-node, or the number of nodes passed from the sub-node to the root node (i.e., the path length) reaches the specified value.
[0064] Step 104, add the binary tree to the insurance claim case recognition model, and return to execute extracting a preset number of target claim cases from the settled claim cases and outlier claim cases until the insurance claim case recognition model contains a second number of binary trees.
[0065] Among them, the second quantity can be 100, 200, 400, etc. The present disclosure does not limit this.
[0066] In some embodiments, the second quantity can be determined according to the first quantity. The smaller the first quantity, the larger the second quantity. In the embodiments of the present disclosure,
[0067] In the embodiments of the present disclosure, a training data set is obtained, where the training data set includes the feature values corresponding to each settled claim case under multiple target features, and the feature values corresponding to each outlier claim case under multiple target features; extract a first quantity of target claim cases from the settled claim cases and outlier claim cases; generate a binary tree according to the feature values and claim amounts corresponding to each target claim case under multiple target features; finally, add the binary tree to the insurance claim case recognition model, and return to execute extracting a preset quantity of target claim cases from the settled claim cases and outlier claim cases until the insurance claim case recognition model contains a second quantity of binary trees. Thereby, an insurance claim case recognition model can be generated quickly and simply, providing support for identifying the claim behavior of a claim case to be settled.
[0068] Figure 2 is a schematic flowchart of an insurance claim case recognition method provided by an embodiment of the present application. As Figure 2 shown, the insurance claim case recognition method may include the following steps:
[0069] Step 201, obtain the feature values corresponding to the claim case to be settled under multiple target features.
[0070] Step 202, input the feature values corresponding to the claim case to be settled under the target features into each binary tree in the insurance claim case recognition model to obtain the target node to which the claim case to be settled belongs in each binary tree.
[0071] In some embodiments, based on the classification rule corresponding to each binary tree, determine the target node to which the claim case to be settled belongs in each binary tree.
[0072] Step 203, determine the path length between the target node and the root node in each binary tree.
[0073] In some embodiments, determine the number of nodes passed from the target node to the root node as the path length between the target node and the root node.
[0074] Step 204, determine the target outlier score corresponding to the claim case to be settled according to the average value of the second quantity of path lengths.
[0075] In some embodiments, the calculation formula of the outlier score can be:
[0076]
[0077] Among them, h(x) represents the path length between the node to which the sample x belongs in each binary tree and the root node. Represents the average value of the second quantity of path lengths. The average path length of constructing a binary tree for m target claim cases. Among them, introducing The purpose is to make Comparable among different insurance claim case recognition models.
[0078] In the embodiments of the present disclosure, the sample x is the claim case to be settled. h(x) is the path length between the target node and the root node in each binary tree.
[0079] Step 205, determine the case type corresponding to the claim case to be settled according to the target outlier score.
[0080] In some embodiments, when the target outlier score is greater than the outlier score threshold, determine that the case type corresponding to the claim case to be settled is an abnormal case; or, when the target outlier score is less than or equal to the outlier score threshold, determine that the case type corresponding to the claim case to be settled is a normal case.
[0081] In some embodiments, the outlier score threshold can be preset. It should be noted that according to the outlier score calculation formula, the outlier score is between 0 and 1. When the outlier score of the sample is close to 1, the path length is very small, and the sample is easily isolated and can be regarded as an outlier. When the outlier score of the sample is less than 0.5, the path length will become larger, and the sample can be regarded as a normal point. If the outlier scores of all samples are around 0.5, then all samples have no abnormality. Therefore, the outlier score can be set to 0.7, 0.8, 0.9, etc. The present disclosure does not limit this.
[0082] In some embodiments, the outlier score threshold can also be determined through the following steps:
[0083] (1) Obtain a training data set, where the training data set includes the feature values corresponding to each settled claim case under multiple target features, and the feature values corresponding to each outlier claim case under multiple target features. Among them, the multiple target features include the claim amount.
[0084] (2) According to the feature values corresponding to each settled claim case under multiple target features and the insurance claim case recognition model, determine the first outlier score corresponding to each settled claim case.
[0085] In the embodiments of the present disclosure, the specific implementation form of determining the first outlier score corresponding to each settled claim case can refer to the specific implementation manner of determining the target outlier score corresponding to the claim case to be settled in the present disclosure.
[0086] In some embodiments, the feature values corresponding to the settled claims under multiple target features can be input into each binary tree in the insurance claim case recognition model to obtain the first node to which the settled claim belongs in each binary tree. Then, the first path length between the first node and the root node in each binary tree is determined, and the first outlier score corresponding to the settled claim is determined according to the average value of the second number of first path lengths.
[0087] (3) According to the feature values corresponding to each outlier claim under multiple target features and the insurance claim case recognition model, determine the second outlier score corresponding to each outlier claim.
[0088] In the embodiments of the present disclosure, for the specific implementation form of determining the second outlier score corresponding to each outlier claim, reference can be made to the specific implementation manner of determining the target outlier score corresponding to the claim to be settled in the present disclosure.
[0089] In some embodiments, the feature values corresponding to the outlier claims under multiple target features can be input into each binary tree in the insurance claim case recognition model to obtain the second node to which the outlier claim belongs in each binary tree. Then, the second path length between the second node and the root node in each binary tree is determined, and the second outlier score corresponding to the settled claim is determined according to the average value of the second number of second path lengths.
[0090] (4) Sort the first outlier score and the second outlier score in ascending order to obtain an outlier score sequence.
[0091] (5) Determine the quantile corresponding to the outlier score sequence at the target quantile as the outlier score threshold.
[0092] In some embodiments, the target quantile can be preset. For example, it can be 0.9, 0.95, etc. The present disclosure does not limit this.
[0093] In some embodiments, if there are 1000 outlier scores in the outlier score sequence and the target quantile is 0.9, then the 900th outlier score in the outlier score sequence is determined as the outlier score threshold.
[0094] In some embodiments, the quantile corresponding to the outlier score sequence at each candidate quantile can also be determined as the candidate threshold corresponding to each candidate quantile. Then, for each candidate threshold, determine the fourth quantity corresponding to the second outlier score greater than the candidate threshold. Finally, determine the target quantile according to the fourth quantity corresponding to each candidate threshold.
[0095] In some embodiments, the candidate quantile corresponding to the candidate threshold with the largest fourth quantity can be determined as the target quantile.
[0096] In some embodiments, the candidate quantile corresponding to the fourth quantity closest to the preset quantity may be determined as the target quantile. In some embodiments, the preset quantity may be half of the third quantity.
[0097] For example, the candidate quantiles include 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, and the third quantity is 100. Then, the fourth quantity corresponding to each candidate quantile can be as shown in Table 1.
[0098] Table 1
[0099]
[0100] As shown in Table 1, when the candidate quantile is 0.9, almost all outlier claim cases can be identified as abnormal cases. When the candidate quantile is 0.95, half of the outlier claim cases can be identified as abnormal cases. The user can determine the target quantile according to the requirements. The present disclosure does not limit this.
[0101] In some embodiments, manual verification can be performed on abnormal cases.
[0102] In the embodiments of the present disclosure, the feature values corresponding to the claim cases to be settled under multiple target features can be obtained first. Then, the feature values corresponding to the claim cases to be settled under the target features are input into each binary tree in the insurance claim case recognition model to obtain the target nodes to which the claim cases to be settled belong in each binary tree. The path lengths between the target nodes and the root nodes in each binary tree are determined. According to the average value of the second quantity of path lengths, the target outlier score corresponding to the claim cases to be settled is determined. According to the target outlier score, the case type corresponding to the claim cases to be settled is determined. Thus, based on the insurance claim case recognition model, the target outlier score of each claim case to be settled can be determined, and based on the target outlier score, the case category corresponding to the claim cases to be settled can be accurately determined. Therefore, in the insurance claim process, each claim case to be settled can be detected, and abnormal cases can be accurately, quickly, and timely identified.
[0103] To implement the above embodiments, the present application also proposes a generating device for an insurance claim case recognition model.
[0104] Figure 3 For the structural schematic diagram of a generating device for an insurance claim case recognition model provided by the embodiments of the present application. As Figure 3 shown, the generating device for the insurance claim case recognition model includes:
[0105] An acquisition module 301, configured to acquire a training data set, where the training data set includes the feature values corresponding to each settled claim under multiple target features, and the feature values corresponding to each outlier claim under multiple target features, and among the multiple target features, there is a claim amount;
[0106] An extraction module 302, configured to extract a first number of target claims from the settled claims and outlier claims;
[0107] A first generation module 303, configured to generate a binary tree according to the feature values and claim amount corresponding to each target claim under multiple target features;
[0108] A second generation module 304, configured to add the binary tree to an insurance claim recognition model, and return to execute extracting a preset number of target claims from the settled claims and outlier claims until the insurance claim recognition model contains a second number of binary trees.
[0109] In some embodiments, it further includes a first determination module, configured to:
[0110] Extract a third number of settled claims from multiple settled claims as claims to be processed;
[0111] Increase the feature value corresponding to each claim to be processed under the claim amount to obtain the feature value corresponding to each claim to be processed under the claim amount;
[0112] Determine the feature values corresponding to each claim to be processed under multiple target features as the feature values corresponding to each outlier claim under multiple target features.
[0113] In some embodiments, it further includes a second determination module, configured to:
[0114] Acquire an initial data set, where the initial data set includes the feature values corresponding to each settled claim under multiple initial features, and the initial features include a claim amount;
[0115] Adopt the principal component analysis method to determine target features from other initial features except the claim amount.
[0116] The generating device of the insurance claim case recognition model provided by this application obtains a training data set, where the training data set includes the feature values corresponding to each settled claim case under multiple target features, and the feature values corresponding to each outlier claim case under multiple target features; extracts the first number of target claim cases from the settled claim cases and outlier claim cases; generates a binary tree according to the feature values and claim amounts corresponding to each target claim case under multiple target features; finally adds the binary tree to the insurance claim case recognition model, and returns to execute extracting a preset number of target claim cases from the settled claim cases and outlier claim cases until the insurance claim case recognition model contains the second number of binary trees. Thus, an insurance claim case recognition model can be generated quickly and simply, providing support for identifying the claim behavior of the claim cases to be settled.
[0117] Figure 4 FIG. 4 is a schematic structural diagram of an insurance claim case recognition device provided by an embodiment of this application. As Figure 4 shown, the insurance claim case recognition device includes:
[0118] An obtaining module 401, configured to obtain the feature values corresponding to the claim case to be settled under multiple target features;
[0119] A processing module 402, configured to input the feature values corresponding to the claim case to be settled under the target features into each binary tree in the insurance claim case recognition model, and obtain the target node to which the claim case to be settled belongs in each binary tree, where the insurance claim case recognition model is generated according to the method of any one of claims 1-3;
[0120] A first determining module 403, configured to determine the path length between the target node and the root node in each binary tree;
[0121] A second determining module 404, configured to determine the target outlier score corresponding to the claim case to be settled according to the average value of the second number of path lengths;
[0122] A third determining module 405, configured to determine the case type corresponding to the claim case to be settled according to the target outlier score.
[0123] In some embodiments, the third determining module 405 is configured to:
[0124] When the target outlier score is greater than the outlier score threshold, determine that the case type corresponding to the claim case to be settled is an abnormal case; or,
[0125] When the target outlier score is less than or equal to the outlier score threshold, determine that the case type corresponding to the claim case to be settled is a normal case.
[0126] In some embodiments, the third determining module 405 is configured to:
[0127] Obtain a training data set, where the training data set includes the feature values corresponding to each settled claim case under multiple target features, and the feature values corresponding to each outlier claim case under multiple target features, where the multiple target features include the claim amount;
[0128] According to the feature values corresponding to each settled claim case under multiple target features and the insurance claim case recognition model, determine the first outlier score corresponding to each settled claim case;
[0129] According to the feature values corresponding to each outlier claim case under multiple target features and the insurance claim case recognition model, determine the second outlier score corresponding to each outlier claim case;
[0130] Sort the first outlier score and the second outlier score in ascending order to obtain an outlier score sequence;
[0131] Determine the quantile corresponding to the outlier score sequence at the target quantile as the outlier score threshold.
[0132] In some embodiments, the third determination module 405 is configured to:
[0133] Determine the quantile corresponding to the outlier score sequence at each candidate quantile as the candidate threshold corresponding to each candidate quantile;
[0134] For each candidate threshold, determine the fourth quantity corresponding to the second outlier score greater than the candidate threshold;
[0135] Determine the target quantile according to the fourth quantity corresponding to each candidate threshold.
[0136] In the embodiments of the present disclosure, the feature values corresponding to the claim case to be settled under multiple target features can be obtained first, and then the feature values corresponding to the claim case to be settled under the target features are input into each binary tree in the insurance claim case recognition model to obtain the target nodes to which the claim case to be settled belongs in each binary tree. Determine the path length between the target node and the root node in each binary tree, and determine the target outlier score corresponding to the claim case to be settled according to the average value of the second quantity of path lengths. Determine the case type corresponding to the claim case to be settled according to the target outlier score. Thus, based on the insurance claim case recognition model, the target outlier score of each claim case to be settled can be determined, and based on the target outlier score, the case category corresponding to the claim case to be settled can be accurately determined, so that in the insurance claim process, each claim case to be settled can be detected, and abnormal cases can be accurately, quickly and timely identified.
[0137] To implement the above embodiments, the present application also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0138] To implement the above embodiments, the present application also provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method provided by the foregoing embodiments when executed by a processor.
[0139] To implement the above embodiments, the present application also provides a computer program product including a computer program, which implements the method provided by the foregoing embodiments when executed by a processor.
[0140] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0141] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and signing an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0142] The present application anticipates providing embodiments that allow users to selectively block the use or access of personal information data. That is, the present disclosure anticipates providing hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.
[0143] In the description of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0144] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0145] Any process or method description represented in a flowchart or otherwise described herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of this application pertain.
[0146] The logic and / or steps represented in a flowchart or otherwise described herein, for example, may be considered as a sequenced list of executable instructions for implementing a logical function, and may be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0147] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0148] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0149] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0150] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for generating an insurance claim case identification model, characterized in that: The following steps are involved: Obtaining a training data set, wherein the training data set includes feature values corresponding to each claim case under multiple target features, and feature values corresponding to each outlier claim case under multiple target features, wherein the multiple target features include the amount of compensation; Extracting a first number of target claim cases from the settled claim cases and the outlier claim cases; Generate a binary tree based on the characteristic values and compensation amounts corresponding to each target claim case under multiple target characteristics; The binary tree is added to the insurance claim case identification model, and the process returns to extract a preset number of target claim cases from the claim cases and the outlier claim cases until the insurance claim case identification model includes a second number of binary trees.
2. The method according to claim 1, characterized in that The method further comprises: Selecting a third number of settled claims cases from the plurality of settled claims cases as pending claims cases; Increase the characteristic value corresponding to each of the pending claims under the compensation amount to obtain the characteristic value corresponding to each of the pending claims under the compensation amount; The characteristic value corresponding to each of the pending claims under multiple target features is determined as the characteristic value corresponding to each outlier claims case under multiple target features.
3. The method according to claim 1, characterized in that The method further comprises: Acquire an initial data set, wherein the initial data set includes feature values corresponding to each of the settled claims under a plurality of initial features, wherein the initial features include the claim amount; The target feature is determined from other initial features except the amount of compensation by using principal component analysis.
4. A method for identifying insurance claims, characterized in that: include: Obtain the feature values corresponding to the pending claims cases under multiple target features; Input the characteristic value corresponding to the pending claim case under the target characteristic into each binary tree in the insurance claim case identification model to obtain the target node to which the pending claim case belongs in each of the binary trees, wherein the insurance claim case identification model is generated according to any one of the methods described in claims 1 to 3; Determine the path length between the target node and the root node in each of the binary trees; Determining a target outlier score corresponding to the pending claim case according to an average value of the second number of path lengths; According to the target outlier score, the case type corresponding to the pending claim case is determined.
5. The method according to claim 4, characterized in that Determining the case type corresponding to the pending claim case according to the target outlier score includes: When the target outlier score is greater than the outlier score threshold, determining that the case type corresponding to the pending claim case is an abnormal case; or, When the target outlier score is less than or equal to the outlier score threshold, it is determined that the case type corresponding to the pending claim case is a normal case.
6. The method according to claim 5, characterized in that The method further comprises: Obtaining a training data set, wherein the training data set includes feature values corresponding to each claim case under multiple target features, and feature values corresponding to each outlier claim case under multiple target features, wherein the multiple target features include the amount of compensation; Determine a first outlier score corresponding to each of the settled claims cases according to the feature values corresponding to each of the settled claims cases under multiple target features and the insurance claim case identification model; Determine a second outlier score corresponding to each of the outlier claims cases according to the feature values corresponding to each of the outlier claims cases under multiple target features and the insurance claims case identification model; Sort the first outlier score and the second outlier score in ascending order to obtain an outlier score sequence; The quantile corresponding to the outlier score sequence under the target quantile is determined as the outlier score threshold.
7. The method according to claim 6, characterized in that The method further comprises: Determine the quantile corresponding to the outlier point sequence under each candidate quantile as the candidate threshold corresponding to each candidate quantile; For each candidate threshold, determining a fourth number corresponding to second outlier scores greater than the candidate threshold; The target percentile is determined according to a fourth quantity corresponding to each of the candidate thresholds.
8. A device for generating an insurance claim case identification model, characterized in that: include: An acquisition module, used to acquire a training data set, wherein the training data set includes feature values corresponding to each claim case under multiple target features, and feature values corresponding to each outlier claim case under multiple target features, wherein the multiple target features include the amount of compensation; An extraction module, configured to extract a first number of target claim cases from the settled claim cases and the outlier claim cases; The first generation module is used to generate a binary tree according to the characteristic values and compensation amounts corresponding to each target claim case under multiple target characteristics; The second generation module is used to add the binary tree to the insurance claim case identification model, and return to execute to extract a preset number of target claim cases from the claimed cases and the outlier claim cases until the insurance claim case identification model contains a second number of binary trees.
9. An insurance claim case identification device, characterized in that: include: An acquisition module is used to obtain the feature values corresponding to the pending claims cases under multiple target features; A processing module, used for inputting the characteristic value corresponding to the pending claim case under the target characteristic into each binary tree in the insurance claim case identification model, and obtaining the target node to which the pending claim case belongs in each of the binary trees, wherein the insurance claim case identification model is generated according to any one of the methods of claims 1 to 3; A first determination module, used to determine the path length between the target node and the root node in each of the binary trees; A second determination module, configured to determine a target outlier score corresponding to the pending claim case according to an average value of a second number of the path lengths; The third determination module is used to determine the case type corresponding to the pending claim case according to the target outlier score.
10. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 3, or the method according to any one of claims 4 to 7.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method according to any one of claims 1 to 3, or the method according to any one of claims 4 to 7.