Training method of car insurance prediction model and car insurance expansion loss identification method
By building a car insurance prediction model and using the car insurance weight dictionary and embedding network to train text vector mapping, the problem of car insurance losses caused by the adjuster's reliance on professional skills is solved, and accurate prediction of the loss amount and resource conservation are achieved.
Patent Information
- Application Number
- CN202510864701.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies have the risk of exaggerating losses during the vehicle damage assessment stage, mainly because the adjuster has high professional skill requirements and subjective judgment, resulting in the assessed damage amount not being consistent with the actual repair amount.
By building a car insurance prediction model, using the car insurance weight dictionary and embedding network, and training text vector mapping processing, an accurate car insurance prediction model is generated, which reduces the dependence on the professional skills of the adjuster and reduces the loss caused by car insurance.
It achieves accurate prediction of the damage amount of accident vehicles, reduces the risk of insurance companies' expanded losses during the vehicle damage assessment stage, and reduces the demand for training data and computing resources.
Smart Images

Figure CN120765397A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of vehicle damage assessment, and in particular to a method for training a vehicle insurance prediction model and a method for identifying vehicle insurance expanded losses. Background Art
[0002] In the field of financial auto insurance, for example, in auto insurance business application scenarios, insurance companies often need to fully and accurately grasp and verify the actual vehicle losses of accident vehicles during the damage assessment stage of the accident vehicle, so as to avoid the insurance company from experiencing expanded auto insurance losses during the damage assessment stage of the vehicle.
[0003] At present, the relevant technology usually determines the approximate vehicle damage amount by having an adjuster during the damage assessment stage of the accident vehicle. However, since this method requires a high level of professional skills from the adjuster and the adjuster has certain subjective judgments when determining the vehicle damage amount, the insurance company is at risk of increased losses from auto insurance during the vehicle damage assessment stage.
[0004] Therefore, the problems existing in related technologies still need to be solved and optimized urgently. Summary of the Invention
[0005] The main purpose of the embodiments of the present application is to propose a training method for a car insurance prediction model and a method for identifying car insurance expanded losses, wherein the training method can provide a car insurance prediction model that can accurately predict and identify the damage amount of an accident vehicle, effectively reducing the professional skill level requirements of the adjuster for vehicle damage assessment, and is conducive to reducing the insurance company's car insurance expanded losses during the vehicle damage assessment stage.
[0006] To achieve the above objectives, a first aspect of an embodiment of the present application provides a method for training a vehicle insurance prediction model and a method for identifying vehicle insurance expanded losses. The vehicle insurance prediction model includes an embedded network, and the method includes:
[0007] Obtaining a motor vehicle insurance training text and a motor vehicle insurance weight dictionary, wherein the motor vehicle insurance weight dictionary includes a plurality of motor vehicle insurance keywords and a keyword weight corresponding to each motor vehicle insurance keyword;
[0008] Inputting the auto insurance training text into a current embedding network for text vector mapping according to the auto insurance weight dictionary to obtain an auto insurance text vector output by the current embedding network, wherein the current embedding network is a first embedding network or a second embedding network obtained by updating the first embedding network, wherein network parameters of the first embedding network are obtained based on keyword weights of a plurality of the auto insurance keywords;
[0009] According to the vehicle insurance text vector, the parameters of the current vehicle insurance prediction model are updated to obtain a trained vehicle insurance prediction model.
[0010] In some embodiments, the first embedding network is obtained by the following steps:
[0011] obtaining an original embedding network of the vehicle insurance prediction model;
[0012] According to the vehicle insurance weight dictionary, the network parameters of the original embedding network are replaced and updated to obtain the first embedding network.
[0013] In some embodiments, the vehicle insurance weight dictionary is obtained by the following steps:
[0014] obtaining a text vector model and a plurality of vehicle insurance sample cases;
[0015] extracting the vocabulary of all the vehicle insurance sample cases to obtain a sample vocabulary table, the sample vocabulary table recording a plurality of vehicle insurance keywords;
[0016] inputting the sample vocabulary table into the text vector model to generate the keyword weight of each vehicle insurance keyword.
[0017] In some embodiments, the second embedding network is obtained by the following steps:
[0018] obtaining the true label of the vehicle insurance training text and the vehicle insurance prediction data corresponding to the vehicle insurance text vector;
[0019] determining a target loss value according to the vehicle insurance prediction data and the true label;
[0020] According to the target loss value, the parameters of the current embedding network are updated to obtain the second embedding network.
[0021] In a second aspect, the embodiments of the present application propose a vehicle insurance prediction model identification method, comprising:
[0022] obtaining a vehicle insurance target text of an accident vehicle;
[0023] inputting the vehicle insurance target text into the vehicle insurance prediction model to perform prediction identification, and obtaining target prediction data;
[0024] performing vehicle insurance extended loss analysis processing on the target prediction data to obtain a vehicle insurance extended loss identification result of the accident vehicle.
[0025] In some embodiments, the vehicle insurance extended loss analysis processing on the target prediction data to obtain the vehicle insurance extended loss identification result of the accident vehicle comprises:
[0026] obtaining a difference threshold and vehicle insurance actual declaration data of the accident vehicle;
[0027] Performing a difference analysis on the actual auto insurance declaration data based on the target prediction data to obtain auto insurance difference data;
[0028] The vehicle insurance expanded loss identification result is obtained according to the difference threshold and the vehicle insurance difference data.
[0029] In some embodiments, performing vehicle insurance expanded loss analysis on the target prediction data to obtain a vehicle insurance expanded loss identification result for the accident vehicle includes:
[0030] Obtain benchmark distribution data;
[0031] Performing data distribution analysis on the target prediction data to obtain prediction distribution data;
[0032] Based on the benchmark distribution data, a divergence analysis is performed on the predicted distribution data to obtain the vehicle insurance expanded loss identification result.
[0033] To achieve the above objectives, a third aspect of an embodiment of the present application provides a system for training a vehicle insurance prediction model, wherein the vehicle insurance prediction model includes an embedded network, and the system:
[0034] A first processing unit obtains a motor vehicle insurance training text and a motor vehicle insurance weight dictionary, wherein the motor vehicle insurance weight dictionary includes a plurality of motor vehicle insurance keywords and a keyword weight corresponding to each motor vehicle insurance keyword;
[0035] a second processing unit, configured to input the auto insurance training text into a current embedding network for text vector mapping based on the auto insurance weight dictionary, thereby obtaining an auto insurance text vector output by the current embedding network, wherein the current embedding network is a first embedding network or a second embedding network obtained by updating the first embedding network, wherein network parameters of the first embedding network are obtained based on keyword weights of a plurality of auto insurance keywords;
[0036] The third processing unit is used to update the parameters of the current vehicle insurance prediction model according to the vehicle insurance text vector to obtain a trained vehicle insurance prediction model.
[0037] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and the processor implements the method of the first aspect or the method of the second aspect when executing the computer program.
[0038] To achieve the above-mentioned purpose, the fifth aspect of the embodiment of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method of the first aspect or the method of the second aspect.
[0039] This application proposes a method for training a vehicle insurance prediction model and a method for identifying vehicle insurance expanded losses. The vehicle insurance prediction model includes an embedding network. The training method obtains vehicle insurance training text and a vehicle insurance weight dictionary, wherein the vehicle insurance weight dictionary includes several vehicle insurance keywords and a keyword weight corresponding to each vehicle insurance keyword. Based on the vehicle insurance weight dictionary, the vehicle insurance training text is input into a current embedding network for text-to-vector mapping processing, thereby obtaining a vehicle insurance text vector output by the current embedding network. The current embedding network is a first embedding network or a second embedding network obtained by updating the first embedding network, wherein the network parameters of the first embedding network are obtained based on the keyword weights of the several vehicle insurance keywords. Based on the vehicle insurance text vector, the parameters of the current vehicle insurance prediction model are updated to obtain a trained vehicle insurance prediction model. This training method, by inputting vehicle insurance training text into the vehicle insurance prediction model for training, enables the trained vehicle insurance prediction model to accurately predict and identify the assessed loss amount of an accident vehicle based on the input vehicle insurance target text, effectively reducing the professional skill level required of adjusters in vehicle loss assessments, thereby helping to reduce insurance companies' risk of expanded losses during vehicle loss assessments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present application or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly expressing some embodiments of the technical solutions of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.
[0041] Figure 1 This is a schematic diagram of the process principle of a method for training a car insurance prediction model provided in an embodiment of the present application;
[0042] Figure 2 yes Figure 1 Schematic diagram of the process principle of step S101 in FIG.
[0043] Figure 3 This is a schematic diagram of the process principle of the first embedded network provided in an embodiment of the present application;
[0044] Figure 4 This is a schematic diagram of the process principle of the second embedded network provided by an embodiment of the present application;
[0045] Figure 5 This is a schematic diagram of the structure of a training system for a vehicle insurance prediction model provided in an embodiment of the present application;
[0046] Figure 6This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0048] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0050] First, let’s analyze some of the terms used in this application:
[0051] An exaggerated loss in motor vehicle insurance refers to an increase in the insurance company's claim amount due to personnel, technical, or other factors during the damage assessment and claims process, resulting in repairs beyond the actual damage scope or unnecessary additional repairs. This can be intentional by the vehicle mechanic to increase revenue, or it can be the result of an inaccurate assessment of the actual damage by the mechanic and / or adjuster.
[0052] At present, the relevant technology usually determines the approximate vehicle damage amount by having an adjuster during the damage assessment stage of the accident vehicle. However, since this method requires a high level of professional skills from the adjuster and the adjuster has certain subjective judgments when determining the vehicle damage amount, the vehicle damage amount determined by the adjuster is often inconsistent with the actual repair amount of the vehicle, which in turn makes the insurance company face the risk of expanded losses from motor vehicle insurance during the vehicle damage assessment stage.
[0053] Some related technologies primarily rely on statistical methods to assess damages for accident vehicles. These methods employ models (such as neural network models or large language models) to learn from a large number of historical vehicle damage assessment cases. The models then analyze the similarities between the current vehicle case input and these historical vehicle damage assessment cases, thereby assessing damages for the accident vehicle. This approach requires a large amount of training data and computing resources, resulting in long training cycles and high training costs.
[0054] In view of this, an embodiment of the present application provides a training method for a motor vehicle insurance prediction model and a method for identifying motor vehicle insurance expanded losses, wherein the training method is performed by inputting motor vehicle insurance training text into the motor vehicle insurance prediction model for training, specifically by inputting the motor vehicle insurance training text into a first embedding network, and / or inputting the motor vehicle insurance training text into a second embedding network obtained by updating the first embedding network, wherein the first embedding network is obtained based on a motor vehicle insurance weight dictionary, which can effectively reduce the training data and computing resources required for model training while ensuring the accuracy of the model in predicting and identifying the motor vehicle insurance target text, thereby reducing the training cycle and training cost of the motor vehicle insurance prediction model, which is beneficial to reducing the motor vehicle insurance expanded losses of insurance companies in the vehicle damage assessment stage.
[0055] The embodiments of the present application provide a method for training a car insurance prediction model and a method for identifying car insurance expanded losses, which are specifically described through the following embodiments.
[0056] The embodiments of the present application provide a method for training a vehicle insurance prediction model and a method for identifying vehicle insurance escalation losses, which can be applied to vehicle insurance business scenarios. In vehicle insurance business scenarios, the model can be trained using the method for training a vehicle insurance prediction model provided by the present application, and then the trained vehicle insurance prediction model and the method for identifying vehicle insurance escalation losses can be used to analyze the vehicle insurance target text of the accident vehicle. This can accurately predict and identify the damage amount of the accident vehicle, effectively reducing the professional skill level required of the adjuster in vehicle damage assessment, and thus helping to reduce the insurance company's escalation losses during the vehicle damage assessment stage.
[0057] The training method for a car insurance prediction model and the method for identifying expanded losses in car insurance provided in the embodiments of the present application can be applied to a terminal, can be applied to a server, and can also be software running on a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet computer, laptop computer, desktop computer, etc.; the server can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements a training method for a car insurance prediction model or an application that implements a method for identifying expanded losses in car insurance, etc., but is not limited to the above forms.
[0058] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0059] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0060] Figure 1 This is an optional flowchart of a method for training a car insurance prediction model provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S103.
[0061] Step S101: Obtaining a car insurance training text and a car insurance weight dictionary, wherein the car insurance weight dictionary includes a plurality of car insurance keywords and a keyword weight corresponding to each car insurance keyword;
[0062] In an embodiment of the present application, the car insurance training text can be text data from historical vehicle damage assessment cases. First, a car insurance professional text library consisting of several historical vehicle damage assessment cases can be established, and then text features can be extracted from the car insurance professional text library and fused in a key-value format to obtain the car insurance training text.
[0063] Specifically, a rule template can be set in advance, and the text features of several historical vehicle damage assessment cases in the professional auto insurance text library can be extracted based on the rule template; alternatively, several historical vehicle damage assessment cases in the auto insurance patent text library can be input into a natural language processing (NLP) model or a large language model (LLM) to extract corresponding text features.
[0064] Among them, the keys in the car insurance training text can be various attributes in historical vehicle damage assessment cases, which can be any one of "case", "vehicle", "brand", "model", "age of vehicle", "driver", "degree of damage", etc., and the values in the car insurance training text can be attribute information corresponding to the attributes. For example, for the attribute "part", its attribute information can be any one of "bumper", "fender", "chassis", etc.; or, for the attribute "degree of damage", its corresponding attribute information can be any one of "slight deformation", "serious damage", "moderate damage", etc. The examples in this application are for illustration only and do not limit the keys and / or values of the car insurance training text. The specific key values of the car insurance training can be flexibly set according to actual conditions. For example, the keys of the car insurance training text can also include "windshield", "vehicle A-pillar", "vehicle B-pillar", etc., and the corresponding attribute information can also be flexibly generated. This application will not go into details here.
[0065] It can be understood that the auto insurance weight dictionary is a collection of several auto insurance keywords and their keyword weights. The auto insurance keywords can be words that appear more frequently in vehicle damage assessment. Specifically, they can be key attributes such as "vehicle", "model", "part", etc., or they can be attribute information such as "minor deformation" and "serious damage", or other keywords such as "cause of vehicle accident" and "accident responsibility division information". The corresponding auto insurance keyword weights can be pre-set or generated through a model.
[0066] In some embodiments, reference Figure 2 The step S101, obtaining the vehicle insurance weight dictionary, includes:
[0067] Step S201, obtaining a text vector model and a plurality of car insurance sample cases;
[0068] Step S202, performing vocabulary extraction on all the car insurance sample cases to obtain a sample vocabulary table, the sample vocabulary table recording a plurality of car insurance keywords;
[0069] Step S203, inputting the sample vocabulary table into the text vector model to generate a keyword weight of each car insurance keyword.
[0070] In the embodiment of the present application, the text vector model can be any one of Word2Vec model, GloVe model, FastText model, ELMo model, Bert model, etc. In this embodiment, the text vector model is taken as the Word2Vec model as an example. The car insurance sample case can be a historical car insurance loss assessment case collected by an insurance company, or a car insurance loss assessment case disclosed on the Internet, etc. There are various ways to obtain the car insurance sample case, which will not be described here.
[0071] It can be understood that for any car insurance sample case, a plurality of car insurance keywords in the car insurance sample case can be obtained. Then, the car insurance keywords of all car insurance sample cases are spliced to generate a sample vocabulary table, and the sample vocabulary table is input into the Word2Vec model to generate a keyword vector of each car insurance keyword in the sample vocabulary table, and the keyword weight of each car insurance keyword is determined based on all keyword vectors.
[0072] It should be noted that after obtaining the keyword vector of all car insurance keywords, in the first implementation, the keyword weight of each car insurance keyword can be obtained by weighting the frequency of occurrence of each car insurance keyword in the sample vocabulary table and the corresponding keyword vector based on the TF-IDF weighting method. Or, in the second implementation, the keyword weight of each car insurance keyword can be obtained by weighting the frequency of occurrence of each car insurance keyword and the keyword vector based on the SIF weighting method.
[0073] In the third implementation, the loss information of the accident vehicle in the sample vocabulary table can also be learned through a neural network (such as a Transformer network) model, which can be the frequency of occurrence, the degree of loss, the range of loss assessment and claim amount, etc. The importance score of each car insurance keyword in the sample vocabulary table is determined based on the attention mechanism; then, the keyword weight of each car insurance keyword is determined based on the importance score and the keyword vector of each car insurance keyword.
[0074] Step S102: Inputting the auto insurance training text into a current embedding network for text vector mapping based on the auto insurance weight dictionary to obtain an auto insurance text vector output by the current embedding network, where the current embedding network is a first embedding network or a second embedding network obtained by updating the first embedding network, wherein the network parameters of the first embedding network are obtained based on the keyword weights of a plurality of auto insurance keywords;
[0075] In an embodiment of the present application, the car insurance prediction model can be trained in a cyclic iterative manner to obtain a trained car insurance prediction model, wherein the car insurance prediction model can specifically be a convolutional neural network (TextCNN) model based on text classification, and the embedding layer (Embedding Layer) of the TextCNN model is recorded as an embedding network.
[0076] Specifically, for any loop iteration process, the car insurance training text of the current loop iteration process can be input into the current embedding network, and the keys and values in the car insurance training text can be mapped into corresponding vectors through the current embedding network to obtain the car insurance text vector.
[0077] It can be understood that if the current loop iteration process is the first loop iteration process in the vehicle insurance prediction model training process, the current embedding network can be the first embedding network, and the network parameters of the first embedding network are obtained by applying the keyword weights in the vehicle insurance weight dictionary; or, if the current loop iteration point process is the second or more loop iteration processes in the vehicle insurance prediction model training process, the current embedding network can still be the first embedding network; or, the current embedding network can be the second embedding network obtained by one or more updates based on the first embedding network.
[0078] Reference Figure 3 In some embodiments, the first embedded network is obtained by the following steps:
[0079] Step S301: obtaining the original embedding network of the vehicle insurance prediction model;
[0080] Step S302: According to the vehicle insurance weight dictionary, the network parameters of the original embedded network are replaced and updated to obtain a first embedded network.
[0081] In the embodiments of the present application, the original embedding network can be an initial embedding network of the vehicle insurance prediction model. Specifically, in the first implementation, the keyword weight of the partial vehicle insurance keywords in the vehicle insurance weight dictionary can be applied to the network parameters of the embedding network of the Word2Vec model to obtain the first embedding network; or in the second implementation, the keyword weight of all vehicle insurance keywords in the vehicle insurance weight dictionary can be applied to the network parameters of the embedding network of the Word2Vec model to obtain the first embedding network.
[0082] In step S103, the vehicle insurance prediction model is updated according to the vehicle insurance text vector to obtain a trained vehicle insurance prediction model.
[0083] In the embodiments of the present application, for any one cycle iteration process, after obtaining the vehicle insurance text vector output by the embedding network in the current cycle iteration process, the vehicle insurance prediction data can be obtained based on the vehicle insurance text vector, and the vehicle insurance prediction model in the current cycle iteration process is updated based on the vehicle insurance prediction data to obtain a trained vehicle insurance prediction model.
[0084] It can be understood that if the vehicle insurance prediction model is a TextCNN model, after obtaining the vehicle insurance text vector, the vehicle insurance text vector can be sequentially input into the convolution layer, the pooling layer and the full connection layer in the TextCNN model to obtain the vehicle insurance prediction data output by the full connection layer of the TextCNN model.
[0085] It should be noted that before the vehicle insurance prediction model is put into use, it needs to be trained to adjust its internal parameters so as to achieve a better prediction effect. Specifically, when training the model, a batch of vehicle insurance training texts can be obtained, and the real labels corresponding to the vehicle insurance training texts are also obtained, which are used to represent the real loss amounts of the vehicle insurance training texts. Then, each vehicle insurance training text and its corresponding real label can be taken as a group of training data, the input data of the model is the vehicle insurance training text, and the vehicle insurance prediction data is output by the model through the prediction of the vehicle insurance training text. After obtaining the vehicle insurance prediction data output by the model, the accuracy of the prediction of the model can be evaluated according to the vehicle insurance prediction data and the real label, so that the parameters of the model are updated.
[0086] Specifically, for machine learning models, the accuracy of the model prediction results can be measured by a loss function (LossFunction). The loss function is defined on a single training data and is used to measure the prediction error of a training data. Specifically, the loss value of the training data is determined by the label of the single training data and the prediction result of the model on the training data. During actual training, a training data set has a lot of training data, so a cost function (CostFunction) is generally used to measure the overall error of the training data set. The cost function is defined on the entire training data set and is used to calculate the average value of the prediction error of all training data, which can better measure the prediction effect of the model. For general machine learning models, based on the aforementioned cost function, plus a regularization term that measures the complexity of the model, it can be used as the objective function of the training. Based on this objective function, the loss value of the entire training data set can be calculated. There are many types of commonly used loss functions, such as 0-1 loss function, square loss function, absolute loss function, logarithmic loss function, cross entropy loss function, smooth L1 loss function, etc., which can all be used as loss functions of machine learning models, which will not be elaborated one by one here. In the embodiment of the present application, any one of the loss functions can be selected to determine the loss value of the training, such as the cross entropy loss function. Based on the training loss, the model parameters are updated using the backpropagation algorithm. After several iterations, a trained auto insurance prediction model is obtained. The specific number of iterations can be pre-set, or training is considered complete when the test set meets the required accuracy.
[0087] It is worth mentioning that for any loop iteration process, the parameter update of the auto insurance prediction model can be to update the remaining network parameters except the network parameters of the embedded network, that is, to freeze the network parameters of the embedded network and update the remaining network parameters of the auto insurance prediction model; or, it can also be to update all network parameters of the auto insurance prediction model. There are many specific implementation methods for parameter updating, which will not be repeated in this application.
[0088] Reference Figure 4 , the second embedding network is obtained by the following steps:
[0089] Step S401: obtaining the true label of the vehicle insurance training text and the vehicle insurance prediction data corresponding to the vehicle insurance text vector;
[0090] Step S402: determining a target loss value based on the vehicle insurance prediction data and the true label;
[0091] Step S403: Update the parameters of the current embedding network according to the target loss value to obtain the second embedding network.
[0092] In the embodiment of the present application, for any one loop iteration process, a target loss value can be determined based on the vehicle insurance prediction data and the true label. The loss function used to determine the target loss value can be a smooth L1 loss function. The target loss value can be specifically expressed as:
[0093]
[0094] in, is the target loss value under the i-th loop iteration process; y i is the true label of the car insurance training text in the i-th loop iteration process; is the car insurance prediction data corresponding to the car insurance training text under the i-th loop iteration process; δ is the threshold for controlling the balance range, and the threshold δ can be flexibly set.
[0095] It can be understood that after obtaining the target loss value under the i-th loop iteration process, the sampling back propagation algorithm can be used to update the parameters of the current embedding network under the i-th loop iteration process based on the target loss value, so as to obtain a second embedding network, which is used as the current embedding network under the next loop iteration process, that is, the current embedding network under the i+1-th loop iteration process.
[0096] The present application is described below with reference to specific embodiments.
[0097] Example 1
[0098] If the auto insurance prediction model is a TextCNN model, the original embedding network can be the initial embedding layer of the TextCNN model; the first embedding network is the updated embedding network obtained by applying the keyword weights of several auto insurance keywords in the auto insurance weight dictionary to the original embedding network; the convolution layer, pooling layer and classification layer of the TextCNN model are collectively referred to as the convolution classification network.
[0099] For the first loop iteration process, the first embedding network can be determined as the current embedding network, that is, the auto insurance training text can be input into the cascaded first embedding network and convolution classification network to obtain the auto insurance prediction data under the first loop iteration process, and the target loss value is calculated based on the auto insurance prediction data and the corresponding true label; then only the target loss value is used to update the convolution classification network of the auto insurance prediction model to obtain a convolution classification network with an update number of 1.
[0100] For the second loop iteration process, the first embedding network can be determined as the current embedding network, and the car insurance training text can be input into the cascaded first embedding network and the convolution classification network with an update number of 1 to obtain the car insurance prediction data under the second loop iteration process, and calculate the target loss value based on the car insurance prediction data and the corresponding true label; then only the target loss value is used to update the convolution classification network of the car insurance prediction model again to obtain a convolution classification network with an update number of 2.
[0101] The third and subsequent iterations are similar to the second iteration, and can be easily deduced from this process. This application will not elaborate on this process. Furthermore, after the vehicle insurance prediction model has been iterated a predetermined number of times or has reached the required accuracy, a trained vehicle insurance prediction model can be obtained.
[0102] Example 2
[0103] If the auto insurance prediction model is a TextCNN model, the original embedding network can be the initial embedding layer of the TextCNN model; the first embedding network is the updated embedding network obtained by applying the keyword weights of several auto insurance keywords in the auto insurance weight dictionary to the original embedding network; the convolution layer, pooling layer and classification layer of the TextCNN model are collectively referred to as the convolution classification network.
[0104] For the first loop iteration process, the first embedding network can be determined as the current embedding network, and the car insurance training text is input into the cascaded first embedding network and convolution classification network to obtain the car insurance prediction data under the first loop iteration process, and the target loss value is calculated based on the car insurance prediction data and the corresponding true label; then the first embedding network and convolution classification network are updated based on the target loss value to obtain the second embedding network and the convolution classification network with an update number of 1.
[0105] For the second loop iteration process, the second embedding network obtained by updating the first loop iteration process can be determined as the current embedding network. The second embedding network is the embedding network obtained by updating the parameters of the first embedding network once; the auto insurance training text is sequentially input into the current embedding network and the convolution classification network with an update number of 1, thereby obtaining the auto insurance prediction data under the second loop iteration process, and the current embedding network and the convolution classification network are updated based on the auto insurance prediction data, to obtain the second embedding network updated under the second loop iteration process and the convolution classification network with an update number of 2.
[0106] For the third loop iteration process, the second embedding network updated in the second loop iteration process can be determined as the current embedding network, which is the embedding network obtained by performing secondary parameter updating on the first embedding network. The specific content of the third loop iteration process and more loop iteration processes is similar to that of the second loop iteration process, which can be simply analogized, and thus will not be described herein.
[0107] It should be noted that the examples of the present application are only for illustration, not for limiting the present application. For example, in the third embodiment, if the total loop iteration times of the vehicle insurance prediction model is N, the first embedding network can be determined as the current embedding network in the first M loop iteration processes; and the second embedding network updated in the previous loop iteration process can be determined as the current embedding network in any one of the (M+1)th loop iteration process to the Nth loop iteration process. The specific content of the specific loop iteration is similar to that of the first embodiment and / or the second embodiment, wherein N and M are positive integers, and M is less than N.
[0108] The embodiments of the present application also provide a vehicle insurance expanded loss identification method, comprising:
[0109] In step S501, a vehicle insurance target text of an accident vehicle is obtained.
[0110] In step S502, the vehicle insurance target text is input into the vehicle insurance prediction model to perform prediction identification, and target prediction data is obtained.
[0111] In step S503, vehicle insurance expanded loss analysis processing is performed on the target prediction data, and a vehicle insurance expanded loss identification result of the accident vehicle is obtained.
[0112] In the embodiments of the present application, a vehicle insurance target text of an accident vehicle can be obtained, which can be a damaged description text of the accident vehicle. The vehicle insurance target text is input into the vehicle insurance prediction model, the loss assessment amount of the accident vehicle is predicted by the vehicle insurance prediction model, the target prediction data is obtained, and the vehicle insurance expanded loss identification result of the accident vehicle is obtained based on the target prediction data.
[0113] In some embodiments, the step S503, the vehicle insurance expanded loss analysis processing is performed on the target prediction data, and the vehicle insurance expanded loss identification result of the accident vehicle is obtained, comprising:
[0114] In step S601, a difference threshold and vehicle insurance actual declaration data of the accident vehicle are obtained.
[0115] In step S602, difference analysis is performed on the vehicle insurance actual declaration data according to the target prediction data, and vehicle insurance difference data is obtained.
[0116] Step S603: Obtain the vehicle insurance expanded loss identification result according to the difference threshold and the vehicle insurance difference data.
[0117] In an embodiment of the present application, the actual declared data of the motor vehicle insurance may be the actual declared amount of the accident vehicle; the difference analysis of step S602 may specifically be to analyze the difference rate between the assessed damage amount indicated by the target prediction data and the actual declared amount indicated by the actual declared data of the motor vehicle insurance, thereby obtaining the motor vehicle insurance difference data; step S603 may be to compare the size relationship between the difference threshold and the motor vehicle insurance difference data, and the specific value of the difference threshold may be flexibly set according to actual conditions, such as any one of 0.03, 0.1, 0.2, etc., thereby obtaining the motor vehicle insurance expanded loss identification result.
[0118] It can be understood that if the auto insurance difference data is less than or equal to the difference threshold, a auto insurance enlarged loss identification result can be generated, indicating that there is no auto insurance enlarged loss; or, if the auto insurance difference data is greater than the difference threshold, a auto insurance enlarged loss identification result can be generated, indicating that there is an auto insurance enlarged loss.
[0119] In some embodiments, the step S503 of performing a vehicle insurance enlarged loss analysis on the target prediction data to obtain a vehicle insurance enlarged loss identification result of the accident vehicle includes:
[0120] Step S701, obtaining reference distribution data;
[0121] Step S702: performing data distribution analysis on the target prediction data to obtain prediction distribution data;
[0122] Step S703: Perform a divergence analysis on the predicted distribution data based on the benchmark distribution data to obtain the vehicle insurance expanded loss identification result.
[0123] In an embodiment of the present application, the benchmark distribution data can be obtained by analyzing the average distribution of the vehicle damage assessment amounts of several case samples; step S702 can be to analyze the probability distribution of the target prediction data in the vehicle damage assessment amounts of all case samples to obtain the predicted distribution data; step S703 can be based on the KL divergence (Kullback-Leibler Divergence, KLD), measuring the difference between the benchmark distribution data and the predicted distribution data, and when the difference between the two benchmark distribution data and the predicted distribution data is less than or equal to a preset divergence range threshold, generating a vehicle insurance enlarged loss identification result indicating that there is no vehicle insurance enlarged loss; otherwise, generating a vehicle insurance enlarged loss identification result indicating that there is a vehicle insurance enlarged loss.
[0124] See also Figure 5The present application also provides a training system for a vehicle insurance prediction model, which can implement the above-mentioned vehicle insurance prediction model training method. The training system includes:
[0125] The first processing unit 801 obtains a vehicle insurance training text and a vehicle insurance weight dictionary, wherein the vehicle insurance weight dictionary includes a plurality of vehicle insurance keywords and a keyword weight corresponding to each vehicle insurance keyword;
[0126] A second processing unit 802 is configured to input the auto insurance training text into a current embedding network for text vector mapping based on the auto insurance weight dictionary, thereby obtaining an auto insurance text vector output by the current embedding network, where the current embedding network is a first embedding network or a second embedding network obtained by updating the first embedding network, wherein network parameters of the first embedding network are obtained based on keyword weights of a plurality of auto insurance keywords;
[0127] The third processing unit 803 is configured to update the parameters of the current vehicle insurance prediction model according to the vehicle insurance text vector to obtain a trained vehicle insurance prediction model.
[0128] The specific implementation of the training system for the vehicle insurance prediction model is basically the same as the specific embodiment of the training method for the vehicle insurance prediction model described above, and will not be repeated here.
[0129] The present application also provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program, implements the aforementioned vehicle insurance prediction model training method or vehicle insurance expanded loss identification method. The electronic device can be any smart terminal, including a tablet computer and an in-vehicle computer.
[0130] See also Figure 6 , Figure 6 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0131] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0132] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the training method of the vehicle insurance prediction model or the vehicle insurance expanded loss identification method of the embodiments of this application.
[0133] Input / output interface 903, used to implement information input and output;
[0134] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0135] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0136] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0137] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned vehicle insurance prediction model training method or the above-mentioned vehicle insurance expanded loss identification method.
[0138] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0139] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0140] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0142] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0143] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0144] It should be understood that, in the application, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases of only A, only B, and A and B existing at the same time, wherein A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent a, b, c, "a and b", "a and c", "b and c", or "a and b and c", wherein a, b, and c can be single or multiple.
[0145] In several embodiments provided in the application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the above units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0146] The units described above as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiment scheme.
[0147] In addition, the functional units in each embodiment of the application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0148] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0149] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for training a car insurance prediction model, characterized in that: The auto insurance prediction model includes an embedded network, and the method includes: Obtaining a motor vehicle insurance training text and a motor vehicle insurance weight dictionary, wherein the motor vehicle insurance weight dictionary includes a plurality of motor vehicle insurance keywords and a keyword weight corresponding to each motor vehicle insurance keyword; According to the auto insurance weight dictionary, the auto insurance training text is input into a current embedding network for text vector mapping processing to obtain an auto insurance text vector output by the current embedding network; the current embedding network is a first embedding network, or a second embedding network obtained by updating the first embedding network; the network parameters of the first embedding network are obtained based on the keyword weights of a plurality of the auto insurance keywords; According to the vehicle insurance text vector, the parameters of the current vehicle insurance prediction model are updated to obtain a trained vehicle insurance prediction model.
2. The training method according to claim 1, characterized in that The first embedding network is obtained by the following steps: Obtaining the original embedding network of the auto insurance prediction model; According to the vehicle insurance weight dictionary, the network parameters of the original embedded network are replaced and updated to obtain a first embedded network.
3. The training method according to claim 1, characterized in that The step of obtaining the vehicle insurance weight dictionary includes: Obtain a text vector model and several auto insurance sample cases; Performing vocabulary extraction on all the auto insurance sample cases to obtain a sample vocabulary table, wherein the sample vocabulary table records a plurality of auto insurance keywords; The sample vocabulary is input into the text vector model to generate weights, and the keyword weight of each of the auto insurance keywords is obtained.
4. The training method according to claim 1, characterized in that The second embedding network is obtained by the following steps: Obtaining the true label of the vehicle insurance training text and the vehicle insurance prediction data corresponding to the vehicle insurance text vector; Determining a target loss value based on the vehicle insurance prediction data and the true label; According to the target loss value, the parameters of the current embedding network are updated to obtain the second embedding network.
5. A method for identifying expanded losses in motor vehicle insurance, characterized in that: The method comprises: Get the insurance target text of the accident vehicle; Inputting the vehicle insurance target text into the vehicle insurance prediction model according to any one of claims 1 to 4 for prediction and recognition to obtain target prediction data; The target prediction data is subjected to a vehicle insurance enlarged loss analysis process to obtain a vehicle insurance enlarged loss identification result of the accident vehicle.
6. The identification method according to claim 5, characterized in that The performing of the vehicle insurance enlarged loss analysis processing on the target prediction data to obtain the vehicle insurance enlarged loss identification result of the accident vehicle includes: Obtaining the difference threshold and the actual automobile insurance declaration data of the accident vehicle; Performing a difference analysis on the actual auto insurance declaration data based on the target prediction data to obtain auto insurance difference data; The vehicle insurance expanded loss identification result is obtained according to the difference threshold and the vehicle insurance difference data.
7. The identification method according to claim 5, characterized in that: The performing of the vehicle insurance enlarged loss analysis processing on the target prediction data to obtain the vehicle insurance enlarged loss identification result of the accident vehicle includes: Obtain benchmark distribution data; Performing data distribution analysis on the target prediction data to obtain prediction distribution data; Based on the benchmark distribution data, a divergence analysis is performed on the predicted distribution data to obtain the vehicle insurance expanded loss identification result.
8. A training system for a car insurance prediction model, characterized in that: The auto insurance prediction model includes an embedded network, and the system includes: A first processing unit obtains a motor vehicle insurance training text and a motor vehicle insurance weight dictionary, wherein the motor vehicle insurance weight dictionary includes a plurality of motor vehicle insurance keywords and a keyword weight corresponding to each motor vehicle insurance keyword; a second processing unit, configured to input the auto insurance training text into a current embedding network for text vector mapping based on the auto insurance weight dictionary, thereby obtaining an auto insurance text vector output by the current embedding network; the current embedding network being a first embedding network or a second embedding network obtained by updating the first embedding network; and network parameters of the first embedding network being obtained based on keyword weights of a plurality of auto insurance keywords; The third processing unit is used to update the parameters of the current vehicle insurance prediction model according to the vehicle insurance text vector to obtain a trained vehicle insurance prediction model.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 4 or the method according to any one of claims 5 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 or the method according to any one of claims 5 to 7 is implemented.