Vehicle intelligent insurance premium pricing method and system, terminal and storage medium

By collecting multi-source data on vehicle intelligent devices and combining machine learning and deep reinforcement learning models, the optimal premium is calculated, which solves the problem that existing auto insurance pricing methods cannot integrate multi-dimensional data, and achieves personalized and dynamically adjusted insurance pricing.

CN120125347APending Publication Date: 2025-06-10SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197938.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing auto insurance pricing methods cannot integrate data from multiple dimensions, lack the ability to personalize and dynamic adjustment, and cannot accurately reflect various risk factors during vehicle use.

Method used

Vehicle operation data, basic data and geographical location data are collected through vehicle intelligent equipment, combined with machine learning algorithms and deep reinforcement learning models, features, configuration weights, and fusion data are extracted, and the optimal premium that comprehensively considers premium pricing strategies, driver risk scores and loss prediction values ​​are finally calculated.

Benefits of technology

It realizes personalized and accurate insurance pricing, can dynamically adjust premiums, accurately reflect the actual risk level of car owners, and improves the scientificity and rationality of insurance pricing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125347A_ABST
    Figure CN120125347A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of insurance pricing, and particularly provides a vehicle intelligent insurance premium pricing method and system, a terminal and a storage medium, and the method comprises the steps: collecting vehicle operation data and basic data based on a vehicle-mounted intelligent device, collecting the vehicle operation data and all basic data, and carrying out the feature extraction, performing weighted fusion on all the features to obtain a unified feature vector; constructing a deep reinforcement learning model based on the target task, and training the deep reinforcement learning model based on the feature vector; and obtaining an insurance premium basic price, calculating a preliminary insurance premium based on the insurance premium basic price in combination with an insurance premium pricing strategy, obtaining a risk adjustment factor based on the driver risk score and the loss prediction value, and adjusting the preliminary insurance premium based on the risk adjustment factor to obtain an optimal insurance premium. The method achieves the regular updating of the pricing, makes full use of the data flow to adjust the pricing, recognizes and controls high-risk events in advance, and guarantees the controllability of the risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of insurance pricing, and in particular relates to a vehicle intelligent premium pricing method, system, terminal and storage medium. Background Art

[0002] With the rapid development of the new energy vehicle market, the intelligence level of vehicles is constantly improving, especially the widespread application of on-board intelligent devices (such as T-Box devices), which enables new energy vehicles to collect and transmit a large amount of dynamic data in real time. These data include but are not limited to the driver's driving behavior, vehicle operating status, real-time traffic flow, environmental conditions and other information.

[0003] When insurance companies assess the risks during vehicle use, these dynamic data provide a richer assessment tool that can accurately reflect the various risk factors during vehicle use. Therefore, how to combine multi-source data with intelligent algorithms to conduct personalized and accurate insurance pricing has become a technical problem that needs to be solved in the field of new energy vehicle insurance.

[0004] Existing auto insurance pricing methods mainly include insurance pricing methods based on traditional statistical models, which rely on fewer data sources, are unable to comprehensively update pricing across multiple dimensions, and lack the ability to personalize and dynamically adjust. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a vehicle intelligent premium pricing method, system, terminal and storage medium to solve the above-mentioned technical problems.

[0006] In a first aspect, the present invention provides a vehicle intelligent premium pricing method, comprising: S1, based on the vehicle-mounted intelligent device, collects vehicle operation data, vehicle basic data and geographic location basic data, obtains driver basic data, obtains external weather data for the period corresponding to the operation data, and the driver uploads all data to the insurance company terminal; S2, extract features from vehicle operation data and all basic data respectively, calculate the importance of each feature based on the machine learning algorithm, assign weights to each feature based on the importance of each feature, and perform weighted fusion of all features to obtain a unified feature vector; S3, building a deep reinforcement learning model based on the target task, and training the deep reinforcement learning model based on the feature vector, wherein the target task includes premium pricing, driver risk assessment, and loss prediction, and the deep reinforcement learning model outputs a premium pricing strategy, a driver risk score, and a loss prediction value; S4. Obtain the basic premium pricing, calculate the preliminary premium based on the basic premium pricing combined with the premium pricing strategy, obtain the risk adjustment factor based on the driver risk score and the loss prediction value, and adjust the preliminary premium based on the risk adjustment factor combined with the external weather data to obtain the optimal premium.

[0007] In an alternative implementation, when constructing the deep reinforcement learning model in step S3, initialize all the parameters of the deep reinforcement learning model. The deep reinforcement learning model specifically includes: A shared layer that extracts general features from the input feature vector through a convolutional layer and a fully connected layer; A task-specific layer that configures an independent fully connected layer for each target task, and each fully connected layer makes decisions for the target task based on the general features; An output layer that configures an independent output layer for each target task, and each output layer belongs to the prediction result of the corresponding target task; In the task-specific layer, set the target loss function for each target task based on the reward of reinforcement learning. The target loss function is expressed as:

[0008] Among them, is the target loss function of target task i, is the immediate reward of target task i, is the state information composed of the input feature vector, is the action information related to the premium pricing strategy, is the current task i in state when taking action the estimated value of the long-term cumulative reward expected to be obtained, and its optimization goal is to minimize the mean square error; Combine the target loss functions of all target tasks to obtain the total loss function and calculate the loss value. When the loss value is greater than the preset threshold, adjust the parameters of the deep reinforcement learning model to obtain the final deep reinforcement learning model; the total loss function is specifically:

[0009] Among them, is the weighting coefficient of task i.

[0010] In an alternative implementation, in step S4, the premium calculation is specifically:

[0011] Among them, i is the vehicle index, is the premium pricing strategy of the i-th vehicle based on the state and is the feature of the vehicle operation data. is the feature of all basic data, is the pricing function optimized by the deep reinforcement learning model, is the preliminary premium of the i-th vehicle;

[0012] Among them, is the optimal premium of the i-th vehicle, is the risk adjustment coefficient, is the adjustment factor for the preliminary premium of the i-th vehicle based on risk assessment, is the external risk coefficient of the i-th vehicle..

[0013] In an alternative implementation, during the training of the deep reinforcement learning model, experiences are composed based on state information, action information, and immediate rewards, and the priorities of experience replay are adjusted based on the TD error of the state, state uncertainty, and risk assessment value of the state. The priority is specifically calculated as: ; Among them, is the TD error when taking action at time t in state , is the weighting coefficient of the TD error, is the uncertainty of state at time t, is the weighting coefficient of state uncertainty, is the risk assessment value of state at time t, is the weighting coefficient of the risk assessment value; Before model training, initialize the priority experience replay pool, collect a preset number of initial experience data, and store the initial experience data in the replay pool based on the calculated priorities; During model training, preferentially select experience data with high priorities for model training, calculate the priorities of new experience data generated during model training, and store them in the replay pool based on the priorities.

[0014] In an alternative implementation, during model training, the model performs exploration operations and exploitation operations, and based on the policy, the exploration rate is reduced in combination with the training steps to balance the exploration operations and exploitation operations. The reduction of the exploration rate is specifically calculated as:

[0015] Among them, is the initial exploration rate, is the decay factor, t is the training step, is the exploration rate when the number of training steps is t.

[0016] In an alternative embodiment, step S2 specifically includes: Using the random forest algorithm, taking all features of the extracted vehicle operation data and all basic data, as well as the corresponding claim probability of each feature as input, performing model training, and during the training process, obtaining the importance of each feature based on the claim probability by calculating the contribution degree of each feature to reducing the sample impurity in all decision trees. The specific calculation is as follows:

[0017] where C is the number of feature categories, is the proportion of samples belonging to class i in node t; Based on feature When splitting node t, the information gain of each feature after splitting is:

[0018] where, is the total number of samples in node t, is the number of samples in the left child node after splitting, is the number of samples in the right child node after splitting, is the left child node, is the right child node; The importance of each feature is obtained by averaging the information gain of the feature in all decision trees

[0019] where, The importance of feature , is the number of decision trees in the random forest, is the set of all nodes of the m-th decision tree in the random forest; Configuring weights for each feature based on the importance degree of each feature is specifically:

[0020] where, is the j-th feature weight, and A is the total number of features.

[0021] In an alternative embodiment, in step S4, after calculating the optimal premium, calculate the deviation between the actual premium and the optimal premium, and use the calculated deviation to adjust the parameters of the deep reinforcement learning model in combination with the gradient descent method, specifically:

[0022] where, is the parameter of the current model at time t, is the parameter updated at time t+1, is the learning rate, is the gradient of the loss function of the j-th feature with respect to the parameter; and based on the deviation, a mechanism that balances the exploration operation and the exploitation operation is used to adjust the weight of each feature in the input feature vector, specifically:

[0023] wherein, is the weight of the j-th feature at time t, is the adjustment step size, is the deviation between the actual premium and the optimal premium of the j-th feature, is the gradient of the loss function of the j-th feature with respect to the feature weight, is the optimal premium.

[0024] In a second aspect, the present invention provides a vehicle intelligent premium pricing system. When the system is implemented, it executes the above-mentioned vehicle intelligent premium pricing method. The system includes: A data acquisition module, which acquires vehicle operation data, vehicle basic data, and geographical location basic data based on an in-vehicle intelligent device, obtains driver basic data, and the driver uploads all the data to the insurance company terminal; A feature fusion module, which respectively extracts features from all the data, calculates the importance degree of each feature based on a machine learning algorithm, configures weights for each feature based on the importance degree of each feature, and performs weighted fusion on all the features to obtain a unified feature vector; A model construction module, which constructs a deep reinforcement learning model based on a target task, and trains the deep reinforcement learning model based on the feature vector. The target task includes premium pricing, driver risk assessment, and loss prediction. The deep reinforcement learning model outputs a premium pricing strategy, a driver risk score, and a loss prediction value; A premium calculation module, which obtains a basic premium pricing, calculates a preliminary premium based on the basic premium pricing and the premium pricing strategy, obtains a risk adjustment factor based on the driver risk score and the loss prediction value, and adjusts the preliminary premium based on the risk adjustment factor to obtain the optimal premium.

[0025] In a third aspect, a terminal is provided, including: A processor and a memory, wherein, the memory is used to store a computer program, the processor is used to call and run the computer program from the memory, so that the terminal executes the above-mentioned method of the terminal.

[0026] Fourthly, a computer-readable storage medium is provided, in which instructions are stored. When the instructions run on a computer, the computer is enabled to execute the methods described in the above aspects.

[0027] The beneficial effects of the present invention are as follows. The vehicle intelligent premium pricing method, system, terminal and storage medium provided by the present invention respectively collect vehicle operation data and all basic data through in-vehicle intelligent devices and external acquisition devices, and perform feature extraction, weight configuration and fusion. Based on the target task, a deep reinforcement learning model is constructed and trained. Finally, the optimal premium that comprehensively considers the premium pricing strategy, driver risk score and loss prediction value can be calculated, and the insurance pricing is dynamically adjusted to ensure that the pricing can timely and accurately reflect the actual risk level of the vehicle owner. According to the feedback signal and the change of the market environment, the pricing strategy is quickly adjusted to ensure that the insurance cost can timely reflect the new risk change, avoid the limitations of the fixed strategy, identify and control high-risk events in advance, and ensure the risk controllability in the dynamic market environment.

[0028] In addition, the design principle of the present invention is reliable and the structure is simple, having a very wide application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the present invention, the drawings required to be used in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 is a schematic flowchart of a vehicle intelligent premium pricing method according to an embodiment of the present invention.

[0031] Figure 2 is a flowchart of deep reinforcement learning modeling according to an embodiment of the present invention.

[0032] Figure 3 is a schematic block diagram of a vehicle intelligent premium pricing system according to an embodiment of the present invention.

[0033] Figure 4 is a schematic structural diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] To enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.

[0036] The following are the keyword explanations of the present invention: T - BOX is the abbreviation of English Telematic BOX, generally referring to the intelligent in - vehicle terminal in the vehicle networking system. The T - BOX can be connected to numerous electronic control units (ECUs) through the in - vehicle CAN bus, such as the engine management system, transmission control unit, etc. It can collect rich vehicle operation data in real - time, covering various information such as vehicle speed, engine speed, coolant temperature, battery power, fault codes, etc. Then, with the help of built - in communication modules, such as 4G, 5G or NB - IoT modules, according to specific communication protocols, it can quickly and stably transmit these data to the remote server, providing basic data support for the remote monitoring, fault diagnosis and performance analysis of the vehicle.

[0037] The vehicle owner can achieve remote control of the vehicle through a mobile application (APP) or other intelligent terminals. This function is extremely convenient. For example, remotely unlocking or locking the vehicle door, without the need for manual operation beside the vehicle, which is convenient for the vehicle owner to let others pick up or place items in special situations; remotely starting or shutting down the engine, in cold winters or hot summers, the vehicle owner can start the vehicle in advance for pre - heating or refrigeration; and can also remotely adjust parameters such as air - conditioner temperature and air volume, enabling the vehicle owner to enjoy a comfortable driving environment when getting in the vehicle.

[0038] The T - BOX is built - in with a high - precision GPS (Global Positioning System) or Beidou satellite positioning module, which can accurately obtain the geographical location information of the vehicle in real - time. This function not only guides the vehicle owner in daily navigation to ensure smooth travel; when the vehicle is stolen unfortunately, the vehicle owner and law enforcement departments can also quickly locate the vehicle based on the real - time position tracking information provided by the T - BOX, increasing the probability of retrieving the vehicle.

[0039] The vehicle intelligent premium pricing method provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the vehicle intelligent premium pricing system runs in the computer device.

[0040] Figure 1 It is a schematic flowchart of a vehicle intelligent insurance premium pricing method according to an embodiment of the present invention. Among them, Figure 1 The execution subject can be a vehicle intelligent insurance premium pricing system. According to different requirements, the order of steps in this flowchart can be changed, and some can be omitted.

[0041] As Figure 1 shown, the method includes: S1, based on the in-vehicle intelligent device, collect vehicle operation data, vehicle basic data, and geographical location basic data, obtain driver basic data, and obtain external weather data for the corresponding period of the operation data. The driver uploads all the data to the insurance company terminal; Use the in-vehicle T-Box device to obtain the operation data of the vehicle in real time through its built-in sensors and data transmission modules, such as vehicle speed, battery power, vehicle fault information, etc. Connect to all basic data sources (such as the traffic management department database, vehicle maintenance record system, etc.) through interfaces, and use API calls or data scraping tools to obtain driver basic data (such as age, driving experience, driving violation records, etc.), vehicle basic data (such as vehicle model, purchase time, maintenance history, etc.), and geographical location basic data (such as current location, traffic conditions in the frequently traveled area, etc.). Users regularly upload the data collected by all devices to the management terminal of the insurance company. The insurance company provides discount services to users who regularly upload real multi-source data. The insurance company obtains the external weather data of the vehicle's historical driving area through the meteorological data interface.

[0042] It provides a comprehensive data basis for subsequent insurance pricing and risk assessment, can more accurately reflect the actual use of the vehicle and the behavior characteristics of the driver, and improves the richness and timeliness of the data compared with the traditional method that only relies on limited static data.

[0043] S2, perform feature extraction on the vehicle operation data and all basic data respectively, calculate the importance of each feature based on machine learning algorithms, configure weights for each feature based on the importance of each feature, and perform weighted fusion on all features to obtain a unified feature vector; For vehicle operation data, time series analysis methods are used to extract features such as the mean and variance of vehicle speed, the consumption rate of battery power, etc.; for driver basic data, classification coding and statistical feature extraction are carried out, such as dividing driving experience into different intervals and counting the proportions of each interval; for vehicle basic data, features such as the service life of the vehicle and the number of repairs are extracted; for geographical location basic data, features such as the regional accident incidence rate and traffic congestion index are extracted. Then, machine learning algorithms such as random forest or XGBoost are used to calculate the importance of each feature through model training and feature importance evaluation, assign weights to each feature according to the importance, and fuse all features by means of weighted summation or concatenation to obtain a unified feature vector.

[0044] Through scientific feature extraction and weight configuration, the most valuable information for insurance pricing and risk assessment can be screened out, the interference of redundant data can be reduced, the quality and effectiveness of the model input data can be improved, enabling the model to capture key information in the data more accurately, thereby enhancing the accuracy of insurance pricing and risk assessment.

[0045] S3. Build a deep reinforcement learning model based on the target tasks, and train the deep reinforcement learning model based on the feature vector. The target tasks include premium pricing, driver risk assessment, and loss prediction. The deep reinforcement learning model outputs a premium pricing strategy, a driver risk score, and a loss prediction value. Build a deep reinforcement learning architecture that includes a shared underlying feature extraction layer and independent task layers. It can optimize multiple insurance business-related tasks simultaneously. Through shared feature extraction and multi-task collaborative learning, the learning efficiency and generalization ability of the model are improved. Compared with traditional single-task models, it can better adapt to the complexity of insurance business and can continuously self-optimize according to market changes and the influx of new data, providing more reliable strategies and predictions for insurance pricing and risk assessment.

[0046] S4. Obtain the basic premium pricing, calculate the preliminary premium based on the basic premium pricing combined with the premium pricing strategy, obtain the risk adjustment factors based on the driver risk score and the loss prediction value, and adjust the preliminary premium based on the risk adjustment factors combined with external weather data to obtain the optimal premium.

[0047] Determine the basic premium pricing according to the basic information of the vehicle (such as vehicle model, purchase price, etc.) and the basic pricing rules of the industry. Input the current state data of the vehicle and driver information into the trained deep reinforcement learning model to obtain the premium pricing strategy, and adjust the basic premium pricing according to this strategy to obtain the preliminary premium. According to the driver risk score and loss prediction value output by the model, calculate the risk adjustment factors using methods such as Value at Risk (VaR) or Conditional Value at Risk (CVaR), and apply the risk adjustment factors and external weather factors to the preliminary premium to obtain the optimal premium.

[0048] It comprehensively considers the impacts of various factors on the insurance premium, realizes the personalized customization and dynamic adjustment of the insurance premium, and can more accurately reflect the actual risk levels of the vehicle and the driver.

[0049] Optionally, as an embodiment of the present invention, step S1 specifically includes: Use a T-Box device to obtain the real-time running data of the vehicle, including but not limited to vehicle speed, fuel consumption, battery status, braking conditions, driving behaviors (such as hard braking, acceleration, etc.), and transmit the data collected by the T-Box device to the cloud platform or local server in real time through the vehicle-mounted communication module; Use an external interface to obtain the driver's personal information and behavior data (such as driving habits, age, gender, etc.), geographical location data (such as current location, surrounding traffic conditions, accident incidence rate, etc.), vehicle type and maintenance status (such as vehicle model, repair history, component health status, etc.), traffic flow and accident frequency data, and the owner's historical claim records in real time through methods such as APIs and data transmission protocols; The vehicle owner integrates all the obtained data and regularly uploads it to the insurance company's management terminal according to the requirements of the insurance company. If the insurance company requires that the running data comes from the T-Box device, then the vehicle owner needs to provide the data with the T-Box device identifier collected by the T-Box device when uploading, ensuring that the vehicle owner himself / herself is not easy to tamper with the data.

[0050] Clean and denoise all the data, remove duplicate data, outliers, and noise to ensure the validity of the data; unify the formats and units of different data sources.

[0051] Convert the data formats of all the data, convert different formats of data (such as GPS coordinates, timestamps, sensor data, etc.) into a unified data format for subsequent processing.

[0052] Fuse all the data, perform time series alignment and fusion on the T-Box data and all the basic data source data to ensure that the data at the same moment can be accurately associated.

[0053] Store the processed data in an efficient data storage system to ensure the accessibility and reliability of the data: To improve the system response speed, cache technologies (such as Redis, Memcached, etc.) can be used to temporarily store frequently used data and reduce the computational burden.

[0054] During the collection and storage process, ensure the security of data and the protection of vehicle owners' privacy. Encrypt sensitive data such as vehicle owners' personal information and driving data to ensure the security of data during transmission and storage. Only authorized users or systems can access relevant data. Protect personal privacy in accordance with relevant regulations (such as GDPR, CCPA, etc.) to ensure that the privacy rights of vehicle owners are not violated during the data collection and use process.

[0055] Optionally, as an embodiment of the present invention, step S2 specifically includes: First, after the data collection and fusion are completed, the data needs to be cleaned, preprocessed, and feature selected to ensure data quality and provide effective input for subsequent modeling; Then, extract features from the vehicle operation data and all basic data respectively. Extract statistical features from time series data such as driving behavior and vehicle status. Specifically, use the sliding window method to extract the average value, maximum value, and minimum value of each time window; Generate new combined features for data with interaction behaviors such as vehicle speed and braking frequency; After the feature extraction is completed, use the random forest algorithm. Take all the features of the extracted vehicle operation data and all basic data as input for model training. During the training process, calculate the importance of each feature through the formula, specifically calculated as:

[0056] where, is the j-th feature, is the impact of feature on the model performance during the i-th training process, is the number of trained models, is the importance of the j-th feature; Configure weights for each feature based on the importance of each feature, specifically as:

[0057] where, is the j-th feature weight, and N is the total number of features; For different features, use the weighted average method for fusion; When calculating the importance of features during model training, use the cross-validation method. Divide the data set into multiple subsets, perform training and validation on different subsets, comprehensively evaluate the model performance, avoid biases caused by data set division, and promptly detect overfitting or underfitting trends.

[0058] Optionally, as an embodiment of the present invention, refer to Figure 2, when constructing the deep reinforcement learning model in step S3, all the parameters of the deep reinforcement learning model are initialized. The deep reinforcement learning model specifically includes: S3-1, Model architecture design Design a reinforcement learning architecture based on the Deep Q-Network (DQN) as the deep reinforcement learning model, which specifically includes: Shared layer, extracting general features from the input feature vector through convolutional layers and fully connected layers; Task-specific layer, configuring an independent fully connected layer for each target task, and each fully connected layer makes decisions for the target task based on the general features; Output layer, configuring an independent output layer for each target task, and each output layer belongs to the prediction result of the corresponding target task; S3-2, Task objective definition Set a target loss function for each target task based on the reward of reinforcement learning in the task-specific layer. The target loss function is expressed as:

[0059] where, is the target loss function of target task i, is the immediate reward of target task i, is the state information composed of the input feature vector, is the action information related to the premium pricing strategy, is the state where the current task i is in when taking the action The estimated value of the long-term cumulative reward expected to be obtained, and its optimization goal is to minimize the mean square error; S3-3, Multi-task objective processing Combine the target loss functions of all target tasks to obtain the total loss function and calculate the loss value. When the loss value is greater than the preset threshold, adjust the parameters of the deep reinforcement learning model to obtain the final deep reinforcement learning model; the total loss function is specifically:

[0060] where, is the weighting coefficient of task i.

[0061] S3-4, Experience replay and priority adjustment During the training of the deep reinforcement learning model, form experiences based on the state information, action information, and immediate rewards, and adjust the priority of the experience replay based on the TD error of the state, state uncertainty, and the risk assessment value of the state. The priority is specifically calculated as: ; where, is the action taken at time t when in state The TD error when taking action is the TD error, is the weighting coefficient of the TD error, is the state at time t of uncertainty, is the weighting coefficient of the state uncertainty, is the state at time t of the risk assessment value, is the weighting coefficient of the risk assessment value; Before model training, initialize the prioritized experience replay pool, collect a preset number of initial experience data, and store the initial experience data in the replay pool based on the calculated priorities; During model training, preferentially select experience data with high priorities for model training, calculate the priorities of the new experience data generated during model training, and store them in the replay pool based on the priorities.

[0062] S3-5, Dynamic Exploration and Priority Adjustment During model training, the model performs exploration operations and exploitation operations. Exploration means that when the model faces an uncertain environment, it tries to take some new actions to discover new information or opportunities, and may even find a better strategy than the current one. Exploitation means that the model selects the currently considered optimal behavior based on existing knowledge to maximize the immediate reward. Based on the strategy, combine the training steps to reduce the exploration rate to balance the exploration operation and the exploitation operation. The specific calculation of reducing the exploration rate is:

[0063] where is the initial exploration rate, is the decay factor, t is the training step, is the exploration rate at training step t.

[0064] S3-6, Distributed Training and Model Optimization To accelerate the training process and use multiple computing nodes for large-scale parallel training, adopt a distributed training framework to achieve distributed training and parameter synchronization. In this way, multiple models can be trained simultaneously, and the training efficiency and convergence speed are improved through global parameter synchronization. Each model updates its parameters for different tasks during training and achieves model optimization through the synchronization of global parameters.

[0065] S3-7, Optimization Objectives and Learning Updates For each task i, its reinforcement learning update rule is:

[0066] Among them, is the learning rate of task i, is the discount factor, is the immediate reward of the current task; In the multi-objective task processing framework, the parameter sharing part of all tasks will be updated through the total loss function to ensure collaborative learning between different tasks.

[0067] S3-8, Joint Optimization of Task and Policy During the training process, the policies between different tasks will be jointly optimized to ensure that the optimization goals of multiple tasks can be achieved collaboratively. Through the model integration method, the policies of all tasks can be combined to select the optimal action policy for each task. For example, for premium pricing, the system can comprehensively score based on multi-dimensional information such as driver behavior and vehicle condition, and then output an optimal insurance price.

[0068] Optionally, as an embodiment of the present invention, in step S4, the pricing model is based on the deep reinforcement learning policy obtained through training (), this policy guides how to select the corresponding action according to the current state (i.e., pricing decision), and the premium calculation is specifically:

[0069] Among them, i is the vehicle index, is the premium pricing policy of the i-th vehicle based on the state ; is the feature of the vehicle operation data, is the feature of all basic data, is the pricing function optimized by the deep reinforcement learning model, is the preliminary premium of the i-th vehicle;

[0070] Among them, is the optimal premium of the i-th vehicle, is the risk adjustment coefficient, is the adjustment factor for the preliminary premium of the i-th vehicle based on risk assessment, is the external risk coefficient of the i-th vehicle.

[0071] Collect a large amount of vehicle accident data under different weather conditions, and count the accident frequency and average loss degree of each weather type to determine the risk coefficient. The weather conditions are divided into categories such as sunny, cloudy, overcast, light rain, moderate rain, heavy rain, rainstorm, light snow, moderate snow, heavy snow, blizzard, strong wind, and dust. For weather conditions like sunny, cloudy, and overcast that have less impact on driving, a relatively small risk coefficient is set; for light rain and light snow weather, since the road surface may be slippery, the coefficient is set slightly larger; for moderate rain, moderate snow, and strong wind weather, the risk increases, and the coefficient also increases accordingly; for heavy rain, heavy snow, and dust weather, which seriously affect visibility and road conditions, the coefficient is further increased; for rainstorm and blizzard weather, the driving risk is extremely high, and the coefficient is the largest.

[0072] Optionally, as an embodiment of the present invention, in step S4, the premium pricing calculation process is specifically as follows: Input feature acquisition: For each vehicle, first collect feature data from multiple data sources (such as T-Box devices, driver behavior data, traffic flow, etc.), and upload it to the insurance company's management terminal to obtain the status.

[0073] Input the status into the reinforcement learning model: Input the status data of the current vehicle into the deep reinforcement learning model, and use the policy in the model to output the pricing action. In this process, the model considers multiple factors, such as driving habits, driving environment, historical claims, etc.

[0074] Policy and risk adjustment: The calculation of insurance premiums is not only based on the pricing action of the current status but also needs to consider risk adjustment factors. For example, methods such as Value at Risk (VaR) or Conditional Value at Risk (CVaR) can be used for adjustment to ensure a full assessment of potential risks during the pricing process.

[0075] Dynamically adjust the insurance premium according to the change of data to ensure that the pricing of each vehicle accurately reflects its risk status.

[0076] Optionally, as an embodiment of the present invention, through data stream processing technology, combined with the exploration-exploitation balance in reinforcement learning, the model can automatically trigger the adjustment of pricing strategies when new data arrives. After calculating the optimal premium, calculate the deviation between the actual premium and the optimal premium, and use the calculated deviation combined with the gradient descent method to adjust the parameters of the deep reinforcement learning model;

[0077] Among them, is the deviation between the actual premium and the optimal premium of the jth feature, is the optimal premium predicted by the model, is the learning rate; The pricing adjustment formula is specifically:

[0078] Among them, is the input feature of the i-th vehicle, and is the current model parameter; The specific calculation for adjusting the model parameters using gradient descent or other optimization algorithms is:

[0079] Among them, is the parameter of the current model at time t, is the updated parameter at time t + 1, is the gradient of the loss function of the j-th feature with respect to the parameter; And based on the deviation, a mechanism that combines the exploration operation and the exploitation operation is used to adjust the weight of each feature in the input feature vector. Specifically:

[0080] Among them, is the weight of the j-th feature at time t, is the adjustment step size, is the gradient of the loss function of the j-th feature with respect to the feature weight.

[0081] Optionally, as an embodiment of the present invention, when performing risk assessment, according to the data streams regularly from multiple data sources (such as driver behavior, geographical location, vehicle type, etc.), multi-dimensional assessment of risks is carried out, and the pricing is dynamically adjusted by calculating the expected loss value (Expected Loss, EL) of the risk, and by combining risk measurement methods such as the value at risk (VaR) or the conditional value at risk (CVaR), the pricing is further optimized.

[0082] In some embodiments, the vehicle intelligent premium pricing system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the vehicle intelligent premium pricing system can be stored in the memory of the computer device and executed by at least one processor to execute (see in detail Figure 1 the description) the functions of vehicle intelligent premium pricing.

[0083] In this embodiment, according to the functions it performs, the vehicle intelligent premium pricing system can be divided into multiple functional modules, such as Figure 3 shown. The functional modules of the system may include: a data acquisition module, a feature fusion module, a model construction module, and a premium calculation module. The modules referred to in the present invention refer to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments. The system includes The data acquisition module collects vehicle operation data, vehicle basic data, and geographical location basic data based on in-vehicle intelligent devices, obtains driver basic data, obtains external weather data for the corresponding period of the operation data, and the driver uploads all the data to the insurance company terminal; The feature fusion module extracts features from the vehicle operation data and all basic data respectively, calculates the importance degree of each feature based on machine learning algorithms, configures weights for each feature based on the importance degree of each feature, and performs weighted fusion on all features to obtain a unified feature vector; The model construction module constructs a deep reinforcement learning model based on the target task, and trains the deep reinforcement learning model based on the feature vector. The target tasks include premium pricing, driver risk assessment, and loss prediction. The deep reinforcement learning model outputs a premium pricing strategy, a driver risk score, and a loss prediction value; The premium calculation module obtains the basic premium pricing, calculates the preliminary premium based on the basic premium pricing combined with the premium pricing strategy, obtains the risk adjustment factors based on the driver risk score and the loss prediction value, and adjusts the preliminary premium based on the risk adjustment factors combined with the external weather data to obtain the optimal premium.

[0084] Through the coordinated operation of the data acquisition module, the feature fusion module, the model construction module, and the premium calculation module, a complete process from multi-source data acquisition to deep reinforcement learning model training and then to accurate premium calculation is realized. The beneficial effects are as follows: it can make full use of in-vehicle and all basic data to comprehensively reflect the vehicle and driver conditions. After feature fusion and model training, tasks such as premium pricing, risk assessment, and loss prediction can be completed simultaneously, and finally the optimal premium considering various factors can be obtained, effectively improving the accuracy, scientificity, and rationality of insurance pricing, and enhancing the risk control ability and market competitiveness of insurance companies.

[0085] Figure 4 The figure is a schematic structural diagram of a terminal provided by an embodiment of the present invention, and this terminal can be used to execute the method for intelligent vehicle premium pricing provided by the embodiment of the present invention.

[0086] Among them, this terminal may include: a processor, a memory, and a communication unit. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0087] Among them, the memory can be used to store the execution instructions of the processor. The memory can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. When the execution instructions in the memory are executed by the processor, the terminal can execute some or all of the steps in the above method embodiments.

[0088] The processor is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and lines. By running or executing software programs and / or modules stored in the memory, and by calling the data stored in the memory, it executes various functions of the electronic terminal and / or processes data. The processor can be composed of an integrated circuit (IC). For example, it can be composed of a single packaged IC, or can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor can include only a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single arithmetic core or can include multiple arithmetic cores.

[0089] The communication unit is used to establish a communication channel so that the storage terminal can communicate with other terminals. It receives user data sent by other terminals or sends user data to other terminals.

[0090] The present invention also provides a computer-readable storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the various embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), etc.

[0091] Therefore, the technical effects that can be achieved by this embodiment can be referred to the description above, and will not be elaborated here.

[0092] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disc, etc., various media that can store program codes, including several instructions for causing a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0093] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the terminal embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0094] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or module can be in an electrical, mechanical or other form.

[0095] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0096] In addition, in each embodiment of the present invention, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0097] Although the present invention has been described in detail by reference to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all fall within the scope of the present invention / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. A vehicle intelligent premium pricing method, characterized in that: The following steps are involved: S1, based on the vehicle-mounted intelligent device, collects vehicle operation data, vehicle basic data and geographic location basic data, obtains driver basic data, obtains external weather data for the period corresponding to the operation data, and the driver uploads all data to the insurance company terminal; S2, extract features from all data, calculate the importance of each feature based on the machine learning algorithm, assign weights to each feature based on the importance of each feature, and perform weighted fusion of all features to obtain a unified feature vector; S3, building a deep reinforcement learning model based on the target task, and training the deep reinforcement learning model based on the feature vector, wherein the target task includes premium pricing, driver risk assessment, and loss prediction, and the deep reinforcement learning model outputs a premium pricing strategy, a driver risk score, and a loss prediction value; S4, obtains the basic premium pricing, calculates the preliminary premium based on the basic premium pricing combined with the premium pricing strategy, obtains the risk adjustment factor based on the driver risk score and the loss prediction value, and adjusts the preliminary premium based on the risk adjustment factor combined with external weather data to obtain the optimal premium.

2. The vehicle intelligent premium pricing method according to claim 1, characterized in that: When constructing the deep reinforcement learning model in step S3, all deep reinforcement learning model parameters are initialized. The deep reinforcement learning model specifically includes: The shared layer extracts common features from the input feature vector through the convolutional layer and the fully connected layer; Task-specific layer: an independent fully connected layer is configured for each target task. Each fully connected layer makes decisions for the target task based on common features. Output layer: an independent output layer is configured for each target task, and each output layer belongs to the prediction result of the corresponding target task; In the task-specific layer, the target loss function is set for each target task based on the reward calculation of reinforcement learning. The target loss function is expressed as: in, is the target loss function for target task i, is the immediate reward of target task i, is the state information composed of the input feature vector, is the action information related to the premium pricing strategy, The current task i is in state Take action when The estimated value of the expected long-term cumulative reward, whose optimization goal is to minimize the mean square error; The target loss functions of all target tasks are combined to obtain the total loss function and calculate the loss value. When the loss value is greater than the preset threshold, the parameters of the deep reinforcement learning model are adjusted to obtain the final deep reinforcement learning model; the total loss function is specifically: in, is the weight coefficient of task i.

3. The intelligent vehicle premium pricing method according to claim 2, characterized in that: In step S4, the premium calculation is specifically as follows: Where i is the vehicle index, is the i-th car based on the state The premium pricing strategy is the characteristic of vehicle operation data, is the characteristic of all basic data, is the pricing function optimized by the deep reinforcement learning model, is the initial premium of the ith car; in, is the optimal premium for the i-th car, is the risk adjustment factor, is the adjustment factor for the initial premium of the i-th vehicle based on risk assessment, is the external risk coefficient of the i-th vehicle.

4. The vehicle intelligent premium pricing method according to claim 2, characterized in that: When training a deep reinforcement learning model, experience is composed based on state information, action information, and immediate rewards. The priority of experience replay is adjusted based on the TD error of the state, the uncertainty of the state, and the risk assessment value of the state. The specific calculation of the priority is: ; in, is in the state at time t Take action when The TD error at is the weighting coefficient of TD error, is the state at time t uncertainty, is the weighting coefficient of state uncertainty, is the state at time t The risk assessment value of is the weighting coefficient of the risk assessment value; Before model training, initialize the priority experience replay pool, collect a preset amount of initial experience data, and store the initial experience data in the replay pool based on the calculated priority; During the model training process, high-priority experience data is preferentially selected for model training. The priority of new experience data generated during the model training process is calculated and stored in the replay pool based on the priority.

5. The vehicle intelligent premium pricing method according to claim 4, characterized in that: During the model training process, the model performs exploration and utilization operations based on The strategy combines the number of training steps to reduce the exploration rate to balance the exploration operation and the utilization operation. The specific calculation of reducing the exploration rate is: in, is the initial exploration rate, is the decay factor, t is the number of training steps, is the exploration rate when the number of training steps is t.

6. The vehicle intelligent premium pricing method according to claim 5, characterized in that: Step S2 specifically includes: The random forest algorithm is used to take the extracted vehicle operation data and all the features of all the basic data, as well as the probability of accident corresponding to each feature as input for model training. During the training process, the importance of each feature is obtained by calculating the contribution of each feature to reducing the sample impurity in all decision trees based on the probability of accident. The specific calculation is: Among them, C is the number of feature categories, is the proportion of samples belonging to category i in node t; Feature-based When node t is split, the information gain of each feature after the split is: in, is the total number of samples of node t, is the number of samples of the left child node after splitting, is the number of samples of the right child node after splitting, is the left child node, is the right child node; The importance of each feature is obtained by averaging the information gain of the feature in all decision trees. in, It is a feature The importance of is the number of decision trees in the random forest, is the set of all nodes of the mth decision tree; Based on the importance of each feature, the weights for each feature are configured as follows: in, is the jth feature weight, and A is the total number of features.

7. The vehicle intelligent premium pricing method according to claim 6, characterized in that: In step S4, after calculating the optimal premium, the deviation between the actual premium and the optimal premium is calculated, and the calculated deviation is combined with the gradient descent method to adjust the parameters of the deep reinforcement learning model, specifically: in, is the parameter of the current model at time t, is the updated parameter at time t+1, is the learning rate, is the gradient of the loss function of the jth feature with respect to the parameter; Based on the deviation, the weight of each feature in the input feature vector is adjusted by combining the mechanism of balanced exploration operation and exploitation operation, specifically: in, is the weight of the jth feature at time t, is the adjustment step size, is the deviation between the actual premium and the optimal premium for the jth feature, is the gradient of the loss function of the jth feature with respect to the feature weight, It is the best premium.

8. A vehicle intelligent premium pricing system, characterized in that: When the system is implemented, the vehicle intelligent premium pricing method according to any one of claims 1 to 7 is executed, and the system includes: The data collection module collects vehicle operation data, vehicle basic data and geographic location basic data based on the on-board intelligent device, obtains the driver's basic data, and obtains the external weather data of the corresponding period of the operation data. The driver uploads all data to the insurance company's management terminal; The feature fusion module extracts features from all data, calculates the importance of each feature based on the machine learning algorithm, assigns weights to each feature based on its importance, and performs weighted fusion of all features to obtain a unified feature vector. A model building module, which builds a deep reinforcement learning model based on a target task, and trains the deep reinforcement learning model based on a feature vector. The target task includes premium pricing, driver risk assessment, and loss prediction. The deep reinforcement learning model outputs a premium pricing strategy, a driver risk score, and a loss prediction value. The premium calculation module obtains the basic premium pricing, calculates the preliminary premium based on the basic premium pricing and the premium pricing strategy, obtains the risk adjustment factor based on the driver risk score and the loss prediction value, and adjusts the preliminary premium based on the risk adjustment factor and external weather data to obtain the optimal premium.

9. A terminal, characterized in that: include: A memory, used for storing a vehicle intelligent premium pricing program; A processor is used to implement the steps of the vehicle intelligent premium pricing method as described in any one of claims 1-7 when executing the vehicle intelligent premium pricing program.

10. A computer-readable storage medium, characterized in that: The readable storage medium stores a vehicle intelligent premium pricing program, and when the vehicle intelligent premium pricing program is executed by the processor, the steps of the vehicle intelligent premium pricing method as described in any one of claims 1-7 are implemented.