Vehicle risk prediction method and device, electronic equipment and storage medium
By analyzing vehicle driving routes and climate data and constructing a random forest model, the problem of ignoring dynamic environmental factors in auto insurance premium assessment was solved, and an accurate assessment of the vehicle's future driving risks was achieved, reducing the insurance company's economic losses.
Patent Information
- Application Number
- CN202511004327.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-10
AI Technical Summary
In existing technologies, the determination of auto insurance premiums relies on static risk factors and ignores dynamic external environmental factors such as natural disasters and severe weather conditions, resulting in inaccurate risk assessments that may cause insurance companies to increase compensation.
By obtaining the vehicle's driving route and status data, analyzing the historical climate data of the preferred driving area, building a random forest model, and combining the vehicle status and climate forecast data, a target risk prediction model is established to achieve an accurate assessment of the vehicle's future driving risks.
Accurately predict the future driving risks of vehicles, reduce the economic losses of insurance companies due to misjudgment of risks, and improve the accuracy of risk assessment.
Smart Images

Figure CN120765401A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology and is applicable to the field of financial technology, and in particular to a vehicle risk prediction method and device, electronic equipment, and storage medium. Background Art
[0002] Currently, risk assessments are typically conducted based on static risk factors related to auto insurance (such as the driver's age, driving history, and vehicle type), and premiums are determined based on these assessments. However, the impact of dynamic external environmental factors (such as natural disasters and adverse weather conditions) on driving risk is ignored, leading to inaccurate adjustments to insurance premiums. This can ultimately underestimate vehicle risk and lead to increased insurance claims. Therefore, accurately predicting vehicle driving risk has become a pressing technical challenge. Summary of the Invention
[0003] The main purpose of the embodiments of the present application is to propose a vehicle risk prediction method and device, electronic equipment and storage medium, aiming to accurately predict the driving risk of a vehicle.
[0004] To achieve the above objectives, a first aspect of an embodiment of the present application provides a vehicle risk prediction method, the method comprising:
[0005] Acquiring sample driving routes and sample vehicle status data of sample vehicles, and performing driving area analysis based on the sample driving routes to obtain preferred driving areas;
[0006] Acquiring historical climate data of the preferred driving area, and performing climate forecasting on the preferred driving area based on the historical climate data to obtain sample climate forecast data;
[0007] constructing a feature vector based on the sample vehicle state data and the sample climate forecast data to obtain a sample vehicle feature vector;
[0008] Performing risk analysis on the sample vehicle feature vector to obtain a sample risk level, and using the sample risk level as a label for the sample vehicle feature vector;
[0009] A random forest model is constructed based on the sample vehicle feature vector and the sample risk level to obtain a target risk prediction model;
[0010] A historical driving route and target vehicle status data of a target vehicle are obtained, and risk prediction is performed on the historical driving route and the target vehicle status data based on the target risk prediction model to obtain a target risk level of the target vehicle.
[0011] In some embodiments, the random forest model construction based on the sample vehicle feature vector and the sample risk level obtains a target risk prediction model, comprising:
[0012] Random sampling is performed on the sample vehicle feature vector to obtain a sample feature vector set;
[0013] A decision tree is constructed according to the sample feature vector set to obtain a risk prediction decision tree;
[0014] The random sampling is performed on the sample vehicle feature vector to obtain a sample feature vector set until the number of risk prediction decision trees reaches a first preset number, and the target risk prediction model is determined based on the preset number of risk prediction decision trees.
[0015] In some embodiments, the decision tree construction according to the sample feature vector set to obtain a risk prediction decision tree comprises:
[0016] A root node is constructed according to the sample feature vector set, and an initial vehicle feature dimension is randomly selected from the feature dimensions of the sample vehicle feature vector;
[0017] Gini impurity calculation is performed on the sample feature vector set according to each initial vehicle feature dimension to obtain vehicle feature Gini data;
[0018] Based on the vehicle feature Gini data, the initial vehicle feature dimension, and the sample feature vector, a node is constructed to obtain a split node, and the split node is a subsequent node of the root node;
[0019] The risk prediction decision tree is determined based on the root node and the split node.
[0020] In some embodiments, the node construction based on the vehicle feature Gini data, the initial vehicle feature dimension, and the sample feature vector to obtain a split node comprises:
[0021] A target vehicle feature dimension is selected from the initial vehicle feature dimension according to the vehicle feature Gini data, and the split node is constructed based on the target vehicle feature dimension;
[0022] A sample feature vector subset is obtained by dividing the sample feature vector set according to the split node;
[0023] The random selection is performed on the feature dimensions of the sample vehicle feature vector to obtain an initial vehicle feature dimension until the number of sample vehicle feature vectors in the sample feature vector subset reaches a second preset number.
[0024] In some embodiments, the risk prediction based on the target risk prediction model on the historical driving route and the target vehicle state data obtains a target risk level, comprising:
[0025] driving area analysis according to the historical driving route obtains a target driving area, and climate prediction is performed on the target driving area to obtain reference climate prediction data;
[0026] According to the vehicle state data, the reference climate prediction data, a feature vector is constructed to obtain a target vehicle feature vector;
[0027] Each risk prediction decision tree of the target risk prediction model is used to evaluate the risk level of the target vehicle feature vector to obtain a plurality of candidate risk levels and the number of prediction decision trees output for each candidate risk level;
[0028] According to the number of prediction decision trees, an intermediate risk level is selected from a plurality of candidate risk levels, and the target risk level is determined based on the intermediate risk level.
[0029] In some embodiments, the target risk level is determined based on the intermediate risk level, comprising:
[0030] In response to a natural disaster warning in the target driving area, a reference risk level is obtained by querying a preset risk level coefficient mapping table according to the natural disaster warning;
[0031] The reference risk level and the intermediate risk level are superimposed to obtain the target risk level.
[0032] In some embodiments, the climate prediction of the historical climate data on the preferred driving area obtains sample climate prediction data, comprising:
[0033] According to a preset first time period, the historical climate data is sliced to obtain at least two historical climate sub-data;
[0034] Each of the historical climate sub-data is identified for disaster events to obtain a disaster occurrence probability of each disaster type in the historical climate sub-data, and the disaster occurrence probability is added to the historical climate sub-data;
[0035] At least two historical climate sub-data are predicted for climate trend based on a long short-term memory network to obtain the sample climate prediction data.
[0036] To achieve the above purpose, a second aspect of an embodiment of the present application proposes a vehicle risk prediction device, the device comprising:
[0037] a first acquisition module, configured to acquire sample driving routes and sample vehicle status data of sample vehicles, and perform driving area analysis based on the sample driving routes to obtain preferred driving areas;
[0038] a second acquisition module, configured to acquire historical climate data of the preferred driving area, and perform climate prediction on the preferred driving area based on the historical climate data to obtain sample climate prediction data;
[0039] a feature vector construction module, configured to construct a feature vector based on the sample vehicle state data and the sample climate prediction data to obtain a sample vehicle feature vector;
[0040] a risk analysis module, configured to perform risk analysis on the sample vehicle feature vector to obtain a sample risk level, and use the sample risk level as a label for the sample vehicle feature vector;
[0041] A risk prediction model construction module is used to construct a random forest model based on the sample vehicle feature vector and the sample risk level to obtain a target risk prediction model;
[0042] The risk prediction module is used to obtain the historical driving route and target vehicle status data of the target vehicle, perform risk prediction on the historical driving route and the target vehicle status data based on the target risk prediction model, and obtain the target risk level of the target vehicle.
[0043] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.
[0044] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.
[0045] The vehicle risk prediction method and device, electronic device and storage medium proposed in this application collect the driving route and status data of the vehicle, and determine the preferred driving area of the vehicle based on the sample driving route, thereby further obtaining the historical climate data of the driving area and performing climate prediction on it, and obtaining prediction data that accurately reflects the future driving environment. Then, the sample vehicle status data reflecting the vehicle's operating status is combined with the predicted sample climate prediction data to construct a comprehensive vehicle feature vector. The feature vector is then trained using a random forest algorithm to establish a target risk prediction model, ultimately achieving an accurate risk level prediction for the target vehicle. As a result, the embodiment of the present application simultaneously considers the static factors of the vehicle's own status and the dynamic environmental factors, achieving an accurate prediction of the vehicle's future driving risk, and solving the problem of inaccurate prediction caused by ignoring environmental changes in the traditional static risk factor analysis method, thereby enabling insurance companies to effectively reduce economic losses caused by misjudgment of risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flow chart of a vehicle risk prediction method provided by an embodiment of the present application;
[0047] Figure 2 yes Figure 1 Flowchart of step S102 in FIG.
[0048] Figure 3 yes Figure 1 Flowchart of step S105 in FIG.
[0049] Figure 4 yes Figure 3 Flowchart of step S302 in FIG.
[0050] Figure 5 yes Figure 4 Flowchart of step S403 in FIG.
[0051] Figure 6 yes Figure 1 Flowchart of step S106 in FIG.
[0052] Figure 7 Schematic diagram of the structure of the vehicle risk prediction device provided in an embodiment of the present application;
[0053] Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0055] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be performed in a manner different from the module division in the device or the sequence in the flowchart. The terms "first", "second", and the like in the specification and claims and the above-described drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application is for the purpose of describing the embodiments of the application only, and is not intended to limit the application.
[0057] First, the meanings of several terms involved in the present application are explained:
[0058] Artificial intelligence (AI): is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science, artificial intelligence aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. The field of research includes robots, language recognition, image recognition, natural language processing and expert systems. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, to perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0059] At present, risk assessment is usually made according to static risk factors related to vehicle insurance (such as the age, gender, driving history, vehicle type, etc. of the driver), and the vehicle insurance premium is determined according to the risk assessment result. However, the influence of dynamic external environmental factors (such as natural disasters and adverse weather conditions) on driving risk is ignored, resulting in inaccurate insurance fee adjustment, which may eventually underestimate the risk of the vehicle and increase the claims of the insurance company. Therefore, how to accurately predict the driving risk of the vehicle has become a technical problem to be solved.
[0060] Based on this, the embodiments of the present application provide a vehicle risk prediction method and device, electronic equipment and storage medium, which aims to accurately predict the driving risk of the vehicle.
[0061] The vehicle risk prediction method and device, electronic equipment and storage medium provided by the embodiments of the present application are specifically explained by the following embodiments, first, the vehicle risk prediction method in the embodiments of the present application is described.
[0062] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0063] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0064] The vehicle risk prediction method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The vehicle risk prediction method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the vehicle risk prediction method, etc., but is not limited to the above forms.
[0065] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0066] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0067] Figure 1 This is an optional flowchart of the vehicle risk prediction method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S106.
[0068] Step S101: Acquire sample driving routes and sample vehicle status data of sample vehicles, and perform driving area analysis based on the sample driving routes to obtain preferred driving areas.
[0069] Step S102 : acquiring historical climate data of the preferred driving area, and performing climate prediction on the preferred driving area based on the historical climate data to obtain sample climate prediction data.
[0070] Step S103 , constructing a feature vector based on the sample vehicle state data and the sample climate prediction data to obtain a sample vehicle feature vector.
[0071] Step S104 , performing risk analysis on the sample vehicle feature vector to obtain a sample risk level, and using the sample risk level as a label for the sample vehicle feature vector.
[0072] Step S105 , constructing a random forest model based on the sample vehicle feature vector and the sample risk level to obtain a target risk prediction model.
[0073] Step S106: Obtain the historical driving route and target vehicle status data of the target vehicle, perform risk prediction on the historical driving route and target vehicle status data based on the target risk prediction model, and obtain the target risk level of the target vehicle.
[0074] In steps S101 to S106 shown in the embodiment of the present application, by collecting the driving route and status data of the vehicle, and determining the preferred driving area of the vehicle based on the sample driving route, the historical climate data of the driving area is further obtained and the climate forecast is performed on it, so as to obtain forecast data that accurately reflects the future driving environment. Then, the sample vehicle status data reflecting the vehicle's operating status is combined with the predicted sample climate forecast data to construct a comprehensive vehicle feature vector. The feature vector is then trained using a random forest algorithm to establish a target risk prediction model, and ultimately an accurate risk level prediction of the target vehicle is achieved. Therefore, the embodiment of the present application simultaneously takes into account the static factors of the vehicle's own state and the dynamic environmental factors, achieves an accurate prediction of the vehicle's future driving risk, and solves the problem of inaccurate prediction caused by ignoring environmental changes in the traditional static risk factor analysis method, thereby enabling insurance companies to effectively reduce economic losses caused by misjudgment of risks.
[0075] In step S101 of some embodiments, the sample vehicle refers to the vehicle used to collect driving data for building the risk prediction model. Its specific information can be obtained through the information registered by the platform user. The sample driving route refers to the driving path information recorded by the sample vehicle during actual driving. The sample vehicle status data refers to the operating status data of the vehicle itself, such as vehicle type, age, braking status, maximum engine speed, and distance traveled. In some embodiments, the driver's driving habits and owner information (such as age, place of residence, and driving history) may also be included.
[0076] Preferred driving areas refer to the geographic areas where sample vehicles frequently travel, which can be provinces, cities, or towns. Cluster analysis algorithms can be used to identify concentrated areas where vehicles frequently travel using the driving route coordinates collected by the vehicles. For example, a month's worth of driving route coordinate data can be spatially clustered using the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm to determine that the sample vehicles' preferred driving area is a specific city.
[0077] In step S102 of some embodiments, see Figure 2 Step S102 may include but is not limited to steps S201 to S203:
[0078] Step S201 , slicing the historical climate data according to a preset first time period to obtain at least two historical climate sub-data.
[0079] Step S202, disaster event identification is performed on each historical climate sub-data to obtain the disaster occurrence probability of each disaster type in the historical climate sub-data, and the disaster occurrence probability is added to the historical climate sub-data.
[0080] Step S203, climate trend prediction is performed on at least two historical climate sub-data based on a long short-term memory network to obtain sample climate prediction data.
[0081] In step S201 of some embodiments, the historical climate data refers to a set of historical values of environmental parameters such as air temperature, precipitation, wind speed, humidity, etc. recorded in a specific time period in the past to characterize the climate conditions of the preferred driving area. This is because the occurrence probability of some natural disasters has certain relevance to the geographical location of the area, for example, the coastal area has a higher probability of typhoon in summer.
[0082] Further, the continuously recorded historical climate data is divided into a fixed time window to generate a plurality of historical climate sub-data with the same length. For example, if the first time period is set to one quarter (3 months), the historical climate data of 12 consecutive months is divided into 4 historical climate sub-data, each with a length of 3 months, representing 4 seasons respectively.
[0083] In step S202 of some embodiments, a convolutional neural network (CNN) classifier can be used to identify the disaster type of the historical climate sub-data based on the numerical features of the air temperature, precipitation, wind speed, etc. of the historical climate sub-data. For example, the CNN model is used to input the air temperature, precipitation, wind speed, etc. of the historical climate sub-data, and output the occurrence probability of disaster types such as rainstorm, flood, typhoon, etc. in the historical climate sub-data, and the corresponding disaster occurrence probability value is added as a new feature dimension to the original historical climate sub-data to form updated historical climate sub-data. For example, the probability of rainstorm is 0.75 and the probability of flood is 0.60 in the historical climate sub-data, then the above two probability values are added to the feature dimension of the historical climate sub-data.
[0084] In step S203 of some embodiments, a long short-term memory network (LSTM) is used for time series prediction, which can effectively capture the sequence features of each climate parameter in the historical climate sub-data and accurately predict the future climate trend. For example, the historical climate sub-data containing air temperature, precipitation, wind speed and disaster occurrence probability is input into the LSTM model, and after model training and learning, the air temperature, precipitation, wind speed change and disaster probability trend of the preferred driving area in the next week are output, thereby obtaining sample climate prediction data.
[0085] The steps S201 to S203 shown in the embodiments of the present application, by slicing the historical climate data in time periods and combining the probability information of disaster events, the recognition and prediction ability of the climate trend prediction model for disaster events is effectively improved, the accurate evaluation of the future climate risk of the driving area is realized, and the accuracy of the future driving risk prediction of the vehicle is improved.
[0086] In step S103 of some embodiments, the parameters in the vehicle state data are numerically processed and normalized to form a vector of uniform length. For example, the average speed, maximum acceleration, emergency braking frequency of the vehicle, and predicted temperature, precipitation, wind speed, etc. are selected as elements to form a 10-dimensional sample vehicle feature vector.
[0087] In step S104 of some embodiments, risk analysis refers to quantifying and classifying the future driving risk of the vehicle according to the feature vector by predetermined risk evaluation criteria. For example, according to the vehicle feature vector, the sample risk level is determined to be low risk, medium risk or high risk, and the risk level is used as the label of the corresponding feature vector. If the vehicle braking frequency of the sample feature vector is high and the natural disaster frequency in the predicted climate is high, it is determined to be high risk.
[0088] In step S105 of some embodiments, please refer to Figure 3 , step S105 can include but is not limited to steps S301 to S303:
[0089] Step S301, randomly sampling the sample vehicle feature vector to obtain a sample feature vector set.
[0090] Step S302, constructing a decision tree according to the sample feature vector set to obtain a risk prediction decision tree.
[0091] Step S303, return to step S301 to randomly sample the sample vehicle feature vector to obtain a sample feature vector set until the number of risk prediction decision trees reaches a first preset number, and determine a target risk prediction model based on the preset number of risk prediction decision trees.
[0092] In step S301 of some embodiments, the sample feature vector machine refers to randomly selecting a subset of feature vector data from the sample vehicle feature vector set with replacement. For example, assuming that there are 1000 sample vehicle feature vectors, 800 are extracted by random sampling with replacement as a sample feature vector set.
[0093] In step S302 of some embodiments, a sample feature vector set is used to construct a decision tree. Further, please refer to Figure 4Step S302 may include but is not limited to steps S401 to S404:
[0094] Step S401 : construct a root node according to the sample feature vector set, and randomly select from the feature dimensions of the sample vehicle feature vectors to obtain initial vehicle feature dimensions.
[0095] Step S402 , performing Gini impurity calculation on the sample feature vector set according to each initial vehicle feature dimension to obtain vehicle feature Gini data.
[0096] Step S403 , constructing nodes based on the vehicle feature Gini data, the initial vehicle feature dimension, and the sample feature vector to obtain split nodes, where the split nodes are subsequent nodes of the root node.
[0097] Step S404: determining a risk prediction decision tree based on the root node and the split nodes.
[0098] In step S401 of some embodiments, the sample feature vector set is set as the root node of the risk prediction decision tree. Several feature dimensions are randomly selected from the feature dimensions of the sample vehicle feature vectors (e.g., vehicle speed, acceleration, braking frequency, engine load, temperature, precipitation probability, etc.), and these feature dimensions are referred to as the initial vehicle feature dimensions.
[0099] In step S402 of some embodiments, for each initial vehicle feature dimension, the sample feature vector set is divided using different thresholds, the Gini impurity of the child nodes after division is calculated, and the feature dimension and corresponding threshold that can reduce the Gini impurity are selected. The vehicle feature Gini data represents the degree of impurity after the sample feature vector set is divided using a certain feature dimension. The lower the value, the higher the purity after division. For example, for the initial vehicle feature dimension of vehicle speed, thresholds of 50km / h and 70km / h are selected for division, and the Gini impurity of each subset is calculated. Finally, the Gini impurity is the smallest under the 50km / h division condition for vehicle speed. The Gini impurity value is the vehicle feature Gini data. Specifically, the implementation method of the Gini impurity calculation can refer to the following analytical formula:
[0100]
[0101] Among them, p k It represents the proportion of category k in the node, that is, the proportion of the number of samples of category k to the total number of samples in the node, and K represents the total number of categories K.
[0102] In step S403 of some embodiments, please refer to Figure 5 Step S403 may also include but is not limited to steps S501 to S503:
[0103] Step S501 : selecting a target vehicle feature dimension from the initial vehicle feature dimension according to the vehicle feature Gini data, and constructing a splitting node based on the target vehicle feature dimension.
[0104] Step S502 : dividing the sample feature vector set into data sets according to the splitting nodes to obtain sample feature vector subsets.
[0105] Step S503 , returning to step 504 , randomly selecting from the feature dimensions of the sample vehicle feature vectors to obtain initial vehicle feature dimensions, until the number of sample vehicle feature vectors in the sample feature vector subset reaches a second preset number.
[0106] In step S501 of some embodiments, the vehicle feature Gini data corresponding to each initial vehicle feature dimension is sorted, and the initial vehicle feature dimension with the lowest Gini impurity is selected as the target vehicle feature dimension. For example, among the initial vehicle feature dimensions of vehicle speed, braking frequency, and precipitation probability, if vehicle speed has the lowest Gini impurity, then vehicle speed is determined as the target vehicle feature dimension.
[0107] In step S502 of some embodiments, dividing the sample feature vector set into a data set according to the splitting node means using the target vehicle feature dimension and its corresponding optimal division threshold to split the original sample feature vector set into multiple sub-data sets. Specifically, based on the optimal division threshold of the target vehicle feature dimension, the sample vehicle feature vectors with eigenvalues higher than the threshold and the sample vehicle feature vectors with eigenvalues lower than or equal to the threshold are divided into different data subsets. For example, the optimal division threshold is determined to be 50km / h based on the target vehicle feature dimension vehicle speed, and the 400 sample vehicle feature vectors with vehicle speeds greater than 50km / h and the 400 sample vehicle feature vectors with vehicle speeds less than or equal to 50km / h in the sample feature vector set are divided into two sample feature vector subsets respectively.
[0108] In some embodiments, in step S503, "until the number of sample vehicle feature vectors in the sample feature vector subset reaches a second predetermined number" refers to continuously and repeatedly randomly extracting initial vehicle feature dimensions and performing splitting until the number of sample vehicle feature vectors contained in the child node is less than or equal to the second predetermined number. For example, when the second predetermined number is set to 20, further splitting of the node is stopped when the number of sample vehicle feature vectors in the last node is 15.
[0109] In steps S501 to S503 shown in the embodiment of the present application, target vehicle feature dimensions are optimized layer by layer based on the vehicle feature Gini data, and split nodes are constructed based on the target vehicle feature dimensions to achieve accurate division of the vehicle feature vector set, thereby constructing a risk prediction decision tree with a rigorous structure, clear nodes and high accuracy.
[0110] In step S404 of some embodiments, the node construction and splitting process is repeated continuously, and the nodes are connected in sequence to form a complete risk prediction decision tree structure with multiple nodes.
[0111] Steps S401 to S404 shown in the embodiment of the present application effectively realize the layer-by-layer splitting of the feature vector set by randomly selecting multiple feature dimensions from the vehicle feature vector set and determining the optimal splitting node based on the Gini impurity, thereby constructing an efficient, accurate and highly generalized risk prediction decision tree.
[0112] In step S303 of some embodiments, the process returns to step S301 until the number of risk prediction decision trees generated meets a first preset number. In this embodiment, the first preset number may be 50, but is not limited thereto.
[0113] In steps S301 to S303 shown in the embodiment of the present application, a high-precision and strong generalization target risk prediction model is established through ensemble learning by randomly sampling vehicle feature vectors multiple times and constructing multiple risk prediction decision trees, thereby significantly improving the accuracy of predicting the vehicle's future driving risks.
[0114] In step S106 of some embodiments, see Figure 6 Step S106 includes but is not limited to steps S601 to S604:
[0115] Step S601 : performing driving area analysis based on historical driving routes to obtain a target driving area, and performing climate prediction on the target driving area to obtain reference climate prediction data.
[0116] Step S602 : constructing a feature vector based on the vehicle status data and the reference climate prediction data to obtain a target vehicle feature vector.
[0117] Step S603 , performing risk level assessment on the target vehicle feature vector through each risk prediction decision tree of the target risk prediction model, obtaining multiple candidate risk levels and outputting the number of prediction decision trees for each candidate risk level.
[0118] Step S604: selecting an intermediate risk level from multiple candidate risk levels according to the number of predicted decision trees, and determining a target risk level based on the intermediate risk level.
[0119] In step S601 of some embodiments, the historical driving route is a series of geographic location information corresponding to the actual driving path of the target vehicle within a certain period of time in the past. The target driving area refers to the concentrated area where the target vehicle frequently travels, obtained by performing spatial cluster analysis on the historical driving route. The target vehicle status data refers to the operating status data of the target vehicle itself, such as vehicle type, age, braking condition, maximum engine speed, and distance traveled. The reference climate forecast data is data that is predicted based on the historical climate data of the target driving area and can reflect the future climate conditions of the area.
[0120] The implementation method of this specific embodiment is the same as that of the embodiment of step S101, and will not be described in detail here.
[0121] In some embodiments, in step S602, the target vehicle feature vector refers to a set of numerical features generated by processing vehicle state data and reference climate forecast data, used to predict the vehicle's risk level. The implementation of this specific embodiment is similar to that of step S102 and is not further described here.
[0122] In step S603 of some embodiments, the target vehicle feature vector is input into each risk prediction decision tree, and each node is judged by its feature threshold, ultimately reaching the leaf node of the decision tree. The candidate risk level output by each decision tree is obtained, and the number of predicted decision trees corresponding to the corresponding risk level is recorded. For example, if the target vehicle feature vector is input into a target risk prediction model consisting of 100 risk prediction decision trees, where 60 decision trees output "medium risk", 30 decision trees output "low risk", and 10 decision trees output "high risk", then the candidate risk levels are "medium risk", "low risk", and "high risk", respectively, and the number of predicted decision trees is 60, 30, and 10, respectively.
[0123] In step S604 of some embodiments, the candidate risk level with the largest number of predicted decision trees is selected as the intermediate risk level based on the number of predicted decision trees. In some embodiments, the intermediate risk level can be directly used as the target risk level. In other embodiments, S604 further includes the following steps:
[0124] In response to a natural disaster warning in the target driving area, the natural disaster warning is queried according to a preset risk level coefficient mapping table to obtain a reference risk level. The reference risk level and the intermediate risk level are then superimposed to obtain a target risk level.
[0125] Specifically, if the target driving area currently has a natural disaster warning type (such as a rainstorm warning), the corresponding risk level coefficient is queried through the risk level coefficient mapping table, and finally the risk level coefficient corresponding to the natural disaster warning is superimposed and calculated with the intermediate risk level to determine the final target risk level. For example, if the intermediate risk level is "medium risk" and the current target driving area has issued a rainstorm warning, the reference risk level corresponding to the rainstorm warning is found to be "high risk" through the risk level coefficient mapping table. After superimposing the two, the final target risk level is raised to "high risk".
[0126] In steps S601 to S604, as illustrated in the present embodiment, a vehicle feature vector accurately reflects the actual driving environment by comprehensively analyzing the vehicle's historical driving routes and vehicle status data, combined with future climate forecasts for the target driving area. Furthermore, an intermediate risk level is determined through a combined prediction of multiple decision trees and counting the number of predicted decision trees. Finally, a target risk level is determined based on this intermediate risk level, significantly improving the accuracy of future driving risk predictions for the vehicle.
[0127] See also Figure 7 The embodiment of the present application further provides a vehicle risk prediction device that can implement the above-mentioned vehicle risk prediction method, and the device includes:
[0128] The first acquisition module 701 is used to acquire sample driving routes and sample vehicle status data of sample vehicles, and perform driving area analysis based on the sample driving routes to obtain preferred driving areas.
[0129] The second acquisition module 702 is configured to acquire historical climate data of the preferred driving area, and perform climate prediction on the preferred driving area based on the historical climate data to obtain sample climate prediction data.
[0130] The feature vector construction module 703 is used to construct a feature vector according to the sample vehicle state data and the sample climate prediction data to obtain the sample vehicle feature vector.
[0131] The risk analysis module 704 is used to perform risk analysis on the sample vehicle feature vector to obtain the sample risk level, and use the sample risk level as a label for the sample vehicle feature vector.
[0132] The risk prediction model construction module 705 is used to construct a random forest model based on the sample vehicle feature vector and the sample risk level to obtain a target risk prediction model.
[0133] The risk prediction module 706 is used to obtain the historical driving route and target vehicle status data of the target vehicle, perform risk prediction on the historical driving route and target vehicle status data based on the target risk prediction model, and obtain the target risk level of the target vehicle.
[0134] The specific implementation of the vehicle risk prediction device is basically the same as the specific embodiment of the above-mentioned vehicle risk prediction method, and will not be repeated here.
[0135] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the vehicle risk prediction method when executing the computer program. The electronic device can be any smart terminal, such as a tablet computer or an in-vehicle computer.
[0136] See also Figure 8 , Figure 8 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0137] The processor 801 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0138] The memory 802 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called by the processor 801 to execute the vehicle risk prediction method of the embodiments of this application.
[0139] Input / output interface 803, used to implement information input and output;
[0140] Communication interface 804, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0141] Bus 805 , which transmits information between various components of the device (e.g., processor 801 , memory 802 , input / output interface 803 , and communication interface 804 );
[0142] The processor 801 , the memory 802 , the input / output interface 803 and the communication interface 804 are connected to each other in communication within the device via a bus 805 .
[0143] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned vehicle risk prediction method is implemented.
[0144] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0145] The vehicle risk prediction method, vehicle risk prediction device, electronic device and storage medium provided in the embodiments of the present application collect the driving route and status data of the vehicle, and determine the preferred driving area of the vehicle based on the sample driving route, so as to further obtain the historical climate data of the driving area and perform climate prediction on it, and obtain prediction data that accurately reflects the future driving environment. Then, the sample vehicle status data reflecting the vehicle's operating status is combined with the predicted sample climate prediction data to construct a comprehensive vehicle feature vector. The feature vector is then trained using a random forest algorithm to establish a target risk prediction model, and ultimately achieve an accurate risk level prediction for the target vehicle. Therefore, the embodiments of the present application simultaneously take into account the static factors of the vehicle's own state and the dynamic environmental factors, achieve an accurate prediction of the vehicle's future driving risk, and solve the problem of inaccurate prediction caused by ignoring environmental changes in the traditional static risk factor analysis method, thereby enabling insurance companies to effectively reduce economic losses caused by misjudgment of risks.
[0146] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0147] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0148] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0149] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0150] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0151] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0152] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0153] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0154] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0155] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0156] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A vehicle risk prediction method, characterized in that: The method comprises: Acquiring sample driving routes and sample vehicle status data of sample vehicles, and performing driving area analysis based on the sample driving routes to obtain preferred driving areas; Acquiring historical climate data of the preferred driving area, and performing climate forecasting on the preferred driving area based on the historical climate data to obtain sample climate forecast data; constructing a feature vector based on the sample vehicle state data and the sample climate forecast data to obtain a sample vehicle feature vector; Performing risk analysis on the sample vehicle feature vector to obtain a sample risk level, and using the sample risk level as a label for the sample vehicle feature vector; A random forest model is constructed based on the sample vehicle feature vector and the sample risk level to obtain a target risk prediction model; A historical driving route and target vehicle status data of a target vehicle are obtained, and risk prediction is performed on the historical driving route and the target vehicle status data based on the target risk prediction model to obtain a target risk level of the target vehicle.
2. The method according to claim 1, characterized in that The random forest model is constructed based on the sample vehicle feature vector and the sample risk level to obtain a target risk prediction model, including: Randomly sampling the sample vehicle feature vectors to obtain a sample feature vector set; Constructing a decision tree based on the sample feature vector set to obtain a risk prediction decision tree; Return to step 1 and randomly sample the sample vehicle feature vectors to obtain a sample feature vector set until the number of risk prediction decision trees reaches a first preset number, and determine the target risk prediction model based on the preset number of risk prediction decision trees.
3. The method according to claim 2, characterized in that The step of constructing a decision tree based on the sample feature vector set to obtain a risk prediction decision tree includes: Constructing a root node based on the sample feature vector set, and randomly selecting from the feature dimensions of the sample vehicle feature vectors to obtain an initial vehicle feature dimension; Calculating the Gini impurity of the sample feature vector set according to each of the initial vehicle feature dimensions to obtain vehicle feature Gini data; Performing node construction based on the vehicle feature Gini data, the initial vehicle feature dimension, and the sample feature vector to obtain a split node, where the split node is a subsequent node of the root node; The risk prediction decision tree is determined based on the root node and the split node.
4. The method according to claim 3, characterized in that The node construction based on the vehicle feature Gini data, the initial vehicle feature dimension, and the sample feature vector to obtain a split node includes: Selecting a target vehicle feature dimension from the initial vehicle feature dimension according to the vehicle feature Gini data, and constructing the splitting node based on the target vehicle feature dimension; Dividing the sample feature vector set into a data set according to the splitting node to obtain a sample feature vector subset; Return to step 100 to randomly select from the feature dimensions of the sample vehicle feature vectors to obtain initial vehicle feature dimensions, until the number of sample vehicle feature vectors in the sample feature vector subset reaches a second preset number.
5. The method according to claim 2, characterized in that The performing risk prediction on the historical driving route and the target vehicle status data based on the target risk prediction model to obtain a target risk level of the target vehicle includes: performing a driving area analysis based on the historical driving route to obtain a target driving area, and performing a climate forecast on the target driving area to obtain reference climate forecast data; Constructing a feature vector based on the vehicle state data and the reference climate prediction data to obtain a target vehicle feature vector; Performing a risk level assessment on the target vehicle feature vector using each risk prediction decision tree of the target risk prediction model to obtain a plurality of candidate risk levels and outputting the number of prediction decision trees for each candidate risk level; An intermediate risk level is selected from the plurality of candidate risk levels according to the number of the prediction decision trees, and the target risk level is determined based on the intermediate risk level.
6. The method according to claim 5, characterized in that The determining the target risk level based on the intermediate risk level includes: In response to a natural disaster warning in the target driving area, querying the natural disaster warning according to a preset risk level coefficient mapping table to obtain a reference risk level; The reference risk level and the intermediate risk level are superimposed to obtain the target risk level.
7. The method according to any one of claims 1 to 6, characterized in that The performing climate forecast for the preferred driving area based on the historical climate data to obtain sample climate forecast data includes: Slicing the historical climate data according to a preset first time period to obtain at least two historical climate sub-data; Performing disaster event identification on each of the historical climate sub-data to obtain a disaster occurrence probability of each disaster type in the historical climate sub-data, and adding the disaster occurrence probability to the historical climate sub-data; The climate trend prediction is performed on at least two of the historical climate sub-data based on a long short-term memory network to obtain the sample climate prediction data.
8. A vehicle risk prediction device, characterized in that: The device comprises: a first acquisition module, configured to acquire sample driving routes and sample vehicle status data of sample vehicles, and perform driving area analysis based on the sample driving routes to obtain preferred driving areas; a second acquisition module, configured to acquire historical climate data of the preferred driving area, and perform climate prediction on the preferred driving area based on the historical climate data to obtain sample climate prediction data; a feature vector construction module, configured to construct a feature vector based on the sample vehicle state data and the sample climate prediction data to obtain a sample vehicle feature vector; a risk analysis module, configured to perform risk analysis on the sample vehicle feature vector to obtain a sample risk level, and use the sample risk level as a label for the sample vehicle feature vector; A risk prediction model construction module is used to construct a random forest model based on the sample vehicle feature vector and the sample risk level to obtain a target risk prediction model; The risk prediction module is used to obtain the historical driving route and target vehicle status data of the target vehicle, perform risk prediction on the historical driving route and the target vehicle status data based on the target risk prediction model, and obtain the target risk level of the target vehicle.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.