New energy vehicle insurance premium evaluation method based on machine learning

By performing feature engineering and machine learning algorithm optimization on multi-dimensional data of new energy vehicles, an accurate premium prediction model is built, and the problems of single evaluation dimensions and large errors of traditional evaluation methods are solved, and efficient and personalized premium evaluation and convenient services are achieved.

CN120471716APending Publication Date: 2025-08-12NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510380158.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The traditional auto insurance premium evaluation method has a single evaluation dimension and lacks the exploration of unique parameters of new energy vehicles, resulting in large evaluation errors and lack of flexibility and personalization.

Method used

By collecting multi-dimensional data of new energy vehicles, including historical driving data, policyholder information and vehicle-specific parameters, data pre-processing and feature engineering are carried out, and combined with machine learning algorithms such as ridge regression, decision trees and random forests, a premium prediction model is built, and hyperparameters are optimized through grid search to develop a web-side prediction platform.

Benefits of technology

It improves the accuracy and flexibility of new energy vehicle insurance premium assessment, can dynamically adjust, provide convenient premium prediction services, and adapt to the rapid changing needs of new energy vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471716A_ABST
    Figure CN120471716A_ABST
Patent Text Reader

Abstract

The invention discloses a new energy vehicle insurance premium evaluation method based on machine learning, mainly solves the problems of single dimension, insufficient data mining and inaccurate prediction of a current vehicle insurance premium evaluation method, and aims to improve the accuracy and efficiency of premium prediction. According to the method, data statistical analysis, data processing, feature engineering and a machine learning algorithm are combined, comprehensive statistical processing is carried out by capturing multi-dimensional data, including historical driving data, insurer information, insurance purchase information, new energy vehicle parameter information and the like, of a new energy vehicle, and key factors influencing insurance premium are identified through correlation analysis. And feature selection and extraction are carried out, and a machine learning algorithm is used to construct an insurance premium prediction model. The model is trained and subjected to hyper-parameter optimization in a parameter grid search mode, and the performance of the prediction model is improved. The invention also relates to development of a universal webpage end program to realize insurance premium prediction, and rapid and convenient insurance premium prediction service can be provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of new energy vehicle insurance and machine learning, and in particular relates to a new energy vehicle insurance premium assessment method based on machine learning. Background Art

[0002] With growing global environmental awareness and the rapid development of new energy technologies, new energy vehicles are becoming a mainstream form of transportation. Compared to traditional fuel vehicles, new energy vehicles offer distinct characteristics in terms of powertrain, performance, and service life, presenting numerous challenges in auto insurance pricing. Traditional auto insurance pricing methods primarily rely on factors such as vehicle make, model, age, and driving record. While these methods offer some degree of protection, they suffer from limitations such as limited assessment dimensions and an inability to accurately capture the unique risks of new energy vehicles. This leads to significant volatility and uncertainty in insurance premiums for new energy vehicles.

[0003] New energy vehicles differ significantly from traditional internal combustion engine vehicles in many ways. Their propulsion systems, batteries, and charging infrastructure differ from those of traditional vehicles. This can lead to different losses, repair difficulties, and risks associated with accidents. The rapid development of new energy vehicle technology and the constant emergence of new technologies make traditional valuation methods difficult to adapt to rapidly changing market demands. Insurance policies, government subsidies, and market demand for new energy vehicles all influence the risk of using these vehicles. These factors are often difficult to reflect in traditional auto insurance pricing methods.

[0004] Existing auto insurance premium assessment methods are mostly rule-based and statistically model-based, considering only simple vehicle and owner information and historical data. They lack in-depth exploration of the unique technical parameters of new energy vehicles, such as battery capacity and battery life. These methods are overly simplistic, lack flexibility and personalization, and the resulting premiums are subject to significant errors. Therefore, a more accurate, intelligent, and personalized auto insurance premium assessment method is needed. Summary of the Invention

[0005] To address the above technical issues, the present invention aims to provide a machine learning-based method for estimating new energy vehicle insurance premiums. This method addresses the limitations of traditional methods, which suffer from a single evaluation dimension and a lack of modeling for parameters unique to new energy vehicles. This method combines historical driving data, policyholder information, policyholder purchase information, and unique parameters of new energy vehicles into a model. Through feature engineering and appropriate machine learning algorithms, the model is trained and optimized to produce a more accurate and flexible model suitable for predicting new energy vehicle premiums.

[0006] The present invention's new energy vehicle insurance premium assessment method based on machine learning adopts the following technical solutions, specifically as follows: Step 1: Collect multi-dimensional data including structural parameters of different new energy vehicles, information about policyholders, and information about policyholders purchasing insurance, and convert them into a universal data format file; Step 2: Preprocess the collected multi-dimensional data, including data cleaning, data merging and segmentation, data encoding, etc., to convert it into numerical data that can be used for model training; Step 3: Identify the key factors affecting new energy vehicle insurance premiums through data correlation coefficient analysis; Step 4: Select and extract model training features to complete feature engineering construction; Step 5: Select appropriate machine learning algorithms to build multiple auto insurance premium prediction models, including ridge regression, decision tree, gradient boosting tree, and random forest; Step 6: Define the value range of common hyperparameters for each model using a hyperparameter grid. Use grid search to find the optimal hyperparameter combination for the model, and train a more accurate and intelligent new energy vehicle insurance premium prediction model. Step 7: Develop a general auto insurance premium prediction web visualization page, integrate the trained premium prediction model into the web page, and users can complete the online auto insurance premium prediction by entering or selecting the corresponding options.

[0007] Step 1 includes: Step 1.1: Collect the structural parameter information of new energy vehicles crawled from the Internet, collect multi-dimensional data of the company's policyholders and policyholders' insurance purchase information; convert the data into a CSV plain text format file.

[0008] Step 2 includes: Step 2.1: Fill missing values in the collected multi-dimensional data and use the average value to fill missing values in the insured person's age column; Step 2.2: Process the abnormal attribute columns in the original dataset. Convert the vehicle registration date to vehicle age, combine commercial insurance and compulsory traffic insurance into vehicle insurance premiums, and separate each vehicle insurance in the vehicle insurance type column into a separate column. Step 2.3: Encode the categorical attribute columns. Encode gender as a label, represented by 0 and 1. One-hot encode the vehicle usage, vehicle type, and accompanying insurance and services, represented by single 1 and residual 0.

[0009] In step 2, the abnormal attribute column processing refers to processing the attribute columns that do not meet the model training requirements; converting the vehicle initial registration date data into vehicle age numerical data to facilitate model training; combining commercial insurance and compulsory traffic insurance into premiums as the dependent variable of the new energy vehicle auto insurance premium prediction model; and dividing the multiple types of insurance in the auto insurance type attribute column into columns for consideration, which can help the model better learn the factors that affect auto insurance premiums.

[0010] In step 2, encoding the categorical attribute columns involves converting data unsuitable for model training into numerical data suitable for model training. For example, the gender columns "male" and "female" are represented by 1 and 0, respectively. For the vehicle usage and vehicle type columns, the selected item is represented by 1, and the others by 0, respectively. For the vehicle insurance type attribute column, the number of insurance types purchased is represented by 1 and 0, respectively.

[0011] Step 3 includes: Step 3.1: Use the Pearson, Spearman, and Kendall rank correlation coefficient analysis methods to perform correlation coefficient analysis on the dependent variable (new energy vehicle insurance premium) and other attributes to obtain the correlation coefficient between each independent variable and the dependent variable; Step 3.2: Based on the evaluation relevance criteria established in advance, delete the attribute columns with low relevance and retain the attribute columns with high relevance.

[0012] In step 3, the correlation coefficient analysis refers to a statistical tool used to measure the correlation between two variables.

[0013] Step 4 includes: Step 4.1: Perform feature engineering based on the highly correlated attributes obtained from the correlation coefficient analysis. Select new vehicle purchase price, actual vehicle value, battery capacity, battery life, driving mode, policyholder age, policyholder gender, vehicle age, number of consecutive years of vehicle insurance coverage, number of vehicle accidents during the consecutive insurance period, and vehicle insurance purchased as features for model training. Step 4.2: Perform z-score normalization on the feature data selected by feature engineering, and save the normalization parameters (mean, standard deviation) for subsequent inverse normalization of auto insurance premium prediction.

[0014] In step 4, feature engineering involves extracting, selecting, and transforming features from the raw data to extract meaningful features that help the machine learning model learn and predict better. Normalization involves scaling the data features to a relatively uniform standard range, ensuring consistent scales across different features. This facilitates subsequent analysis and modeling, improving model effectiveness and training efficiency.

[0015] Step 5 includes: Step 5.1: Select a commonly used machine learning algorithm for predicting unknown data, divide the standardized data into a training set and a test set, build initial auto insurance premium prediction models for each, and observe their performance on the test set; Step 5.2: By observing their performance on the test set, preliminarily screen out models that are more suitable for new energy vehicle insurance premium prediction.

[0016] Step 6 includes: Step 6.1: Identify the machine learning models used to predict auto insurance premiums and determine the hyperparameters that affect the performance and generalization ability of each model. Step 6.2: Define a hyperparameter grid for each model, set the range of variation of the model hyperparameters, and perform grid search through GridSearchCV to determine the optimal parameter combination of the model; Step 6.3: Use the optimal model parameter combination for training to obtain the optimal new energy vehicle insurance premium prediction model and save it as a .joblib file for loading and use in general web application development.

[0017] In step 6, the grid search means that the algorithm will traverse the given hyperparameter grid, select a value for each hyperparameter of the model to combine and train, and evaluate its performance in cross-validation to find the best hyperparameter combination of the model, thereby determining the optimal model.

[0018] Step 7 includes: Step 7.1: Develop a front-end web visualization page using Vue, use Axios to implement cross-domain communication between the front-end and back-end to complete page jumps, and use JS to implement user interaction with the page; Step 7.2: Integrate the optimal auto insurance premium prediction model into the web backend through the Flask framework, monitor and process frontend requests, and establish a connection between the frontend and backend. Beneficial effects

[0019] The present invention provides a method for evaluating new energy vehicle insurance premiums based on machine learning. It establishes a multi-dimensional data model for new energy vehicles, and then trains and optimizes the model through machine learning algorithms to obtain a new energy vehicle insurance premium prediction model. This method can not only improve the accuracy of vehicle insurance pricing, but also realize dynamic adjustment and optimization, and has broad application prospects.

[0020] This method integrates historical driving data, vehicle policyholder information, vehicle insurance information, and unique technical parameters of new energy vehicles to conduct comprehensive modeling and deeply analyze the correlation between various features and premiums. Through feature engineering and advanced machine learning algorithms, a more accurate and intelligent model is constructed that can better meet the needs of auto insurance premium forecasting for new energy vehicles, effectively addressing the problems of traditional evaluation methods, which are simple, single-dimensional, and have large prediction errors. In addition, the present invention has also developed a supporting web-based new energy vehicle insurance premium forecasting platform to provide more convenient and efficient services. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 This is an overall flow chart of a method for evaluating new energy vehicle insurance premiums based on machine learning according to an embodiment of the present invention; Figure 2 A schematic diagram of the machine learning model training and optimization process in an embodiment of the present invention; Figure 3 A schematic diagram of the process of implementing the prediction of new energy vehicle insurance premiums by introducing the Flask project into the model of an embodiment of the present invention; Figure 4 2 is a system architecture diagram in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The present invention will be described in further detail below with reference to the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention. Based on the present invention, other embodiments obtained by ordinary technicians in this field without making any creative work are all within the scope of protection of the present invention.

[0024] like Figure 1 As shown, this embodiment provides a method for evaluating new energy vehicle insurance premiums based on machine learning, which includes: Step 1: Collect multi-dimensional data including the parameter structure of different new energy vehicles, information about the insured and the insurance purchased by the insured, and convert it into a universal data format file; Step 1.1: Crawl information on the parameter structure and unique technical parameters of new energy vehicles from the Internet, and collect information on policyholders and insurance purchases from companies; This includes the new car's purchase price, reference actual value, power, battery capacity, battery life, drive mode, range, charging system, approved passenger capacity, approved load capacity, curb weight, number of consecutive years of insurance coverage, number of claims during the consecutive coverage period, handling fee ratio, self-pricing discount factor, no-claim discount factor, no-claim discount factor level, owner's age, owner's gender, purchase date, vehicle usage, vehicle type, and insurance type. Convert the data to a CSV plain text file.

[0025] Step 2: Integrate and process the collected multi-dimensional data, perform data cleaning, data merging and segmentation, data encoding, etc., and convert it into a data format suitable for model training; Step 2.1: Combine the new energy vehicle configuration information crawled from the Internet and the insured and insurance purchase information collected by the company into the same Excel file, and convert it into a CSV plain text file; Step 2.2: Open PyCharm, read the CSV plain text file, and fill in the missing values of the collected multi-dimensional data. Fill the missing values of the insured person's age column with the average age. Step 2.3: Process the abnormal attribute columns in the original dataset. Convert the vehicle registration date to vehicle age, combine commercial insurance and compulsory traffic insurance into auto insurance premiums, and divide auto insurance types into multiple auto insurance attribute columns, including additional medical expense liability insurance (third party liability insurance) and compulsory traffic liability insurance. Step 2.4: Encode the categorical attribute columns. Encode gender as a label, represented by 0 and 1. One-hot encode the vehicle usage, vehicle type, and vehicle insurance attribute columns, represented by single 1 and residual 0. Step 3: Determine the key factors affecting auto insurance premiums through correlation coefficient analysis; Step 3.1: Open PyCharm and use the Pearson, Spearman, and Kendall correlation coefficient analysis functions in the Scipy library to perform correlation coefficient analysis on the dependent variable, auto insurance premium, and other attributes. Obtain the correlation coefficient value between each independent variable and auto insurance premium. Step 3.2: Analyze the three sets of correlation coefficient analysis results and select attributes with significant correlation based on the pre-established evaluation criteria. These attributes include new vehicle purchase price, actual reference value, power, battery capacity, drive mode, battery life, approved passenger capacity, range, charging system, number of consecutive years of insurance coverage, number of accidents during the consecutive insurance period, owner age, owner gender, purchase date, compulsory vehicle insurance, vehicle usage, vehicle type, and insurance type.

[0026] Step 4: Select and extract model training features to complete feature engineering construction; Step 4.1: Based on the results of the correlation coefficient analysis, first determine the new car purchase price, actual vehicle value, power, continuous insurance years, number of claims during the continuous insurance period, vehicle usage nature, vehicle type, etc. as part of the feature engineering; Step 4.2: Select the insured's important information, such as age, gender, insurance type, and vehicle age, and add them to the feature engineering selection as features for model training; Step 4.3: Specific technical parameters of new energy vehicles, such as battery capacity, battery life, range, driving mode, and charging system, are also converted into features for model training; Step 4.4: Perform z-score normalization on the feature data selected by feature engineering and save the normalization parameters (mean, standard deviation) to a file for use in inverse normalization for auto insurance premium prediction.

[0027] Step 5: Select an appropriate machine learning algorithm to build a car insurance premium prediction model, including ridge regression, decision tree, gradient boosting tree, and random forest; Step 5.1: Create a project using PyCharm, using Python 3.7 as the interpreter. Download the pandas and scikit-learn libraries. Divide the processed data into a training set and a test set with a ratio of 4:1. Train the linear regression, ridge regression, decision tree, gradient boosted tree, random forest, and LSTM models. First, obtain an initial auto insurance premium prediction model and observe its performance on the test set. Step 5.2: By observing their performance on the test set, we preliminarily screen out models that are more suitable for predicting auto insurance premiums. These models include ridge regression, decision tree, gradient boosting tree, and random forest.

[0028] Step 6: Define the common hyperparameter value ranges of the machine learning model through the hyperparameter grid, use GridSearchCV to perform grid search to find the optimal parameter combination of the model, train the model based on the optimal model parameter combination, and perform cross-validation through KFold to evaluate the model's generalization ability and prevent overfitting, ultimately obtaining the optimal new energy vehicle insurance premium prediction model; Step 6.1: Determine the machine learning model for predicting new energy vehicle insurance premiums, including ridge regression, decision tree, gradient boosting tree, and random forest. For the assessment of new energy vehicle insurance premiums, open Pycharm and set the hyperparameters of ridge regression, including alpha (regularization strength parameter), to control the model's fit to the training data; set the hyperparameters of the decision tree, including max_depth (maximum depth), min_samples_split (minimum number of samples required for internal node re-division), and min_samples_leaf (minimum number of samples required for leaf nodes); set the hyperparameters of the gradient boosting tree, including n_estimators (number of weak classifiers (trees)), max_depth (maximum depth of each tree), min_samples_split (minimum number of samples for each tree node re-division), min_samples_leaf (minimum number of samples for leaf nodes), and learning_rate (learning rate); set the hyperparameters of the random forest, including n_estimators (number of trees in the forest), max_depth (maximum depth of each tree), min_samples_split (minimum number of samples required for internal node division of the decision tree), min_samples_leaf (minimum number of samples required for leaf nodes), max_features (maximum number of features considered when splitting each tree), and bootstrap (enable self-sampling).

[0029] Step 6.2: Set the value range of the hyperparameter alpha of the ridge regression model to 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 10, 100; set the value range of the hyperparameter max_depth of the decision tree model to None, 5, 10, and 15, the value of min_samples_split to 2, 5, and 10, and the value of min_samples_leaf to 1, 2, and 4; set the hyperparameter n_estimators of the gradient boosting tree model to 50, 100, 200, 300, l The values of earning_rate include 0.01, 0.1, and 0.2, the values of max_depth include 3, 4, and 5, the values of min_samples_split include 2, 3, 4, 5, and 10, and the values of min_samples_leaf include 1, 2, 3, 4, and 5. The values of the random forest hyperparameter n_estimators include 50, 100, and 200, the values of max_depth include None, 10, 20, and 30, the values of min_samples_split include 2, 5, and 10, and the values of min_samples_leaf include 1, 2, 3, 4, and 5.

[0030] Step 6.3: Add KFold cross validation to the model. Here, set 5-fold cross validation to evaluate the generalization ability of the model and prevent overfitting; set the indicators used to evaluate the performance of the model in regression problems: root mean square error (MSE), root mean square error (RMSE), and coefficient of determination (R). 2 , according to the coefficient of determination R 2 To select the optimal model.

[0031] In one embodiment, the model training and optimization process of step 6 is as follows Figure 2 The specific steps are as follows: Save the acquired data as a CSV text file; Process data in CSV text files, including data cleaning and completion, data merging and segmentation, and data encoding; Through correlation coefficient analysis, we can obtain attributes that are highly correlated with premiums; Perform feature selection, extraction, and conversion to complete feature engineering and obtain feature data for model training and testing; Divide the data into training sets and test sets, select an appropriate machine learning algorithm, set the range of hyperparameter values, train on the training set, and make predictions on the test set to evaluate model performance. At the same time, continuously adjust the model hyperparameters to find the optimal parameter combination for the model.

[0032] In regression problems, mean square error (MSE) and root mean square error (RMSE) are usually used to measure the accuracy of the model, and the coefficient of determination (R) is used to measure the accuracy of the model. 2 To evaluate the degree of fit of the model. The indicators for evaluating model performance are MSE, RMSE and R 2 The calculation formula is as follows:

[0033]

[0034]

[0035] The mean square error (MSE) is used to measure the difference between the predicted value and the actual value. It calculates the average of the squares of the prediction errors and is more sensitive to larger errors. It represents the true value, i.e. the real premium value; represents the predicted value, that is, the new energy vehicle insurance premium value predicted by the model; n represents the total number of sample data.

[0036] The root mean square error (RMSE) has the same unit as the original data and is more interpretable. represents the true value; represents the predicted value; n represents the total number of sample data.

[0037] Coefficient of determination Indicates the degree of fit between the model prediction value and the actual value. The value is between 0 and 1. The larger the value, the better the fit of the model to the data. represents the true value; represents the predicted value; represents the mean of the true value y, that is, the mean of the actual car insurance premium; n represents the total number of samples.

[0038] Step 7: Develop a general web visualization page for new energy vehicle insurance premium prediction and integrate the trained premium prediction model into the web page.

[0039] Step 7.1: Develop a front-end web page using Vue. Use Axios to implement cross-domain front-end and back-end communication to navigate between pages. Use JavaScript to enable user interaction with the page. The front-end web page developed in this example uses the Element UI framework. Use its component library to design front-end components, including navigation bars, forms, buttons, icons, labels, and images.

[0040] Step 7.2: Integrate the optimal auto insurance premium prediction model into the web backend using the Flask framework to monitor and process frontend requests, establishing a connection between the frontend and backend. Using PyCharm, download the flask and flask_cors libraries, create the app.py file, and create view functions within the file that map to the frontend request URLs. Within the view functions, process the data sent from the frontend, call the model to predict the premium value, perform denormalization, and return the final prediction result to the frontend page for display.

[0041] In one embodiment, the process of introducing the model into the Flask project in step 7 to realize the prediction of new energy vehicle insurance premium is as follows: Figure 3 As shown, the steps are as follows: First save the trained model file and the standardized parameter file during the data standardization process; Load the model file and standardized parameter file into the Flask project; In the view function, the model file and the standardized parameter file are used to standardize the received data, and the model is called for prediction. The predicted premium is de-standardized, and the result is returned to the front-end page.

[0042] The present invention provides a method for evaluating auto insurance premiums for new energy vehicles based on machine learning. The tools used are not limited to those provided by the present invention, and other related tools can also implement the steps of the present invention. Although the embodiments of the present invention have been described before, they are only used to illustrate the technical solutions and main features of the present invention, but are not used to limit the present invention. Once those skilled in the art know the basic creative concepts, they can make additional changes and modifications to these embodiments, or make equivalent replacements for some of the technical features therein. Any changes or replacements that can be easily thought of by any person skilled in the art within the scope of the technology disclosed by the present invention are covered by the protection scope of the present invention.

Claims

1. A method for evaluating new energy vehicle insurance premiums based on machine learning, characterized in that: The steps include: Step 1: Collect multi-dimensional data including structural parameters of different new energy vehicles, information about policyholders, and information about policyholders purchasing insurance, and convert it into a universal data format file; Step 2: Preprocess the collected multi-dimensional data, including data cleaning, data merging and segmentation, data encoding, etc., to convert the data into numerical data suitable for model training; Step 3: Identify the key factors affecting new energy vehicle insurance premiums through correlation coefficient analysis; Step 4: Based on the key factors, select and extract model training features to complete the construction of feature engineering; Step 5: Select appropriate machine learning algorithms to build multiple auto insurance premium prediction models, including ridge regression, decision tree, gradient boosting tree, and random forest; Step 6: Define the value range of common hyperparameters for each auto insurance premium prediction model through the hyperparameter grid, use the GridSearchCV method to perform grid search to find the optimal hyperparameter combination of the model, and further optimize and train the most suitable auto insurance premium prediction model; Step 7: Develop a general auto insurance premium prediction web visualization page, integrate the trained premium prediction model into the web page, and users can complete the online auto insurance premium prediction by entering or selecting the corresponding options.

2. The method according to claim 1, characterized in that The step 1 is specifically as follows: Step 1.1: Collect the structural parameter information of new energy vehicles crawled from the Internet, collect multi-dimensional data of the company's policyholders and the policyholders' insurance purchase information, and merge the data into a CSV plain text format file.

3. The method according to claim 1, characterized in that The step 2 is specifically as follows: Step 2.1: Fill missing values in the collected multi-dimensional data and use the average value to fill missing values in the insured person's age column; Step 2.2: Process the abnormal attribute columns in the original dataset. Convert the vehicle registration date to vehicle age, combine commercial insurance and compulsory traffic insurance into vehicle insurance premiums, and divide the vehicle insurance type column into multiple vehicle insurance attribute columns. Step 2.3: Encode the categorical attribute columns. Encode gender as a label, represented by 0 and 1. One-hot encode the vehicle usage, vehicle type, and accompanying insurance and services, represented by single 1 and residual 0.

4. The method according to claim 1, wherein The step 3 is specifically as follows: Step 3.1: Select Pearson, Spearman, and Kendall rank correlation coefficient analysis methods to perform correlation coefficient analysis on the dependent variable and other attributes respectively, and obtain the correlation coefficient between each independent variable and the dependent variable; Step 3.2: Analyze the three sets of correlation coefficient analysis results. According to the established criteria, select attributes with high correlation and consider them as input features for model training.

5. The method according to claim 1, wherein The step 4 is specifically as follows: Step 4.1: Perform feature engineering based on the highly correlated attributes obtained from the correlation coefficient analysis. Select new vehicle purchase price, actual vehicle value, battery capacity, battery life, driving mode, policyholder age, policyholder gender, vehicle age, number of years of continuous vehicle insurance coverage, number of vehicle accidents during the continuous insurance coverage period, vehicle type, vehicle usage, and insurance purchase as features for model training. Step 4.2: Perform z-score standardization on the feature data selected after feature engineering, convert data of different scales and units such as vehicle price and owner age into data with the same standard scale, and save the data with the same standard scale.

6. The method according to claim 1, characterized in that The step 5 is specifically as follows: Step 5.1: Select a machine learning algorithm for predicting unknown data. Divide the data with the same standard scale into a training set and a test set. Build initial auto insurance premium prediction models for each set and observe their performance on the test set. Step 5.2: By observing their performance on the test set, preliminarily screen out models that are more suitable for auto insurance premium prediction.

7. The method according to claim 1, characterized in that The step 6 is specifically as follows: Step 6.1: Select a machine learning model suitable for new energy vehicle insurance premium prediction and determine the key hyperparameters that affect the performance and generalization ability of each model; Step 6.2: Define a hyperparameter grid search space for each machine learning model, set the range of variation of each hyperparameter, and perform grid search using the GridSearchCV method to determine the optimal parameter combination of the machine learning model during training; Step 6.3: Train the model based on the optimal hyperparameter combination to obtain the optimal auto insurance premium prediction model and save it for subsequent use.

8. The method according to claim 1, characterized in that The step 7 is specifically as follows: Step 7.1: Use Vue to develop the front-end web visualization page, use Axios to implement front-end and back-end communication to complete page jumps, and use JS to realize user interaction with the page; Step 7.2: Integrate the optimal auto insurance premium prediction model into the web backend through the Flask framework, monitor and process requests from the frontend, and implement data transmission and interaction between the frontend and backend.