New energy vehicle safety assessment method based on large language model
Through fine-tuning of the Qwen-7B-Chat model and multi-index fusion method, combined with the entropy value method and the coefficient of variation method, a new energy vehicle safety assessment method was constructed, which solved the problem of difficulty in comprehensively evaluating the overall safety of the vehicle in the existing technology, and achieved efficient and accurate safety assessment.
Patent Information
- Application Number
- CN202510217793.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-04
AI Technical Summary
The existing new energy vehicle safety assessment methods lack an overall evaluation framework that comprehensively combines data from various vehicle systems, making it difficult to accurately evaluate the overall safety of the vehicle.
The Qwen-7B-Chat model is adjusted by using the fine-tuning technology of large language model, and combined with the entropy value method and the coefficient of variation method, a vehicle safety assessment algorithm is constructed, and the front-end page is constructed for user interaction through the Gradio framework, and safety scores and comprehensive scores of various vehicle configurations are generated.
It improves the accuracy and efficiency of safety assessment of new energy vehicles, reduces computing resource consumption and time costs, and can fully reflect the safety performance of the vehicle.
Smart Images

Figure CN120256950A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle configuration management and safety assessment, and particularly relates to a vehicle safety assessment method based on large language model fine-tuning and multi-index fusion. Background Art
[0002] With the development of the vehicle industry and the popularization of new energy vehicles, the complexity and diversity of vehicle configurations are increasing continuously. How to effectively manage and evaluate the safety of new energy vehicles has become an important issue. The current safety assessment methods for new energy vehicles usually focus on the safety performance assessment of specific components of new energy vehicles, lacking a comprehensive assessment framework that combines the data of various vehicle systems. Although this fragmented research method helps to deeply analyze the safety performance of a certain component, it is difficult to comprehensively grasp the overall safety condition of the vehicle.
[0003] In recent years, large language models have made remarkable progress in the fields of natural language processing and data analysis. By fine-tuning these models, efficient processing of specific tasks can be achieved. At the same time, multi-index fusion methods such as the entropy value method and the coefficient of variation method have been widely used in multi-attribute decision analysis and can effectively evaluate the comprehensive influence of multiple factors.
[0004] However, no new energy vehicle safety assessment framework combining large language models has been proposed yet. Therefore, the present invention aims to provide a new energy vehicle safety assessment method based on large language models and multi-index fusion to improve the accuracy and efficiency of assessment. Summary of the Invention
[0005] Aiming at the above technical problems, the present invention aims to provide a new energy vehicle safety assessment method based on large language models, which effectively combines large language models and multi-index fusion methods, thereby realizing the efficient management and safety assessment of new energy vehicle configuration information and improving the accuracy and efficiency of assessment.
[0006] The specific technical solution of the present invention is as follows:
[0007] A new energy vehicle safety assessment method based on large language models, comprising the following steps:
[0008] Step 1: Obtain the original data of the new energy vehicle;
[0009] Step 2: Preprocess and clean the data, set scoring rules, pre-score the vehicle safety, and generate a question-and-answer type instruction fine-tuning data set;
[0010] Step 3: Adjust the Qwen-7B-Chat large language model using the LoRA fine-tuning method and combine the adjusted model weights with the original model weights;
[0011] Step 4: Set up a MySQL database, design database tables, and store relevant information about vehicle configurations;
[0012] Step 5: Adopt the entropy value method and the coefficient of variation method, and fuse the two in proportion to construct a vehicle safety assessment algorithm;
[0013] Step 6: Use the Gradio framework to build a front-end page, obtain the vehicle model information input by the user, query the detailed vehicle configuration in the database, splice it to the user input, call the fine-tuned model to generate safety scores for various vehicle configurations, and finally call the vehicle safety assessment algorithm to calculate the comprehensive score of the vehicle.
[0014] Further, the specific steps of Step 1 are as follows:
[0015] Step 1.1, obtain vehicle brand information from a specified URL list and write the brand ID into an Excel file;
[0016] Step 1.2, obtain detailed information about sub-brands from a given URL list and write the sub-brand ID into an Excel file;
[0017] Step 1.3, obtain vehicle details from a given vehicle ID list and write the vehicle details into an Excel file.
[0018] Further, the specific steps of Step 2 are as follows:
[0019] Step 2.1, read the Excel file and define attributes that contain important vehicle features;
[0020] Step 2.2, clean the values of the attributes;
[0021] Step 2.3, for each row of data, extract the cleaned attribute values to obtain vehicle configuration information, and generate scores for different safety scoring dimensions according to multiple rules to obtain the vehicle's safety configuration information;
[0022] Step 2.4, export the processed data into a json file.
[0023] Further, Step 3 includes:
[0024] Step 3.1, prepare a training dataset and apply the LoRA technique to adjust the Qwen-7B-Chat large language model;
[0025] Step 3.2, merge the adapted parameters obtained after adjustment with the parts of the weights in the original Qwen-7B-Chat large language model that violate the modification to form an optimized Qwen-7B-Chat large language model.
[0026] Further, step 4 is specifically as follows:
[0027] Step 4.1, set up a MySQL database, which is used to store and manage vehicle configuration information;
[0028] Step 4.2, design database tables, and at least the fields of vehicle ID, model, level, energy type, and battery type need to be included in the database tables;
[0029] Step 4.3, import the cleaned vehicle configuration information into the designed database tables.
[0030] Further, step 5 is specifically as follows:
[0031] Step 5.1, read the json file obtained in step 2.4 and extract the scoring data;
[0032] Step 5.2, process the vehicle's safety configuration information using the entropy method to calculate the weight of each safety configuration item;
[0033] Step 5.3, process the vehicle's complete configuration information using the coefficient of variation method to calculate the coefficient of variation of each safety configuration;
[0034] Step 5.4, fuse the weights obtained by the entropy method and the weights obtained by the coefficient of variation method according to a predetermined ratio to generate a comprehensive weight.
[0035] Further, the specific steps for processing the vehicle's safety configuration information using the entropy method include:
[0036] Perform min-max standardization on the safety configuration information and calculate the information entropy of each configuration item:
[0037]
[0038] where x ij represents the standardized value of the i-th sample on the j-th configuration, P ij is the probability after standardization, and E j is the information entropy of the j-th configuration item;
[0039] Calculate the entropy weight of each configuration item according to the information entropy:
[0040] d j = 1 - E j ;
[0041]
[0042] where d j is the entropy weight of the j-th configuration item, w jis the normalized entropy weight of the j-th configuration item, and m is the total number of configuration items.
[0043] Further, the coefficient of variation method is used to process the complete configuration information of the vehicle. The specific steps for calculating the coefficient of variation of each safety configuration include: calculating the standard deviation and mean of each configuration item; the coefficient of variation CV j is calculated by the following formula:
[0044]
[0045] where σ j is the standard deviation of the j-th configuration item, and μ j is the mean of the j-th configuration item.
[0046] Further, the specific steps for fusing the weights obtained by the entropy value method and the weights obtained by the coefficient of variation method in a predetermined ratio to generate a comprehensive weight include:
[0047] Set the proportion of the entropy value method weight to be α, and the proportion of the coefficient of variation method weight to be β, and α + β = 1; the comprehensive weight W j is calculated by the following formula:
[0048]
[0049] Further, step 6 is specifically as follows:
[0050] Step 6.1, import the required dependency libraries, load the pre-trained model and tokenizer, and establish a connection with the database;
[0051] Step 6.2, query the detailed configuration information corresponding to the vehicle model input by the user in the MySQL database and splice it with the user input;
[0052] Step 6.3, call the adjusted Qwen-7B-Chat large language model, input the spliced vehicle configuration data, and generate the safety scores of each vehicle configuration;
[0053] Step 6.4, according to the generated safety scores, call the vehicle safety assessment algorithm to calculate the comprehensive score of the vehicle;
[0054] Step 6.5, use Gradio to create a Web page to listen for user input and generate responses.
[0055] Beneficial effects:
[0056] An embodiment of the present invention provides a new energy vehicle safety assessment method based on large language models and multi-index fusion. By using the LoRA fine-tuning technique to fine-tune the Qwen-7B-Chat large language model, the adaptability and processing efficiency of the model for new energy vehicle configuration management and safety assessment tasks are significantly improved, and the consumption of computing resources and time costs are reduced. The entropy method and the coefficient of variation method are used for multi-index fusion, comprehensively considering the importance of multiple configuration items and the degree of data dispersion, improving the accuracy and reliability of vehicle safety assessment, and being able to comprehensively reflect the safety performance of the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flowchart of the new energy vehicle safety assessment method combined with a large language model according to an example of the present invention;
[0058] Figure 2 It is a flowchart of the safety assessment algorithm according to an example of the present invention;
[0059] Figure 3 It is a flowchart of building a Gradio front-end page according to an example of the present invention;
[0060] Figure 4 It is a system architecture diagram according to an example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0062] As Figure 1 shown, an embodiment of the present invention provides a new energy vehicle safety assessment method based on a large language model. According to the vehicle information input by the user, the detailed configuration of the vehicle is queried from the database, and together with the prompt words, it is input into the large language model Qwen-7B to generate a score for the corresponding configuration, and then the comprehensive scoring algorithm is called to calculate the overall score of the corresponding vehicle model. A web page is built using Gradio to interact with the user, thereby implementing a web page project for new energy vehicle safety assessment based on a large language model.
[0063] In this example, the large language model is Qwen-7B, which is a large language model launched by Alibaba Cloud. It aims to provide strong support for various natural language processing tasks, including but not limited to text generation, dialogue understanding, sentiment analysis, etc. What makes Qwen-7B special is that it has approximately 7 billion parameters, with powerful learning and expression capabilities, enabling it to capture subtle differences and complex patterns in language. Qwen-7B is based on the Transformer architecture and incorporates the latest technological improvements such as efficient attention mechanisms and normalization methods, enhancing the model's performance and stability. Qwen-7B supports fine-tuning and can be further trained on specific tasks, so it can be used for new energy vehicle safety assessment.
[0064] In this example, the web-based user interface uses Gradio, which is a lightweight, open-source Python library designed specifically for quickly building user interfaces for machine learning and data processing models. Gradio allows us to build interactive demonstration applications in the browser without writing any front-end code. Gradio supports a variety of input and output components, enabling us to easily convert complex models or data processing functions into user-friendly web application interfaces. By encapsulating the model as a Python function and connecting it to the Gradio interface, users can input data in the web page and view the model's output without any plugins.
[0065] As Figure 1 shown, the specific steps of a new energy vehicle safety assessment method based on a large language model provided by this application are as follows:
[0066] Step 1: Utilize the official dataset provided by the cooperative enterprise and combine it with publicly available Internet data to collect and supplement data; use web crawler technology to crawl the vehicle configuration data publicly available on the official websites of the automotive industry.
[0067] Step 1.1, Use the get_car_pingpai module to obtain vehicle brand information from the specified URL list and write the brand ID to a file.
[0068] Step 1.2, Use the get_car_pp_sub module to obtain detailed information about sub-brands from the given URL list and write the sub-brand ID to a file.
[0069] Step 1.3, Use the get_car_detail module to obtain vehicle details from the given vehicle ID list and write the vehicle details to an Excel file.
[0070] Step 2: Preprocess and clean the data, set scoring rules, pre-rate the vehicle safety, and generate a Q&A type instruction fine-tuning dataset;
[0071] Step 2.1: Read the excel file of vehicle data and define attributes that contain important vehicle features.
[0072] Step 2.2: Use the clean_value module to clean the attribute values, remove special symbols and unify the null value format.
[0073] Step 2.3: For each row of data, extract the attribute values and generate scores for different safety scoring dimensions according to multiple rules.
[0074] Step 2.4: Export the processed data into a json file.
[0075] Step 3: Adopt the LoRA fine-tuning method to fine-tune the Qwen-7B-Chat large language model, and combine the fine-tuned model weights with the original model weights;
[0076] Step 3.1: Prepare the training data and apply the known LoRA technology to fine-tune the Qwen-7B-Chat large language model.
[0077] Step 3.2: Merge the adaptive parameters obtained after fine-tuning with the weights of the parts that violate the modification in the original model to form the final optimized model.
[0078] Step 4: Set up a MySQL database, design database tables, and store relevant information about vehicle configurations;
[0079] Step 4.1: Set up a MySQL database for storing and managing vehicle configuration information.
[0080] Step 4.2: Design database tables, which should at least include fields such as vehicle id, model, level, energy type, battery type, etc.
[0081] Step 4.3: Import the vehicle configuration information into the designed data table and perform necessary data verification and cleaning.
[0082] Step 5: Adopt the entropy method and the coefficient of variation method, and fuse the two in proportion to construct a vehicle safety assessment algorithm.
[0083] Step 5.1: Read the train.json file obtained in Step 2.4 and extract the vehicle scoring data.
[0084] Step 5.2: Use the entropy method to process the vehicle's safety configuration information and calculate the weight of each safety configuration item.
[0085] Step 5.3: Use the coefficient of variation method to process the vehicle's complete configuration information and calculate the coefficient of variation of each safety configuration.
[0086] Step 5.4: Integrate the weights obtained by the entropy method and the weights obtained by the coefficient of variation method in a predetermined ratio to generate a comprehensive weight.
[0087] Step 6: Use the Gradio framework to build a front-end page, obtain the vehicle model entered by the user, query the vehicle's detailed configuration in the database, splice it to the user input, call the fine-tuned model to generate the safety scores of various vehicle configurations, and finally call the vehicle safety assessment algorithm to calculate the comprehensive score of the vehicle.
[0088] As Figure 1 shown, this embodiment provides a new energy vehicle safety assessment method based on a large language model, which includes:
[0089] Step 1.1: Use the Pycharm development software to create a Python project named llm with a project interpreter of 3.8. Create a file named dataFetch.py under the llm folder. Define headers in the dataFetch.py file to simulate browser behavior, define a global variable car_info to store the list of vehicle brand names, and define a global variable car_id to store the list of vehicle brand URLs. Define a function get_car_pingpai in the dataFetch.py file. Traverse each element in the car_list list, construct the URL to access each vehicle brand page, send an HTTP GET request using the requests library to obtain the web page content, use BeautifulSoup to parse the HTML content and find the elements containing the vehicle brand name and URL, extract the vehicle brand name and URL, and store them in the car_info and car_id lists respectively. Finally, write the URLs of the extracted vehicle brands to the 123.txt file.
[0090] Step 1.2: Define a function get_car_pp_sub. Traverse each element in the url list, construct the URL to access each model under the brand, send an HTTP GET request using the Session object of the request library, use BeautifulSoup to parse the HTML content, and find the elements containing the model information. Finally, extract the relevant model IDs and write them to the 234.txt file.
[0091] Step 1.3, Define the get_car_detail function. Traverse each element in the car_id list, construct the URL to access the details page of each car model, send an HTTP GET request using the requests library to obtain the web page content, parse the HTML content using BeautifulSoup, find the elements containing the car model configuration information, then extract the car model configuration information. After processing special characters, use the pandas library to save the information to the Excel file 234.xlsx.
[0092] Step 2.1, Create a dataprocess.py file in the llm folder. Use the read_excel method in the pandas library to read the Excel file of vehicle data. Define the list of attributes attributes that need to be extracted, and define an empty list output_data to store each processed piece of data.
[0093] Step 2.2, Define the clean_value function to remove null values and special symbols such as "-", "--", "---", "●" from the data.
[0094] Step 2.3, Traverse each row of data and score the vehicle configuration according to different rules. The specific scoring rules are as follows:
[0095] (1) Body structure score: Models with vehicle types of "Large SUV", "Large car", "Mid - large SUV" get 10 points; models with vehicle types of "Medium car", "Medium pickup truck", "Medium MPV", "Medium SUV" get 8 points; models with vehicle types of "Compact car", "Compact SUV", "Compact MPV" get 6 points; others get 4 points.
[0096] (2) Power system safety score: Energy type of "Pure electric" gets 6 points, energy type of "Plug - in hybrid" gets 4 points, others get 2 points. On this basis, obtain the motor horsepower value, and add 1 point for every 80 horsepower, with a maximum addition of 4 points.
[0097] (3) Charging safety score: Battery type of "Lithium iron phosphate battery" gets 9 points, battery type of "Lithium iron phosphate battery + Nickel - cobalt - manganese battery" gets 7 points, battery type of "Lithium - ion battery" gets 6 points, battery type of
[0098] "Nickel - cobalt - manganese battery" gets 5 points, battery type of "Lead - acid battery" gets 3 points, others get 1 point.
[0099] (4) Driving assistance system score: If the model is equipped with "Driving assistance image", add 2 points; if it is equipped with "L3 - level assisted driving"
[0100] Add 3 points, add 2 points for equipped with "Level 2 assisted driving", add 1 point for equipped with "Level 1 assisted driving", equipped with "
[0101] Cruise system" add 2 points, equipped with "Lane keeping assist system" add 2 points, no points added in other cases.
[0102] (5) Braking system scoring: Add 1 point if the vehicle model is equipped with "parking brake", add 2 points for equipped with "ABS anti-lock braking", add 2 points for equipped with "Brakeforce distribution", add 2 points for equipped with "Active braking", add 2 points for equipped with "Regenerative braking energy", no points added in other cases.
[0103] (6) Safety assistance system scoring: Add 2 points if the vehicle model is equipped with "Active safety warning system", equipped with "Blind spot monitoring"
[0104] Add 2 points, add 2 points for equipped with "Driver fatigue warning", add 2 points for equipped with "Low-speed driving warning sound", equipped with "
[0105] Wading induction system" add 2 points, no points added in other cases.
[0106] (7) Night driving safety scoring: Add 5 points if the vehicle model is equipped with "Night vision system", add 2 points for equipped with "Adaptive high and low beam", add 2 points for equipped with "Automatic headlight", add 1 point in other cases.
[0107] (8) Airbag system safety scoring: Add 2 points if the vehicle model is equipped with "Front airbag", equipped with "
[0108] "Side airbag", "Side curtain airbag", "Front knee airbag", "Passenger seat cushion airbag", "Center airbag", "Rear forward airbag", "Rear inflatable seat belt", "Rear seat anti-slip airbag" add 1 point, no points added in other cases.
[0109] (9) Tire and suspension system safety scoring: Add 2 points if the vehicle model is equipped with "Safety tire", equipped with "Aluminum alloy wheels"
[0110] Add 1 point, add 1 point for equipped with "Tire pressure monitoring system", add 3 points if the vehicle drive type is "Three-motor four-wheel drive", add 2 points for drive type of "Dual-motor four-wheel drive", no points added in other cases.
[0111] (10) Electronic stability control system: Add 6 points if the vehicle model is equipped with "Electronic stability program", add 3 points for equipped with "Traction control", add 1 point in other cases.
[0112] Step 2.4, construct each processed data into JSON format and store it in the train.json file.
[0113] Step 3.1, download the Qwen-7B-Chat source code and the finetune.py file through the download channels provided by Tongyi Qianwen official, import them into the project, and install the peft code library and the required dependency libraries in the requirements.txt file through pip. Set the correct model, data, and output paths, and then use the finetune.py script to fine-tune the Qwen-7B-Chat model.
[0114] Step 3.2, use the AutoTokenizer class to load the tokenizer and save it to a new path. Then, load the fine-tuned Qwen-7B-Chat model through the AutoPeftModelForCausalLM class, and call the merge_and_unload() method to merge the LoRA fine-tuning weights into the original model to generate a complete merged model. Finally, store the merged model in the output path according to the set maximum shard size.
[0115] Step 4.1, set up a MySQL database with the database name car.
[0116] Step 4.2, design the database table information_car, which needs to include fields such as vehicle ID, model, level, energy type, battery type, etc.
[0117] Step 4.3, import the vehicle configuration information into the designed data table and perform necessary data verification and cleaning.
[0118] The flowchart for constructing the vehicle safety assessment algorithm is as Figure 2 shown:
[0119] Step 5.1, read the train.json file obtained in Step 2.4, extract the vehicle scoring data, and convert it into a numpy array.
[0120] Step 5.2, use the entropy method to process the vehicle safety configuration information and calculate the weight of each safety configuration item. The specific calculation process is as follows: perform min-max standardization on the safety configuration information; calculate the information entropy of the configuration items in 5.1 through the following formula:
[0121]
[0122] where x ij represents the standardized value of the i-th sample on the j-th configuration, P ij is the probability after standardization, and E j is the information entropy of the j-th configuration item; calculate the entropy weight of each configuration item according to the information entropy:
[0123] d j = 1 - Ej
[0124]
[0125] Step 5.3: Process the complete configuration information of the vehicle using the coefficient of variation method, and calculate the coefficient of variation of each safety configuration. The specific calculation process is as follows: Calculate the standard deviation and mean of each configuration item; the coefficient of variation CV j is calculated by the following formula:
[0126]
[0127] where σ j is the standard deviation of the j-th configuration item, and μ j is the mean of the j-th configuration item.
[0128] Step 5.4: Fuse the weights obtained by the entropy value method and the weights obtained by the coefficient of variation method in a predetermined ratio to generate a comprehensive weight. The specific calculation process is as follows: Set the weight ratio of the entropy value method as α, the weight ratio of the coefficient of variation method as β, and α + β = 1; the comprehensive weight W j is calculated by the following formula
[0129]
[0130] Build a Gradio front-end page, and the flowchart for realizing interaction with users is as Figure 3 shown as follows:
[0131] Step 6.1: Import the required dependency libraries, load the pre-trained model and tokenizer, and establish a connection to the database.
[0132] These dependency libraries are related to the code, and the code needs to use these dependency libraries. Specifically, these dependency libraries are as follows:
[0133] os: This library is used to interact with the operating system, such as handling file paths, environment variables, etc.
[0134] argparse: Used to parse command-line arguments for system configuration and management.
[0135] gradio: This library is used to build a graphical interface for interacting with users, enabling users to easily interact with the system and obtain the prediction results of the model.
[0136] mdtex2html: Used to convert Markdown-formatted text to HTML format for text display processing in the system.
[0137] torch: The PyTorch library is a framework for deep learning that provides functions such as model training and inference. In the system of this patent, it is used to load pre-trained models and perform inference tasks.
[0138] transformers: A library provided by Hugging Face for loading pre-trained natural language processing models and tokenizers.
[0139] mysql.connector: Used to connect and interact with the MySQL database, allowing the system to obtain vehicle configuration information from the database and perform security score calculations.
[0140] In this embodiment, create an app.py file in the llm folder to create the front-end page. Import relevant libraries such as Gradio and Transformer, use the mysql.connector library to establish a connection with the MySQL database, load the pre-trained model and tokenizer, and set the device to GPU.
[0141] Step 6.2, use Gradio's Blocks to create a web page, including an input box, a button, and a chat record display area, define the button functions, including submitting user input, regenerating, and clearing the history. Define the _launch_demo(args, model, tokenizer, config) function to start the Gradio service and listen for user input and generate responses.
[0142] Step 6.3, define the fetch_from_database(user_input) function, which receives the user input, uses fuzzy query to query the detailed configuration information of the vehicle model that matches the user input from the database and assigns it to result, define the field name list field_name to store the vehicle configuration information that needs to be extracted, traverse each row and column of result,
[0143] remove the unnecessary characters and concatenate the field names and corresponding data into a string.
[0144] Step 6.4: Define the predict(_query, _chatbot, _task_history) function, which receives the user input, chat history, and historical tasks. In this function, call the fetch_from_database(user_input) function in Step 6.2 to obtain the processed vehicle configuration data, concatenate it with the user input and the prompt, and form a rich query string. Use the model fine-tuned in Step 3.2 to input the rich query string to generate a safety score response for the vehicle configuration, check whether the response meets the format requirements, and throw an exception if it does not meet the format requirements, or return the safety score response if it meets the format requirements.
[0145] Step 6.5: Call the weights.py module in Step 5.1 to obtain the weights of each configuration item for calculating the comprehensive score. Define the calculate_comprehensive_score(response, weights, expected_dimensions) function, split the input response by line into dimensions and scores, and store them in the scores_dit dictionary. According to the dimensions in the expected_dimensions list, obtain the corresponding scores from the scores_dict, multiply the scores by the corresponding weights in the weights list, and sum them to obtain the comprehensive score.
[0146] The system architecture diagram is as Figure 4 shown, specifically including:
[0147] (1) The PC accesses the Web system;
[0148] (2) Gradio is in the display layer and is responsible for the front-end UI;
[0149] (3) The business logic layer mainly includes the large language model Qwen-7B to generate scores for various vehicle configurations and the comprehensive scoring algorithm to calculate the weights of each configuration item;
[0150] (4) Msql.connector is used as the data access layer to establish a connection with the MySQL database;
[0151] (5) The database uses MySQL to store specific vehicle configuration information.
[0152] The present invention provides a new energy vehicle safety assessment method for large language models. The tools used are not limited to those provided by the present invention, and other related tools can also implement the steps of the present invention. Although the embodiments of the present invention have been described previously, they are only used to illustrate the technical solutions and main features of the present invention, but not to limit the present invention. Once those skilled in the art learn the basic creative concepts, they can make additional changes and modifications to these embodiments, or make equivalent replacements for some of the technical features. Any changes or replacements that are easily conceivable by those skilled in the art within the technical scope disclosed by the invention within the principles of the present invention are covered by the protection scope of the present invention.
Claims
1. A new energy vehicle safety assessment method based on large language models, characterized in that, It includes the following steps: Step 1: Obtain the original data of new energy vehicles; Step 2: Preprocess and clean the data, set a scoring rule, pre-score the vehicle safety, and generate a Q&A type instruction fine-tuning dataset; Step 3: Use the LoRA fine-tuning method to adjust the Qwen-7B-Chat large language model, and combine the adjusted model weights with the original model weights; Step 4: Build a MySQL database, design database tables, and store the relevant information of vehicle configurations; Step 5: Adopt the entropy method and the coefficient of variation method, and fuse the two in proportion to construct a vehicle safety assessment algorithm; Step 6: Use the Gradio framework to build a front-end page, obtain the vehicle model information input by the user, query the vehicle detailed configuration in the database, splice it to the user input, call the fine-tuned model, generate the safety scores of various vehicle configurations, and finally call the vehicle safety assessment algorithm to calculate the comprehensive score of the vehicle.
2. The method according to claim 1, wherein The specific content of Step 1 includes: Step 1.1, Obtain vehicle brand information from the specified URL list, and write the brand ID into an Excel file; Step 1.2, Obtain the detailed information of sub-brands from the given URL list, and write the sub-brand ID into an Excel file; Step 1.3, Obtain vehicle details from the given vehicle ID list, and write the vehicle details into an Excel file.
3. The method according to claim 1, characterized in that The specific content of Step 2 is: Step 2.1, Read the excel file, and define the attributes that contain important vehicle features; Step 2.2, Clean the values of the attributes; Step 2.3, For each row of data, extract the cleaned attribute values to obtain vehicle configuration information, and generate scores for different safety scoring dimensions according to multiple rules to obtain the vehicle's safety configuration information; Step 2.4, Export the processed data into a json file.
4. The method according to claim 1, wherein Step 3 includes: Step 3.1, Prepare a training dataset, and apply the LoRA technology to adjust the Qwen-7B-Chat large language model; Step 3.2, Merge the adapted parameters obtained after adjustment with the weights of the parts that violate the modification in the original Qwen-7B-Chat large language model to form an optimized Qwen-7B-Chat large language model.
5. The method according to claim 1, characterized in that, The specific content of Step 4 is: Step 4.1, Build a MySQL database, which is used to store and manage vehicle configuration information; Step 4.2, Design database tables, and at least the fields of vehicle id, model, level, energy type, and battery type need to be included in the database tables; Step 4.3, Import the cleaned vehicle configuration information into the designed database tables.
6. The method according to claim 1, wherein The specific content of Step 5 is: Step 5.1, Read the json file obtained in Step 2.4 and extract the scoring data; Step 5.2, Use the entropy method to process the vehicle's safety configuration information and calculate the weight of each safety configuration item; Step 5.3, Use the coefficient of variation method to process the vehicle's complete configuration information and calculate the coefficient of variation of each safety configuration; Step 5.4, fuse the weights obtained by the entropy method and the weights obtained by the coefficient of variation method according to a predetermined ratio to generate a comprehensive weight.
7. The method according to claim 6, characterized in that, The specific steps for processing the vehicle's safety configuration information using the entropy method include: Perform min-max normalization on the safety configuration information and calculate the information entropy of each configuration item: where x ij represents the normalized value of the i-th sample on the j-th configuration, and P ij is the probability after normalization, and E j is the information entropy of the j-th configuration item; Calculate the entropy weight of each configuration item based on the information entropy: d j = 1 - E j ; where d j is the entropy weight of the j-th configuration item, and w j is the normalized entropy weight of the j-th configuration item, and m is the total number of configuration items.
8. The method according to claim 6, wherein The coefficient of variation method is used to process the complete configuration information of the vehicle. The specific steps for calculating the coefficient of variation of each safety configuration include: calculating the standard deviation and mean of each configuration item; the coefficient of variation CV j It is calculated by the following formula: where σ j is the standard deviation of the j-th configuration item, and μ j is the mean of the j-th configuration item.
9. The method according to claim 6, wherein The specific steps for fusing the weights obtained by the entropy method and the weights obtained by the coefficient of variation method according to a predetermined ratio to generate a comprehensive weight include: Set the weight ratio of the entropy value method as α, the weight ratio of the coefficient of variation method as β, and α + β = 1; the comprehensive weight W j is calculated by the following formula:
10. The method according to claim 1, wherein The specific content of step 6 is: Step 6.1, import the required dependency libraries, load the pre-trained model and tokenizer, and establish a connection with the database; Step 6.2, query the detailed configuration information corresponding to the vehicle model input by the user in the MySQL database and splice it with the user input; Step 6.3, call the adjusted Qwen-7B-Chat large language model, input the spliced vehicle configuration data, and generate the safety scores for each vehicle configuration; Step 6.4, according to the generated safety scores, call the vehicle safety assessment algorithm to calculate the comprehensive score of the vehicle; Step 6.5, use Gradio to create a Web page to listen for user input and generate responses.