Methods, devices, electronic equipment and storage media for generating vehicle parts-to-vehicle ratio coefficients
By constructing a vehicle parts atlas and detecting outliers, the problem of insufficient data coverage and dynamic factor consideration in the calculation of the vehicle parts-to-vehicle ratio in existing technologies has been solved, achieving efficient and accurate calculation of the vehicle parts-to-vehicle ratio and supporting insurance pricing and vehicle cost assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEOPLE'S INSURANCE COMPANY OF CHINA
- Filing Date
- 2026-01-22
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, the calculation of the vehicle parts-to-vehicle ratio relies on manual statistics and traditional weighted algorithms, which have problems such as limited data coverage, low processing efficiency, and insufficient consideration of dynamic factors, making it difficult to meet the industry's needs for real-time performance, accuracy, and large-scale coverage.
By constructing a vehicle parts map using historical vehicle repair claims data and a pre-trained parts big data model, outlier detection is performed to obtain target parts price data. Based on the vehicle parts map, a complete parts list and parts-to-vehicle ratio coefficient for each vehicle model are determined.
It significantly improves the accuracy and efficiency of calculating the vehicle parts-to-vehicle ratio, and enhances the calculation coverage, accuracy, and timeliness, making it suitable for insurance pricing and vehicle cost assessment.
Smart Images

Figure CN122134407A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device and storage medium for generating the vehicle parts-to-vehicle ratio coefficient. Background Technology
[0002] In the insurance industry, the parts-to-vehicle ratio is a crucial indicator for measuring the economics of vehicle repairs, widely used in insurance pricing, claims management, and industry regulation. Users need to use the parts-to-vehicle ratio to quantitatively evaluate the repair costs of different vehicle models, thereby determining differentiated insurance rates. For example, in motor vehicle insurance pricing, a higher parts-to-vehicle ratio typically indicates a higher risk of claims, and insurance companies will adjust premiums to match this risk level. Furthermore, the parts-to-vehicle ratio is also an important reference indicator for consumers when purchasing a vehicle, helping the public assess the long-term operating costs.
[0003] In related technologies, the calculation of the vehicle parts-to-vehicle ratio relies on manual statistics and traditional weighted algorithms, which suffers from limited data coverage, low processing efficiency, and insufficient consideration of dynamic factors. For example, insurance companies need to rely on parts price data provided by OEMs or authorized dealers, but data collection cycles are long, manual data processing costs are high, and it is difficult to cover the price differences of a large number of vehicle models and regions in the market. With the rapid expansion of vehicle models and parts types, traditional methods can no longer meet the industry's needs for real-time, accurate, and large-scale coverage.
[0004] Therefore, there is an urgent need for a method to generate the vehicle parts-to-vehicle ratio coefficient in order to solve the above-mentioned technical problems. Summary of the Invention
[0005] This application provides a method, apparatus, electronic device, and storage medium for generating the vehicle parts-to-vehicle ratio coefficient, thereby improving the accuracy and efficiency of the vehicle parts-to-vehicle ratio coefficient calculation.
[0006] In a first aspect, embodiments of this application provide a method for generating the vehicle parts-to-vehicle ratio coefficient, including:
[0007] Based on historical vehicle repair claims data and pre-trained parts big language model, construct a vehicle parts atlas;
[0008] Anomaly detection processing is performed based on the historical vehicle repair claims data and external spare parts price data to obtain the target spare parts price data;
[0009] Based on the vehicle parts map, determine the complete parts list for each vehicle model;
[0010] Based on the complete list of spare parts for each vehicle model, the price data of the target spare parts, and the price of the whole vehicle for each model, determine the vehicle parts-to-whole ratio coefficient for each model.
[0011] In an optional embodiment of this application, the construction of a vehicle parts atlas based on historical vehicle repair claims data and a pre-trained parts large language model includes: extracting parts name data, price data, and vehicle model data from historical vehicle repair claims data; inputting the parts name data into the pre-trained parts large language model to obtain standard parts name data; constructing a multi-level parts name classification system based on the standard parts name data; and constructing a structured mapping relationship between vehicle models and parts as a vehicle parts atlas based on the multi-level parts name classification system, price data, vehicle model data, and pre-stored standard parts data.
[0012] In an optional embodiment of this application, the pre-trained parts large language model is obtained through model training, wherein the model training includes: acquiring parts repair data processing and conversion data, and generating a training dataset that conforms to the format required by the initial large language model based on the parts repair data processing and conversion data; and training the initial large language model using a multi-turn dialogue format based on the training dataset to obtain the pre-trained parts large language model.
[0013] In an optional embodiment of this application, the step of performing outlier detection processing based on the historical vehicle repair claim data and external spare parts price data to obtain target spare parts price data includes: performing price extraction processing based on the historical vehicle repair claim data to obtain claim parts price data; performing preliminary identification processing based on the claim parts price data and external spare parts price data to obtain multiple sets of intermediate spare parts price data after removing invalid price data; performing outlier detection and removal processing on the multiple sets of intermediate spare parts price data to obtain multiple sets of target spare parts price data; and performing aggregation processing based on the multiple sets of target spare parts price data to obtain target spare parts price data.
[0014] In an optional embodiment of this application, the preliminary identification processing based on the claimed parts price data and external parts price data to obtain multiple sets of intermediate parts price data after removing invalid price data includes: determining the mean, standard deviation, median, interpolation, and dataset size of the parts price data based on the claimed parts price data and external parts price data; determining the deviation value between the parts price data points and the mean based on the claimed parts price data, external parts price data, and the mean and standard deviation of the parts price data; and determining the deviation value between the parts price data points and the mean based on the dataset size, a pre-selected significance level, and a preset verification threshold. The system uses a value table to determine a target threshold value; based on the target threshold value and the deviation degree value, it identifies outliers in the claimed parts price data and external parts price data; it replaces the outliers in the claimed parts price data and external parts price data with the mean, median, or interpolation value, and then jumps to the step of determining the mean, standard deviation, median, interpolation, and dataset size of the parts price data based on the claimed parts price data and external parts price data, until there are no outliers in the claimed parts price data and external parts price data, to obtain multiple sets of intermediate parts price data after removing invalid price data.
[0015] In an optional embodiment of this application, the step of obtaining multiple sets of target spare parts price data by performing outlier detection and removal processing on multiple sets of intermediate spare parts price data includes: constructing an array from the multiple sets of intermediate spare parts price data to obtain a spare parts price feature array, wherein the spare parts price feature array includes multiple elements, each element corresponding to a set of intermediate spare parts price data; inputting the spare parts price feature array into a pre-trained isolated forest model to obtain the outlier scores of each element in the spare parts price feature array; determining the outlier elements in the spare parts price feature array based on the outlier scores of each element in the spare parts price feature array and a preset outlier ratio estimate; and removing the outlier elements in the spare parts price feature array to obtain multiple sets of target spare parts price data.
[0016] In an optional embodiment of this application, the vehicle parts-to-vehicle ratio coefficient for each vehicle model is determined based on the complete parts list, target parts price data, and vehicle price for each vehicle model. This includes: determining the total price of vehicle parts for each vehicle model based on the complete parts list and target parts price for each vehicle model; and determining the vehicle parts-to-vehicle ratio coefficient for each vehicle model based on the total price of vehicle parts and vehicle price for each vehicle model.
[0017] Secondly, embodiments of this application provide an apparatus for generating a vehicle parts-to-vehicle ratio, comprising:
[0018] The graph construction module is used to build a vehicle parts graph based on historical vehicle repair claims data and a pre-trained parts big language model.
[0019] The abnormal data processing module is used to perform outlier detection processing based on the historical car repair claims data and external spare parts price data to obtain the target spare parts price data;
[0020] The list generation module is used to determine the complete list of spare parts for each vehicle model based on the vehicle model spare parts map;
[0021] The parts-to-vehicle ratio determination module is used to determine the parts-to-vehicle ratio for each vehicle model based on the complete list of spare parts for each model, the price data of the target spare parts, and the price of the whole vehicle for each model.
[0022] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0023] The memory stores computer-executed instructions;
[0024] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0025] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0026] This application provides a method, apparatus, electronic device, and storage medium for generating the vehicle parts-to-vehicle ratio coefficient. It constructs a vehicle parts map based on historical vehicle repair and claims data and a pre-trained parts large language model, then processes abnormal data to obtain target parts price data. Next, using the vehicle parts map, it determines a complete standard parts table for each vehicle model, and based on the target parts price data and the vehicle price for each model, it derives the final vehicle parts-to-vehicle ratio coefficient for each model. This significantly improves the coverage, accuracy, and timeliness of the calculated vehicle models, achieving the technical effect of improving the accuracy and efficiency of vehicle parts-to-vehicle ratio coefficient calculation. Attached Figure Description
[0027] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0028] Figure 1 A schematic diagram illustrating a method for generating the vehicle parts-to-vehicle ratio provided in this application;
[0029] Figure 2 This application provides a schematic diagram of a process for generating a vehicle-to-whole ratio coefficient.
[0030] Figure 3 A schematic diagram of a device for generating the vehicle parts-to-vehicle ratio provided in this application;
[0031] Figure 4 This is a schematic diagram of the structure of an electronic device provided in this application.
[0032] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0033] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0034] In related technologies, data collection of parts and vehicle prices published by OEMs or authorized users is done manually. Data entry, deduplication, and data cleaning all require manual operation, resulting in a time-consuming and error-prone process that is difficult to scale up and update in real time. Faced with the rapid expansion of vehicle models and parts types, traditional manual processes can no longer meet the demands for efficient and accurate calculations.
[0035] Furthermore, in related technologies, the calculation of the vehicle parts-to-whole ratio lacks the identification and processing of dynamic factors such as parts loss rate, regional price differences, and parts price changes. Existing methods are mostly based on static prices, ignoring price fluctuations of parts in different regions and channels, as well as dynamic factors such as parts loss rate and parts price changes. This makes it difficult to reflect these changes in a timely manner, resulting in calculation results that lag far behind the actual market situation. This affects the accuracy and timeliness of the calculation results, and consequently, the accuracy and timeliness of pricing vehicle insurance for car repair and claims users based on the existing vehicle parts-to-whole ratio and common parts burden index.
[0036] Based on the aforementioned technical issues, this application provides a method for generating the vehicle parts-to-vehicle ratio coefficient. This method constructs a vehicle parts atlas using historical vehicle repair claims data and a pre-trained parts large language model. It leverages a large amount of existing historical parts claims price data and parts price data from external data providers. Combining multiple processes such as the mapping relationship between parts and vehicle models, parts standard name normalization and classification, anomaly data processing, and parts benchmark price verification, this method significantly improves the coverage, accuracy, timeliness, and industry applicability of the calculated vehicle models. The new method enhances the accuracy and efficiency of calculating the vehicle parts-to-vehicle ratio coefficient.
[0037] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0038] Figure 1 This application provides a schematic diagram of a scenario for generating the vehicle parts-to-vehicle ratio coefficient. Figure 1 As shown, the specific application scenarios of this application include:
[0039] User 101, client 102, and server 103. User 101 can be a staff member responsible for car repair claims or managing car repair claims data, and reviewing relevant documents. Client 102 can be a hardware device that responds to user 101's operations to complete corresponding tasks; for example, client 102 can be a personal computer, mobile phone, all-in-one machine, or other terminal device with car repair claims data display or manual gesture response capabilities. Server 103 can receive historical car repair claims data from other terminals and spare parts price data from external data providers, and execute the relevant steps of the car parts-to-vehicle ratio generation method to generate the car parts-to-vehicle ratio coefficient. This coefficient, used to complete document content review, is ideally displayed on client 101. This provides a solid foundation for car repair claims staff to calculate and compare the parts-to-vehicle ratio coefficient for each car model, improving the efficiency and accuracy of car repair claims.
[0040] Figure 2 This is a schematic diagram illustrating the process of generating the vehicle-to-parts ratio coefficient, as provided in an embodiment of this application.
[0041] like Figure 2 As shown, the entity responsible for generating the vehicle parts-to-vehicle ratio can be... Figure 1 The server 103 shown can also be other hardware devices with the same function, and this embodiment does not limit it here.
[0042] like Figure 2As shown in the embodiments of this application, the method for generating the vehicle parts-to-vehicle ratio includes the following steps:
[0043] S201: Construct a vehicle parts atlas based on historical vehicle repair claims data and a pre-trained parts big language model.
[0044] In this embodiment, historical vehicle repair claims data can be vehicle-related data from the parts repair data already entered into the insurance service user backend, such as vehicle model, parts name, and parts price. The vehicle model parts map can be a structured mapping relationship between vehicle models and parts.
[0045] Based on the above embodiments, in an optional embodiment of this application, step S201 includes:
[0046] S201a: Extract parts name data, price data, and vehicle model data based on historical vehicle repair claims data.
[0047] In this embodiment, historical vehicle repair claim data includes parts names from diverse sources, along with corresponding price data for each part and vehicle model data for which these parts belong. In this embodiment, open-source natural language processing models can be used to extract parts names, including vehicle model and price, from historical vehicle repair claim data.
[0048] S201b: Input the spare parts name data into the pre-trained spare parts large language model to obtain standard spare parts name data.
[0049] In this embodiment, the pre-trained parts large language model can be trained on an artificial intelligence model by processing and transforming parts repair data to generate a training dataset that conforms to the format required by the large language model. This pre-trained parts large language model can intelligently standardize messy parts names into standard parts names.
[0050] Specifically, in an optional embodiment of this application, the pre-trained accessory large language model is obtained through model training, wherein model training includes:
[0051] b1: Obtain and process spare parts repair data, and generate a training dataset that meets the format requirements of the initial large language model based on the processed and transformed spare parts repair data.
[0052] In this embodiment, the data processing and transformation of spare parts repair data can be done by processing or transforming historical repair claim data to generate training data that conforms to the format required by the initial large language model. This training data serves as the training dataset, which is in JSON format. Each line of JSON represents a piece of training data, where the JSON data consists of a list (array) of messages, each message being an object: system (system settings), user (user input), and assistant (assistant / model output). The system sets the identity and capabilities of the large model, helping it understand its own positioning and task scope. The user input is a jumbled list of spare parts names, and the assistant is the model's response based on the user input and the provided standard names of the spare parts.
[0053] b2: Based on the training dataset, the initial large language model is trained using a multi-turn dialogue format to obtain a pre-trained accessory large language model.
[0054] In this embodiment, the aforementioned data is organized into a multi-turn dialogue format for training a large language model. This allows the pre-trained large language model for auto parts to learn how to understand user input and provide more accurate responses regarding auto parts in specific scenarios defined by the system. Each line of JSON data represents a complete dialogue turn. Using a multi-turn dialogue input format during training facilitates the model's learning of contextual understanding and task execution capabilities.
[0055] In an optional embodiment of this application, the specific training process includes: first, loading the original pre-trained weights of the large language model. These pre-trained weights have been trained on a large-scale text corpus and have learned rich language knowledge and general capabilities. Then, using a lightweight supervised fine-tuning method based on Lora (LORA), the model is trained using the aforementioned multi-turn dialogue data. During model training, the model can convert non-standard accessory names into standard names. In the supervised fine-tuning process, data with the role as "user" is used as cues, and data with the role as "assistant" is used as labels for supervised learning. The model uses these cue-response pairs to update the weights of the large language model. The loss function uses the cross-entropy loss function as the optimization objective. The cross-entropy loss function measures the difference between the model's predicted results and the true labels, aiming to maximize the accuracy of the assistant's response. The optimizer algorithm uses the AdamW algorithm to update the model parameters. The learning rate is an important hyperparameter that controls the magnitude of model parameter updates.
[0056] S201c: Construct a multi-level parts name classification system based on standard parts name data, and construct a structured mapping relationship between vehicle models and parts based on the multi-level parts name classification system, price data, vehicle model data, and pre-stored standard parts data, as a vehicle model parts map.
[0057] In this embodiment, based on the processing of the pre-trained parts large language model, the messy parts naming data is converted into standard parts name data, and a multi-level parts name classification system is constructed on this basis. For example, in this embodiment, a three-level parts classification system is established, categorizing each part's original equipment (OE) code into 15 major categories, 147 subcategories, and 9000 standard parts names. Simultaneously, data from official parts catalogs of some automotive OEMs and 4S stores, as well as 4S store parts lists, are integrated for supplementation and cross-validation to ensure the completeness and timeliness of the structured mapping and classification of vehicle models and parts, constructing a structured vehicle-parts mapping relationship covering tens of thousands of vehicle models and 20 million parts. This mapping relationship can be a structured data table, with indexes between elements for mutual access.
[0058] In this embodiment, by constructing a vehicle parts and components map, the accuracy of matching vehicle models with parts and components and the level of automated indexing are greatly improved, providing data support for subsequent calculation of the vehicle parts-to-whole ratio coefficient for each vehicle model and horizontal data comparison.
[0059] S202: Perform outlier detection processing based on historical vehicle repair claims data and external spare parts price data to obtain target spare parts price data.
[0060] In this embodiment, the standard OE code for auto parts is used as the main index. Based on massive historical auto parts claim price data from automotive repair and claims service users, as well as external auto parts price data, a multi-dimensional auto parts price information database is constructed, including OEM suggested retail prices, dealership quotes, third-party market prices, and historical average repair prices. However, outliers still exist, requiring outlier detection processing. Outliers are replaced with normal values, and the resulting outlier-free auto parts price data is used as the target auto parts price data. In this embodiment, outliers in automotive auto parts prices can be initially identified using the Grubbs test. Then, the Isolation Forest algorithm further detects and processes outliers in the Grubbs test-cleaned auto parts price data, improving the quality and accuracy of the final target auto parts price data.
[0061] Based on the above embodiments, in an optional embodiment of this application, step S202 includes:
[0062] S202a: Price extraction and processing are performed based on historical vehicle repair claim data to obtain claim parts price data.
[0063] In this embodiment, relevant data on spare parts prices are extracted from historical vehicle repair claim data using natural language processing (NLP) tools or an artificial intelligence entity with specific data extraction functions to obtain claim parts price data.
[0064] S202b: Based on the claim parts price data and external spare parts price data, preliminary identification and processing are performed to obtain multiple sets of intermediate spare parts price data after removing invalid price data.
[0065] In this embodiment, the preliminary identification process can be the process of calculating and judging abnormal price data through the Grubbs test method.
[0066] Specifically, in an optional embodiment of this application, step S202b specifically includes:
[0067] Step B1: Based on the claimed parts price data and external parts price data, determine the mean, standard deviation, median, interpolation, and dataset size of the spare parts price data.
[0068] Step B2: Based on the mean and standard deviation of the claimed parts price data, external parts price data, and spare parts price data, determine the degree of deviation between the spare parts price data points and the mean.
[0069] Step B3: Determine the target critical value based on the dataset size, pre-selected significance level values, and the preset validation critical value table.
[0070] Step B4: Based on the target threshold and deviation level, identify outliers in the claims component price data and external component price data.
[0071] Step B5: Replace outliers in the claims parts price data and external parts price data with the mean, median, or interpolation, and then proceed to the step of determining the mean, standard deviation, median, interpolation, and dataset size of the parts price data based on the claims parts price data and external parts price data, until there are no outliers in the claims parts price data and external parts price data, so as to obtain multiple sets of intermediate parts price data after removing invalid price data.
[0072] In this embodiment, the price data of claimed parts and external parts are calculated by data processing software or by calling artificial intelligence tools with calculation functions to obtain the mean, standard deviation, median, interpolation and dataset size of the parts price data.
[0073] In this embodiment, for each price in the dataset, its deviation value Z is calculated. The formula for calculating the Z value is: Z = (price - mean) / standard deviation.
[0074] In this embodiment, the critical value of the Grubbs test is found based on the Grubbs test critical value table, according to the size of the dataset and the selected significance level of 0.05.
[0075] In this embodiment, after outliers are detected, they are removed from the dataset or replaced with other values such as the mean, median, or interpolation. This results in multiple sets of intermediate spare parts price data after removing invalid price data. Since the Grubbs test can only identify one outlier at a time, multiple iterations are required to clean up all outliers. In each iteration, the mean, standard deviation, and Z-value need to be recalculated, and the above operations of identifying, processing, and replacing outliers are repeated until no more outliers remain in the dataset.
[0076] S202c: Perform outlier detection and removal on multiple sets of intermediate spare parts price data to obtain multiple sets of target spare parts price data.
[0077] In an optional embodiment of this application, step S202c specifically includes:
[0078] Step c1: Perform array construction processing on multiple sets of intermediate spare parts price data to obtain a spare parts price feature array, where the spare parts price feature array includes multiple elements, each element corresponding to a set of intermediate spare parts price data.
[0079] Step c2: Input the component price feature array into the pre-trained isolated forest model to obtain the outlier scores of each element in the component price feature array.
[0080] Step c3: Determine the abnormal elements in the component price feature array based on the outlier scores of each element in the component price feature array and the preset outlier ratio estimate.
[0081] Step c4: Clear the abnormal elements in the component price feature array to obtain multiple sets of target component price data.
[0082] In this embodiment, after initial cleaning and detection using the Grubbs test, a two-dimensional price feature array is constructed, for example: [[price1], [price2], ..., [price11]]. Then, the Isolation Forest algorithm can be used to further detect and process outliers in this two-dimensional price feature array.
[0083] In this embodiment, the process of outlier detection and removal for multiple sets of intermediate spare parts price data can be achieved using an isolated forest algorithm model to further detect and process outliers in prices. This ultimately yields multiple sets of target spare parts price data.
[0084] In this embodiment, the pre-trained isolated forest model can be an initialized open-source isolated forest model trained using pre-prepared automotive parts price data. Initializing the open-source isolated forest model involves setting key parameters, such as: setting the number of isolated trees in the forest (n_estimators, e.g., n_estimators=40); setting the estimated proportion of outliers in the dataset (contamination), which can be used to set the threshold of the decision function and influence the identification of outliers; here, contamination=0.1; and setting the random number seed (random_state) to control the randomness of the model and ensure the repeatability of the results.
[0085] In this embodiment, outlier scoring is performed first. For each price in the dataset, a trained Isolation Forest model is used to calculate its outlier score. The lower the score, the easier it is for the price to be isolated, and therefore the more likely it is to be an outlier. Then, outlier determination is performed. Based on the outlier scores and the contamination parameter, the contamination parameter is used as an estimate of the outlier proportion, and samples with scores lower than the quantile corresponding to this proportion are selected as outliers.
[0086] In this embodiment, the isolated forest algorithm is used to further detect and process outliers in the automotive parts price data after the initial cleaning by Grubbs test through the above steps, which greatly improves the accuracy of the basic parts price data.
[0087] S202d: Aggregate multiple sets of target spare parts price data to obtain target spare parts price data.
[0088] In this embodiment, based on the cleaned price data of spare parts from multiple sources, the median aggregation algorithm can be used as the final target price data of the spare parts, which improves the representativeness and robustness of the target price data and provides a solid data foundation for the subsequent calculation of the spare parts to whole vehicle ratio.
[0089] S203: Based on the vehicle parts map, determine the complete parts list for each vehicle model.
[0090] In this embodiment, based on the vehicle model and parts map, a complete parts list corresponding to each vehicle model is calculated. For example, the names of standard parts that have a mapping relationship with vehicle model A are extracted to generate a complete parts list.
[0091] S204: Determine the vehicle parts-to-vehicle ratio coefficient for each vehicle model based on the complete parts list, target parts price data, and vehicle price for each model.
[0092] In this embodiment, the price data of each part in the complete parts list corresponding to each model is extracted from the target parts price data and added together. Finally, the ratio is calculated based on the whole vehicle price of each model to obtain the vehicle parts-to-whole ratio coefficient for each model.
[0093] Based on the above embodiments, in an optional embodiment of this application, step S204 specifically includes:
[0094] S204a: Determine the total price of vehicle parts for each vehicle model based on the complete parts list and target parts prices for each model.
[0095] S204b: Determine the vehicle parts-to-vehicle ratio coefficient for each vehicle model based on the total price of vehicle parts and the price of the whole vehicle.
[0096] In this embodiment, the total price of vehicle parts for each model can be obtained through cumulative calculation. In this embodiment, the vehicle price can be a pre-set value such as the vehicle's selling price or the market average price. In this embodiment, the vehicle parts-to-vehicle ratio refers to the ratio of the total price of vehicle parts to the vehicle price or the vehicle's selling price. Therefore, by calculating the ratio of the total price of vehicle parts to the vehicle price for each model, the vehicle parts-to-vehicle ratio for each model can be obtained.
[0097] In summary, the method for generating the vehicle parts-to-vehicle ratio provided in this application constructs a vehicle parts atlas using historical vehicle repair and claims data and a pre-trained parts large language model. Then, it performs anomaly data processing to obtain target parts price data. Next, using the vehicle parts atlas, it determines a complete standard parts table for each vehicle model. Finally, using the target parts price data and the vehicle price for each model, it derives the final vehicle parts-to-vehicle ratio for each model. This significantly improves the coverage, accuracy, timeliness, and industry applicability of the calculated vehicle models, enhancing the accuracy and efficiency of the vehicle parts-to-vehicle ratio calculation.
[0098] Figure 3 A schematic diagram of a device for generating the vehicle parts-to-vehicle ratio provided in this application is shown below. Figure 3 As shown, the vehicle parts-to-vehicle ratio generation device provided in this embodiment includes: a map construction module 301, an abnormal data processing module 302, a list generation module 303, and a parts-to-vehicle ratio determination module 304.
[0099] Among them, the graph construction module 301 is used to construct a vehicle parts graph based on historical vehicle repair claims data and pre-trained parts big language model;
[0100] The abnormal data processing module 302 is used to perform outlier detection processing based on historical automobile repair claim data and external spare parts price data to obtain target spare parts price data;
[0101] The list generation module 303 is used to determine the complete list of spare parts for each vehicle model based on the vehicle model spare parts map;
[0102] The parts-to-vehicle ratio determination module 304 is used to determine the parts-to-vehicle ratio of each vehicle model based on the complete parts list, target parts price data, and vehicle price of each model.
[0103] In an optional embodiment of this application, the graph construction module 301 is specifically used for: extracting spare parts name data, price data, and vehicle model data based on historical vehicle repair and claims data; inputting the spare parts name data into a pre-trained spare parts large language model to obtain standard spare parts name data; constructing a multi-level spare parts name classification system based on the standard spare parts name data; and constructing a structured mapping relationship between vehicle models and spare parts as a vehicle spare parts graph based on the multi-level spare parts name classification system, price data, vehicle model data, and pre-stored standard spare parts data.
[0104] In an optional embodiment of this application, the graph construction module 301 is further specifically used to: acquire spare parts repair data processing and conversion data, and generate a training dataset that conforms to the format required by the initial large language model based on the spare parts repair data processing and conversion data; and train the initial large language model using a multi-turn dialogue format based on the training dataset to obtain a pre-trained spare parts large language model.
[0105] In an optional embodiment of this application, the abnormal data processing module 302 is specifically used for: performing price extraction processing based on historical automobile repair claim data to obtain claim parts price data; performing preliminary identification processing based on claim parts price data and external spare parts price data to obtain multiple sets of intermediate spare parts price data after removing invalid price data; performing outlier detection and removal processing on the multiple sets of intermediate spare parts price data to obtain multiple sets of target spare parts price data; and performing aggregation processing based on the multiple sets of target spare parts price data to obtain target spare parts price data.
[0106] In an optional embodiment of this application, the abnormal data processing module 302 is further specifically used for: determining the mean, standard deviation, median, interpolation, and dataset size of the spare parts price data based on the claimed spare parts price data and the external spare parts price data; determining the deviation value between the spare parts price data points and the mean based on the claimed spare parts price data, the external spare parts price data, and the mean and standard deviation of the spare parts price data; determining a target threshold value based on the dataset size, a pre-selected significance level value, and a preset verification threshold value table; determining outliers in the claimed spare parts price data and the external spare parts price data based on the target threshold value and the deviation value; replacing the outliers in the claimed spare parts price data and the external spare parts price data with the mean, median, or interpolation, and jumping to the step of determining the mean, standard deviation, median, interpolation, and dataset size of the spare parts price data based on the claimed spare parts price data and the external spare parts price data, until there are no outliers in the claimed spare parts price data and the external spare parts price data, so as to obtain multiple sets of intermediate spare parts price data after removing invalid price data.
[0107] In an optional embodiment of this application, the abnormal data processing module 302 is further specifically used for: performing array construction processing on multiple sets of intermediate spare parts price data to obtain a spare parts price feature array, wherein the spare parts price feature array includes multiple elements, each element corresponding to a set of intermediate spare parts price data; inputting the spare parts price feature array into a pre-trained isolated forest model to obtain the outlier score of each element in the spare parts price feature array; determining the abnormal elements in the spare parts price feature array based on the outlier score of each element in the spare parts price feature array and a preset outlier ratio estimate; and clearing the abnormal elements in the spare parts price feature array to obtain multiple sets of target spare parts price data.
[0108] In an optional embodiment of this application, the parts-to-vehicle ratio determination module 304 is specifically used to: determine the total price of vehicle parts for each vehicle model based on the complete parts list and target parts price for each vehicle model; and determine the vehicle parts-to-vehicle ratio for each vehicle model based on the total price of vehicle parts and the price of the whole vehicle.
[0109] This embodiment provides a vehicle parts-to-vehicle ratio generation device, which can execute the method provided in the above-described method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0110] Figure 4 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 4As shown, the electronic device 40 provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the electronic device 40 further includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.
[0111] In a specific implementation, at least one processor 401 executes computer execution instructions stored in memory 402, causing at least one processor 401 to perform the above-described method.
[0112] The specific implementation process of processor 401 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0113] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0114] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0115] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0116] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0117] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0118] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0119] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0120] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0121] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0122] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0123] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0124] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0125] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for generating the vehicle parts-to-vehicle ratio coefficient, characterized in that, include: Based on historical vehicle repair claims data and pre-trained parts big language model, construct a vehicle parts atlas; Anomaly detection processing is performed based on the historical vehicle repair claims data and external spare parts price data to obtain the target spare parts price data; Based on the vehicle parts map, determine the complete parts list for each vehicle model; Based on the complete list of spare parts for each vehicle model, the price data of the target spare parts, and the price of the whole vehicle for each model, determine the vehicle parts-to-whole ratio coefficient for each model.
2. The method according to claim 1, characterized in that, The construction of the vehicle parts atlas based on historical vehicle repair claims data and a pre-trained parts big language model includes: Extract spare parts name data, price data, and vehicle model data from historical auto repair claims data; The part name data is input into the pre-trained part large language model to obtain standard part name data; A multi-level parts name classification system is constructed based on the standard parts name data. Based on the multi-level parts name classification system, price data, vehicle model data, and pre-stored standard parts data, a structured mapping relationship between vehicle models and parts is constructed as a vehicle model parts map.
3. The method according to claim 2, characterized in that, The pre-trained accessory large language model is obtained through model training, wherein the model training includes: Acquire and process spare parts repair data, and generate a training dataset that conforms to the format required by the initial large language model based on the processed and transformed spare parts repair data. The initial large language model is trained using a multi-turn dialogue format based on the training dataset to obtain a pre-trained accessory large language model.
4. The method according to claim 1, characterized in that, The step of performing outlier detection processing based on the historical vehicle repair claims data and external spare parts price data to obtain the target spare parts price data includes: Price extraction processing is performed based on the historical vehicle repair claims data to obtain the price data of the claimed parts; Based on the claimed parts price data and external spare parts price data, preliminary identification processing is performed to obtain multiple sets of intermediate spare parts price data after removing invalid price data; Outlier detection and removal are performed on multiple sets of intermediate spare parts price data to obtain multiple sets of target spare parts price data. The target spare parts price data is obtained by aggregating the multiple sets of target spare parts price data.
5. The method according to claim 4, characterized in that, The preliminary identification process based on the claimed parts price data and external spare parts price data yields multiple sets of intermediate spare parts price data after removing invalid price data, including: Based on the claimed parts price data and external parts price data, determine the mean, standard deviation, median, interpolation, and dataset size of the spare parts price data; Based on the mean and standard deviation of the claimed parts price data, external parts price data, and spare parts price data, determine the degree of deviation between the spare parts price data points and the mean. The target threshold value is determined based on the dataset size, pre-selected significance level values, and a preset validation threshold value table. Based on the target threshold and the deviation value, outliers in the claimed parts price data and external parts price data are determined; Replace the outliers in the claimed parts price data and external parts price data with the mean, median, or interpolation, and then proceed to the step of determining the mean, standard deviation, median, interpolation, and dataset size of the parts price data based on the claimed parts price data and external parts price data, until there are no outliers in the claimed parts price data and external parts price data, so as to obtain multiple sets of intermediate parts price data after removing invalid price data.
6. The method according to claim 4, characterized in that, The process of detecting and removing outliers from multiple sets of intermediate spare parts price data yields multiple sets of target spare parts price data, including: The multiple sets of intermediate spare parts price data are processed to form an array to obtain a spare parts price feature array, wherein the spare parts price feature array includes multiple elements, and each element corresponds to a set of intermediate spare parts price data; The component price feature array is input into a pre-trained isolated forest model to obtain the outlier scores of each element in the component price feature array; Based on the outlier scores of each element in the component price feature array and the preset outlier ratio estimate, the outlier elements in the component price feature array are determined. Abnormal elements in the component price feature array are cleared to obtain multiple sets of target component price data.
7. The method according to claim 1, characterized in that, Based on the complete parts list for each vehicle model, target parts price data, and the vehicle price for each model, determine the parts-to-vehicle ratio for each model, including: Based on the complete parts list and target parts prices for each vehicle model, determine the total price of vehicle parts for each vehicle model; The parts-to-vehicle ratio coefficient for each vehicle model is determined based on the total price of vehicle parts and the price of the whole vehicle.
8. A device for generating the vehicle parts-to-vehicle ratio, characterized in that, include: The graph construction module is used to build a vehicle parts graph based on historical vehicle repair claims data and a pre-trained parts big language model. The abnormal data processing module is used to perform outlier detection processing based on the historical car repair claims data and external spare parts price data to obtain the target spare parts price data; The list generation module is used to determine the complete list of spare parts for each vehicle model based on the vehicle model spare parts map; The parts-to-vehicle ratio determination module is used to determine the parts-to-vehicle ratio for each vehicle model based on the complete list of spare parts for each model, the price data of the target spare parts, and the price of the whole vehicle for each model.
9. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.