Vehicle diagnosis report processing method and device
Through standardized processing and annotation label vector generation, and combining historical vehicle diagnostic reports to train diagnostic models, the inaccurate diagnosis results caused by engineer experience dependence is solved, and high accuracy and reliability of vehicle failure prediction is achieved.
Patent Information
- Application Number
- CN202510282827.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-22
AI Technical Summary
The annotated information in the existing vehicle diagnostic report depends on the engineer's experience and habits, resulting in low credibility in the diagnosis results and affecting the accuracy of vehicle failure prediction.
By obtaining vehicle diagnostic reports in real time, standardized processing and annotation label vector generation, combining historical vehicle diagnostic reports to train diagnostic models, predict vehicle failure risks, and using incremental learning and data cleaning technologies to ensure data quality and model adaptability.
Improve the accuracy of vehicle failure prediction, ensure data reliability and model performance, promptly detect and handle vehicle failures, and avoid abnormal data interference.
Smart Images

Figure CN120356272A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle fault diagnosis, and particularly to a method and device for processing vehicle diagnostic reports. Background Art
[0002] With the widespread use of automobiles, in order to provide better user experience to users, the data stream of a vehicle can be monitored through a diagnostic tool to understand the working status of each system and component in the vehicle in real time, discover and handle potential problems of the vehicle in a timely manner. In the case where a fault has occurred, the fault location can be accurately determined, thereby improving the repair efficiency and quality and ensuring that the vehicle maintains a good operating state. In the process of initially generating a vehicle diagnostic report through a diagnostic tool, professional engineers are required to annotate the content and generate a complete diagnostic report in combination with the monitored vehicle data.
[0003] However, since the annotation information is relatively dependent on the experience and habits of engineers, and the description habits of each engineer are different, relatively subjective descriptions may occur, which has a certain impact on the data processing and analysis of vehicle diagnostic reports, resulting in a low credibility of diagnostic results. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method and device for processing vehicle diagnostic reports, which can fuse annotations and vehicle data, avoid the influence of single data on decision-making, and improve the accuracy of vehicle fault prediction.
[0005] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for processing a vehicle diagnostic report is provided, including: obtaining in real time a vehicle diagnostic report provided by a diagnostic tool, where the vehicle diagnostic report includes vehicle data and annotations for the vehicle data; performing standardization processing on the vehicle data to obtain standardized data, and generating an annotation tag vector for the annotations; according to the standardized data and the annotation tag vector, using a preset diagnostic model to predict a fault probability, and according to the output result of the diagnostic model, predicting whether the vehicle has a fault risk, where the diagnostic model is trained through historical vehicle diagnostic reports within a first preset historical time period.
[0006] Optionally, the performing standardization processing on the vehicle data to obtain standardized data includes: for each type of vehicle data of the first dimension data type in the vehicle diagnostic report, performing the following operations: determining the mean and standard deviation corresponding to the vehicle data, where the mean and the standard deviation are calculated based on the historical vehicle data corresponding to the first dimension data type; based on the mean and the standard deviation, standardizing the vehicle data belonging to the first dimension data type to obtain standardized data.
[0007] Optionally, the above method further includes: obtaining a historical vehicle diagnosis report within a first preset historical time period; the historical vehicle diagnosis report includes historical vehicle data and historical annotations for the historical vehicle data; performing standardization processing on the historical vehicle data to obtain standardized historical data, and generating a historical annotation label vector for the historical annotations; iteratively training a big data model according to the standardized historical data and the historical annotation label vector to obtain a diagnosis model.
[0008] Optionally, the performing standardization processing on the historical vehicle data to obtain standardized historical data includes: for historical vehicle data belonging to the same first dimension data type, calculating the mean and standard deviation of the historical vehicle data; standardizing each piece of the historical vehicle data based on the mean and standard deviation to obtain standardized historical data.
[0009] Optionally, the generating a historical annotation label vector for the historical annotations includes: for each of the historical annotations, performing: matching the annotation keywords in the annotation keyword library with the historical annotation; in response to the number of occurrences of the annotation keyword in the historical annotation being greater than or equal to a first preset number threshold, setting the initial label vector corresponding to the historical annotation to 1; in response to the number of occurrences of the annotation keyword in the historical annotation being less than the first preset number threshold, setting the initial label vector corresponding to the historical annotation to 0; in response to there being more than one type of annotation keyword in the historical annotation, merging the initial label vectors associated with each annotation keyword to generate a multi-dimensional historical annotation label vector for the historical annotation, where each dimension of the historical annotation label vector corresponds to a historical keyword; in response to there being only one type of annotation keyword in the historical annotation, using the initial label vector associated with the annotation keyword as the historical annotation label vector for the historical annotation.
[0010] Optionally, the above method further includes: obtaining a historical vehicle diagnosis report within a second preset historical time period; constructing a historical annotation set by integrating the historical annotations in the historical vehicle diagnosis report within the second preset historical time period; using the historical annotations with the number of occurrences greater than a second preset number threshold in the historical annotation set as keywords to construct an annotation keyword library.
[0011] Optionally, the iteratively training a big data model according to the standardized historical data and the historical annotation label vector to obtain a diagnosis model includes: classifying the historical annotation label vector according to the second dimension data type of the standardized historical data; using the historical annotation label vectors belonging to the same class in the classification result to generate a type feature vector for each second dimension data type; training the big data model according to multiple such type feature vectors.
[0012] Optionally, the above method further includes: for each of the above historical annotation tag vectors, using the above historical annotation tag vector and its corresponding standardized historical data to determine the feature weight of the above historical annotation tag vector. The above generating a type feature vector for each second-dimension data type includes: for the historical annotation tag vectors belonging to the same second-dimension data type, performing weighted fusion based on the feature weights corresponding to each historical annotation tag vector to obtain a feature vector.
[0013] Optionally, the above determining the feature weight of the above historical annotation tag vector includes: for each of the above historical annotation tag vectors, using the above historical annotation tag vector and its corresponding standardized data to calculate the importance score of the above historical annotation tag vector; according to the importance score of the above historical annotation tag vector, calculating its corresponding feature weight.
[0014] Optionally, the above method further includes: screening the above historical annotation tag vectors belonging to the same class and their corresponding standardized data according to the above importance score. The above calculating its corresponding feature weight according to the importance score of the above historical annotation tag vector includes: calculating the corresponding feature weight according to the importance scores of each target historical annotation tag vector belonging to the same class after screening.
[0015] Optionally, before the above standardizing the above historical vehicle data to obtain standardized historical data, the above method further includes: for the historical vehicle data of each second-dimension data type and its corresponding historical annotation, marking abnormal historical vehicle data and abnormal historical annotation to determine the number of valid data corresponding to the above second-dimension data type; according to the number of valid data of each of the above second-dimension data types and the total number of data of each of the above second-dimension data types, analyzing whether to start a data cleaning program for the above historical vehicle data and its corresponding historical annotation.
[0016] Optionally, the above analyzing whether to start a data cleaning program for the above historical vehicle data and its corresponding historical annotation includes: according to the number of valid data of each of the above second-dimension data types and the total number of data of each of the above second-dimension data types, calculating the data quality score of all the above historical vehicle data and its corresponding historical annotation; in response to the above data quality score being less than a first preset score threshold, triggering data cleaning for all the above historical vehicle data and its corresponding historical annotation;
[0017] Or,
[0018] For each of the above second - dimension data types, calculate the data quality score corresponding to the second - dimension data type according to the number of valid data of the second - dimension data type and the total number of data of the second - dimension data type; in response to the data quality score corresponding to the second - dimension data type being less than the second preset score threshold, trigger data cleaning for the historical vehicle data and historical annotations corresponding to the second - dimension data type.
[0019] Optionally, the above - mentioned method further includes: in response to the newly obtained vehicle diagnostic report reaching a preset data volume, calculate the data quality score of the newly obtained vehicle diagnostic report; in response to the data quality score of the newly obtained vehicle diagnostic report being greater than or equal to the third preset score threshold, trigger incremental learning to update the parameters of the diagnostic model.
[0020] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided a vehicle diagnostic report processing device, including: an acquisition module, configured to acquire in real - time a vehicle diagnostic report provided by a diagnostic tool, where the vehicle diagnostic report includes vehicle data and annotations for the vehicle data; a processing module, configured to perform normalization processing on the vehicle data to obtain normalized data, and generate an annotation label vector for the annotations; a prediction module, configured to perform a fault probability prediction using a preset diagnostic model according to the normalized data and the annotation label vector, and predict whether there is a fault risk for the vehicle according to the output result of the diagnostic model, where the diagnostic model is trained by historical vehicle diagnostic reports within a first preset historical time period.
[0021] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided an electronic device for vehicle diagnostic report processing.
[0022] An electronic device for vehicle diagnostic report processing according to an embodiment of the present invention includes: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the one or more processors to implement a vehicle diagnostic report processing method according to an embodiment of the present invention.
[0023] To achieve the above object, according to still another aspect of the embodiments of the present invention, there is provided a computer - readable storage medium.
[0024] A computer - readable storage medium according to an embodiment of the present invention stores a computer program, and when the program is executed by a processor, it implements a vehicle diagnostic report processing method according to an embodiment of the present invention.
[0025] One embodiment of the above invention has the following advantages or beneficial effects: By analyzing the obtained vehicle diagnostic report using a diagnostic model trained with historical vehicle diagnostic reports to predict whether there is a fault risk in the vehicle. Since the data input into the diagnostic model comes from the vehicle, these data can relatively truthfully reflect the vehicle condition. And the data input into the diagnostic model is the annotation label vector generated by standardizing the vehicle data and adding annotations, which fuses the annotations and vehicle data, avoids the influence of single data on decision-making, improves the data quality of the data input into the diagnostic model, ensures the high reliability of the data used in the model input and training process, and improves the accuracy of vehicle fault prediction.
[0026] At the same time, data cleaning ensures the overall quality of the data, ensures the stability of the data distribution, and avoids the interference of abnormal data on model training.
[0027] By incrementally learning to dynamically adjust model parameters, dynamic feature weighting, etc., the model performance is significantly improved, making the model pay more attention to important features, and updating the model parameters according to the newly obtained vehicle diagnostic report, enabling the diagnostic model to adapt to new data in a timely manner and maintaining good model performance.
[0028] By setting a medium-risk probability threshold and a high-risk probability combined with an alarm mechanism, risk early warning is carried out according to the fault probability predicted by the model to ensure the timely discovery and handling of vehicle faults.
[0029] The further effects of the above non-conventional optional methods will be described below in combination with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0031] Figure 1 is a schematic flowchart of vehicle diagnostic report processing according to an embodiment of the present invention;
[0032] Figure 2 is a schematic flowchart of standardization processing according to an embodiment of the present invention;
[0033] Figure 3 is a schematic flowchart of generating annotation label vectors according to an embodiment of the present invention;
[0034] Figure 4 is a schematic flowchart of establishing a diagnostic model according to an embodiment of the present invention;
[0035] Figure 5 is a schematic flowchart of generating historical annotation label vectors according to an embodiment of the present invention;
[0036] Figure 6It is a schematic flow chart of constructing an annotation keyword library according to an embodiment of the present invention;
[0037] Figure 7 It is a schematic flow chart of training a big data model according to an embodiment of the present invention;
[0038] Figure 8 It is a schematic flow chart of establishing a diagnosis model according to another embodiment of the present invention;
[0039] Figure 9 It is a schematic flow chart of a vehicle diagnosis report processing method according to another embodiment of the present invention;
[0040] Figure 10 It is a schematic structural diagram of a vehicle diagnosis report processing device according to an embodiment of the present invention;
[0041] Figure 11 It is an exemplary system architecture diagram to which an embodiment of the present invention can be applied;
[0042] Figure 12 It is a schematic structural diagram of a computer system of a server suitable for implementing an embodiment of the present invention. Detailed implementation manners
[0043] The following makes an explanation of exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0044] It should be noted that, without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0045] Figure 1 It is a schematic diagram of the main steps of a vehicle diagnosis report processing method according to an embodiment of the present invention. As Figure 1 shown, the vehicle diagnosis report processing method of the embodiment of the present invention mainly includes the following steps S101 to step S103:
[0046] Step S101, obtaining in real time a vehicle diagnosis report provided by a diagnostic tool, where the vehicle diagnosis report includes vehicle data and annotations for the vehicle data.
[0047] A diagnostic tool is pre-installed in the vehicle to monitor vehicle data in real time and generate a vehicle diagnostic report after being annotated by a vehicle engineer. Vehicle data refers to the data related to vehicle performance generated by the vehicle that is monitored in real time by the diagnostic tool. According to the second-dimensional data type, the above vehicle data may include, but is not limited to, battery (Battery Management System, BMS) data, engine data, charging data, inverter data, etc. The second-dimensional data type can be further divided to obtain vehicle data of the first-dimensional data type. For example, battery data includes, but is not limited to, battery cell temperature, State of Health (SOH) of the battery, State Of Charge (SOC) of the battery, voltage, discharge rate, etc. Engine data includes, but is not limited to, motor speed, torque, power, etc. Charging data includes, but is not limited to, charging voltage, current, charging stage identifier, etc. Inverter data includes, but is not limited to, inverter output voltage, inverter current, inverter efficiency, etc.
[0048] The annotation for vehicle data is the vehicle data annotated by a vehicle engineer based on the vehicle data. For example, when the battery cell temperature is high, the vehicle engineer may annotate "BMS temp high"; when there is a spike in the inverter current, the vehicle engineer may annotate "INV current spike". It can be understood that each annotation may only be for one piece of vehicle data or may be for multiple pieces of vehicle data.
[0049] Step S102, perform standardization processing on the above vehicle data to obtain standardized data, and generate an annotation label vector for the above annotation;
[0050] Before using the diagnostic model to predict the vehicle fault probability, in order to make different data have the same scale and eliminate the difference in data dimensions, the vehicle data of the same first-dimensional data type can be standardized. Valuable information is extracted from the annotation by generating an annotation label vector, and then it is transformed into a structured label vector.
[0051] Step S103, according to the above standardized data and the above annotation label vector, use a preset diagnostic model to predict the fault probability, and according to the output result of the above diagnostic model, predict whether there is a fault risk for the above vehicle. The above diagnostic model is trained through historical vehicle diagnostic reports within a first preset historical time period.
[0052] Use the diagnostic model to output the predicted fault probability according to the standardized data and the annotation label vector, so as to predict whether there is a fault risk for the vehicle through the fault probability.
[0053] Specifically, a medium-risk probability threshold and a high-risk probability threshold are preset. When the failure probability output by the diagnosis model is less than or equal to the medium-risk probability threshold, it is determined that the failure risk of the current vehicle is low; when the failure probability output by the diagnosis model is less than or equal to the high-risk probability threshold and greater than the medium-risk probability threshold, it is determined that the failure risk of the current vehicle is medium; when the failure probability output by the diagnosis model is greater than the high-risk probability threshold, it is determined that the failure risk of the current vehicle is high.
[0054] Corresponding warning prompts can be made according to the prediction results of the failure risk.
[0055] Furthermore, the method of warning prompt can be executed through a smart contract or a control strategy. Among them, a smart contract means that when the contract terms are met, the smart contract will be automatically executed. For example, contract clause 1 is whether the failure probability is less than or equal to the high-risk probability threshold and greater than the medium-risk probability threshold. When the failure probability is less than or equal to the high-risk probability threshold and greater than the medium-risk probability threshold, the vehicle failure risk is medium risk, and the smart contract is automatically executed to give a medium-risk failure warning; contract clause 2 is whether the failure probability is greater than the high-risk probability threshold. When the failure probability is greater than the high-risk probability threshold, the vehicle failure risk is high risk, and the smart contract is automatically executed to give a high-risk failure warning.
[0056] Even further, a medium-risk occurrence threshold for the failure probability less than or equal to the high-risk probability threshold and greater than the medium-risk probability threshold, and a high-risk occurrence threshold for the number of times the failure probability is greater than the high-risk probability threshold can also be set. For example, contract clause 3 is whether the failure probability is greater than the high-risk probability threshold and whether the consecutive number of times greater than the high-risk probability threshold reaches the high-risk occurrence threshold. When this is met, the smart contract is automatically executed to give a high-risk failure warning; when the failure probability is less than the high-risk probability threshold, the smart contract will not be triggered; when the failure probability is greater than the high-risk probability threshold but the consecutive number of times greater than the high-risk threshold is less than the high-risk occurrence threshold, the smart contract will not be triggered. It can be understood that the consecutive number of times the failure probability is greater than the high-risk probability threshold reaching the high-risk occurrence threshold means that the failure probability output by the diagnosis model is continuously greater than the high-risk probability threshold, and the number of times greater than the high-risk probability threshold reaches the high-risk occurrence threshold, without any situation where the failure probability is less than or equal to the high-risk probability threshold in between.
[0057] During the warning process, vehicle overall fault warning can be performed, or the frequency of keyword occurrences in the annotation can be analyzed. If the fault probability indicates that the vehicle is in medium or high risk, it can be determined whether there are keywords in the annotation whose occurrence times exceed the keyword frequency threshold when the medium or high risk occurs. If so, vehicle fault warning for the system corresponding to the keyword can be performed according to the keyword. As an example, when the fault probability indicates that the vehicle is in medium risk for three consecutive cycles, analyzing the annotation shows that the keyword "INV freq shake" appears continuously in the above three cycles, then inverter fault warning can be performed.
[0058] In an alternative embodiment, the above-mentioned standardization processing of the above-mentioned vehicle data to obtain standardized data includes:
[0059] As Figure 2 shown, for the vehicle data of each first-dimensional data type in the above-mentioned vehicle diagnosis report, the following operations are performed:
[0060] Step S201, determining the mean value and standard deviation corresponding to the above-mentioned vehicle data, where the above-mentioned mean value and the above-mentioned standard deviation are calculated based on the historical vehicle data corresponding to the above-mentioned first-dimensional data type;
[0061] Step S202, based on the above-mentioned mean value and the above-mentioned standard deviation, standardize the vehicle data belonging to the above-mentioned first-dimensional data type to obtain standardized data.
[0062] The first-dimensional data type refers to the smallest data type level of division. For example, the battery cell temperature, remaining power, etc. in battery data, and the motor speed, torque, etc. in engine data. For the vehicle data of each first-dimensional data type, numerical calculations are performed through the mean value and standard deviation to obtain the standardized data corresponding to each vehicle data of the first-dimensional data type, as shown in Equation 1.
[0063]
[0064] Among them, DN(x) represents the standardized data corresponding to the vehicle data; x represents the vehicle data; μ represents the mean value of the historical vehicle data corresponding to the first-dimensional data type of the vehicle data; σ represents the standard deviation of the historical vehicle data corresponding to the first-dimensional data type of the vehicle data. It should be noted that both the mean value and the standard deviation are the mean value and standard deviation of the historical vehicle data calculated according to the historical vehicle data during the model training process.
[0065] As an example, for the vehicle data with the first-dimensional data type of rotational speed, the mean value of the historical vehicle data corresponding to the rotational speed is 0, and the standard deviation is 1. Furthermore, the standardized data of the vehicle data can be calculated through the above Equation 1.
[0066] Further, after calculating the standardized data, in order to avoid the influence of abnormal data on the prediction accuracy of the diagnostic model, the abnormal data therein can be adjusted to the normal range or the abnormal data can be removed. The abnormal data can be standardized data greater than or equal to ±5σ. As an example, the calculated standardized data of the battery cell temperature is 6, which is greater than 5 times the standard deviation. Therefore, this standardized data can be removed or adjusted to 5 to make the distribution of the standardized data stable.
[0067] In an alternative embodiment, generating the annotation label vector for the above-mentioned annotation includes:
[0068] As Figure 3 shown, the following steps S301 to S303 are performed for each annotation:
[0069] Step S301, matching the annotation keywords in the annotation keyword library with the above-mentioned annotation;
[0070] Step S302, in response to the number of occurrences of the above-mentioned annotation keyword in the above-mentioned annotation being greater than or equal to the first preset number threshold, setting the initial label vector corresponding to the above-mentioned annotation to 1; in response to the number of occurrences of the above-mentioned annotation keyword in the above-mentioned annotation being less than the first preset number threshold, setting the initial label vector corresponding to the above-mentioned annotation to 0;
[0071] Step S303, in response to the number of types of annotation keywords existing in the above-mentioned annotation being greater than 1, merging the initial label vectors associated with each annotation keyword to generate a multi-dimensional annotation label vector for the above-mentioned annotation, where each dimension of the annotation label vector corresponds to a keyword; in response to the number of types of annotation keywords existing in the above-mentioned annotation being equal to 1, using the initial label vector associated with the above-mentioned annotation keyword as the annotation label vector for the above-mentioned annotation.
[0072] The annotation keyword library includes a number of annotation keywords. The annotation keywords in the annotation keyword library can be sequentially matched with the annotation, and when the annotation keyword appears in the annotation, the number of occurrences of the annotation keyword is calculated.
[0073] In the case where the number of occurrences is greater than or equal to the first preset number threshold, setting the initial label vector corresponding to the annotation to 1; in the case where the number of occurrences is less than the first preset number threshold and not 0, setting the initial label vector corresponding to the annotation to 0; in the case where no annotation keyword is matched in the annotation, identifying the annotation as an abnormally low-value annotation and removing it without generating a corresponding initial label vector.
[0074] If multiple types of annotation keywords appear in an annotation, the initial label vectors associated with each annotation keyword are merged to generate a multi-dimensional annotation label vector corresponding to the annotation; if only one type of annotation keyword appears in an annotation, the initial label vector corresponding to the annotation keyword is used as the annotation label vector of the annotation.
[0075] As an example, when the number of occurrences of "BMS temp high" in annotation 1 is 2 times, equal to the first preset number threshold of 2 times, the initial label vector associated with the annotation keyword corresponding to this annotation is 1; at the same time, the number of occurrences of "BMS temp continues to rise" in annotation 1 is 1 time, less than the first preset number threshold of 2 times, so the initial label vector associated with the annotation keyword corresponding to this annotation is 0. Since two types of annotation keywords appear in annotation 1, the initial label vectors associated with these two types of annotation keywords can be merged to generate a multi-dimensional annotation label vector [1, 0] corresponding to annotation 1, and each dimension of the annotation label vector corresponds to an annotation keyword.
[0076] Furthermore, in the process of generating a multi-dimensional annotation label vector, the order of multiple initial label vectors in the multi-dimensional annotation label vector can be sorted according to the pre-set first dimension data type order. As an example, the battery cell temperature takes precedence over the battery temperature change rate which takes precedence over the battery power. When the above three types of annotation keywords exist simultaneously in an annotation, the initial label vectors are merged in the above order to generate a multi-dimensional annotation label vector.
[0077] In an alternative embodiment, the above method further includes: as Figure 4 shown, the establishment of the diagnostic model includes the following steps S401 to step S403:
[0078] Step S401, obtaining a historical vehicle diagnostic report within a first preset historical time period; the above historical vehicle diagnostic report includes historical vehicle data and historical annotations for the above historical vehicle data;
[0079] Step S402, performing standardization processing on the above historical vehicle data to obtain standardized historical data, and generating historical annotation label vectors for the above historical annotations;
[0080] Step S403, iteratively training a big data model according to the above standardized historical data and the above historical annotation label vectors to obtain a diagnostic model.
[0081] Perform data processing on the historical vehicle data and historical annotations in the historical vehicle diagnosis report, and use the above formula (1) to generate the standardized historical data corresponding to the historical vehicle data. Similar to steps S301 to S303, generate the historical annotation label vector corresponding to the historical annotation, which will not be elaborated here.
[0082] Based on the standardized historical data and the historical annotation label vector, and using the failure probability as the expected output, perform iterative training on the big data model. After the training is completed, a diagnosis model is obtained. As an example, the big data model can adopt any model that can implement the vehicle diagnosis method of the embodiments of the present invention, and no specific limitation is made here.
[0083] In an optional embodiment, before the above-mentioned historical vehicle data is standardized to obtain the standardized historical data, the above method further includes:
[0084] For the historical vehicle data of each second-dimensional data type and its corresponding historical annotation, mark the abnormal historical vehicle data and abnormal historical annotation to determine the number of valid data corresponding to the above second-dimensional data type;
[0085] According to the number of valid data of each of the above second-dimensional data types and the total number of data of each of the above second-dimensional data types, analyze whether to start the data cleaning program for the above historical vehicle data and its corresponding historical annotation.
[0086] After obtaining the historical vehicle diagnosis report, since the amount of historical vehicle data and historical annotation data involved is large, there may be some low-quality historical vehicle data and historical annotations. To avoid the impact of these low-quality data on the prediction accuracy of the failure probability, it is possible to determine the abnormal historical vehicle data and abnormal historical annotations existing in the historical vehicle data of each second-dimensional data type and the corresponding historical annotation, and then determine the number of valid data. According to the proportion of the number of valid data in the total number of data, remove the low-quality data through data cleaning.
[0087] Among them, for each historical vehicle data, the historical vehicle data with a missing value ratio exceeding the preset missing ratio can be used as the abnormal historical vehicle data; for each historical annotation, the historical annotation with a null value or a length less than the preset character length can be used as the abnormal historical annotation.
[0088] As an example, when the data type of the second dimension is battery data, the battery data includes 4 items of historical vehicle data and 2 items of historical annotations. The missing value ratio of one item of historical vehicle data exceeds the preset missing ratio by 5%. Therefore, this item of historical vehicle data is abnormal historical vehicle data. The missing values of the remaining 3 items of historical vehicle data do not exceed the preset missing ratio. The length of one item of historical annotation is less than the preset character length by 5 characters. Therefore, this historical annotation is abnormal historical annotation, and the length of the remaining one item of historical annotation does not exceed the preset character length. Therefore, the total number of data items in the battery data is 6, and the number of valid data items is 4.
[0089] In an alternative embodiment, whether to start the data cleaning process for the above historical vehicle data and its corresponding historical annotations includes:
[0090] For each of the above data types of the second dimension, calculate the data quality score corresponding to the data type of the second dimension according to the number of valid data items and the total number of data items of the data type of the second dimension.
[0091] In response to the data quality score corresponding to the data type of the second dimension being less than the second preset score threshold, trigger data cleaning for the historical vehicle data and historical annotations corresponding to the data type of the second dimension.
[0092] For each data type of the second dimension, it is possible to determine whether to perform data cleaning on the historical vehicle data and historical annotations of the data type of the second dimension according to the proportion of the number of valid data items in the historical vehicle data and historical annotations of the data type of the second dimension to the total number of data items of the data type of the second dimension. That is, each data cleaning is for the historical vehicle data and historical annotations of one data type of the second dimension. As shown in Equation 2 below,
[0093]
[0094] where DQS 单 represents the data quality score of the historical vehicle data and historical annotations corresponding to the data type of the second dimension; C represents the number of valid data items in the historical vehicle data and historical annotations corresponding to the data type of the second dimension; M represents the total number of data items of the historical vehicle data and historical annotations corresponding to the data type of the second dimension.
[0095] When the data quality scores of the historical vehicle data and historical annotations of the calculated second - dimension data type are greater than or equal to the second preset score threshold, it indicates that the data quality of the historical vehicle data and historical annotations corresponding to the second - dimension data type is relatively high, and there is no need for data cleaning; when the data quality scores of the historical vehicle data and historical annotations of the calculated second - dimension data type are less than the second preset score threshold, it indicates that the data quality of the historical vehicle data and historical annotations corresponding to the second - dimension data type is relatively low. To avoid the impact of low - quality data on the accuracy of fault probability prediction, data cleaning for the historical vehicle data and historical annotations corresponding to the second - dimension data type can be triggered. During the data cleaning process, low - quality data in the historical vehicle data and historical annotations of the second - dimension data type is removed, including abnormal historical vehicle data and abnormal historical annotations.
[0096] As an example, continuing with the above example, when the second - dimension data type is battery data, C i is 4, M i is 6, and the data quality score of the battery data is calculated using Equation 2, and the calculated DQS 单 i is 66.7%.
[0097] By performing data quality scoring on the data of each second - dimension data type, when performing data cleaning, it is possible to more specifically clean the data of the second - dimension data type with a lower data quality score, without having to traverse all the data, reducing the computational volume and data processing volume.
[0098] Furthermore, the above analysis of whether to start the data cleaning program for the above - mentioned historical vehicle data and its corresponding historical annotations includes:
[0099] Calculate the data quality scores of all historical vehicle data and its corresponding historical annotations based on the number of valid data of each of the above - mentioned second - dimension data types and the total number of data of each of the above - mentioned second - dimension data types;
[0100] In response to the above data quality score being less than the first preset score threshold, trigger data cleaning for all the above - mentioned historical vehicle data and its corresponding historical annotations.
[0101] The data quality scores of all historical vehicle data and historical annotations can be calculated based on the number of valid data and the total number of data of each second - dimension data type in all historical vehicle diagnostic reports, combined with the number of second - dimension data types. As shown in Equation 3 below,
[0102]
[0103] where DQS 全The data quality score representing all historical vehicle data and historical annotations; N represents the number of data types in the second dimension; M i represents the total number of data of the historical vehicle data and historical annotations corresponding to the i-th data type in the second dimension; C i represents the number of valid data of the historical vehicle data and historical annotations corresponding to the i-th data type in the second dimension.
[0104] When the data quality score of all historical vehicle data and historical annotations is greater than or equal to the first preset score threshold, it indicates that the data quality of all historical vehicle data and historical annotations is relatively high and no data cleaning is required; when the data quality score of all historical vehicle data and historical annotations is less than the first preset score threshold, it indicates that the data quality of all historical vehicle data and historical annotations is relatively low, triggering data cleaning for all historical vehicle data and historical annotations. During the data cleaning process, low-quality data in all historical vehicle data and historical annotations is removed, including abnormal historical vehicle data and abnormal historical annotations.
[0105] It should be noted that the first preset score threshold and the second preset score threshold can be the same or different.
[0106] As an example, in all historical vehicle data and all historical annotations of the historical vehicle diagnostic reports within the first preset historical time period, there are a total of four data types in the second dimension, namely battery data, engine data, charging data, and inverter data. The first data type in the second dimension is battery data, with the number of valid data being 4 and the total number of data being 6, and the data quality score of the battery data being 66.7%; the second data type in the second dimension is engine data, with the number of valid data being 2 and the total number of data being 4, and the data quality score of the engine data being 50%; the third data type in the second dimension is charging data, with the number of valid data being 7 and the total number of data being 7, and the data quality score of the charging data being 100%; the fourth data type in the second dimension is inverter data, with the number of valid data being 9 and the total number of data being 10, and the data quality score of the inverter data being 90%. Based on the data quality scores corresponding to the above four data types in the second dimension, the calculated data quality score of all historical vehicle data and historical annotations is 76.7%.
[0107] By performing data quality scoring on all historical vehicle data and historical annotations, the data quality of all historical vehicle data and historical annotations can be obtained, comprehensively evaluating the data quality of all historical databases. In the case of low data quality, low-quality data in all data is removed to reduce the impact of low-quality data and ensure the rationality of subsequent feature weighting and the accuracy of failure probability.
[0108] In an alternative embodiment, the above-mentioned standardization of the historical vehicle data to obtain standardized historical data includes:
[0109] For historical vehicle data belonging to the same first-dimensional data type, calculate the mean and standard deviation of the historical vehicle data;
[0110] Based on the above mean and standard deviation, standardize each of the above historical vehicle data to obtain standardized historical data.
[0111] It should be noted that the process of standardizing historical vehicle data is similar to the process of standardizing vehicle data above, and the calculation is also carried out using Equation 1, which will not be elaborated here.
[0112] In an alternative embodiment, the above-mentioned generation of historical annotation label vectors for the above historical annotations includes:
[0113] As Figure 5 shown, the following steps S501 to S503 are performed for each of the above historical annotations:
[0114] Step S501, match the annotation keywords in the annotation keyword library with the above historical annotations;
[0115] Step S502, in response to the number of occurrences of the above annotation keywords in the above historical annotation being greater than or equal to the first preset number threshold, set the initial label vector corresponding to the above historical annotation to 1; in response to the number of occurrences of the above annotation keywords in the above historical annotation being less than the first preset number threshold, set the initial label vector corresponding to the above historical annotation to 0;
[0116] Step S503, in response to the number of types of annotation keywords existing in the above historical annotation being greater than 1, merge the initial label vectors associated with each annotation keyword to generate a multi-dimensional historical annotation label vector of the above historical annotation, where each dimension of the historical annotation label vector corresponds to a historical keyword; in response to the number of types of annotation keywords existing in the above historical annotation being equal to 1, use the initial label vector associated with the above annotation keyword as the historical annotation label vector of the above historical annotation.
[0117] It should be noted that the process of generating historical annotation label vectors for historical annotations is similar to the process of generating annotation label vectors for annotations above, and it is necessary to match the annotation keywords in the annotation keyword library with historical annotations, which will not be elaborated here.
[0118] In an alternative embodiment, the above method further includes: a method for constructing an annotation keyword library, as Figure 6 shown, specifically including the following steps S601 to S603:
[0119] Step S601: Obtain the historical vehicle diagnostic reports within the second preset historical time period;
[0120] Step S602: Construct a historical annotation set by integrating the historical annotations in the above-mentioned historical vehicle diagnostic reports within the second preset historical time period;
[0121] Step S603: Use the historical annotations that appear more than the second preset frequency threshold in the above-mentioned historical annotation set as keywords to construct an annotation keyword library.
[0122] Integrate all the historical annotations involved in the historical vehicle diagnostic reports within the preset second historical time period to construct a historical annotation set. And calculate the number of occurrences of each historical annotation in all the historical annotations, so as to use the historical annotations with relatively more occurrences as keywords to construct an annotation keyword library.
[0123] Furthermore, the historical annotation set can be continuously updated based on the obtained vehicle diagnostic reports, so as to continuously incorporate the annotations in the new vehicle diagnostic reports into the historical annotation set, further expand the annotation keyword library, realize the automatic update and expansion of the annotation keyword library, and provide a real-time data basis for vehicle diagnostic report processing.
[0124] In an alternative embodiment, as Figure 7 shown, the diagnostic model obtained by iteratively training the big data model according to the above-mentioned standardized historical data and the above-mentioned historical annotation tag vectors includes the following steps S701 to S703:
[0125] Step S701: Classify the above-mentioned historical annotation tag vectors according to the second-dimensional data type of the above-mentioned standardized historical data;
[0126] Step S702: Use the historical annotation tag vectors belonging to the same class in the classification result to generate type feature vectors for each second-dimensional data type;
[0127] Step S703: Train the big data model according to multiple above-mentioned type feature vectors.
[0128] Since the vehicle diagnostic reports contain multiple data sources, by fusing the historical annotation tag vectors from the same data source according to the second-dimensional data type, type feature vectors can be generated for the historical annotation tag vectors of each second-dimensional data type respectively, so as to fuse the type feature vectors to obtain a fusion result, and train the big data model based on the fusion result, so that the model has the ability to distinguish faults of different second-dimensional data types. As shown in Equation 4 below,
[0129] F d = F1 + F2 + …… + FN Formula 4
[0130] Wherein, F d represents the fusion result; F1 represents the type feature vector corresponding to the first second - dimension data type; F2 represents the type feature vector corresponding to the second second - dimension data type... F N represents the type feature vector corresponding to the Nth second - dimension data type; N represents the number of second - dimension data types. To prevent the data of a certain second - dimension data type from dominating the decision - making, the type feature vectors of each second - dimension data type can be fused and then the big data model can be trained to make full use of the information of multi - source data and enhance the model's ability to detect and diagnose faults.
[0131] Furthermore, the above - mentioned method further includes:
[0132] For each of the above - mentioned historical annotation label vectors, using the above - mentioned historical annotation label vector and its corresponding standardized historical data, determine the feature weight of the above - mentioned historical annotation label vector;
[0133] The above - mentioned generation of type feature vectors for each second - dimension data type includes:
[0134] For historical annotation label vectors belonging to the same second - dimension data type, perform weighted fusion based on the feature weights corresponding to each historical annotation label vector to obtain the type feature vector. As shown in Formula 5,
[0135]
[0136] Wherein, F i represents the type feature vector of the ith second - dimension data type; W h represents the feature weight corresponding to the hth historical annotation label vector in the ith second - dimension data type; D i represents the hth historical annotation label vector in the ith second - dimension data type; R represents the number of historical annotation label vectors participating in the calculation of the type feature vector of the ith second - dimension data type.
[0137] Among them, the feature weight of each historical annotation label vector can be set in advance or calculated according to the importance degree of the historical annotation label vector in all historical annotation label vectors of its corresponding second - dimension data type. Set the upper limit of the feature weight of each historical annotation label vector in advance to avoid a certain historical annotation label vector or its corresponding second - dimension data type from dominating the decision - making. When it is indicated in the historical annotation that the historical vehicle data of a certain second - dimension data type may have problems and the number of problem occurrences increases, the feature weight of this historical annotation label vector can be appropriately increased while ensuring that the feature weight does not exceed the upper limit of the feature weight.
[0138] By fusing different historical annotation tag vectors, a more representative and decision-making valuable type feature vector is obtained, and a big data model is trained through the fusion results of the type feature vectors of multiple second-dimensional data types, improving the reliability of the model.
[0139] Furthermore, determining the feature weights of the above historical annotation tag vectors includes:
[0140] For each of the above historical annotation tag vectors, using the above historical annotation tag vector and its corresponding standardized data, calculate the importance score of the above historical annotation tag vector;
[0141] According to the importance score of the above historical annotation tag vector, calculate its corresponding feature weight.
[0142] For each historical annotation tag vector, calculate its corresponding importance score through the following formula 6,
[0143] Q h =λ1·I h +λ2·R h +λ3·A h Formula 6
[0144] where, Q h represents the importance score of the h-th historical annotation tag vector; I h represents the average gain in improving the model prediction accuracy of the historical annotation corresponding to the h-th historical annotation tag vector within a preset third historical time period; λ1 represents the weight coefficient of I h ; R h represents the correlation between the h-th historical annotation tag vector and the historical vehicle data corresponding to it; λ2 represents the weight coefficient of R h ; A h represents the availability ratio of the historical vehicle data corresponding to the h-th historical annotation tag vector, that is, among the historical vehicle data of the preset historical data volume corresponding to the h-th historical annotation tag vector, the proportion of valid historical vehicle data. Among them, valid historical vehicle data refers to other historical vehicle data except abnormal historical vehicle data. As an example, the preset historical data volume can be the most recent 100 pieces of historical vehicle data.
[0145] I h measures the positive contribution of the h-th historical annotation tag vector to the model prediction accuracy during training within a preset third historical time period, and its value range is between [0,1]. I hIt can be provided by a feature importance ranking algorithm, which comprehensively considers the influence degree of the h-th historical annotation label vector on the prediction result accuracy during the model training process. Among them, the preset third historical time period can be the past month, the past week, the past three days, etc., and no specific limitation is made here.
[0146] Calculate the Pearson correlation coefficient R between the h-th historical annotation label vector and the expected corresponding historical vehicle data. h It can be positively correlated with the above Pearson correlation coefficient, or a Pearson correlation coefficient threshold can be set. When the Pearson correlation coefficient is greater than the Pearson correlation coefficient threshold, h take the first correlation value; when the Pearson correlation coefficient is less than or equal to the Pearson correlation coefficient threshold, h take the second correlation value. As an example, the Pearson correlation coefficient threshold can be set to 0.5, and the value of h is as follows,
[0147]
[0148] where b is the Pearson correlation coefficient between the h-th historical annotation label vector and the expected corresponding historical vehicle data.
[0149] A h can be the actual proportion of valid historical vehicle data among the historical vehicle data of the preset historical data volume corresponding to the h-th historical annotation label vector calculated; or it can be a proportion determined according to the above actual proportion. When A h is the proportion determined according to the above actual proportion, determine A according to whether the actual proportion is greater than the preset actual proportion threshold h . As an example, set the actual proportion threshold to 95%, and the value of A h is as follows,
[0150]
[0151] where c represents the actual proportion of valid historical vehicle data among the historical vehicle data of the preset historical data volume corresponding to the h-th historical annotation label vector.
[0152] λ1, λ2 and λ3 can be set according to the actual situation and adjusted based on experience during the model training process. The sum of the above three weight coefficients should be 1. As an example, λ1 can be 0.5, λ2 can be 0.3, and λ3 can be 0.2.
[0153] After calculating the importance scores of the historical annotation label vectors, normalization can be performed using the softmax function to obtain the feature weights of each historical annotation label vector among all the historical annotation label vectors of its corresponding second-dimensional data type.
[0154] In an alternative embodiment, the above method further includes:
[0155] Screen the above historical annotation label vectors belonging to the same class and their corresponding standardized data according to the above importance scores.
[0156] Specifically, to avoid the low-value historical annotation label vectors from affecting the accuracy of the model, the departure-time annotation label vectors for calculating the type feature vectors of this second-dimensional data type can be screened according to the proportion of the importance scores of the historical annotation label vectors in the importance scores of all the historical annotation label vectors of its corresponding second-dimensional data type, to determine the target historical annotation label vectors participating in the calculation, and the feature weights are determined based on the importance scores of the target historical annotation label vectors.
[0157] Preset a proportion threshold of the importance scores. When the proportion of the importance score of a historical annotation label vector in the importance scores of all the historical annotation label vectors of its corresponding second-dimensional data type is less than the proportion threshold of the importance scores, this historical annotation label vector is regarded as a low-value historical annotation label vector and is excluded from participating in the calculation of the type feature vectors of its corresponding second-dimensional data type. The other historical annotation label vectors in this second-dimensional data type are used as the target historical annotation label vectors.
[0158] Furthermore, an importance score update period threshold can be set. Whenever the proportion of the importance score of a historical annotation label vector in the importance scores of all historical annotation label vectors of its corresponding second - dimension data type is less than the importance score proportion threshold, continuously monitor the importance score of this historical annotation label vector to determine whether the proportion of the importance score of the historical annotation label vector in the importance scores of all historical annotation label vectors of its corresponding second - dimension data type is less than the importance score proportion threshold within the period indicated by the importance score update period threshold. If it is less in all cases, then eliminate this historical annotation label vector. As an example, set the importance score proportion threshold to 0.05 and the importance score update period threshold to 3. In the first update period, the proportion of the importance score of the historical annotation label vector in the importance scores of all historical annotation label vectors of its corresponding second - dimension data type is 0.02. In the second update period, the proportion is 0.04. In the third update period, the proportion is 0.04. It can be seen that starting from the first update period, the proportion of the importance score of this historical annotation label vector in the importance scores of all historical annotation label vectors of its corresponding second - dimension data type is less than 0.05 for three consecutive periods. Therefore, this historical annotation label vector is regarded as a low - value historical annotation label vector and is eliminated.
[0159] In addition, during the process of selecting the target historical annotation label vector, the correlation between this historical annotation label vector and the expected corresponding historical vehicle data can also be referred to. If the correlation is lower than the correlation threshold, then correspondingly increase the importance score proportion threshold corresponding to this historical annotation label vector. That is to say, if the correlation between the historical annotation label vector and the expected corresponding historical vehicle data is low, then it is required that the proportion of the importance score of this historical annotation label vector in the importance scores of all historical annotation label vectors of its corresponding second - dimension data type is higher in order to retain this historical annotation label vector.
[0160] By adjusting the importance score proportion threshold and the importance score update period threshold, adjust the number of target historical annotation label vectors selected from the historical annotation label vectors belonging to the same second - dimension data type, so as to control the historical annotation label vectors in each second - dimension data type within a reasonable scale.
[0161] Furthermore, calculating the corresponding feature weight according to the importance score of the above - mentioned historical annotation label vector includes:
[0162] According to the importance scores of each target historical annotation label vector belonging to the same category screened out, calculate its corresponding feature weight. Refer to Equation 7 below,
[0163]
[0164] where P represents the number of target historical annotation label vectors in the second-dimensional data type of the i-th item; Q j represents the importance score of the j-th target historical annotation label vector in the second-dimensional data type of the i-th item.
[0165] In an alternative embodiment, the above method further includes:
[0166] In response to the newly acquired vehicle diagnostic report reaching a preset data volume, calculate the data quality score of the newly acquired vehicle diagnostic report;
[0167] In response to the data quality score of the above newly acquired vehicle diagnostic report being greater than or equal to the third preset score threshold, trigger incremental learning and update the parameters of the above diagnostic model.
[0168] Since the vehicle's diagnostic tool continuously generates new vehicle diagnostic reports, and since the newly acquired vehicle diagnostic reports may have certain differences from the historical vehicle diagnostic reports, therefore, after the newly acquired vehicle diagnostic reports reach a certain scale, the parameters of the diagnostic model can be updated through incremental learning. As shown in Equation 8 below,
[0169]
[0170] where M t+1 represents the model parameters of the updated diagnostic model; M t represents the model parameters of the current diagnostic model; η represents the learning rate, and the value range is 0.01 to 00.05; y t represents the current actual failure probability; represents the failure probability output by the current diagnostic model. Among them, the learning rate can be dynamically adjusted based on the model error and is positively correlated with the model error. As an example, when the model error of the diagnostic model decreases, the learning rate can be reduced and adjusted to half of the previous value; when the model error of the diagnostic model increases, the learning rate can be increased and adjusted to 1.5 times the previous value.
[0171] By calculating the error between the actual failure probability (true value) and the failure probability (predicted value) output by the model, and under the adjustment of the learning rate, the current model parameters are corrected to achieve iterative update of the model parameters, enabling the model to continuously fit new data, improving the prediction accuracy, and enhancing the model performance.
[0172] In addition, when the data quality score of the newly obtained vehicle diagnostic report is less than the third preset score threshold, even if the newly obtained vehicle diagnostic report reaches the preset data volume, incremental learning is not triggered, thus avoiding interference of low-quality data on the model.
[0173] According to the vehicle diagnostic report processing method of the embodiments of the present invention, by analyzing the obtained vehicle diagnostic report with a diagnostic model trained by historical vehicle diagnostic reports to predict whether there is a fault risk in the vehicle, through standardization processing and annotation tag vector generation, annotations and vehicle data can be fused, avoiding the influence of single data on decision-making, improving the data quality, ensuring high reliability of the data used in the model input and training process, and improving the accuracy of vehicle fault prediction.
[0174] Meanwhile, the overall quality of the data is ensured through data cleaning, ensuring stable data distribution and avoiding interference of abnormal data on model training.
[0175] By dynamically adjusting model parameters through incremental learning, dynamic feature weighting, etc., the model performance is significantly improved, enabling the model to pay more attention to important features, and updating model parameters according to the newly obtained vehicle diagnostic report, enabling the diagnostic model to adapt to new data in a timely manner and maintaining good model performance.
[0176] By setting a medium risk probability threshold and a high risk probability combined with an alarm mechanism, risk early warning is carried out according to the fault probability predicted by the model to ensure timely discovery and handling of vehicle faults.
[0177] The construction method of the diagnostic model is explained below through a specific embodiment.
[0178] As Figure 8 shown, the construction method of the diagnostic model of the embodiments of the present invention includes the following steps S801 to step S812:
[0179] Step S801, obtain historical vehicle diagnostic reports within a first preset historical time period; the above historical vehicle diagnostic reports include historical vehicle data and historical annotations for the above historical vehicle data. Exemplarily, the following table provides part of the content of the vehicle diagnostic report.
[0180]
[0181] Step S802, for each of the above second-dimension data types, calculate the data quality score corresponding to the second-dimension data type according to the effective data quantity and the total data quantity of the second-dimension data type;
[0182] Step S803: In response to the data quality score corresponding to the above second-dimensional data type being less than the second preset score threshold, trigger data cleaning for the historical vehicle data and historical annotations corresponding to the above second-dimensional data type, and remove the abnormal historical vehicle data in the historical vehicle data of the second-dimensional data type and the abnormal historical annotations in the historical annotations;
[0183] Step S804: After data cleaning, for the historical vehicle data of each first-dimensional data type, calculate the mean and standard deviation of the above historical vehicle data, and based on the above mean and the above standard deviation, standardize the historical vehicle data belonging to the above first-dimensional data type to obtain standardized data;
[0184] Step S805: After data cleaning, for each of the above historical annotations, match the annotation keywords in the annotation keyword library with the above historical annotations; in response to the number of occurrences of the above annotation keywords in the above historical annotations being greater than or equal to the first preset number threshold, set the initial label vector corresponding to the above historical annotation to 1; in response to the number of occurrences of the above annotation keywords in the above historical annotations being less than the first preset number threshold, set the initial label vector corresponding to the above historical annotation to 0; in response to the situation that the above historical annotation does not match any of the annotation keywords, identify the above historical annotation as an abnormal low-value annotation and eliminate it without generating a corresponding initial label vector;
[0185] Among them, the construction method of the annotation keyword library includes: obtaining the historical vehicle diagnostic reports within the second preset historical time period; comprehensively constructing a historical annotation set from the historical annotations in the above historical vehicle diagnostic reports within the second preset historical time period; using the historical annotations with the number of occurrences greater than the second preset number threshold in the above historical annotation set as keywords to construct an annotation keyword library;
[0186] Step S806: In response to the number of types of annotation keywords existing in the above historical annotations being greater than 1, merge the initial label vectors associated with each annotation keyword to generate a multi-dimensional historical annotation label vector for the above historical annotation, where each dimension of the historical annotation label vector corresponds to a historical keyword; in response to the number of types of annotation keywords existing in the above historical annotations being equal to 1, use the initial label vector associated with the above annotation keyword as the historical annotation label vector of the above historical annotation;
[0187] Step S807: Classify the above historical annotation label vectors according to the second-dimensional data type of the above standardized historical data;
[0188] Step S808: For each second-dimensional data type, for each of the above historical annotation label vectors, calculate the importance score of the above historical annotation label vector using the above historical annotation label vector and its corresponding standardized historical data;
[0189] Step S809, screening the above historical annotation label vectors and their corresponding standardized data belonging to the same class according to the above importance scores, and calculating the corresponding feature weight according to the importance score of each target historical annotation label vector belonging to the same class screened out;
[0190] Step S810, for target historical annotation label vectors belonging to the same second dimensional data type, weighted fusion is performed based on the feature weight corresponding to each target historical annotation label vector to obtain a type feature vector of the second dimensional data type;
[0191] Step S811, fusing type feature vectors of different second-dimensional data types to obtain a fusion result, performing big data model training based on the fusion result, and obtaining a diagnosis model after the training is completed;
[0192] Step S812, in response to the newly acquired vehicle diagnostic report reaching a preset data amount, calculate the data quality score of the newly acquired vehicle diagnostic report, and when the data quality score of the newly acquired vehicle diagnostic report is greater than or equal to a third preset score threshold, trigger incremental learning and update the parameters of the diagnostic model.
[0193] According to the method for constructing a vehicle diagnostic model in an embodiment of the present invention, the acquired vehicle diagnostic report can be analyzed by using a diagnostic model trained with historical vehicle diagnostic reports to predict whether the vehicle has a risk of failure. Through standardization processing and annotation label vector generation, annotations and vehicle data can be fused to avoid the influence of single data on decision-making, improve data quality, ensure the high reliability of data used in the model input and training process, and improve the accuracy of vehicle fault prediction.
[0194] At the same time, data cleaning ensures the overall quality of the data, ensures stable data distribution, and avoids interference of abnormal data on model training.
[0195] Through incremental learning, dynamic adjustment of model parameters, dynamic feature weighting, etc., the model performance is significantly improved, making the model pay more attention to important features, and updating the model parameters according to the newly acquired vehicle diagnostic report, so that the diagnostic model can adapt to new data in a timely manner and maintain good model performance.
[0196] By setting medium-risk probability thresholds and high-risk probability combined with an alarm mechanism, risk warnings are issued based on the failure probability predicted by the model to ensure that vehicle failures are discovered and handled in a timely manner.
[0197] The vehicle diagnostic report processing method is explained below through a specific embodiment.
[0198] like Figure 9As shown in the figure, the vehicle diagnosis report processing method according to the embodiment of the present invention includes the following steps S901 to step S911:
[0199] Step S901, obtain the vehicle diagnosis report provided by the diagnostic tool in real time. The vehicle diagnosis report includes vehicle data and annotations for the vehicle data.
[0200] Step S902, for the vehicle data of each first - dimension data type in the vehicle diagnosis report, determine the mean value and standard deviation corresponding to the vehicle data, where the mean value and the standard deviation are calculated based on the historical vehicle data corresponding to the first - dimension data type.
[0201] Step S903, standardize the vehicle data belonging to the first - dimension data type based on the mean value and the standard deviation to obtain standardized data.
[0202] Step S904, for each annotation, match the annotation keywords in the annotation keyword library with the annotation; in response to the number of occurrences of the annotation keyword in the annotation being greater than or equal to the first preset number threshold, set the initial label vector corresponding to the annotation to 1; in response to the number of occurrences of the annotation keyword in the annotation being less than the first preset number threshold, set the initial label vector corresponding to the annotation to 0.
[0203] Step S905, in response to the number of types of annotation keywords existing in the annotation being greater than 1, merge the initial label vectors associated with each annotation keyword to generate a multi - dimensional annotation label vector for the annotation, where each dimension of the annotation label vector corresponds to a keyword; in response to the number of types of annotation keywords existing in the annotation being equal to 1, use the initial label vector associated with the annotation keyword as the annotation label vector of the annotation.
[0204] Step S906, classify the annotation label vectors according to the second - dimension data type of the standardized data.
[0205] Step S907, for each second - dimension data type, for each of the above - mentioned annotation label vectors, calculate the importance score of the annotation label vector by using the annotation label vector and its corresponding standardized data.
[0206] Step S908, calculate the corresponding feature weight according to the importance score of the selected target annotation label vector.
[0207] Here, in the process of using the diagnostic model to predict the fault probability, there is no need to screen the annotation label vectors, and the corresponding feature weights can be directly calculated according to the importance scores of each annotation label vector.
[0208] Step S909: For the annotation label vectors belonging to the same second - dimension data type, perform weighted fusion based on the feature weights corresponding to each annotation label vector to obtain the type feature vector of the second - dimension data type;
[0209] Step S910: Fuse the type feature vectors to obtain a fusion result, and obtain the fault probability through a diagnostic model according to the fusion result;
[0210] Step S911: Determine the vehicle fault risk according to the comparison result of the fault probability with the medium - risk probability threshold and the high - risk probability threshold, and perform fault warning.
[0211] According to the vehicle diagnostic report method of the embodiment of the present invention, by using a diagnostic model trained with historical vehicle diagnostic reports to analyze the obtained vehicle diagnostic report, it is predicted whether the vehicle has a fault risk. Through standardization processing and annotation label vector generation, annotations and vehicle data can be fused, avoiding the influence of single data on decision - making, improving data quality, ensuring high reliability of the data used in the model input and training process, and improving the accuracy of vehicle fault prediction.
[0212] By setting a medium - risk probability threshold and a high - risk probability combined with an alarm mechanism, risk warning is performed according to the fault probability predicted by the model to ensure timely discovery and handling of vehicle faults.
[0213] Figure 10 It is a schematic diagram of the main modules of the vehicle diagnostic report processing device according to the embodiment of the present invention. As Figure 10 shown, the vehicle diagnostic report processing device 1000 according to the embodiment of the present invention includes: an acquisition module 1001, configured to obtain in real time a vehicle diagnostic report provided by a diagnostic tool, where the vehicle diagnostic report includes vehicle data and annotations for the vehicle data; a processing module 1002, configured to perform standardization processing on the vehicle data to obtain standardized data, and generate annotation label vectors for the annotations; a prediction module 1003, configured to perform fault probability prediction using a preset diagnostic model according to the standardized data and the annotation label vectors, and predict whether the vehicle has a fault risk according to the output result of the diagnostic model, where the diagnostic model is trained with historical vehicle diagnostic reports within a first preset historical time period.
[0214] In an optional embodiment of the present invention, the processing module 1002 is further configured to perform the following operations on the vehicle data of each first - dimension data type in the vehicle diagnostic report: determine the mean and standard deviation corresponding to the vehicle data, where the mean and the standard deviation are calculated based on the historical vehicle data corresponding to the first - dimension data type; based on the mean and the standard deviation, standardize the vehicle data belonging to the first - dimension data type to obtain standardized data.
[0215] In an alternative embodiment of the present invention, the vehicle diagnostic report processing device 1000 further includes: a training module, configured to: obtain historical vehicle diagnostic reports within a first preset historical time period; the historical vehicle diagnostic reports include historical vehicle data and historical annotations for the historical vehicle data; perform normalization processing on the historical vehicle data to obtain normalized historical data, and generate a historical annotation tag vector for the historical annotations; iteratively train a big data model based on the normalized historical data and the historical annotation tag vector to obtain a diagnostic model.
[0216] In an alternative embodiment of the present invention, the training module is further configured to: for historical vehicle data belonging to the same first-dimensional data type, calculate the mean and standard deviation of the historical vehicle data; normalize each of the historical vehicle data based on the mean and standard deviation to obtain normalized historical data.
[0217] In an alternative embodiment of the present invention, the training module is further configured to perform the following for each of the historical annotations: match the annotation keywords in the annotation keyword library with the historical annotation; in response to the number of occurrences of the annotation keyword in the historical annotation being greater than or equal to a first preset number threshold, set the initial tag vector corresponding to the historical annotation to 1; in response to the number of occurrences of the annotation keyword in the historical annotation being less than the first preset number threshold, set the initial tag vector corresponding to the historical annotation to 0; in response to there being more than one type of annotation keyword in the historical annotation, merge the initial tag vectors associated with each annotation keyword to generate a multi-dimensional historical annotation tag vector for the historical annotation, where each dimension of the historical annotation tag vector corresponds to a historical keyword; in response to there being only one type of annotation keyword in the historical annotation, use the initial tag vector associated with the annotation keyword as the historical annotation tag vector for the historical annotation.
[0218] In an alternative embodiment of the present invention, the vehicle diagnostic report processing device 1000 further includes: a construction module, configured to: obtain historical vehicle diagnostic reports within a second preset historical time period; comprehensively analyze the historical annotations in the historical vehicle diagnostic reports within the second preset historical time period to construct a historical annotation set; use the historical annotations that appear more than a second preset number threshold in the historical annotation set as keywords to construct an annotation keyword library.
[0219] In an alternative embodiment of the present invention, the above training module is further configured to: classify the above historical annotation label vectors according to the second-dimensional data type of the above standardized historical data; use the historical annotation label vectors belonging to the same class in the classification result to generate type feature vectors for each second-dimensional data type; and train the big data model according to multiple types of the above type feature vectors.
[0220] In an alternative embodiment of the present invention, the above vehicle diagnostic report processing device 1000 further includes: a calculation module, configured to: for each of the above historical annotation label vectors, determine the feature weight of the above historical annotation label vector by using the above historical annotation label vector and its corresponding standardized historical data.
[0221] The above training module is further configured to: perform weighted fusion on the historical annotation label vectors belonging to the same second-dimensional data type based on the feature weight corresponding to each historical annotation label vector to obtain a feature vector.
[0222] In an alternative embodiment of the present invention, the above calculation module is further configured to: for each of the above historical annotation label vectors, calculate the importance score of the above historical annotation label vector by using the above historical annotation label vector and its corresponding standardized data; and calculate the corresponding feature weight according to the importance score of the above historical annotation label vector.
[0223] In an alternative embodiment of the present invention, the above calculation module is further configured to: screen the above historical annotation label vectors belonging to the same class and their corresponding standardized data according to the above importance score; the step of calculating the corresponding feature weight according to the importance score of the above historical annotation label vector includes: calculating the corresponding feature weight according to the importance score of each target historical annotation label vector belonging to the same class after screening.
[0224] In an alternative embodiment of the present invention, the above vehicle diagnostic report processing device 1000 further includes: a cleaning module, configured to: for the historical vehicle data of each second-dimensional data type and its corresponding historical annotation, mark the abnormal historical vehicle data and abnormal historical annotation to determine the effective data quantity corresponding to the above second-dimensional data type; and analyze whether to start a data cleaning program for the above historical vehicle data and its corresponding historical annotation according to the effective data quantity of each second-dimensional data type and the total data quantity of each second-dimensional data type.
[0225] In an alternative embodiment of the present invention, the above cleaning module is further configured to: calculate the data quality score of all the historical vehicle data and its corresponding historical annotation according to the effective data quantity of each second-dimensional data type and the total data quantity of each second-dimensional data type.
[0226] In response to the above data quality score being less than the first preset score threshold, trigger data cleaning for all the above historical vehicle data and its corresponding historical annotations.
[0227] In an alternative embodiment of the present invention, the above cleaning module is further configured to: for each of the above second-dimensional data types, calculate the data quality score corresponding to the second-dimensional data type according to the number of valid data of the second-dimensional data type and the total number of data of the data type; in response to the data quality score corresponding to the second-dimensional data type being less than the second preset score threshold, trigger data cleaning for the historical vehicle data and historical annotations corresponding to the second-dimensional data type.
[0228] In an alternative embodiment of the present invention, the above vehicle diagnosis report processing device 1000 further includes: an update module, configured to: in response to the newly acquired vehicle diagnosis reports reaching a preset data volume, calculate the data quality score of the newly acquired vehicle diagnosis reports;
[0229] In response to the data quality score of the newly acquired vehicle diagnosis reports being greater than or equal to the third preset score threshold, trigger incremental learning to update the parameters of the above diagnosis model.
[0230] According to the vehicle diagnosis report processing device of the embodiment of the present invention, by analyzing the acquired vehicle diagnosis reports using a diagnosis model trained with historical vehicle diagnosis reports, predicting whether there is a fault risk for the vehicle, through standardization processing and annotation label vector generation, the annotations and vehicle data can be fused, avoiding the influence of single data on decision-making, improving the data quality, ensuring high reliability of the data used in the model input and training process, and improving the accuracy of vehicle fault prediction.
[0231] At the same time, by data cleaning, the overall quality of the data is guaranteed, the data distribution is ensured to be stable, and the interference of abnormal data on model training is avoided.
[0232] By dynamically adjusting model parameters through incremental learning, dynamic feature weighting, etc., the model performance is significantly improved, making the model pay more attention to important features, and updating the model parameters according to the newly acquired vehicle diagnosis reports, enabling the diagnosis model to adapt to new data in a timely manner and maintaining good model performance.
[0233] By setting a medium risk probability threshold and a high risk probability combined with an alarm mechanism, risk early warning is carried out according to the fault probability predicted by the model to ensure timely discovery and handling of vehicle faults.
[0234] Figure 11 An exemplary system architecture 1100 is shown to which the vehicle diagnosis report processing method or vehicle diagnosis report processing device of the embodiment of the present invention can be applied.
[0235] As shown Figure 11 in FIG. 1, the system architecture 1100 may include vehicle terminals 1101, 1102, 1103, a network 1104, and a server 1105. The network 1104 serves as a medium for providing communication links between the vehicle terminals 1101, 1102, 1103 and the server 1105. The network 1104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0236] The vehicle terminals 1101, 1102, 1103 interact with the server 1105 via the network 1104 to receive or send data, etc. Various communication client applications may be installed on the vehicle terminals 1101, 1102, 1103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0237] The server 1105 may be a server that provides various services, such as a back-end management server that supports the vehicle diagnostic reports sent by the vehicle terminals 1101, 1102, 1103. The back-end management server may analyze and process the obtained vehicle diagnostic report data, and feedback the processing results (such as whether there is a fault risk in the vehicle) to the vehicle terminals.
[0238] It should be noted that the vehicle diagnostic report processing method provided by the embodiments of the present invention is generally executed by the server 1105. Correspondingly, the vehicle diagnostic report processing device is generally disposed in the server 1105.
[0239] It should be understood Figure 11 that the numbers of vehicle terminals, networks, and servers in FIG. 1 are merely illustrative. According to actual needs, there may be any number of vehicle terminals, networks, and servers.
[0240] Next, referring to Figure 12 FIG. 2, which shows a schematic structural diagram of a computer system 1200 of a server suitable for implementing the embodiments of the present invention. Figure 12 The server shown in FIG. 2 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0241] As shown Figure 12 in FIG. 2, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 1202 or the programs loaded from the storage section 1208 into the random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the computer system 1200 are also stored. The CPU 1201, ROM 1202, and RAM 1203 are connected to each other via a bus 1204. The input / output (I / O) interface 1205 is also connected to the bus 1204.
[0242] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as required. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 1210 as required so that a computer program read therefrom is installed into the storage section 1208 as required.
[0243] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by a central processing unit (CPU) 1201, the above-described functions defined in the system of the present invention are performed.
[0244] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0245] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0246] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes an acquisition module, a processing module, and a prediction module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the acquisition module can also be described as "a module that obtains in real time a vehicle diagnosis report provided by a diagnosis tool, and the vehicle diagnosis report includes vehicle data and annotations for the vehicle data".
[0247] As another aspect, the present invention further provides a computer-readable medium, which can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: obtaining in real time a vehicle diagnosis report provided by a diagnosis tool, where the vehicle diagnosis report includes vehicle data and annotations for the vehicle data; performing a normalization process on the vehicle data to obtain normalized data, and generating an annotation tag vector for the annotations; according to the normalized data and the annotation tag vector, using a preset diagnosis model to predict a failure probability, and according to the output result of the diagnosis model, predicting whether there is a failure risk for the vehicle, and the diagnosis model is trained by historical vehicle diagnosis reports within a first preset historical time period.
[0248] According to the technical solution of the embodiments of the present invention, by analyzing the obtained vehicle diagnosis report using a diagnosis model trained with historical vehicle diagnosis reports to predict whether there is a failure risk for the vehicle, through normalization processing and annotation tag vector generation, the annotations and vehicle data can be fused, avoiding the influence of single data on decision-making, improving the data quality, ensuring high reliability of the data used in the model input and training process, and improving the accuracy of vehicle failure prediction.
[0249] At the same time, the overall quality of the data is ensured through data cleaning, ensuring the stability of the data distribution and avoiding the interference of abnormal data on model training.
[0250] By incrementally learning to dynamically adjust model parameters, dynamic feature weighting, etc., the model performance is significantly improved, making the model pay more attention to important features, and updating the model parameters according to newly obtained vehicle diagnosis reports, enabling the diagnosis model to adapt to new data in a timely manner and maintaining good model performance.
[0251] By setting a medium risk probability threshold and a high risk probability combined with an alarm mechanism, risk early warning is performed according to the failure probability predicted by the model to ensure timely discovery and handling of vehicle failures.
[0252] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A vehicle diagnostic report processing method, characterized in that, Including: Obtaining in real time a vehicle diagnostic report provided by a diagnostic tool, where the vehicle diagnostic report includes vehicle data and annotations for the vehicle data; Performing normalization processing on the vehicle data to obtain normalized data, and generating an annotation tag vector for the annotations; According to the normalized data and the annotation tag vector, using a preset diagnostic model to predict the fault probability, and according to the output result of the diagnostic model, predicting whether there is a fault risk for the vehicle, where the diagnostic model is trained through historical vehicle diagnostic reports within a first preset historical time period.
2. The vehicle diagnosis report processing method according to claim 1, wherein The performing normalization processing on the vehicle data to obtain normalized data includes: For the vehicle data of each first-dimensional data type in the vehicle diagnostic report, perform the following operations: Determine the mean value and standard deviation corresponding to the vehicle data, where the mean value and the standard deviation are calculated based on historical vehicle data corresponding to the first-dimensional data type; Based on the mean value and the standard deviation, perform normalization on the vehicle data belonging to the first-dimensional data type to obtain normalized data.
3. The vehicle diagnosis report processing method according to claim 1, characterized in that, The method further includes: Obtaining historical vehicle diagnostic reports within a first preset historical time period; the historical vehicle diagnostic reports include historical vehicle data and historical annotations for the historical vehicle data; Performing normalization processing on the historical vehicle data to obtain normalized historical data, and generating a historical annotation tag vector for the historical annotations; Iteratively training a big data model according to the normalized historical data and the historical annotation tag vector to obtain a diagnostic model; Preferably, the performing normalization processing on the historical vehicle data to obtain normalized historical data includes: For historical vehicle data belonging to the same first-dimensional data type, calculate the mean value and standard deviation of the historical vehicle data; Based on the mean value and the standard deviation, perform normalization on each piece of historical vehicle data to obtain normalized historical data.
4. The vehicle diagnosis report processing method according to claim 3, wherein The generating a historical annotation tag vector for the historical annotations includes: For each of the historical annotations, perform: Match the annotation keywords in the annotation keyword library with the historical annotation; In response to the number of times the annotation keyword appears in the historical annotation being greater than or equal to a first preset number threshold, set the initial tag vector corresponding to the historical annotation to 1; in response to the number of times the annotation keyword appears in the historical annotation being less than the first preset number threshold, set the initial tag vector corresponding to the historical annotation to 0; In response to the number of types of annotation keywords existing in the historical annotation being greater than 1, merge the initial tag vectors associated with each annotation keyword to generate a multi-dimensional historical annotation tag vector for the historical annotation, where each dimension of the historical annotation tag vector corresponds to a historical keyword; in response to the number of types of annotation keywords existing in the historical annotation being equal to 1, use the initial tag vector associated with the annotation keyword as the historical annotation tag vector for the historical annotation; Preferably, the method further includes: Obtaining historical vehicle diagnostic reports within a second preset historical time period; Construct a historical annotation set by synthesizing the historical annotations in the historical vehicle diagnostic reports within the second preset historical time period; Use the historical annotations that appear more than the second preset frequency threshold in the historical annotation set as keywords to construct an annotation keyword library.
5. The vehicle diagnosis report processing method according to claim 3, characterized in that The iterative training of the big data model based on the standardized historical data and the historical annotation label vectors to obtain a diagnostic model includes: Classify the historical annotation label vectors according to the second-dimensional data type of the standardized historical data; Use the historical annotation label vectors belonging to the same class in the classification result to generate a type feature vector for each second-dimensional data type; Train the big data model according to multiple types of the type feature vectors; Preferably, the method further includes: In response to the newly acquired vehicle diagnostic reports reaching a preset data volume, calculate the data quality score of the newly acquired vehicle diagnostic reports; In response to the data quality score of the newly acquired vehicle diagnostic reports being greater than or equal to the third preset score threshold, trigger incremental learning to update the parameters of the diagnostic model.
6. The vehicle diagnosis report processing method according to claim 5, wherein The method further includes: For each of the historical annotation label vectors, use the historical annotation label vector and its corresponding standardized historical data to determine the feature weight of the historical annotation label vector; The generating a type feature vector for each second-dimensional data type includes: For the historical annotation label vectors belonging to the same second-dimensional data type, perform weighted fusion based on the feature weights corresponding to each historical annotation label vector to obtain a type feature vector.
7. The vehicle diagnosis report processing method according to claim 6, wherein, The determining the feature weight of the historical annotation label vector includes: For each of the historical annotation label vectors, use the historical annotation label vector and its corresponding standardized data to calculate the importance score of the historical annotation label vector; Calculate the corresponding feature weight according to the importance score of the historical annotation label vector.
8. The vehicle diagnosis report processing method according to claim 7, characterized in that The method further includes: Screen the historical annotation label vectors belonging to the same class and their corresponding standardized data according to the importance score; The calculating the corresponding feature weight according to the importance score of the historical annotation label vector includes: Calculate the corresponding feature weight according to the importance scores of each target historical annotation label vector belonging to the same class after screening.
9. The vehicle diagnosis report processing method according to claim 3, characterized in that, Before the standardization process of the historical vehicle data to obtain standardized historical data, the method further includes: For the historical vehicle data of each second-dimensional data type and its corresponding historical annotation, mark the abnormal historical vehicle data and abnormal historical annotation to determine the effective data quantity corresponding to the second-dimensional data type; Analyze whether to start a data cleaning program for the historical vehicle data and its corresponding historical annotation according to the effective data quantity of each second-dimensional data type and the total data quantity of each second-dimensional data type; Preferably, the analyzing whether to start a data cleaning program for the historical vehicle data and its corresponding historical annotation includes: Calculate the data quality scores of all historical vehicle data and their corresponding historical annotations according to the number of valid data of each of the second - dimension data types and the total number of data of each of the second - dimension data types; in response to the data quality score being less than the first preset score threshold, trigger data cleaning for all the historical vehicle data and their corresponding historical annotations. Alternatively, for each of the second - dimension data types, calculate the data quality score corresponding to the second - dimension data type according to the number of valid data of the second - dimension data type and the total number of data of the second - dimension data type; in response to the data quality score corresponding to the second - dimension data type being less than the second preset score threshold, trigger data cleaning for the historical vehicle data and historical annotations corresponding to the second - dimension data type.
10. A vehicle diagnostic report processing device, characterized in that, Including: An acquisition module, configured to acquire in real time a vehicle diagnostic report provided by a diagnostic tool, where the vehicle diagnostic report includes vehicle data and annotations for the vehicle data; A processing module, configured to perform standardization processing on the vehicle data to obtain standardized data, and generate an annotation label vector for the annotations; A prediction module, configured to predict the fault probability using a preset diagnostic model according to the standardized data and the annotation label vector, and predict whether the vehicle has a fault risk according to the output result of the diagnostic model, where the diagnostic model is trained by historical vehicle diagnostic reports within a first preset historical time period.