Violation anomaly identification method for periodic environmental inspection data of in-use vehicles

The classification prediction model constructed by the random forest model and ROC curve, combined with vehicle information and detection methods, identify violations and cheating behaviors in regular environmental inspections of motor vehicles, and achieve complete abnormal identification of environmental inspection data in use of vehicles, reduce the impact of outliers, adapt to the strict requirements of different regions, and is easy to promote.

CN115828178BActive Publication Date: 2025-08-15SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211447137.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2025-08-15
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively identify and supervise illegal cheating behaviors in regular environmental inspections of motor vehicles, resulting in untrue test results, affecting the effectiveness of environmental inspections and the difficulty of supervision.

Method used

A random forest model and ROC curve are used to construct a classification prediction model, and combined with vehicle information parameters and detection methods, the first inspection passes and unqualified data are identified, abnormal data is identified through changes in the detection station, detection line and vehicle information, and different thresholds are set for localized adjustments.

Benefits of technology

It realizes complete abnormal identification of environmental inspection data of vehicles in use, reduces the impact of outliers, solves multi-classification problems, adapts to the strict requirements of different regions, and does not require additional equipment and complex models, making it easy to promote.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115828178B_ABST
    Figure CN115828178B_ABST
Patent Text Reader

Abstract

In view of the limitations of the existing technology, the present invention proposes a method for identifying violations and anomalies in the periodic environmental protection inspection data of in-use vehicles. Starting from whether the first inspection is qualified or not, the present invention uses different methods to identify abnormal data according to the actual characteristics of different first inspection result data. The identification data range is complete and the method is operational. For the first inspection qualified vehicle data that accounts for the vast majority of the data volume, the random forest model is used to identify anomalies in the environmental protection inspection data of in-use vehicles, which can effectively reduce the influence of outliers, solve multi-classification problems, and has strong generalization ability. The establishment of the anomaly recognition model can be completed by only using the historical data of the environmental protection inspection of in-use vehicles. There is no need to add additional detection equipment or establish a complex physical model, and it is easy to implement. Different thresholds can be set according to the strictness of anomaly recognition in different regions, and local adjustments can be made to facilitate promotion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic environment big data analysis, and in particular to a method for identifying violation anomalies in periodic environmental protection inspection data of in-use vehicles. Background Art

[0002] To effectively reduce motor vehicle exhaust emissions and mitigate the environmental and human health hazards of vehicle exhaust, national requirements for regular environmental inspections of in-use vehicles have been introduced, along with relevant inspection standards. However, due to the varying levels of expertise and capabilities among market-based motor vehicle exhaust inspection agencies in my country, and the lack of a unified and effective regulatory strategy, as well as the fact that these agencies are often profit-driven, some, in pursuit of maximum economic returns, disregard motor vehicle environmental inspection management requirements and violate national inspection standards. These agencies, through temporary modifications to substandard vehicles, cheating during the exhaust inspection process, and tampering with inspection results, allow high-emission vehicles to pass exhaust inspections. Various methods of testing violations continue to emerge, and the problem of testing fraud persists despite repeated crackdowns. The primary reason for this is the lack of a method for identifying fraudulent in-use vehicle periodic environmental inspection data, making it difficult for regulatory authorities to effectively screen out irregular and irregular inspection data.

[0003] With the rapid growth of motor vehicle ownership, the number of vehicles awaiting inspection, inspection agencies, testing stations, and inspection lines will increase in the future. This will also increase the number of daily inspection tasks, inspection process records, and inspection data collection and compilation, posing significant challenges to the supervision of motor vehicle exhaust inspections. Efficiently overseeing environmental inspections for in-use vehicles, preventing irregularities and fraud during the inspection process, ensuring authentic and standardized data collection, and promoting the healthy and orderly development of the motor vehicle exhaust inspection industry are key challenges that urgently need to be addressed.

[0004] The Chinese invention application, published on February 28, 2020, describes a method, device, and storage medium for detecting anomalies in vehicle repair reimbursement claims. This method uses a fusion of isolation forest and PCA models to detect anomalies, such as false claims, to avoid the errors associated with using the isolation forest model alone. The unsupervised machine learning method incorporates a PCA semantic parsing step to enhance the parseability of the detection results. Furthermore, the unsupervised machine learning method avoids manual labeling. However, the aforementioned approach cannot be applied to identifying fraudulent vehicle annual inspections, and therefore, existing technologies still have certain limitations. Summary of the Invention

[0005] To ensure that motor vehicle environmental inspection results can truly reflect the status of motor vehicle exhaust emissions and guarantee the high-quality implementation of regular environmental inspections, the present invention proposes a method for identifying violations and anomalies in the regular environmental inspection data of in-use vehicles. This prediction method can effectively identify abnormal and falsified data, providing a basis for motor vehicle exhaust inspection supervision. The technical solution adopted by the present invention is:

[0006] A method for identifying violation anomalies in periodic environmental inspection data of in-use vehicles includes the following steps:

[0007] S1, obtaining the periodic environmental protection inspection data of the in-use vehicle to be identified;

[0008] S2, determining whether the in-use vehicle periodic environmental protection inspection data is the first inspection qualified vehicle data or the first inspection unqualified vehicle data: if the first inspection qualified vehicle data, then go to step S3; if the first inspection unqualified vehicle data, then go to step S5;

[0009] S3, inputting the vehicle information parameters and detection methods in the in-use vehicle periodic environmental protection inspection data into a random forest model obtained by modeling and tuning a preset training set and validation set, to obtain a passing score probability for the in-use vehicle periodic environmental protection inspection data;

[0010] S4, based on the qualified score probability, using a classification prediction model constructed based on the ROC curve, and using a preset warning level and level threshold to identify abnormal data in the periodic environmental protection inspection data of the in-use vehicle;

[0011] S5, selecting re-inspection records from the periodic environmental protection inspection data of the in-use vehicles, i.e., records of vehicles that were previously unqualified but ultimately passed the inspection;

[0012] S6, identifying abnormal data by checking whether the test method, test station, test line, and vehicle information in the qualified test record and the re-inspection record of the in-use vehicle periodic environmental protection inspection data are changed, provided that the in-use vehicle periodic environmental protection inspection data are for the same vehicle;

[0013] S7: Combine steps S4 and S6 to output the violation anomaly identification result.

[0014] Compared with the existing technology, the present invention starts from whether the first inspection is qualified or not, and uses different methods to identify abnormal data according to the actual characteristics of different first inspection result data. The identification data range is complete and can be implemented. For the first inspection qualified vehicle data that accounts for the vast majority of the data volume, the random forest model is used to identify abnormalities in the environmental protection inspection data of in-use vehicles, which can effectively reduce the impact of outliers, solve multi-classification problems, and has strong generalization ability. The establishment of the abnormality recognition model can be completed by using only the historical data of the environmental protection inspection of in-use vehicles. There is no need to add additional detection equipment or establish a complex physical model, and it is easy to implement. Different thresholds can be set according to the strictness of abnormality recognition in different regions, and localized adjustments can be made to facilitate promotion.

[0015] As a preferred solution, the training set and the validation set are obtained by the following processing method:

[0016] Obtaining sample data from periodic environmental inspections of motor vehicles; deleting duplicate data from the sample data; and removing vehicle data from the sample data with a cumulative mileage greater than 600,000 kilometers or a maximum gross mass greater than 50 tons as outlier data;

[0017] Calibrate the vehicle brand field in the sample data; calculate the vehicle age based on the difference between the vehicle environmental inspection time and the vehicle registration date in the sample data; select vehicle information parameters related to environmental inspection in the sample data, and extract the initial inspection historical data of the vehicle information parameters, inspection methods, and inspection results in the sample data to form an initial inspection data set;

[0018] Integer coding is used to perform numerical conversion on the parameters in the preliminary inspection data set, and the preliminary inspection data set is divided into a training set and a validation set.

[0019] As a preferred solution, the random forest model is obtained in the following way:

[0020] The training set is randomly resampled with replacement to generate a sampling set: the input of the training set is vehicle information parameters and detection method parameters, wherein the vehicle information parameters include license plate color, vehicle type, maximum gross mass, usage nature, fuel type, vehicle brand, engine model, emission standard, engine displacement, cumulative mileage, and vehicle age; the detection methods include dual idle, steady-state operating condition, transient operating condition, simple transient operating condition, free acceleration, and loaded deceleration; the input of the training set is expressed as i=1,2,...,N, where N is the number of training set samples and L is the total number of vehicle information parameters and detection method parameters. The output of the training set is the test category probability, expressed as y i , i=0,1, 0 means unqualified, 1 means qualified;

[0021] The random forest algorithm is used to model the sample set: a feature random selection mechanism is adopted, the maximum number of features of a single decision tree is K=12, k (k<=K) features are randomly selected from all features, the best segmentation attributes are selected as nodes to establish a decision tree, the size of K remains unchanged during the growth of the decision tree, and the training results are output;

[0022] According to the training results, the area under the ROC curve (AUC) value is used as the evaluation indicator. Using the validation set, the Bayesian parameter tuning method is used to tune the hyperparameters of the random forest model to determine the optimal hyperparameters.

[0023] Furthermore, the random forest algorithm is used to model the sample set in the following way:

[0024] The class weight is calculated using the training sample size: that is, the more samples of a certain type, the lower the weight, and the fewer samples, the higher the weight; the weight of the unqualified class: Eligible Category Weights: Where N is the total number of samples in the training set, n is the number of qualified samples in the training set, and Nn is the number of unqualified samples;

[0025] Randomly select k (k <= K) features, and use the obtained sampling subset as the root node. Calculate the Gini coefficient of the feature for the sampling subset: among all possible features and all possible split points, select the feature with the smallest weighted Gini index and the corresponding split point as the optimal feature and optimal split point. Based on the optimal feature and optimal split point, generate leaf nodes and split the subsets to ensure that each subset is correctly assigned to a leaf node.

[0026] The above steps are recursively called for leaf nodes until one of the following conditions is met: the weighted Gini index of the training set is less than the predetermined threshold; or there are no more features; or the number of samples in the node is less than the predetermined threshold. Once the conditions are met, this round of training is completed and a decision tree is generated.

[0027] Repeat the above steps to complete the construction of all decision trees, and perform arithmetic average calculation to obtain the output result of the input on the entire random forest.

[0028] Furthermore, the Gini coefficient of the sampled subset of feature pairs is calculated as follows:

[0029]

[0030] X represents the training set of the node, which has two classes: qualified and unqualified; N i is the training subset of X that belongs to the i-th category;

[0031] Note that the training set X can be divided into X1, X2, ..., X according to feature A. k, k parts, then under the condition of feature A, the weighted Gini index of set X is;

[0032]

[0033] Furthermore, determining the hyperparameters includes the following process:

[0034] The area under the ROC curve (AUC) value is used as the evaluation index to construct a replacement function for random forest;

[0035] Define the six hyperparameters of the model: the number of decision trees n_estimators, the maximum depth of the decision tree max_depth, the minimum number of samples in a leaf node min_samples_leaf, the condition in_samples_split that limits the further division of the subtree, the maximum number of features max_features, and the division standard criterion when splitting the node, to determine the hyperparameter search space of the random forest;

[0036] Repeat the following steps until the maximum number of iterations is reached. After reaching the maximum number of iterations, the hyperparameter combination with the largest AUC value is selected as the optimal hyperparameter of the model:

[0037] Using the validation set, a random hyperparameter combination is applied to the random forest model to obtain the AUC value evaluation index score; considering the prior knowledge of the previous hyperparameter combination, the Bayesian theorem is used to estimate the posterior distribution of the surrogate function, and then the next sampled hyperparameter combination is selected based on the distribution.

[0038] As a preferred solution, the classification prediction model includes four levels of classification prediction models constructed with recall rates of 80%, 85%, 90% and 95% as thresholds respectively;

[0039] In step S4, according to the severity of the judgment of the vehicle environmental protection test results, four classification prediction models are combined to form a five-level warning of the possibility of failure of the environmental protection test of the in-use vehicle:

[0040] If only 80% of the classification prediction models predict failure, the warning level is 1; if a higher-level classification prediction model predicts failure, the warning level is based on the higher level; if no classification prediction model predicts failure, the warning level is 0; the higher the sample warning level, the higher the probability that the test result will fail.

[0041] As a preferred solution, in step S6, abnormal data is identified by:

[0042] If the inspection line, inspection station or inspection method changes in the inspection record of the same vehicle, it will be considered as abnormal data;

[0043] Compare multiple test records of the same vehicle to see if its engine rated power has changed. If so, it is determined to be abnormal data.

[0044] The present invention also provides the following:

[0045] A storage medium stores a computer program, which, when executed by a processor, implements the steps of the aforementioned method for identifying violations and anomalies in periodic environmental protection inspection data of in-use vehicles.

[0046] A computer device includes a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, the steps of the aforementioned method for identifying violations and anomalies in periodic environmental protection inspection data of in-use vehicles are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 A flowchart of a method for identifying violations and anomalies in periodic environmental inspection data of in-use vehicles provided by an embodiment of the present invention;

[0048] Figure 2 A schematic diagram of a method for constructing and optimizing a random forest model and a method for constructing a classification prediction model used in an embodiment of the present invention;

[0049] Figure 3 An example of an ROC curve drawn when building a classification prediction model according to an embodiment of the present invention;

[0050] Figure 4 This is an example of an early warning regarding qualified data of the first inspection result according to an embodiment of the present invention;

[0051] Figure 5 This is an example of multiple detection records in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;

[0053] It should be clear that the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the embodiments of the present application.

[0054] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the embodiments of the present application. The singular forms "a," "the," and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0055] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.

[0056] In addition, in the description of this application, unless otherwise specified, "plurality" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship. The present invention is further described below with reference to the accompanying drawings and examples.

[0057] In order to solve the limitations of the prior art, this embodiment provides a technical solution, which will be further described below in conjunction with the accompanying drawings and embodiments.

[0058] Example 1

[0059] Please refer to Figure 1 A method for identifying violation anomalies in periodic environmental inspection data of in-use vehicles includes the following steps:

[0060] S1, obtaining the periodic environmental protection inspection data of the in-use vehicle to be identified;

[0061] S2, determining whether the in-use vehicle periodic environmental protection inspection data is the first inspection qualified vehicle data or the first inspection unqualified vehicle data: if the first inspection qualified vehicle data, then go to step S3; if the first inspection unqualified vehicle data, then go to step S5;

[0062] S3, inputting the vehicle information parameters and detection methods in the in-use vehicle periodic environmental protection inspection data into a random forest model obtained by modeling and tuning a preset training set and validation set, to obtain a passing score probability for the in-use vehicle periodic environmental protection inspection data;

[0063] S4, based on the qualified score probability, using a classification prediction model constructed based on a receiver operating characteristic (ROC) curve, and using a preset warning level and level threshold to identify abnormal data in the periodic environmental protection inspection data of the in-use vehicle;

[0064] S5, selecting re-inspection records from the periodic environmental protection inspection data of the in-use vehicles, i.e., records of vehicles that were previously unqualified but ultimately passed the inspection;

[0065] S6, identifying abnormal data by checking whether the test method, test station, test line, and vehicle information in the qualified test record and the re-inspection record of the in-use vehicle periodic environmental protection inspection data are changed, provided that the in-use vehicle periodic environmental protection inspection data are for the same vehicle;

[0066] S7: Combine steps S4 and S6 to output the violation anomaly identification result.

[0067] Compared with the existing technology, the present invention starts from whether the first inspection is qualified or not, and uses different methods to identify abnormal data according to the actual characteristics of different first inspection result data. The identification data range is complete and can be implemented. For the first inspection qualified vehicle data that accounts for the vast majority of the data volume, the random forest model is used to identify abnormalities in the environmental protection inspection data of in-use vehicles, which can effectively reduce the impact of outliers, solve multi-classification problems, and has strong generalization ability. The establishment of the abnormality recognition model can be completed by using only the historical data of the environmental protection inspection of in-use vehicles. There is no need to add additional detection equipment or establish a complex physical model, and it is easy to implement. Different thresholds can be set according to the strictness of abnormality recognition in different regions, and localized adjustments can be made to facilitate promotion.

[0068] As a preferred embodiment, the training set and the validation set are obtained by the following processing method:

[0069] Obtaining sample data from periodic environmental inspections of motor vehicles; deleting duplicate data from the sample data; and removing vehicle data from the sample data with a cumulative mileage greater than 600,000 kilometers or a maximum gross mass greater than 50 tons as outlier data;

[0070] Calibrate the vehicle brand field in the sample data; calculate the vehicle age based on the difference between the vehicle environmental inspection time and the vehicle registration date in the sample data; select vehicle information parameters related to environmental inspection in the sample data, and extract the initial inspection historical data of the vehicle information parameters, inspection methods, and inspection results in the sample data to form an initial inspection data set;

[0071] Integer coding is used to perform numerical conversion on the parameters in the preliminary inspection data set, and the preliminary inspection data set is divided into a training set and a validation set.

[0072] Specifically, during the above process, the vehicle brand field in the environmental inspection data can be calibrated with reference to the vehicle brand on the Autohome website. The vehicle age is calculated by taking the difference between the environmental inspection time and the vehicle registration date: vehicle age = [(vehicle environmental inspection time - vehicle registration date) / 365] + 1.

[0073] During the laboratory phase, this example obtained over 1.7 million pieces of initial environmental inspection data from a specific city's motor vehicles. License plate color, vehicle type, maximum gross mass, usage, fuel type, vehicle brand, engine model, emission standard, engine displacement, cumulative mileage, vehicle age, inspection method, and final determination result data were selected as analysis targets. The final determination result was used as output, and the remaining 12 variables were used as input to construct a dataset and perform preprocessing. After data preprocessing, a dataset with a sample size of approximately 1.2 million pieces of data was obtained. Each parameter in the dataset was converted using integer encoding. Furthermore, to ensure a balance between the "pass" and "fail" environmental inspection data across the training, validation, and test sets, the entire initial inspection dataset was first divided into a pass dataset and a fail dataset. These two datasets were then randomly assigned to the training, validation, and test sets in an 8:1:1 ratio. The datasets were then merged to form a final training set of 960,000 pieces, a validation set of 120,000 pieces, and a test set of 120,000 pieces.

[0074] As a preferred embodiment, please refer to Figure 2 , the random forest model is obtained in the following way:

[0075] The training set is randomly resampled with replacement to generate a sampling set: the input of the training set is vehicle information parameters and detection method parameters, wherein the vehicle information parameters include license plate color, vehicle type, maximum gross mass, usage nature, fuel type, vehicle brand, engine model, emission standard, engine displacement, cumulative mileage, and vehicle age; the detection methods include dual idle, steady-state operating condition, transient operating condition, simple transient operating condition, free acceleration, and loaded deceleration; the input of the training set is expressed as i=1,2,...,N, where N is the number of training set samples and L is the total number of vehicle information parameters and detection method parameters. The output of the training set is the test category probability, expressed as y i , i=0,1, 0 means unqualified, 1 means qualified;

[0076] The random forest algorithm is used to model the sample set: a feature random selection mechanism is adopted, the maximum number of features of a single decision tree is K=12, k (k<=K) features are randomly selected from all features, the best segmentation attributes are selected as nodes to establish a decision tree, the size of K remains unchanged during the growth of the decision tree, and the training results are output;

[0077] According to the training results, the area under the ROC curve (AUC) value is used as the evaluation indicator. Using the validation set, the Bayesian parameter tuning method is used to tune the hyperparameters of the random forest model to determine the optimal hyperparameters.

[0078] Specifically, in the laboratory stage, this embodiment randomly reuses a sub-training data set that constitutes a single decision tree in the random forest, and uses the same sampling method for 960,000 times to obtain 960,000 different sampling sets.

[0079] Furthermore, the random forest algorithm is used to model the sample set in the following manner:

[0080] The class weight is calculated using the training sample size: that is, the more samples of a certain type, the lower the weight, and the fewer samples, the higher the weight; the weight of the unqualified class: Eligible Category Weights: Where N is the total number of samples in the training set, n is the number of qualified samples in the training set, and Nn is the number of unqualified samples;

[0081] Randomly select k (k <= K) features, and use the obtained sampling subset as the root node. Calculate the Gini coefficient of the feature for the sampling subset: among all possible features and all possible split points, select the feature with the smallest weighted Gini index and the corresponding split point as the optimal feature and optimal split point. Based on the optimal feature and optimal split point, generate leaf nodes and split the subsets to ensure that each subset is correctly assigned to a leaf node.

[0082] The above steps are recursively called for leaf nodes until one of the following conditions is met: the weighted Gini index of the training set is less than the predetermined threshold; or there are no more features; or the number of samples in the node is less than the predetermined threshold. Once the conditions are met, this round of training is completed and a decision tree is generated.

[0083] Repeat the above steps to complete the construction of all decision trees, and perform arithmetic average calculation to obtain the output result of the input on the entire random forest.

[0084] Specifically, in the laboratory stage, this embodiment trained 960,000 decision trees from 960,000 sample sets, output 960,000 training results, and calculated the arithmetic average to obtain the output result of the sample input on the entire random forest.

[0085] Furthermore, the Gini coefficient of the sampled subset of feature pairs is calculated as follows:

[0086]

[0087] X represents the training set of the node, which has two classes: qualified and unqualified; N i is the training subset of X that belongs to the i-th category;

[0088] Note that the training set X can be divided into X1, X2, ..., X according to feature A. k , k parts, then under the condition of feature A, the weighted Gini index of set X is;

[0089]

[0090] Furthermore, determining the hyperparameters includes the following process:

[0091] The area under the ROC curve (AUC) value is used as the evaluation index to construct a replacement function for random forest;

[0092] Define the six hyperparameters of the model: the number of decision trees n_estimators, the maximum depth of the decision tree max_depth, the minimum number of samples in a leaf node min_samples_leaf, the condition in_samples_split that limits the further division of the subtree, the maximum number of features max_features, and the division standard criterion when splitting the node, to determine the hyperparameter search space of the random forest;

[0093]

[0094]

[0095] Repeat the following steps until the maximum number of iterations is reached. After reaching the maximum number of iterations, the hyperparameter combination with the largest AUC value is selected as the optimal hyperparameter of the model:

[0096] Using the validation set, a random hyperparameter combination is applied to the random forest model to obtain the AUC value evaluation index score; considering the prior knowledge of the previous hyperparameter combination, the Bayesian theorem is used to estimate the posterior distribution of the surrogate function, and then the next sampled hyperparameter combination is selected based on the distribution.

[0097] max_depth min_samples_leaf min_samples_split n_estimators max_features criterion 19 21 114 160 4 gini

[0098] As a preferred embodiment, the classification prediction model includes four levels of classification prediction models constructed with recall rates of 80%, 85%, 90% and 95% as thresholds respectively;

[0099] In step S4, according to the severity of the judgment of the vehicle environmental protection test results, four classification prediction models are combined to form a five-level warning of the possibility of failure of the environmental protection test of the in-use vehicle:

[0100] If only 80% of the classification prediction models predict failure, the warning level is 1; if a higher-level classification prediction model predicts failure, the warning level is based on the higher level; if no classification prediction model predicts failure, the warning level is 0; the higher the sample warning level, the higher the probability of failure in the test result, and this part of the data should be paid special attention to.

[0101] Specifically, in the laboratory stage, this embodiment constructs a classification prediction model in the following manner to form different warning levels:

[0102] The vehicle information parameters and detection method of the test set are used as input, and the predicted category probability of the random forest model is output, that is, the qualified score probability; the predicted probability (score) of the random forest model of the test set is removed from duplicate values and sorted in descending order; please refer to Figure 3 , using the predicted probability as the threshold to calculate the recall rate (the proportion of qualified samples predicted to be qualified) and specificity (the proportion of unqualified samples predicted to be unqualified), and use the specificity as the horizontal axis and the recall rate as the vertical axis to draw the ROC curve;

[0103] Considering the practical difficulty of regular environmental testing and supervision for in-use vehicles, four classification prediction models were constructed based on the strictness of judging abnormal data violations, selecting different thresholds corresponding to recall rates of 80%, 85%, 90%, and 95%. Specifically, Model 1 predicts 80% of qualified samples as qualified, Model 2 predicts 85% of qualified samples as qualified, Model 3 predicts 90% of qualified samples as qualified, and Model 4 predicts 95% of qualified samples as qualified.

[0104]

[0105] For data that fail the initial inspection, it is considered that there is no abnormal violation and no warning will be processed. For data that pass the initial inspection, according to the severity of the judgment of the vehicle environmental protection inspection results, 4 classification prediction models are combined to form a 5-level warning for the possibility of failure of the environmental protection inspection of in-use vehicles. If only Model 1 predicts failure, the warning level is 1; if a higher-level classification prediction model predicts failure, the warning level is based on the higher level, that is, Model 2 and Model 3 predict failure at the same time, the warning level is 3; if no classification prediction model predicts failure, the warning level is 0; Figure 4 Taking the results as an example, the higher the warning level, the higher the probability of unqualified test results, and this part of the data should be paid special attention to.

[0106] As a preferred embodiment, in step S6, abnormal data is identified by:

[0107] If the inspection line, inspection station or inspection method changes in the inspection record of the same vehicle, it will be considered as abnormal data;

[0108] Compare multiple test records of the same vehicle to see if its engine rated power has changed. If so, it is determined to be abnormal data.

[0109] Specifically, if a vehicle fails the initial inspection, it will normally need to undergo continuous maintenance and re-inspections until it passes the exhaust gas test before it can pass the annual inspection. Therefore, for vehicles that failed the initial inspection, the environmental inspection data is filtered to identify those that have undergone re-inspections, that is, those that failed the inspection and ultimately passed the inspection.

[0110] Taking the regular environmental protection inspection data of in-use vehicles in a certain city as an example, we screened out 143,416 records of vehicles that had undergone and passed re-inspections, including 118,010 vehicles;

[0111] We screened out vehicles that failed the initial inspection, along with their corresponding re-inspection records, from the regular environmental inspection data for in-use vehicles. We performed data pre-processing, such as deleting duplicate data. We used the vehicle identification number to confirm that the vehicle was the same before and after the re-inspection. We then conducted a statistical analysis of changes to information such as the inspection method, inspection station name, engine power, and engine model. We identified 2,633 records with changes, representing 1.84% of the total vehicle population.

[0112] For multiple inspection records of the same vehicle (same VIN), please refer to Figure 5, comparing the detection methods and detection station information, it was found that the number of records in which only the detection station information was changed was the largest, with a total of 1,876 records, accounting for 71.25%; the number of records in which only the detection method was changed was also relatively large, with a total of 261 records, accounting for 9.91%. If the records of changes in all detection stations are considered, the total number reaches 2,081, accounting for as high as 79.04%, involving 2,064 vehicles. If the records of changes in all detection methods are considered, the total number reaches 362, accounting for 13.75%, of which 133 records (involving 107 vehicles) were changed from the loaded deceleration method to the opaque smoke method. Since the opaque smoke method does not have NOX detection requirements, it is a more relaxed detection method than the loaded deceleration method, and vehicles are more likely to pass. Therefore, this type of data on changes in detection lines, detection stations or detection methods deserves special attention;

[0113] By comparing multiple test records of the same vehicle, we can see if its engine rated power has changed. If it has, it is considered abnormal data. Through practice, we found that under the premise of the engine model not changing, we found 42 records of engine power changes, of which 6 records belong to the loading and deceleration method, and 2 records have the phenomenon of explicitly writing down the engine power. The specific situation is as follows:

[0114] ① Vehicle ***43055, with yellow license plate color, had its engine power registered as 247kW at 2:22:00 PM on July 10, 2020, but only registered as 88kW at 4:31:00 PM on July 10, 2020;

[0115] ② For vehicle ***LW139 with a blue license plate, the engine power was registered as 68kW at 9:42:00 on May 5, 2019, but was only registered as 22kW at 17:19:00 on July 1, 2020.

[0116] Example 2

[0117] A storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for identifying violations and anomalies in periodic environmental protection inspection data of in-use vehicles in embodiment 1.

[0118] Example 3

[0119] A computer device includes a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, the steps of the method for identifying violations and anomalies in periodic environmental protection inspection data of in-use vehicles in Example 1 are implemented.

[0120] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A method for identifying violations and anomalies in periodic environmental inspection data of in-use vehicles, characterized by: The following steps are involved: S1, obtaining the periodic environmental protection inspection data of the in-use vehicle to be identified; S2, determining whether the in-use vehicle periodic environmental protection inspection data is the first inspection qualified vehicle data or the first inspection unqualified vehicle data: if the first inspection qualified vehicle data, then go to step S3; if the first inspection unqualified vehicle data, then go to step S5; S3, inputting the vehicle information parameters and detection methods in the in-use vehicle periodic environmental protection inspection data into a random forest model obtained by modeling and tuning a preset training set and validation set, to obtain a passing score probability for the in-use vehicle periodic environmental protection inspection data; The random forest algorithm is used to model the sample set in the following way: The class weight is calculated using the training sample size: that is, the more samples of a certain type, the lower the weight, and the fewer samples, the higher the weight; the weight of the unqualified class: Eligible Category Weights: Where N is the total number of samples in the training set, n is the number of qualified samples in the training set, and Nn is the number of unqualified samples; Randomly select k features, k <= K, and use the obtained sampling subset as the root node. Calculate the Gini coefficient of the feature for the sampling subset: among all possible features and all possible split points, select the feature with the smallest weighted Gini index and the corresponding split point as the optimal feature and optimal split point. Based on the optimal feature and optimal split point, generate leaf nodes and split the subsets to ensure that each subset is correctly assigned to a leaf node. The above steps are recursively called for leaf nodes until one of the following conditions is met: the weighted Gini index of the training set is less than the predetermined threshold; or there are no more features; or the number of samples in the node is less than the predetermined threshold. Once the conditions are met, this round of training is completed and a decision tree is generated. Repeat the above steps to complete the construction of all decision trees, and perform arithmetic average calculation to obtain the output result of the input on the entire random forest; The Gini coefficient of a sampled subset of feature pairs is calculated as follows: X represents the training set of the node, which has two classes: qualified and unqualified; N i is the training subset of X that belongs to the i-th category; Note that the training set X can be divided into X1, X2, ..., X according to feature A. k , k parts, then under the condition of feature A, the weighted Gini index of set X is; S4, based on the qualified score probability, using a classification prediction model constructed based on the ROC curve, and using a preset warning level and level threshold to identify abnormal data in the periodic environmental protection inspection data of the in-use vehicle; S5, selecting re-inspection records from the periodic environmental protection inspection data of the in-use vehicles, i.e., records of vehicles that were previously unqualified but ultimately passed the inspection; S6, identifying abnormal data by checking whether the test method, test station, test line, and vehicle information in the qualified test record and the re-inspection record of the in-use vehicle periodic environmental protection inspection data are changed, provided that the in-use vehicle periodic environmental protection inspection data are for the same vehicle; S7: Combine steps S4 and S6 to output the violation anomaly identification result.

2. The method for identifying violations and anomalies in periodic environmental inspection data of in-use vehicles according to claim 1 is characterized in that: The training set and validation set are obtained by the following processing method: Obtaining sample data from periodic environmental inspections of motor vehicles; deleting duplicate data from the sample data; and removing vehicle data from the sample data with a cumulative mileage greater than 600,000 kilometers or a maximum gross mass greater than 50 tons as outlier data; Calibrate the vehicle brand field in the sample data; calculate the vehicle age based on the difference between the vehicle environmental inspection time and the vehicle registration date in the sample data; Selecting vehicle information parameters related to environmental inspection in the sample data, extracting the initial inspection historical data of the vehicle information parameters, inspection methods, and inspection results in the sample data to form an initial inspection data set; Integer coding is used to perform numerical conversion on the parameters in the preliminary inspection data set, and the preliminary inspection data set is divided into a training set and a validation set.

3. The method for identifying violations and anomalies in periodic environmental inspection data of in-use vehicles according to claim 1 is characterized in that: The random forest model is obtained in the following way: The training set is subjected to random resampling with replacement to generate a sampling set: the input of the training set is vehicle information parameters and detection method parameters, wherein the vehicle information parameters include license plate color, vehicle type, maximum gross mass, usage nature, fuel type, vehicle brand, engine model, emission standard, engine displacement, cumulative mileage, and vehicle age; the detection methods include dual idle, steady-state operating condition, transient operating condition, simple transient operating condition, free acceleration, and loaded deceleration; The input of the training set is represented as i=1,2,...,N, where N is the number of training set samples, L is the total number of vehicle information parameters and detection method parameters, x represents the training set samples; the output of the training set is the test category probability, expressed as y i , i=0,1, 0 means unqualified, 1 means qualified; The random forest algorithm is used to model the sample set: a feature random selection mechanism is adopted, the maximum number of features of a single decision tree is K=12, k (k<=K) features are randomly selected from all features, the best segmentation attributes are selected as nodes to establish a decision tree, the size of K remains unchanged during the growth of the decision tree, and the training results are output; According to the training results, the area under the ROC curve (AUC) value is used as the evaluation indicator. Using the validation set, the Bayesian parameter tuning method is used to tune the hyperparameters of the random forest model to determine the optimal hyperparameters.

4. The method for identifying violations and anomalies in periodic environmental inspection data of in-use vehicles according to claim 3 is characterized in that: Determining hyperparameters involves the following process: The area under the ROC curve (AUC) value is used as the evaluation index to construct a replacement function for random forest; Define the six hyperparameters of the model: the number of decision trees n_estimators, the maximum depth of the decision tree max_depth, the minimum number of samples in a leaf node min_samples_leaf, the condition in_samples_split that limits the further division of the subtree, the maximum number of features max_features, and the division standard criterion when splitting the node, to determine the hyperparameter search space of the random forest; Repeat the following steps until the maximum number of iterations is reached. After reaching the maximum number of iterations, the hyperparameter combination with the largest AUC value is selected as the optimal hyperparameter of the model: Using the validation set, a random hyperparameter combination is applied to the random forest model to obtain the AUC value evaluation index score; considering the prior knowledge of the previous hyperparameter combination, the Bayesian theorem is used to estimate the posterior distribution of the surrogate function, and then the next sampled hyperparameter combination is selected based on the distribution.

5. The method for identifying violations and anomalies in periodic environmental inspection data of in-use vehicles according to claim 1 is characterized in that: The classification prediction model includes four levels of classification prediction models constructed with recall rates of 80%, 85%, 90% and 95% as thresholds respectively; In step S4, according to the severity of the judgment of the vehicle environmental protection test results, four classification prediction models are combined to form a five-level warning of the possibility of failure of the environmental protection test of the in-use vehicle: If only 80% of the classification prediction models predict failure, the warning level is 1; if a higher-level classification prediction model predicts failure, the warning level is based on the higher level; If no classification prediction model predicts that the sample is unqualified, the warning level is 0; the higher the sample warning level, the higher the probability that the test result is unqualified.

6. The method for identifying violations and anomalies in periodic environmental inspection data of in-use vehicles according to claim 1 is characterized in that: In step S6, abnormal data is identified by: If the inspection line, inspection station or inspection method changes in the inspection record of the same vehicle, it will be considered as abnormal data; Compare multiple test records of the same vehicle to see if its engine rated power has changed. If so, it is determined to be abnormal data.

7. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method for identifying violations and anomalies in the periodic environmental protection inspection data of in-use vehicles as described in any one of claims 1 to 6 are implemented.

8. A computer device, characterized in that: The method comprises a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor, wherein when the computer program is executed by the processor, the steps of the method for identifying violations and anomalies in the periodic environmental inspection data of in-use vehicles as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Coal type identification method based on random forest

    CN111797883A

  • Credit card default fraud identification method based on RF-DBSCAN algorithm

    CN112001788A