Method and system for intelligently analyzing authenticity of motor vehicle exhaust emission inspection data
By receiving and standardizing multi-source data, constructing a correlation graph of testing institutions, and using LSTM and spiking neural networks to analyze vehicle emission trajectories, the problems of data silos and collaborative anomaly identification in motor vehicle exhaust emission testing are solved, and more accurate analysis of the authenticity of test data is achieved.
Patent Information
- Application Number
- CN202511056899.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing motor vehicle exhaust emission testing suffers from data silos, making it difficult to conduct cross-institutional and cross-period correlation analysis. It also fails to effectively identify data deviations and collaborative abnormal behaviors among testing institutions, resulting in a lack of a global perspective in analyzing the authenticity of data.
By receiving multi-source data, including motor vehicle exhaust emission test data, vehicle maintenance data, and test equipment calibration logs, and after standardization processing, a correlation graph of testing institutions is constructed, and an isolated forest model is used to identify abnormal subgraphs. Combined with LSTM model and spiking neural network, vehicle emission trajectories are analyzed to generate early warning information on the authenticity of test data.
It achieves integrated analysis of data across the entire domain, enabling the identification of group data deviations in testing agencies and abnormal vehicle emission trajectories, thereby improving the accuracy of identifying genuine and fake testing data and the long-term effectiveness of the system.
Smart Images

Figure CN120875908A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of motor vehicle emission supervision technology, specifically involving an intelligent analysis system for the authenticity of exhaust gas testing based on regional full-volume inspection data. By integrating periodic motor vehicle inspection data within a province or city, it uses artificial intelligence technology to identify systemic authenticity issues such as data tampering and equipment calibration anomalies by testing agencies. Background Technology
[0002] Motor vehicle emissions testing is a crucial means of environmental protection and air quality control, playing a significant role in reducing vehicle pollutant emissions and improving air quality. However, existing motor vehicle emissions testing methods suffer from data inaccuracies, primarily manifested in the following aspects: The problem of data silos is severe. Individual test data is analyzed in isolation, failing to correlate data across institutions or time periods, resulting in a lack of a holistic perspective in analyzing the authenticity of test data. Traditional exhaust emission test data analysis methods primarily evaluate single test results, failing to effectively utilize historical data and data from multiple institutions for correlation analysis, making it difficult to detect data falsification.
[0003] Regional anomalies are often concealed, making it difficult for traditional methods to detect collective data deviations among testing institutions. When multiple testing institutions employ similar methods to verify data authenticity, the lack of comprehensive data comparison and analysis capabilities often prevents traditional regulatory approaches from identifying such coordinated anomalies, leading to regulatory blind spots.
[0004] Therefore, an intelligent analysis system is needed that can integrate data from across the entire domain and identify regional anomalies to effectively solve the problem of the authenticity of motor vehicle exhaust emission test data. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for intelligent analysis of the authenticity of motor vehicle exhaust emission test data. Through multi-source data fusion, intelligent analysis models and dynamic knowledge bases, it can effectively identify the authenticity of motor vehicle exhaust emission test data and provide decision support for environmental regulatory departments.
[0006] To achieve the above objectives, this invention provides a method for intelligent analysis of the authenticity of motor vehicle exhaust emission test data, comprising: The system receives vehicle exhaust emission test data, combines it with vehicle maintenance data and test equipment calibration logs to obtain multi-source test data, and performs standardization processing on the multi-source test data to generate standardized multi-source test data. The vehicle exhaust emission test data includes the vehicle VIN code, test time, pollutant concentration value, and test equipment serial number. Based on the standardized multi-source detection data, a detection agency association graph is constructed, and an isolated forest model is used to detect abnormal subgraphs in the detection agency association graph to obtain the agency collaborative anomaly identification results. Based on the trained first neural network model, the target vehicle information from the standardized multi-source detection data is input into the first neural network model to generate a theoretical emission curve. The time warping distance between the measured data of the target vehicle and the theoretical emission curve is calculated to obtain the vehicle trajectory anomaly analysis results. The first neural network model includes an LSTM model and / or a spiking neural network. Based on the results of the collaborative anomaly identification of the aforementioned institutions and the results of the anomaly analysis of the vehicle trajectory, a warning message for the authenticity of the detection data is generated.
[0007] Optionally, based on the standardized multi-source detection data, a detection agency association graph is constructed, and an isolated forest model is used to detect abnormal subgraphs in the detection agency association graph to obtain agency collaborative anomaly identification results, including: The standardized multi-source detection data is compressed using a deep autoencoder to generate a low-dimensional representation vector of the detection mechanism; Comparative learning is performed on the low-dimensional representation vector, and data from the same vehicle at different detection agencies are used as positive sample pairs for training to obtain the normal detection mode vector; Based on the low-dimensional representation vector and the normal detection mode vector, calculate the Mahalanobis distance between each detection institution and its historical detection data, and screen potential abnormal detection institutions. A dynamic prediction model for the emission characteristics of the target vehicle is constructed using variational Bayesian inference, and the emission characteristic distribution of the target vehicle is generated. Based on the emission characteristic distribution and the potential anomaly detection agencies, calculate the anomaly score of newly added emission data in the standardized multi-source detection data, and identify the anomaly detection agency group.
[0008] Optionally, the step of performing comparative learning on the low-dimensional representation vector, using data from the same vehicle at different detection agencies as positive sample pairs for training, to obtain a normal detection pattern vector, includes: Based on the emission data of the target vehicle from different testing institutions, data similarity features are extracted to construct a comparative learning sample set; Gaussian noise and a random mask are added to the contrastive learning sample set to generate an enhanced training dataset; The time-series features of the augmented training dataset are processed using a self-attention mechanism to obtain sequence encodings. Based on the sequence encoding, the contrast loss of positive and negative pairs of samples is calculated, and the parameters of the isolated forest model are updated. Based on the parameters of the isolated forest model, the low-dimensional representation vector is mapped to the normal detection pattern vector space to obtain the normal detection pattern vector.
[0009] Optionally, the step of constructing a dynamic prediction model for the emission characteristics of the target vehicle using variational Bayesian inference, and generating the emission characteristic distribution of the target vehicle, includes: Collect historical emission data of the target vehicle and construct an emission feature training set; Based on the emission feature training set, emission concentration trend change features and periodic fluctuation features are extracted; A dynamic prediction model of vehicle maintenance records and emission characteristic changes was established using variational Bayesian inference, and the emission characteristic abrupt change points were analyzed. The anomaly detection threshold is adjusted based on the emission characteristic mutation points to generate an adaptive prediction interval, which is used to characterize the emission characteristic distribution of the target vehicle.
[0010] Optionally, based on the trained first neural network model, the target vehicle information from the standardized multi-source detection data is input into the first neural network model to generate a theoretical emission curve, and the time warping distance between the measured data of the target vehicle and the theoretical emission curve is calculated to obtain the vehicle trajectory anomaly analysis result; wherein, the first neural network model includes an LSTM model and / or a spiking neural network; the first neural network model using a spiking neural network includes: A spiking neural network is constructed to convert the standardized multi-source detection data into a pulse sequence; Calculate the number of causal segments in the pulse sequence under different spiking neural network parameter configurations, and select the optimal initialization parameters; The spiking neural network is trained using a time-series backpropagation algorithm, and the neuron parameters are dynamically adjusted to obtain the trained spiking neural network. Analyze the pulse propagation path and activation mode in the trained spiking neural network to extract emission time-series features; The theoretical emission curve is generated based on the emission time-series characteristics.
[0011] Optionally, training the spiking neural network using the time-series backpropagation algorithm includes: The gradient of the pulse generation function of the spiking neural network is calculated using the surrogate gradient method. Based on the gradient of the function, a loss function is constructed that includes the pulse sequence prediction error term and the causal segment number regularization term; The parameters of the spiking neural network are optimized using the loss function while maintaining the number of causal segments; The parameters of the spiking neural network are quantized and the connections are sparsified to compress the model size of the spiking neural network. Based on preset evaluation metrics, the detection accuracy of the compressed spiking neural network is verified on a test dataset.
[0012] Optionally, the step of analyzing the pulse propagation path and activation mode in the trained spiking neural network and extracting emission time-series features includes: Extract the neuron activation sequences of the spiking neural network to identify sensitive regions where changes in input cause significant changes in output; Locate the boundaries of causal segments corresponding to the sensitive regions and mark potential anomalous data points; Analyze the emission parameter combinations of the potential abnormal data points to determine the abnormal triggering conditions; Using the aforementioned abnormal triggering conditions as a benchmark, emission time-series features are extracted from the potential abnormal data points.
[0013] Optionally, the standardization process performed on the multi-source detection data to generate standardized multi-source detection data includes: A physical emission model is built based on vehicle type and engine parameters to generate theoretical emission baseline data; Collect real-time OBD monitoring data and roadside remote sensing data of the vehicle to establish statistical benchmark data; Extract the timestamp information from the multi-source detection data, the OBD real-time monitoring data, and the roadside remote sensing data to construct a time alignment optimization target; Based on the aforementioned time alignment optimization objective, a genetic algorithm is used to optimize the time offset parameters to obtain time-aligned multi-source detection data. Based on the time-aligned multi-source detection data, the theoretical emission baseline data, and the statistical baseline data, a unified monitoring sequence for vehicle emission behavior is constructed to generate standardized multi-source detection data.
[0014] Optionally, the step of using a genetic algorithm to optimize the time offset parameters to obtain time-aligned multi-source detection data includes: Construct a chromosome encoding scheme to represent the time offset and scaling factor of the multi-source data; Construct a fitness function based on the mutual information loss of the multi-source detection data to evaluate the time alignment effect; Chromosomes are selected according to the fitness function, and crossover and mutation operations are performed on the selected chromosomes. The selected chromosome is iteratively optimized until the time alignment accuracy reaches the preset accuracy threshold requirement. The optimal time offset parameter is then output to obtain the time-aligned multi-source detection data.
[0015] This invention also provides a system for intelligent analysis of the authenticity of motor vehicle exhaust emission test data, comprising: The data receiving module is used to receive vehicle exhaust emission detection data, combine it with vehicle maintenance data and testing equipment calibration logs to obtain multi-source detection data, and perform standardization processing on the multi-source detection data to generate standardized multi-source detection data; wherein, the vehicle exhaust emission detection data includes vehicle VIN code, detection time, pollutant concentration value and testing equipment serial number; The organization analysis module is used to construct an association graph of the detection organizations based on the standardized multi-source detection data, and to use the isolated forest model to detect abnormal subgraphs in the association graph of the detection organizations to obtain the results of organization collaboration anomaly identification. The trajectory analysis module is used to input the target vehicle information from the standardized multi-source detection data into the first neural network model based on the trained first neural network model, generate a theoretical emission curve, calculate the time warping distance between the measured data of the target vehicle and the theoretical emission curve, and obtain the vehicle trajectory anomaly analysis results; wherein, the first neural network model includes an LSTM model and / or a spiking neural network; The early warning module is used to generate early warning information on the authenticity of detection data based on the results of the collaborative anomaly identification of the institution and the results of the vehicle trajectory anomaly analysis.
[0016] The beneficial effects of this invention are as follows: 1. This invention establishes a global data lake, integrating multi-source data such as data from all provincial and municipal testing institutions, vehicle maintenance data, and equipment calibration logs, breaking down data silos and providing a data foundation for global analysis; 2. This invention uses the isolated forest model to detect abnormal subgraphs in the association graph of the detection agencies, which can effectively identify the collective data deviation of the detection agencies and discover collaborative abnormal behavior.
[0017] 3. This invention utilizes LSTM models and / or spiking neural networks to analyze vehicle emission trajectories, and combined with vehicle maintenance records, it can accurately identify abnormal fluctuations in detection data, thereby improving the accuracy of identifying genuine and fake behaviors. 4. This invention continuously learns new true / false patterns through a dynamic knowledge base, enabling dynamic responses to constantly changing true / false strategies and improving the long-term effectiveness of the system. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0019] Figure 1 A flowchart of the intelligent analysis method for verifying the authenticity of motor vehicle exhaust emission test data provided in an embodiment of the present invention; Figure 2 A flowchart of the collaborative anomaly identification method for detection agencies provided in an embodiment of the present invention; Figure 3 A flowchart of a vehicle emission anomaly detection method based on a spiking neural network provided in an embodiment of the present invention; Figure 4 This is a structural diagram of the intelligent analysis system for verifying the authenticity of motor vehicle exhaust emission test data provided in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0021] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0022] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Example 1 This embodiment provides a method for intelligent analysis of the authenticity of motor vehicle exhaust emission test data, such as... Figure 1 As shown, it includes: Step S1: Receive vehicle exhaust emission detection data, combine it with vehicle maintenance data and testing equipment calibration logs to obtain multi-source detection data, and standardize the multi-source detection data to generate standardized multi-source detection data; wherein, the vehicle exhaust emission detection data includes vehicle VIN code, detection time, pollutant concentration value and testing equipment serial number; In this embodiment, step S1 primarily involves core vehicle exhaust emission testing data, including the vehicle's VIN code, testing time, pollutant concentration values (such as CO, HC, NOx, etc.), and the testing equipment serial number. Simultaneously, vehicle maintenance data (such as M-station maintenance records) and testing equipment calibration logs (including calibration time, calibration personnel, sensor parameters, etc.) are also acquired. Furthermore, multi-source data can be accessed, including video surveillance data from testing institutions during the exhaust emission testing process, vehicle remote sensing data, heavy-duty diesel vehicle OBD remote monitoring data, and vehicle roadside inspection and on-site testing data, to obtain multi-source testing data. This data is then standardized to generate standardized multi-source testing data, providing a data foundation for subsequent analysis.
[0025] Specifically, the multi-source detection data is standardized to generate standardized multi-source detection data, including: A physical emission model is built based on vehicle type and engine parameters to generate theoretical emission baseline data; Collect real-time OBD monitoring data and roadside remote sensing data of the vehicle to establish statistical benchmark data; Extract the timestamp information from the multi-source detection data, the OBD real-time monitoring data, and the roadside remote sensing data to construct a time alignment optimization target; Based on the aforementioned time alignment optimization objective, a genetic algorithm is used to optimize the time offset parameters to obtain time-aligned multi-source detection data. Based on the time-aligned multi-source detection data, the theoretical emission baseline data, and the statistical baseline data, a unified monitoring sequence for vehicle emission behavior is constructed to generate standardized multi-source detection data.
[0026] Among them, a genetic algorithm is used to optimize the time offset parameter to obtain time-aligned multi-source detection data, including: Design a chromosome encoding scheme to represent the time offset and scaling factor of the multi-source data; Construct a chromosome encoding scheme to represent the time offset and scaling factor of the multi-source data; Construct a fitness function based on the mutual information loss of the multi-source detection data to evaluate the time alignment effect; Chromosomes are selected according to the fitness function, and crossover and mutation operations are performed on the selected chromosomes. The selected chromosome is iteratively optimized until the time alignment accuracy reaches the preset accuracy threshold requirement. The optimal time offset parameter is then output to obtain the time-aligned multi-source detection data.
[0027] In Example 1, standardizing the multi-source detection data is a fundamental step in the entire analysis process. Specifically, the process of standardizing the multi-source detection data to generate standardized multi-source detection data includes several key steps.
[0028] First, the system constructs a physical emission model based on vehicle type and engine parameters to generate theoretical emission benchmark data. This physical model considers factors such as vehicle displacement, fuel type, engine technology level, and service life, and establishes emission levels for various pollutants (CO, HC, NOx, etc.) under ideal conditions using thermodynamic and combustion kinetic principles. The physical model can provide theoretical emission predictions for different types and usage conditions of motor vehicles; these values serve as the first reference benchmark for assessing the reasonableness of actual emission data.
[0029] Secondly, the system collects real-time OBD monitoring data and roadside remote sensing data from the vehicle to establish statistical benchmark data. The OBD system continuously records key emission control parameters such as engine operating status, catalytic conversion efficiency, and oxygen sensor operation, while the roadside remote sensing equipment captures instantaneous emission levels during actual vehicle operation using infrared and ultraviolet spectroscopy analysis. The system uses these two types of data as the vehicle's emission performance under real-world operating conditions, and establishes the measured distribution characteristics of vehicle emissions through statistical methods, forming statistical benchmark data.
[0030] Then, the system extracts the timestamp information from the detection data, the OBD data, and the remote sensing data to construct a time alignment optimization objective. Because the three types of data differ in their acquisition frequency, recording methods, and time stamping mechanisms, direct comparison and analysis would result in temporal misalignment. The system transforms the time alignment problem into an optimization objective: finding the optimal time mapping parameters so that data from different sources can accurately correspond in the time dimension. Specifically, the system aims to maximize the mutual information between aligned data sequences, ensuring that different data sources reflect the same physical process.
[0031] Next, based on the aforementioned time alignment optimization objective, the system employs a genetic algorithm to optimize the time offset parameters, achieving time alignment of multi-source data. As an evolutionary computation method, the genetic algorithm is particularly suitable for solving this type of multi-parameter optimization problem. The specific process of optimizing the time offset parameters using the genetic algorithm includes five key steps: The first step involves designing a chromosome encoding scheme to represent the time offset and scaling factor of the multi-source detection data. Each chromosome contains multiple sets of genes, corresponding to time mapping parameters between different data sources, including the overall offset (e.g., system clock differences of several hours or days) and the time scaling factor (to handle inconsistent sampling frequencies). The encoding uses real numbers to ensure the continuity and accuracy of the parameter search space.
[0032] The second step involves constructing a fitness function based on the mutual information loss of the multi-source detection data to evaluate the time alignment effect. This function calculates the mutual information gain between the data sources after alignment, while also considering the smoothness and consistency constraints of the data alignment. The fitness function design also incorporates a penalty term to avoid signal distortion caused by excessive time stretching or compression.
[0033] Third, the system selects a subset of chromosomes for crossover and mutation operations based on the fitness function. The selection strategy combines roulette wheel selection with elite retention to ensure the inheritance of high-quality solutions while maintaining population diversity. The crossover operation uses arithmetic crossover, weighting the time parameters of the two parent generations; the mutation operation dynamically adjusts the mutation rate according to the algorithm's iteration stages, maintaining high exploratory activity in the early stages and improving local fine-grained search capabilities in the later stages.
[0034] The fourth step involves iteratively optimizing the selected chromosomes until the time alignment accuracy meets the requirements. During the iteration process, the system continuously evaluates the fitness value of the best individual in the population. The algorithm terminates when the improvement in the best fitness is below a preset threshold for several consecutive generations (usually set to 15-20 generations), or when the maximum number of iterations (usually 50-100 generations) is reached. As the iteration progresses, the chromosome population gradually converges to the parameter combination that best preserves the time correspondence of multi-source data.
[0035] Fifth, the system outputs the optimal time offset parameters to establish the temporal correspondence of the multi-source data. After optimization, the system applies the optimal chromosome-encoded parameter set to perform temporal remapping on the original data, generating a time-consistent multi-source data stream. This alignment enables emission data from different sources to be compared and fused for analysis within the same time frame.
[0036] Finally, based on the time-aligned multi-source data, theoretical emission benchmark data, and statistical benchmark data, the system constructs a unified monitoring sequence for vehicle emission behavior to generate standardized multi-source detection data. The system integrates the three types of benchmark data (theoretical model, OBD monitoring, and remote sensing measurements) into a unified standardized sequence using a weighted fusion algorithm. The fusion process considers the reliability, completeness, and timeliness of each data source, generating the best estimate and its confidence interval for the emission status at each time point. Standardization processing also includes data normalization, missing value handling, and outlier labeling to ensure data quality for subsequent analysis. Through this series of processes, the original multi-source heterogeneous data is transformed into structured, time-aligned, standardized multi-source detection data, providing a reliable data foundation for subsequent anomaly detection and authenticity analysis.
[0037] Step S2: Based on the standardized multi-source detection data, construct a detection agency association graph, and use the isolated forest model to detect abnormal subgraphs in the detection agency association graph to obtain standardized multi-source detection data of agency collaborative anomaly identification results; In step S2, a correlation graph of testing institutions is constructed based on standardized multi-source detection data. Nodes in the graph represent testing institutions, and edges represent the correlation strength between institutions. The correlation strength is calculated by weighting factors such as equipment homogeneity, geographical proximity, and data distribution similarity. Then, an isolated forest model is used to detect anomalous subgraphs in the correlation graph and identify groups of testing institutions that may exhibit coordinated anomalous behavior.
[0038] Step S3: Establish a first neural network model, input the target vehicle information from the standardized multi-source detection data into the first neural network model, generate a theoretical emission curve, calculate the time warping distance between the measured data of the target vehicle and the theoretical emission curve, and obtain the vehicle trajectory anomaly analysis result. The first neural network model includes an LSTM model and / or a spiking neural network. In step S3, based on the trained first neural network model, the target vehicle information from the standardized multi-source detection data is input into the first neural network model to generate a theoretical emission curve. The time warp distance between the measured data of the target vehicle and the theoretical emission curve is calculated to obtain the vehicle trajectory anomaly analysis result. The first neural network model includes an LSTM model and / or a spiking neural network. By analyzing the historical trajectory of vehicle emission data and combining it with maintenance records, abnormal fluctuations in the detection data can be identified.
[0039] Step S4: Based on the results of the mechanism collaboration anomaly identification and the results of the vehicle trajectory anomaly analysis, generate early warning information on the authenticity of the detection data.
[0040] In step S4, the system comprehensively considers the results of the identification of abnormal collaborative data between the testing agencies and the analysis results of abnormal vehicle trajectories to generate a warning message for the authenticity of the test data. When a data deviation is detected in the testing agency's collaborative data, or when there is an unreasonable improvement in the vehicle emission data (such as an improvement exceeding 45% without corresponding maintenance records), the system will trigger a warning to remind regulatory authorities to pay attention to the authenticity of the relevant test data.
[0041] Example 2 Building upon Example 1, this example details a specific method for constructing an association graph of detection institutions based on the standardized multi-source detection data, employing an isolated forest model to detect anomalous subgraphs within the association graph, and obtaining the results of collaborative anomaly identification of institutions. Figure 2 As shown, it includes: S2.1. The standardized multi-source detection data is compressed using a deep autoencoder to generate a low-dimensional representation vector of the detection mechanism; First, in step S2.1, the system utilizes a deep autoencoder to perform dimensionality reduction and compression on the standardized multi-source detection data. A deep autoencoder is a special neural network structure consisting of an encoder and a decoder, capable of mapping high-dimensional data to a low-dimensional latent space while retaining key information. For the detection agency data, the system constructs a 5-layer autoencoder network. The input layer receives multi-dimensional feature vectors (typically 72-128 dimensions) containing pollutant concentration distribution, equipment calibration frequency, and detection pass rate. The intermediate encoding layers progressively compress the data to a typical 128-dimensional low-dimensional representation space. The training process uses reconstruction loss as the optimization objective, employs the Adam optimizer for parameter updates, and sets the learning rate to 0.001. After 300 training cycles, the autoencoder can losslessly compress the original high-dimensional data into low-dimensional vectors containing key features. These low-dimensional representation vectors capture the essence of the characteristics of each detection agency's data, providing computational efficiency and dimensionality reduction advantages for subsequent analysis.
[0042] S2.2. Perform comparative learning on the low-dimensional representation vector, and use the data of the same vehicle in different detection agencies as positive sample pairs for training to obtain the normal detection mode vector; Next, in step S2.2, the system performs contrastive learning on these low-dimensional representation vectors to learn detection patterns that distinguish between normal and abnormal patterns. The core idea of contrastive learning is to learn effective representations by comparing the similarity and differences between samples. In this application, the system treats data from the same vehicle at different detection agencies as positive sample pairs (should be similar), while data from significantly different vehicles are treated as negative sample pairs (should be different). In this way, the model learns the common features of normal detection patterns and can identify abnormal behaviors that deviate from these commonalities.
[0043] Before model training, the system performed comprehensive preprocessing on the raw data to ensure data quality and representativeness. The preprocessing process included key steps such as handling missing values, identifying and adjusting outliers, data standardization, and balancing imbalanced datasets.
[0044] To address the missing value problem, the system first analyzes the data missing patterns, distinguishing between random and systematic missing values. For randomly missing data points, the system uses interpolation methods based on time series characteristics for imputation. Specifically, linear interpolation is used for short-term missing values (intervals less than 3 time points), spline interpolation is used for medium-term missing values (intervals 3-7 time points), and long-term missing values are imputed by combining data patterns from similar vehicles. The system sets a missing value tolerance threshold, discarding samples with consecutive missing values exceeding a preset threshold (usually 20% of the sequence length) to avoid introducing excessive uncertainty. Furthermore, the system records the location information of missing values as auxiliary features to help the model identify differences in data reliability.
[0045] Outlier identification is a crucial step in ensuring model stability. The system employs an improved Z-score method combined with domain knowledge for anomaly detection. Traditional Z-score methods assume a normal distribution, which may not be suitable for the skewed distribution of emissions data. Therefore, the system introduces a robust Z-score variant based on quantiles. Specifically, the system first groups the data by vehicle type and year, then calculates the median and interquartile range within each group. Data points deviating from the median by more than 1.5 times the interquartile range are marked as potential anomalies. For detected outliers, the system does not simply delete or replace them, but performs secondary verification using knowledge from vehicle maintenance records, testing equipment calibration logs, and other domains. Verified outliers are appropriately adjusted (usually limited to reasonable boundaries) rather than completely removed to retain potentially useful information. Extreme values that cannot be verified (such as those deviating from the mean by more than 5 standard deviations) are deleted to avoid model training bias.
[0046] Data standardization is a necessary step in training neural network models. Considering the varying dimensions and distribution characteristics of different pollutant concentrations (such as CO, HC, and NOx), the system employs multiple standardization strategies. For features with approximately normal distributions, Z-score standardization is used; for skewed distributions, a logarithmic transformation is performed before standardization; and for features with clear upper and lower limits, Min-Max scaling is used to scale them to the [0,1] interval. Standardization parameters (such as mean, standard deviation, maximum and minimum values) are calculated based on the training set and stored for subsequent testing and inference phases. To enhance the model's generalization ability, the system also introduces a mini-batch standardization layer during training to dynamically adjust the feature distribution and mitigate internal covariate bias.
[0047] To address the data imbalance issue, the system employs a hybrid strategy combining stratified sampling and synthetic sample generation. First, vehicles are stratified based on key attributes such as vehicle type, year, and emission standard to ensure that the proportion of samples in each stratum in the training set is consistent with the distribution of the target application scenario. For rare categories with insufficient sample size (such as specific high-emission models or vehicles meeting new emission standards), the system uses an improved SMOTE (Synthetic Minority Oversampling) algorithm to generate synthetic samples. This algorithm generates new samples by interpolating between k-nearest neighbors in the feature space, enhancing the model's ability to identify rare categories. Furthermore, the system employs a balanced batch construction strategy to ensure that each training batch contains diverse vehicle types and emission modes, preventing the model from being frequently dominated by a single category during training.
[0048] Temporal feature engineering is crucial for improving model performance. The system constructs multi-scale time window features, including short-term (hourly), medium-term (daily), and long-term (monthly) features, to capture changes in emission patterns across different time scales. For each time window, the system extracts statistical features (mean, variance, quantiles, etc.), trend features (linear trend coefficients, seasonality indices, etc.), and frequency domain features (Fourier transform coefficients, wavelet energy, etc.). The system also constructs domain-knowledge-based interactive features, such as ratios between different pollutants and ratios of the same pollutant under different loads. These features reflect engine combustion status and aftertreatment system efficiency, providing the model with more discriminative information.
[0049] Finally, the system implements a rigorous training-validation-test data splitting strategy to ensure the reliability of model evaluation. Data splitting employs a time-series-aware approach, rather than simple random sampling, to avoid information leakage. Specifically, the system divides the data into three parts in chronological order: the first 70% is used for training, the middle 15% for validation and hyperparameter tuning, and the last 15% for final performance evaluation. This splitting method more realistically reflects the model's performance in real-world applications, especially its adaptability to time-evolving emission characteristics and fraudulent patterns.
[0050] The system employs a multi-level, multi-dimensional evaluation framework to comprehensively measure model performance, ensuring its reliability and effectiveness in practical applications. The evaluation index system covers four key dimensions: classification accuracy, ranking quality, probability calibration, and business value, providing comprehensive guidance for model selection and optimization.
[0051] In terms of classification accuracy, the system first calculates a standard confusion matrix, including precision, accuracy, recall, and F1 score. Considering the specific needs of detecting genuine and counterfeit data, the system places particular emphasis on balancing precision (ensuring a low false positive rate) and recall (ensuring a high detection rate), using a weighted F1 score (typically with a target value ≥ 0.85) for comprehensive evaluation. The weighting is dynamically adjusted based on business needs; in the early stages of regulatory enforcement, precision may be prioritized to establish system credibility, while as the system matures, recall may be more emphasized to improve regulatory coverage. The system also calculates stratified performance indicators for different types of vehicles and testing institutions, identifying the model's strengths and weaknesses for specific subgroups, providing direction for subsequent optimization.
[0052] Ranking quality assessment focuses on the model's ability to accurately rank anomalies, which is particularly important for resource-constrained regulatory enforcement. The system uses the Receiver Operating Characteristic (ROC) curve and its area under the curve (AUC) to evaluate the model's discriminative ability at different decision thresholds. Considering the significant imbalance in the ratio of positive to negative samples in real / false detection (typically, the proportion of real and false data does not exceed 10%), the system places greater emphasis on the precision-recall (PR) curve and its area under the curve (PR-AUC), a more sensitive performance metric when handling imbalanced datasets. Furthermore, the system calculates ranking correlation coefficients (such as the Spearman correlation coefficient) to assess the consistency between the model's predicted anomaly scores and expert judgments of anomaly severity, ensuring that high-risk cases receive priority attention.
[0053] Probabilistic calibration evaluation ensures that the probability scores output by the model are meaningful; that is, among samples with a predicted probability of p, the proportion actually belonging to the positive class should be close to p. The system uses reliability plots and expected calibration error (ECE) to measure the quality of probabilistic calibration. To improve calibration performance, the system applies techniques such as temperature scaling or binning calibration during the model post-processing stage to convert the original model output into well-calibrated probability scores. Good probabilistic calibration enables regulators to develop intervention strategies based on confidence levels; for example, investigations can be initiated directly for cases with a predicted probability exceeding 90%, while cases in the 70%-90% range may require additional evidence.
[0054] The evaluation of emissions forecasting tasks employs regression metrics. The system calculates the root mean square error (RMSE) to assess forecast accuracy, the mean absolute percentage error (MAPE) to assess relative error levels, and uses the coefficient of determination (R²) to measure the model's ability to explain variance. Considering the need to quantify the uncertainty in emissions forecasting, the system additionally evaluates the quality of the forecast interval, including forecast interval coverage (PICP) and average forecast interval width (MPIW). While maintaining high coverage (typically ≥90%), the system strives for narrower forecast intervals to provide an accurate and reliable forecast range.
[0055] The model's decision threshold is determined using a cost-sensitive optimization method. The system first quantifies the operational costs of false positives and false negatives through expert interviews and historical case analysis. False positive costs include unnecessary investigative resource consumption and potential reputational damage to testing agencies, while false negative costs include increased environmental pollution risks and loss of regulatory credibility. Based on these cost estimates, the system constructs a cost matrix and then uses grid search or Bayesian optimization methods to find the decision threshold that minimizes the expected total cost on the validation set. Furthermore, the system supports a multi-threshold strategy, categorizing prediction results into three types: "highly suspicious," "requiring further investigation," and "likely normal," providing a tiered response framework for regulatory enforcement.
[0056] The system's time stability assessment ensures the model remains effective over long-term operation. Using a sliding window validation method, the system evaluates model performance on test sets at different time intervals, identifying potential performance degradation. When a significant performance decline is detected (e.g., a drop of more than 5 percentage points in key metrics), the system triggers a model update process. Performance degradation analysis also helps identify emerging true and false patterns that may not yet be fully learned by the current model.
[0057] Finally, the system evaluated the effectiveness of different model architectures and feature combinations through comparative experiments. Benchmark models included traditional machine learning methods (such as random forests and gradient boosting trees) and deep learning methods (such as CNNs, LSTMs, and spiking neural networks), and their performance on various metrics was compared through cross-validation. Experimental results show that for collaborative anomaly detection tasks involving detection agencies, the graph neural network-based model performs optimally; while for vehicle trajectory anomaly analysis, LSTMs combined with attention mechanisms and spiking neural networks have significant advantages in capturing temporal features. These comparative results not only guided the final model selection but also verified the scientific validity and advanced nature of the system's technical solution.
[0058] Specifically, the low-dimensional representation vector undergoes comparative learning, using data from the same vehicle at different detection institutions as positive sample pairs for training, to obtain a normal detection mode vector, including: S2.2.1 Based on the emission data of the target vehicle from different testing institutions, extract data similarity features and construct a comparative learning sample set; The comparative learning process in step S2.2 can be further divided into five sub-steps. In S2.2.1, the system extracts data similarity features and constructs a comparative learning sample set based on the emission data of the target vehicle from different testing institutions. Specifically, the system selects test records of the same vehicle from different institutions within 30 days as positive sample pairs, ensuring that these samples are sufficiently close in time and that their emission characteristics should not have significant differences. For each pair of data, the system calculates the relative deviation and consistency of fluctuation trends of each pollutant index, constructing a feature vector that reflects the intrinsic similarity of the samples. The final comparative learning sample set contains a large number of positive sample pairs and randomly combined negative sample pairs, typically maintaining a 1:4 positive-to-negative sample ratio to handle class imbalance problems.
[0059] S2.2.2 Add Gaussian noise and a random mask to the contrastive learning sample set to generate an enhanced training dataset; In S2.2.2, the system applies data augmentation techniques to the contrastive learning sample set, adding Gaussian noise and a random mask to generate an enhanced training dataset. The standard deviation of the Gaussian noise is set to 10% of the standard deviation of the original data, and the random mask sets the feature values to zero with a 15% probability. These augmentation techniques simulate data fluctuations and missing values that may occur during actual detection, improving the robustness of the model. The size of the augmented training set is typically 3-5 times that of the original sample set, providing the model with richer training materials.
[0060] S2.2.3 The time-series features of the augmented training dataset are processed using a self-attention mechanism to obtain sequence encoding; In the next step, S2.2.3, the system employs a self-attention mechanism to process the time-series features of the augmented training dataset, obtaining sequence encoding. The self-attention mechanism can capture long-range dependencies in sequence data, making it particularly suitable for processing temporal patterns in emission data. The system implements a self-attention layer with four attention heads, each independently learning different temporal relationship patterns. The input sequence is first augmented with temporal information through positional encoding, and then undergoes self-attention computation to obtain a context-aware representation. The final sequence encoding integrates key patterns and contextual information from the time series, providing a high-quality data representation for subsequent comparative learning.
[0061] S2.2.4 Based on the sequence encoding, calculate the contrast loss of positive and negative pairs of samples, and update the parameters of the isolated forest model; In step S2.2.4, the system calculates the contrastive loss between positive and negative sample pairs based on sequence encoding and updates the parameters of the Isolation Forest model accordingly. The contrastive loss function uses a modified InfoNCE formula, encouraging positive sample pairs to be close together in the representation space, while negative sample pairs are kept far apart. Specifically, the system uses cosine similarity to measure the distance between samples and controls the sensitivity of the loss function using a temperature parameter (usually set to 0.07). The contrastive loss directly guides the update of the Isolation Forest model parameters, enabling the model to tightly cluster normal pattern samples while effectively isolating abnormal samples.
[0062] S2.2.5 Based on the parameters of the isolated forest model, the low-dimensional representation vector is mapped to the normal detection pattern vector space to obtain the normal detection pattern vector.
[0063] Finally, in step S2.2.5, the system maps the low-dimensional representation vector to the normal detection pattern vector space based on the updated isolated forest model parameters, obtaining the normal detection pattern vector. This mapping process essentially projects the original low-dimensional representation onto the learned normal pattern manifold, making the representation of the normal detection data more consistent and compact. The mapping function is implemented using a multilayer perceptron, containing two hidden layers (64 and 32 neurons respectively) and an output layer, with LeakyReLU used as the non-linear activation function. Through this mapping, the system obtains a vector representation of the normal detection pattern, providing a benchmark for subsequent anomaly detection.
[0064] S2.3. Based on the low-dimensional representation vector and the normal detection mode vector, calculate the Mahalanobis distance between each detection institution and its historical detection data, and screen potential abnormal detection institutions; In step S2.3, the system calculates the Mahalanobis distance between each detection agency and its historical detection data based on the low-dimensional representation vector and the normal detection pattern vector, thus filtering out potentially anomalous detection agencies. Compared to Euclidean distance, Mahalanobis distance considers the correlation and scale differences between features, and can more accurately measure the degree of anomalousness in multidimensional space. The system establishes a historical data covariance matrix for each detection agency and calculates the standardized distance between new detection data and historical normal patterns. When the Mahalanobis distance of a detection agency is significantly higher than its historical mean (usually set to the mean plus three standard deviations), the agency is marked as potentially anomalous.
[0065] S2.4. A dynamic prediction model for the emission characteristics of the target vehicle is constructed using variational Bayesian inference, and the emission characteristic distribution of the target vehicle is generated; In step S2.4, the system employs variational Bayesian inference to construct a dynamic prediction model of the target vehicle's emission characteristics, generating an emission characteristic distribution. The variational Bayesian method combines the uncertainty quantification of Bayesian inference with the computational efficiency of variational inference, generating probability distributions rather than point estimates, thus more comprehensively representing the uncertainty of the prediction results. The system uses a conditional variational autoencoder architecture, taking historical vehicle emission data, vehicle model information, and maintenance records as conditional inputs to learn the conditional probability distribution of emission characteristics. Model training uses variational lower bound (ELBO) as the optimization objective, while balancing reconstruction error and KL divergence to ensure the accuracy and diversity of the generated distribution. After training, the model can generate reliable distribution predictions of the future emission characteristics of a given vehicle, including mean and variance information.
[0066] A dynamic prediction model for the emission characteristics of the target vehicle is constructed using variational Bayesian inference, generating the emission characteristic distribution of the target vehicle, including: S2.4.1 Collect historical emission data of the target vehicle and construct an emission feature training set; In step S2.4.1, the system first collects historical emissions data of the target vehicle to construct an emissions feature training set. During this process, the system associates all historical inspection records of the vehicle with its VIN code, including records from different time points and different inspection agencies. The collected data includes the concentrations of various pollutants (CO, HC, NOx, PM, etc.), inspection time, mileage, engine operating parameters, etc. The system also pays special attention to the temporal distribution of the data to ensure that the training set covers different stages of vehicle use (such as the new car period, break-in period, stable period, aging period, etc.). For vehicles with insufficient data, the system supplements the data with data from vehicles of the same model and year as prior knowledge to improve the model's generalization ability. The final training set typically contains 10-20 inspection records of the target vehicle, forming a longitudinal dataset spanning from several months to several years.
[0067] S2.4.2 Based on the emission feature training set, extract the emission concentration trend change features and periodic fluctuation features; In step S2.4.2, the system extracts trend change features and periodic fluctuation features of pollutant concentrations based on the emission feature training set. The trend change features reflect the long-term evolution pattern of vehicle emissions over usage time and mileage. The system uses a non-parametric regression method (such as Local Weighted Regression LOESS) to fit the long-term trend line of the emission data. This trend line smooths short-term fluctuations and reflects the true emission performance degradation curve. For the periodic fluctuation features, the system applies wavelet transform and Fourier analysis to decompose the periodic components in the emission data, identifying seasonal changes (such as increased emissions in winter and relative stability in summer) and detecting periodic effects. The system pays particular attention to the fluctuation spectrum features of the emission data, using them as an important basis for anomaly detection. Through this step, the original time-series emission data is converted into a structured feature representation containing trend components, periodic components, and residual components.
[0068] S2.4.3 A dynamic prediction model of vehicle maintenance records and emission characteristic changes is established using variational Bayesian inference, and the emission characteristic abrupt change points are analyzed and obtained. Step S2.4.3 is the core of the entire prediction model. It uses variational Bayesian inference to establish a dynamic prediction model of vehicle maintenance records and changes in emission characteristics, and analyzes abrupt changes in emission characteristics. The technical implementation of this step is particularly critical and requires detailed explanation.
[0069] Variational Bayesian inference is an approximate Bayesian inference method that approximates the posterior distribution by optimizing the variational lower bound (ELBO), combining the uncertainty quantification capability of Bayesian methods with the computational efficiency of deep learning. In this system, variational Bayesian inference is used to build a probabilistic graphical model containing latent variables to capture the complex relationship between vehicle emission characteristics, maintenance records, and test data.
[0070] Specifically, the system constructs a dynamic Bayesian network, which includes the following key variables: Observed variables: pollutant concentration measurements at each time point, vehicle mileage, and maintenance records (including maintenance time and maintenance items). Latent variables: State variables that cannot be directly observed, such as the actual emissions of the vehicle, engine health, and catalytic converter efficiency. Parameter variables: model parameters such as emission state transition probability, maintenance impact coefficient, and observation noise variance. The system first defines an emission state transition model, describing how a vehicle's emission state evolves over time and mileage without maintenance intervention. This model uses nonlinear state-space equations to correlate the current state with the previous state, cumulative mileage increments, and vehicle characteristic parameters. The model specifically considers the correlations between different pollutants; for example, HC and CO emissions typically change synchronously, while NOx may exhibit different patterns.
[0071] Next, the system establishes an impact model on emissions status for maintenance events. Each maintenance event is modeled as an intervention in emissions status, with the magnitude of the impact depending on the maintenance type, the relevance of the maintenance item to the emissions system, and the quality of the maintenance. For example, replacing an oxygen sensor has a direct impact on Lambda values and related emissions, while replacing fluids has a relatively indirect impact on emissions. The system learns typical impact patterns for different maintenance types from historical data, forming a prior knowledge base of maintenance impacts.
[0072] Finally, the system defines an observation model that describes the probability of observing a specific detection value under a given emission state. This model takes into account sources of uncertainty such as detection error, equipment accuracy, and environmental factors, and typically uses a Gaussian observation model with state-dependent noise.
[0073] After integrating the three components mentioned above, the system constructs a complete probabilistic graphical model. Then, a variational inference algorithm is used to approximate the posterior distribution. Specifically, the system uses the mean-field variational inference method, assuming that the posterior distribution can be decomposed into the product of the independent distributions of each latent variable, and then iteratively optimizes the parameters of these distributions. The optimization objective is to maximize the lower bound of evidence (ELBO), i.e. ELBO = E[log p(observed data|latent variables)] - KL[q(latent variables)||p(latent variables)] Where q (latent variable) is the variational approximation of the posterior distribution, and p (latent variable) is the prior distribution. The optimization process uses the stochastic variational inference (SVI) algorithm to maximize ELBO through the stochastic gradient ascent method.
[0074] After model training, the system uses the learned variational posterior distribution to analyze abrupt changes in emission characteristics. An abrupt change is a point in time where emission characteristics undergo a significant change, which could be a positive change due to normal maintenance or an unreasonable change due to abnormal data. The system identifies abrupt changes by calculating the rate of change of the relative entropy (KL divergence) of the emission state posterior distribution over time. When the change in KL divergence exceeds a threshold (usually set to three standard deviations of historical fluctuations), that point in time is marked as an abrupt change.
[0075] For each identified abrupt change, the system further analyzes its plausibility. The system queries maintenance records within a 30-day window before and after the change to assess the causal consistency between the maintenance items and the emission changes. For example, if an abrupt change shows a significant reduction in HC and CO emissions, and there is a catalytic converter replacement record, this change is considered a plausible maintenance effect. Conversely, if there is a significant emission improvement (e.g., exceeding 45%) but no corresponding maintenance support, the abrupt change is marked as highly suspicious, possibly indicating the falsification of the test data.
[0076] The system also utilizes another advantage of Bayesian inference—the posterior prediction distribution—to assess the plausibility of new detection data. Given a vehicle's historical state and maintenance records, the system can calculate a predicted distribution of emission levels at future points in time, including the mean and uncertainty interval. New detection data that significantly deviates from this predicted distribution (e.g., exceeding the 95% prediction interval) is flagged as potentially anomaly.
[0077] S2.4.4 Adjust the anomaly determination threshold according to the emission characteristic mutation point to generate an adaptive prediction interval, which is used to characterize the emission characteristic distribution of the target vehicle.
[0078] In step S2.4.4, the system adjusts the anomaly detection threshold based on the identified abrupt change points to generate an adaptive prediction interval, which is used to characterize the emission feature distribution of the target vehicle. Unlike traditional fixed threshold methods, the innovation of this system lies in the fact that the threshold is dynamically adjusted based on the vehicle's historical fluctuation patterns, maintenance records, and seasonal factors.
[0079] Specifically, the system first calculates a baseline prediction interval for emission characteristics based on the posterior predicted distribution obtained from a variational Bayesian model. This interval is typically the posterior mean ± 2 times the posterior standard deviation, corresponding to approximately 95% confidence. Then, the system adjusts this baseline interval based on identified reasonable abrupt changes (abrupt changes supported by repairs). If there are abrupt changes resulting from multiple effective repairs in the historical data, the system increases the width of the prediction interval to accommodate potential rapid improvements in vehicle performance; conversely, if the vehicle's historical performance is stable, the prediction interval is narrowed accordingly to improve detection sensitivity.
[0080] The system also considers the impact of seasonal factors on the prediction range. For example, under low-temperature conditions in winter, the concentration of certain pollutants in exhaust gases will naturally increase, and the system will appropriately relax the judgment thresholds for these seasonally sensitive indicators. Similarly, vehicles from different years and with different emission standards will also be subject to differentiated prediction range strategies.
[0081] The resulting adaptive prediction range is a dynamic boundary that changes over time. It considers both the inherent evolution of vehicle emission characteristics and the influence of external factors such as maintenance events and seasonal changes. This prediction range is directly used for subsequent anomaly detection—when new detection data falls within the range, it is considered normal fluctuation; when data deviates significantly from the prediction range (especially with unreasonable sudden improvements) and there is no corresponding maintenance support, the system will trigger a data authenticity warning.
[0082] By employing this dynamic prediction model based on variational Bayesian inference, the system can distinguish between reasonable emission improvements resulting from normal maintenance and suspicious data anomalies, significantly improving the accuracy of true / false alarm detection and reducing the false alarm rate. This method is particularly suitable for handling complex systems such as motor vehicle emissions, which are influenced by multiple factors and exhibit high uncertainty.
[0083] S2.5. Based on the emission characteristic distribution and the potential anomaly detection agencies, calculate the anomaly score of the newly added emission data in the standardized multi-source detection data, and identify the anomaly detection agency group.
[0084] Finally, in step S2.5, based on the generated emission characteristic distribution, the system calculates the anomaly score of newly added emission data in the standardized multi-source detection data, identifying anomaly detection agency groups. The system calculates the log-likelihood of the new detection data under the predicted distribution and converts it into a standardized anomaly metric represented by Z-scores. For each detection agency, the system aggregates the anomaly scores of all its detected vehicles to form an agency-level anomaly score. Through spectral clustering algorithms, the system further identifies detection agency groups with similar anomaly patterns, which are likely to exhibit coordinated true / false behavior. Ultimately, the system outputs a list of anomaly detection agency groups, including agency name, anomaly score, number of related vehicles, and description of typical anomaly patterns, providing regulatory authorities with intuitive and clear early warning information.
[0085] Through the above steps, the system can effectively identify collaborative anomalies among testing agencies, especially systemic patterns of authenticity that are difficult to detect from data from a single agency. This anomaly detection framework based on deep learning and statistical models provides more precise technical support for emissions regulation.
[0086] In this embodiment, the Isolation Forest algorithm for detecting anomalous subgraphs is the core technology for identifying collaborative anomalies among detection agencies. This method first constructs an association graph network of each detection agency, and then applies an improved Isolation Forest algorithm to identify clusters of anomalous agencies. The specific implementation steps are as follows: The input data includes historical testing data feature vectors from each testing institution (containing multi-dimensional features such as pass rate, pollutant concentration distribution, and equipment calibration frequency) and inter-institutional relationship information (such as equipment supplier relationships, geographical proximity, and personnel mobility). First, the system constructs an institution relationship graph, where nodes represent testing institutions and edges represent the strength of the relationship between institutions. The relationship strength is calculated using weighted factors, including equipment homogeneity (35%), geographical proximity (25%), and data distribution similarity (40%).
[0087] The second step involves extracting subgraph features from the constructed association graph, including topological properties such as subgraph density, average node degree, and KL divergence within the subgraph, forming a subgraph feature vector. The system randomly selects features and constructs random partitioning trees, recording the path length required for a node to be isolated in each partitioning process. This step is repeated to construct approximately 100 random trees, forming an isolated forest model.
[0088] The third step is to calculate the anomaly score for each subgraph. The anomaly score is calculated based on the average isolated path length; the shorter the path, the easier it is to be isolated, and the higher the probability of an anomaly. The system sets a threshold (usually the 75th percentile of the global average score) and marks subgraphs with scores above the threshold as potential anomaly groups.
[0089] The output results are abnormal association groups of testing institutions, including a list of institutions within the group, anomaly scores, and descriptions of anomaly characteristics (such as the degree of deviation in pollutant concentration distribution, synchronicity anomalies, etc.). The system also generates a visual association diagram, using color intensity to represent the degree of anomaly, making it easier for regulatory personnel to intuitively understand the network structure and scope of impact of abnormal institutions.
[0090] Example 3 Based on Example 1, this example details the specific implementation method for establishing a first neural network model, inputting the target vehicle information from the standardized multi-source detection data into the first neural network model, generating a theoretical emission curve, calculating the time warp distance between the measured data of the target vehicle and the theoretical emission curve, and obtaining the vehicle trajectory anomaly analysis results. When the first neural network model is an LSTM model, the LSTM prediction model is used to construct the theoretical emission curve of the vehicle and identify abnormal trajectories by comparing it with actual detection data. The detailed implementation scheme of this model is as follows: During the model building phase, the system collects historical vehicle emissions data as the training set, including basic vehicle information (vehicle model, fuel type, engine displacement, initial registration date) and data from previous tests (concentrations of various pollutants, test dates, and mileage). The LSTM network adopts a three-layer structure: the input layer receives an 8-dimensional feature vector, the hidden layer contains 64 neurons, and the output layer predicts the concentration values of four major pollutants (CO, HC, NOx, and PM). The model training parameters are set as follows: batch size 32, initial learning rate 0.001 dynamically adjusted using the Adam optimizer, training epochs 200, and early stopping threshold of no improvement in loss on the validation set after 10 consecutive epochs.
[0091] After model training, for the target vehicle, the system inputs its basic information and historical data to generate a theoretical curve showing emissions changing over time. Simultaneously, actual detection records for the vehicle are collected to form a measured curve. The system applies the Dynamic Time Warping (DTW) algorithm to calculate the similarity distance between the two curves. The DTW algorithm allows for flexible matching of curves along the time axis and can handle situations where detection time is uneven. The DTW distance calculation uses Euclidean distance as the basic metric, and the window width is set to 10% of the sequence length to balance computational efficiency and matching accuracy.
[0092] The system also integrates a vehicle maintenance record database to extract time points of maintenance events related to emissions. By comparing the temporal correlation between maintenance events and abrupt changes in the emission curve, the system determines whether emission improvements are reasonably supported by maintenance. The system defines an emission abrupt change as a point where the emission value improvement exceeds three times the historical standard deviation of fluctuation. For each abrupt change, the system queries for relevant maintenance records within a 30-day window before and after it. If no records are found and the improvement exceeds 45%, it is marked as a highly suspicious data anomaly.
[0093] Finally, the system outputs a vehicle emissions trajectory analysis report, including a comparison chart of theoretical and measured curves, DTW distance values, markers of abnormal abrupt changes, and confidence scores. The report also includes a comparative analysis of emissions trajectories with vehicles of the same model and year, providing a more comprehensive basis for anomaly assessment and offering regulatory authorities accurate support for identifying the authenticity of test data.
[0094] Example 4 In Example 4, the Spiking Neural Network (SNN) serves as the primary implementation of the first neural network model, providing a novel theoretical foundation and technical approach for anomaly detection in motor vehicle emission data. Unlike traditional neural networks, SNN simulates the pulse coding and temporal dynamics characteristics of biological neurons, enabling it to more accurately capture temporal dependencies and causal relationships in emission data. The implementation steps of this method are detailed below.
[0095] Based on Example 1, when the first neural network model is a spiking neural network, a first neural network model is established, and the target vehicle information from the standardized multi-source detection data is input into the first neural network model to generate a theoretical emission curve, such as... Figure 3 As shown, it includes: S3.1 Construct a spiking neural network to convert the standardized multi-source detection data into a spiking sequence; In step S3.1, the system constructs a spiking neural network and converts standardized multi-source detection data into spiking sequences. The construction of the spiking neural network includes determining the network structure and neuron model. The system adopts a multi-layer architecture, including an input layer, two hidden layers, and an output layer. The number of neurons in the input layer corresponds to the feature dimensions of the standardized data (typically 12-16 features, including pollutant concentrations, vehicle state parameters, etc.). The first hidden layer contains 64 neurons, the second hidden layer contains 32 neurons, and the output layer corresponds to the predicted pollutant emission indicators (typically 4-6 indicators). Each neuron uses the Leaky Integrate-and-Fire (LIF) model, a widely used biological neuron simulation model whose membrane potential dynamics are described by the following differential equation: τ(dV / dt) = -(VV rest ) + R·I(t) Where V is the membrane potential, V rest V is the resting potential, τ is the membrane time constant, R is the membrane resistance, and I(t) is the input current. When the membrane potential V exceeds the threshold V... th At this time, the neuron generates a pulse and resets the membrane potential to V. reset The system is configured with different parameter values for neurons at different levels, such as τ=5ms for input layer neurons and τ=10ms for hidden layer neurons, in order to optimize information processing capabilities at different time scales.
[0096] Data transformation is a crucial step in the application of spiking neural networks. The system employs a time-coding method to convert continuous emission data into pulse sequences. Specifically, for each feature dimension, the system maps its numerical value to the time delay or frequency of pulse occurrence. For example, higher pollutant concentrations correspond to higher pulse frequencies or earlier occurrence times. The system uses Gaussian tuning curves to ensure a smooth mapping from numerical values to pulse codes while preserving the relative relationships of the data. For time-series data, the system applies a sliding window technique, converting data samples within each time window into a pulse sequence, with 50% overlap between adjacent windows to ensure temporal continuity.
[0097] S3.2 Calculate the number of causal segments in the pulse sequence under different spiking neural network parameter configurations, and select the optimal initialization parameters; In step S3.2, the system calculates the number of causal segments in the pulse sequence under different parameter configurations and selects the optimal initialization parameters. A causal segment is a region in the SNN input domain where the output pulse time exhibits local Lipschitz continuity relative to the input pulse time and network parameters. This concept is crucial for understanding and optimizing SNN behavior because the number of causal segments directly relates to the network's expressive power and learning efficiency.
[0098] The system first designs a grid search strategy in the parameter space to systematically explore key parameters (such as neuron threshold, membrane time constant, and initial range of synaptic weights). For each parameter configuration, the system randomly samples a subset of the training data (typically 20% of the entire set), observes the network response through forward simulation, and analyzes the relationship between input perturbations and output changes. Specifically, the system applies a small perturbation (e.g., a random change of ±5%) to each input sample and records the variation pattern of the output pulse time. Continuous variation regions are identified as individual causal segments, while output abrupt changes are marked as causal segment boundaries. In this way, the system can quantify the number of causal segments generated by the network under different parameter configurations.
[0099] Analysis shows a significant positive correlation between the number of causal fragments and the network training success rate. The more causal fragments generated by the parameter configuration, the faster the network converges and the better its generalization performance in subsequent training. Based on this finding, the system selects the parameter set that generates the most causal fragments as the network initialization scheme. Typical optimal configurations include a moderate neuron threshold distribution (e.g., following a Gaussian distribution with a mean of 0.5 and a standard deviation of 0.1), a large membrane time constant (τ = 10⁻¹⁵ ms), and a small initial range of synaptic weights (e.g., a uniform distribution of [-0.1, 0.1]).
[0100] S3.3 The spiking neural network is trained using the time-series backpropagation algorithm, and the neuron parameters are dynamically adjusted to obtain the trained spiking neural network; Step S3.3 is a crucial step in training the spiking neural network. The system uses the temporal backpropagation algorithm to train the network and dynamically adjust the neuron parameters. This step consists of five sub-steps, which are described in detail below.
[0101] The process of training the spiking neural network using a time-series backpropagation algorithm includes: S3.3.1 The gradient of the pulse generation function of the spiking neural network is calculated using the surrogate gradient method; In S3.3.1, the system utilizes a surrogate gradient method to calculate the gradient of the heaviside step function (Hercules step function) in a spiking neural network. The backpropagation algorithm used in traditional neural network training cannot be directly applied to SNNs because the Hercules step function is non-differentiable at the threshold. To address this issue, the system employs a surrogate gradient method, using a differentiable function to approximate the gradient of the Hercules step function. Specifically, the system selects the following surrogate function: g(x) = 1 / γ * max(0, 1-|xV th | / γ) Where γ is a hyperparameter controlling the approximation accuracy, typically set to 0.1. This function in V th The system provides a smooth transition in the vicinity, enabling efficient gradient propagation while maintaining consistent behavior with the original impulse function far from the threshold region. It implements an automatic differentiation framework that records the computational graph during forward propagation and uses a surrogate function to compute gradients during the backpropagation phase.
[0102] S3.3.2 Based on the gradient of the function, construct a loss function that includes the pulse sequence prediction error term and the causal segment number regularization term; In S3.3.2, the system constructs a loss function that includes a pulse sequence prediction error term and a causal segment quantity regularization term. The loss function is the objective of training optimization and directly affects the feature representations learned by the network. The system-designed loss function comprises two main components: L = L pred + λ*L caus Where L pred This is the prediction error term, which calculates the difference between the predicted output and the target output. Since the output is a pulse sequence, the system uses van Rossum distance to measure the similarity between two pulse sequences; this metric considers the temporal position and number of pulses. caus This is the causal segment number regularization term, which encourages the network to maintain a high number of causal segments, thus enhancing its generalization ability. This term is achieved by estimating the number of causal segments under the current parameters and maximizing this number. λ is a hyperparameter that balances the two terms, and is typically set to a value between 0.1 and 0.5.
[0103] S3.3.3 Optimize the parameters of the spiking neural network using the loss function while maintaining the number of causal segments; In S3.3.3, the system optimizes the spiking neural network parameters using a constructed loss function while maintaining the number of causal segments. The optimization process employs an Adam optimizer with an adaptive learning rate, initially set to 0.001 and dynamically adjusted based on training progress. To maintain the number of causal segments, the system evaluates the current network's causal segments after each training epoch and introduces a dynamic regularization strategy. If the number of causal segments decreases significantly, the system increases L... caus The weight λ of each term is adjusted accordingly; conversely, it is slightly reduced to balance prediction accuracy and network stability. The training process employs a mini-batch strategy with a batch size of 32 and a total training cycle of 100-150, including an early stopping mechanism to prevent overfitting.
[0104] S3.3.4 Perform parameter quantization and connection sparsification on the spiking neural network to compress the model size of the spiking neural network; In S3.3.4, the system performs parameter quantization and connection sparsification on the trained spiking neural network to compress the model size. This step aims to reduce model complexity, making it suitable for deployment in resource-constrained environments. Parameter quantization converts the continuous floating-point representation of neuron parameters and synaptic weights into discrete representations, such as 8-bit fixed-point numbers. The system employs a training-aware quantization method, considering the impact of quantization errors on network performance during the quantization process and performing fine-tuning compensation. Connection sparsification reduces network complexity by pruning synaptic connections that contribute less. The system uses an amplitude-based pruning strategy, retaining the top 60-80% of connections with the largest absolute weight values and setting the rest to zero. After pruning, the system performs a short retraining (typically 10 epochs) to restore network performance. Through these techniques, the final model size is compressed to approximately 25% of the original version, significantly reducing computational and storage requirements.
[0105] S3.3.5 Based on preset evaluation metrics, verify the detection accuracy of the compressed spiking neural network on a test dataset.
[0106] In S3.3.5, the system validates the detection accuracy of the compressed spiking neural network on a test dataset. The system uses an independent test set (typically 20-30% of the total data) to evaluate model performance, ensuring that compression does not significantly reduce detection accuracy. Evaluation metrics include mean squared error of emission prediction, anomaly detection precision, recall, and F1 score. Experiments show that the optimized compressed model maintains over 95% of the original detection accuracy while significantly improving inference efficiency. The system also performs sensitivity analysis to test the model's robustness under different noise levels and anomaly types, ensuring reliability in real-world applications.
[0107] S3.4 Analyze the pulse propagation path and activation mode in the trained spiking neural network to extract emission time-series features; Step S3.4 is the model interpretation and feature extraction stage, which systematically analyzes the pulse propagation path and activation mode in the trained spiking neural network and extracts emission time-series features. This step includes four sub-steps, which are described in detail below.
[0108] Analyze the pulse propagation path and activation mode in the trained spiking neural network to extract emission time-series features, including: S3.4.1 Extract the neuron activation sequence of the spiking neural network and identify sensitive regions where changes in input cause significant changes in output; In S3.4.1, the system extracts the neuronal activation sequences of a spiking neural network to identify sensitive regions where input changes cause significant output changes. The system first records detailed activity of all neurons in the network as they process input data, including membrane potential changes and pulse occurrence times. Then, the system applies systematic perturbations to the input data and observes the changing patterns of neuronal activity. By calculating the perturbation sensitivity matrix, the system quantifies the degree of influence of each input feature on the activity of each neuron. High-sensitivity regions typically correspond to important decision boundaries or feature transition points, which are key to understanding network behavior. The system also uses principal component analysis to reduce the dimensionality of the activity sequences, revealing the main activity patterns and neuronal population dynamics.
[0109] S3.4.2 Locate the boundary of the causal segment corresponding to the sensitive area and mark potential abnormal data points; In S3.4.2, the system locates the causal segment boundaries corresponding to sensitive regions and marks potential anomalous data points. Causal segment boundaries are regions in the input space that cause significant changes in the output impulse pattern, typically manifested as decision thresholds or mode transition points. The system identifies these boundaries by analyzing the local Lipschitz constant of the input-output mapping. Specifically, the system performs fine-grained sampling of the input space, calculates the rate of change of output from adjacent sample points, and marks regions where the rate of change suddenly increases as causal segment boundaries. Interestingly, the system finds that true anomalous data points often lie near these boundaries because they represent transitions from normal to anomalous patterns. Based on this finding, the system focuses its analysis on data points located near causal segment boundaries, marking them as potential anomalous points for further evaluation.
[0110] S3.4.3 Analyze the emission parameter combinations of the potential abnormal data points to determine the abnormal triggering conditions; In S3.4.3, the system analyzes the emission parameter combinations of potential anomaly data points to determine anomaly triggering conditions. For each marked potential anomaly point, the system performs in-depth feature analysis to identify which specific emission parameter combinations trigger anomaly determination. The system first extracts the raw values of each emission indicator and their deviations relative to the reference baseline. Then, it applies decision trees and rule extraction algorithms to extract interpretable conditional rules from the decision logic of the spiking neural network. For example, the system may find that specific combinations such as "when the HC value is 40% below the baseline and the CO value is 35% below the baseline, while NOx only decreases by 5-10%" are highly correlated with anomaly determination. These rules not only improve the interpretability of anomaly detection but also provide regulatory authorities with specific anomaly pattern characteristics, facilitating actual law enforcement.
[0111] S3.4.4 Using the aforementioned abnormal triggering conditions as a benchmark, extract the emission time-series features from the potential abnormal data points.
[0112] In S3.4.4, the system uses anomaly triggering conditions as a benchmark to extract emission time-series features from potential anomaly data points. Based on the anomaly triggering conditions identified in the previous step, the system further analyzes the performance patterns of these conditions over time. The system performs time-series modeling on the data from the anomaly data points and the time windows before and after them, extracting time features including mutation rate, fluctuation frequency, and autocorrelation. These features can distinguish between reasonable emission changes (such as gradual improvement after maintenance) and suspicious data anomalies (such as sudden and large jumps in detection values). The system pays particular attention to the temporal synergistic relationships between emission parameters, such as the relative time series of changes in the concentrations of different pollutants. These features are of great value in identifying subtle true and false patterns (such as adjusting only some parameters).
[0113] S3.5 Generate the theoretical emission curve based on the emission time-series characteristics.
[0114] Finally, in step S3.5, the system generates a theoretical emission curve based on the extracted emission time-series features. The system integrates the various features extracted in the preceding steps to construct a theoretical prediction model for vehicle emissions. This model considers basic vehicle characteristics (model, year, mileage, etc.), historical emission trajectories, maintenance records, and time-series features extracted from a spiking neural network, comprehensively generating a theoretical curve showing the change of vehicle pollutant emissions over time. The theoretical curve is represented by point estimates (most probable values) and confidence intervals, reflecting the uncertainty of the prediction. For each pollutant, the system generates an independent theoretical curve, while also considering the correlation constraints between pollutants. These theoretical curves are directly used for anomaly detection—the system compares the vehicle's measured emission data with the theoretical curves, calculates the Dynamic Time Warping (DTW) distance, and quantifies the difference between the two. When the difference significantly exceeds a preset threshold (usually three standard deviations of historical fluctuations), the system triggers an anomaly warning.
[0115] This spiking neural network-based approach enables the system to capture subtle temporal patterns and causal relationships in emissions data, providing more accurate anomaly detection capabilities than traditional neural networks, particularly excelling in handling complex true and false patterns and considering causal relationships. Simultaneously, causal fragment analysis and spiking path tracing enhance the system's interpretability, allowing regulatory authorities to clearly understand the basis for anomaly determinations and improving the scientific rigor and persuasiveness of regulatory enforcement.
[0116] Spiking Neural Network (SNN) causal segment analysis technology provides a novel theoretical foundation and implementation path for anomaly detection in motor vehicle emission data. Based on the temporal encoding characteristics of SNNs, this technology can accurately capture temporal dependencies and causal relationships in emission data. The implementation process first constructs an SNN model architecture suitable for emission data, including an input layer, several hidden layers, and an output layer. Each neuron employs a Leaky Integrate-and-Fire (LIF) model, capable of accumulating input until a threshold is reached to generate a pulse output. The system converts the emission data time series into a pulse sequence, mapping pollutant concentration values to pulse time or pulse frequency through rate encoding or time encoding.
[0117] During the SNN initialization phase, the system employs a causal segment maximization strategy. A causal segment refers to a region in the SNN input domain where the output pulse time exhibits local Lipschitz continuity relative to the input pulse time and network parameters. The system first performs a grid search in the parameter space to evaluate the number of causal segments generated by the network under different initialization schemes. Specifically, the system randomly samples a subset of training data, calculates the piecewise linearity of the network response under different parameter configurations, and selects the parameter set that generates the most causal segments as the initialization scheme. Practice shows that the greater the number of causal segments, the stronger the network's expressive power and learning potential, which is highly correlated with the observed training success rate.
[0118] The training process employs an improved version of the Temporal Backpropagation (STBP) algorithm. Since the impulse function is non-differentiable, the system uses a surrogate gradient method to approximate the gradient of the impulse generation function. To enhance the network's ability to capture long-term dependencies in emission data, the system dynamically adjusts the leakage and threshold parameters of neurons during training, enabling the network to handle both rapidly changing and slowly accumulating emission features simultaneously. The training objective function includes a prediction error term and a causal segment regularization term; the latter encourages the network to maintain a high number of causal segments, thus preserving good generalization ability.
[0119] During the model deployment phase, the system applies the trained SNN to detect anomalies in real-time emissions data. After inputting historical vehicle emissions data, the SNN output pulse sequence is decoded into emissions trend predictions and anomaly scores. The system pays particular attention to regions where small changes in the input cause significant changes in the output; these regions typically correspond to the boundaries of causal segments and are also high-risk areas for potential anomalies. By analyzing the pulse propagation paths and activation patterns in the SNN, the system can provide interpretable evidence for anomaly detection, indicating which specific combinations of emissions parameters triggered the anomaly determination.
[0120] To adapt to hardware resource constraints, the system implemented quantization and sparsification of the SNN model. Neuron parameters and synaptic weights were discretized into 8-bit fixed-point representations, and pruning techniques were used to remove connections that contributed little to the prediction. The final model size was compressed to 25% of the original version while maintaining a detection accuracy of over 95%. This allows the technology to be deployed on edge computing devices, supporting offline anomaly detection scenarios.
[0121] Example 5 This embodiment provides a system for intelligent analysis of the authenticity of motor vehicle exhaust emission test data, such as... Figure 4 As shown, it includes: The data receiving module is used to receive vehicle exhaust emission detection data, combine it with vehicle maintenance data and testing equipment calibration logs to obtain multi-source detection data, and perform standardization processing on the multi-source detection data to generate standardized multi-source detection data; wherein, the vehicle exhaust emission detection data includes vehicle VIN code, detection time, pollutant concentration value and testing equipment serial number; The organization analysis module is used to construct an association graph of the detection organizations based on the standardized multi-source detection data, and to use the isolated forest model to detect abnormal subgraphs in the association graph of the detection organizations to obtain the results of organization collaboration anomaly identification. The trajectory analysis module is used to input the target vehicle information from the standardized multi-source detection data into the first neural network model based on the trained first neural network model, generate a theoretical emission curve, calculate the time warping distance between the measured data of the target vehicle and the theoretical emission curve, and obtain the vehicle trajectory anomaly analysis results; wherein, the first neural network model includes an LSTM model and / or a spiking neural network; The early warning module is used to generate early warning information on the authenticity of detection data based on the results of the collaborative anomaly identification of the institution and the results of the vehicle trajectory anomaly analysis.
[0122] The system provided in this embodiment corresponds to the method provided in the foregoing embodiments. The specific implementation of each module can be referred to the corresponding steps in the foregoing method embodiments, and will not be repeated here.
[0123] Those skilled in the art will understand that the above-described embodiments are merely illustrative of the principles of the invention and are not intended to limit the possible implementations of the invention. Modifications and improvements can be made to the embodiments without departing from the principles and scope of the invention, and such modifications and improvements are also considered to fall within the protection scope of the invention.
[0124] The embodiments of this application employ the following core technologies: Adaptive Long-Term Embedding Technology Adaptive long-term embedding technology improves the accuracy of anomaly detection by establishing long-term behavioral patterns of vehicles and testing agencies. This technology first constructs a multi-dimensional embedding space, mapping vehicle emission characteristics and testing agency behavior patterns to a unified vector space. The system collects historical vehicle testing data, including pollutant concentration values from multiple tests, testing agency information, and testing time, etc., and compresses the high-dimensional data into low-dimensional embedding vectors through a deep autoencoder network, typically with 128 dimensions. These embedding vectors can capture the inherent patterns and evolutionary laws of vehicle emission characteristics.
[0125] During the data denoising stage, the system employs a contrastive learning framework, treating data from the same vehicle across different trusted testing institutions as positive sample pairs and clearly abnormal data as negative samples, training the model to learn the ability to distinguish between normal and abnormal patterns. The system uses a self-attention mechanism to process time-series data, modeling the time dependence of emissions data. Simultaneously, it enhances the model's robustness to noise by adding Gaussian noise and random masks for data augmentation. For each embedding vector, the system calculates its Mahalanobis distance to the set of historical embedding vectors, identifying and filtering out anomalous noise points while retaining a reliable subset of historical data.
[0126] The enhanced recommendation module, based on the purified embedding space, establishes a dynamic prediction model for the emission characteristics of each vehicle. The system employs variational Bayesian inference to generate a distribution prediction of the future emission characteristics of vehicles and constructs confidence intervals. When new detection data enters the system, the deviation of its embedding vector from the predicted distribution is quantified as an anomaly score. The system further correlates external factors such as vehicle maintenance records and changes in regional pollution control policies to adjust the anomaly detection threshold and reduce the false alarm rate.
[0127] The adaptive learning mechanism enables the system to continuously update the embedding model from new data. The system periodically (typically monthly) updates the embedding parameters using incremental learning methods and determines whether to adopt the new parameters by comparing performance on the validation set. To address the concept drift problem, the system maintains multiple versions of the embedding model for different time windows and uses ensemble methods to integrate the prediction results of multiple versions, improving the system's adaptability to long-term trends.
[0128] Cross-modal time fusion technology Cross-modal time fusion technology integrates multi-source heterogeneous data to establish a comprehensive monitoring framework for vehicle emission behavior, significantly improving the accuracy of identifying genuine and false emission patterns. This technology first constructs a multi-reference simulation system, modeling vehicle emission behavior as a multi-dimensional time series, considering multiple data sources including emission detection data, real-time vehicle OBD monitoring data, roadside remote sensing data, and environmental conditions. The system builds dual reference benchmarks based on a physical emission model and a data-driven model. The physical model establishes theoretical emission curves based on factors such as vehicle type, age, and engine parameters, while the data-driven model establishes statistical benchmarks based on large-scale real-world samples.
[0129] Time alignment is a key challenge in cross-modal fusion. This system employs an adaptive time warping method based on a genetic algorithm to address the issues of inconsistent sampling frequencies and timestamp offsets across different data sources. This method formalizes the time alignment problem as an optimization problem, with the objective function being to minimize the mutual information loss of the aligned data sequences. The genetic algorithm uses a specific chromosome encoding scheme to represent time offset and scaling parameters, iteratively optimizing these parameters through evolutionary operations (selection, crossover, and mutation). The system specifically designs an initial population generation strategy based on domain knowledge, significantly accelerating the convergence process. After approximately 50 generations of evolution, the algorithm typically finds a near-globally optimal time alignment scheme, ensuring precise temporal correspondence between different modalities.
[0130] The data fusion layer employs a hierarchical approach to process multimodal information. Low-level fusion uses an attention mechanism to weightedly integrate features from various data sources; mid-level fusion focuses on extracting complementary information and correlation patterns between modalities; and high-level fusion focuses on decision-level integration. The system introduces a variational autoencoder framework to establish a latent space mapping between modalities, enabling comparison and fusion of different modalities in a unified representation space. To address the issue of missing modalities, the system implements a modality completion mechanism based on generative adversarial networks, capable of inferring a reasonable distribution of missing modalities from existing modalities.
[0131] The anomaly detection module performs multi-scale analysis based on the fused multimodal time series. The system simultaneously detects anomaly patterns at three time scales: short-term (single detection), medium-term (periodic changes), and long-term (aging trends). Short-term analysis focuses on instantaneous anomalies during a single detection process, medium-term analysis focuses on the periodic fluctuation patterns of detected values, and long-term analysis tracks the degradation curve of vehicle emission performance over time. The system uses a combination of an improved Long Short-Term Memory (LSTM) network and wavelet transform to decompose the trend, seasonality, and random components in the time series, applying specific anomaly detection strategies to different components.
[0132] The system implements an active learning framework to continuously optimize model performance. By periodically updating the model from the subset with the most information from high-confidence samples, the system can adapt to external factors such as changes in emission standards and upgrades to detection equipment. Simultaneously, the system maintains an adversarial example library, recording historically discovered true and false patterns, and uses these samples to strengthen the model's robustness. In actual deployment, the system undergoes a comprehensive update quarterly, while daily operation employs incremental learning for continuous fine-tuning to maintain the model's adaptability and sensitivity.
[0133] Validation experiments show that cross-modal time fusion technology improves the accuracy of true / false detection from 78% to 92% compared to traditional methods, while reducing the false positive rate by 40%. This technology demonstrates significant advantages, especially in handling complex scenarios (such as emission fluctuations due to seasonal variations) and subtle true / false behaviors (such as fine-tuning only for specific pollutants), providing regulatory authorities with a more reliable decision support tool.
[0134] It should be noted that those skilled in the art can make various modifications and variations to this invention without departing from the spirit and scope of this invention. If such modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include such modifications and variations.
[0135] This disclosure also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program performs the steps of the method described above for intelligent analysis of the authenticity of motor vehicle exhaust emission test data. The storage medium can be either volatile or non-volatile computer-readable storage.
[0136] Furthermore, this disclosure also provides a computer program product, which stores a computer program. When the computer program is run by a processor, it executes the steps of a method for intelligent analysis of the authenticity of motor vehicle exhaust emission test data provided in any of the above embodiments of this disclosure. For details, please refer to the above method embodiments, which will not be repeated here.
[0137] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0138] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices and apparatuses described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0139] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0140] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0141] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0142] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A method for intelligent analysis of the authenticity of motor vehicle exhaust emission test data, characterized in that, include: The system receives vehicle exhaust emission test data, combines it with vehicle maintenance data and test equipment calibration logs to obtain multi-source test data, and performs standardization processing on the multi-source test data to generate standardized multi-source test data. The vehicle exhaust emission test data includes the vehicle VIN code, test time, pollutant concentration value, and test equipment serial number. Based on the standardized multi-source detection data, a detection agency association graph is constructed, and an isolated forest model is used to detect abnormal subgraphs in the detection agency association graph to obtain the agency collaborative anomaly identification results. Based on the trained first neural network model, the target vehicle information from the standardized multi-source detection data is input into the first neural network model to generate a theoretical emission curve. The time warping distance between the measured data of the target vehicle and the theoretical emission curve is calculated to obtain the vehicle trajectory anomaly analysis results. The first neural network model includes an LSTM model and / or a spiking neural network. Based on the results of the collaborative anomaly identification of the aforementioned institutions and the results of the anomaly analysis of the vehicle trajectory, a warning message for the authenticity of the detection data is generated.
2. The method according to claim 1, characterized in that, Based on the standardized multi-source detection data, a detection agency association graph is constructed. An isolated forest model is then used to detect anomalous subgraphs within this graph, yielding agency collaborative anomaly identification results, including: The standardized multi-source detection data is compressed using a deep autoencoder to generate a low-dimensional representation vector of the detection mechanism; Comparative learning is performed on the low-dimensional representation vector, and data from the same vehicle at different detection agencies are used as positive sample pairs for training to obtain the normal detection mode vector; Based on the low-dimensional representation vector and the normal detection mode vector, calculate the Mahalanobis distance between each detection institution and its historical detection data, and screen potential abnormal detection institutions. A dynamic prediction model for the emission characteristics of the target vehicle is constructed using variational Bayesian inference, and the emission characteristic distribution of the target vehicle is generated. Based on the emission characteristic distribution and the potential anomaly detection agencies, calculate the anomaly score of newly added emission data in the standardized multi-source detection data, and identify the anomaly detection agency group.
3. The method according to claim 2, characterized in that, The step of performing comparative learning on the low-dimensional representation vector, using data from the same vehicle at different detection agencies as positive sample pairs for training, to obtain a normal detection pattern vector includes: Based on the emission data of the target vehicle from different testing institutions, data similarity features are extracted to construct a comparative learning sample set; Gaussian noise and a random mask are added to the contrastive learning sample set to generate an enhanced training dataset; The time-series features of the augmented training dataset are processed using a self-attention mechanism to obtain sequence encodings. Based on the sequence encoding, the contrast loss of positive and negative pairs of samples is calculated, and the parameters of the isolated forest model are updated. Based on the parameters of the isolated forest model, the low-dimensional representation vector is mapped to the normal detection pattern vector space to obtain the normal detection pattern vector.
4. The method according to claim 3, characterized in that, The dynamic prediction model for the emission characteristics of the target vehicle is constructed using variational Bayesian inference, generating the emission characteristic distribution of the target vehicle, including: Collect historical emission data of the target vehicle and construct an emission feature training set; Based on the emission feature training set, emission concentration trend change features and periodic fluctuation features are extracted; A dynamic prediction model of vehicle maintenance records and emission characteristic changes was established using variational Bayesian inference, and the emission characteristic abrupt change points were analyzed. The anomaly detection threshold is adjusted based on the emission characteristic mutation points to generate an adaptive prediction interval, which is used to characterize the emission characteristic distribution of the target vehicle.
5. The method according to claim 1, characterized in that, The first neural network model, based on the completed training, inputs the target vehicle information from the standardized multi-source detection data into the first neural network model to generate a theoretical emission curve, calculates the time warp distance between the measured data of the target vehicle and the theoretical emission curve, and obtains the vehicle trajectory anomaly analysis results; wherein, the first neural network model includes an LSTM model and / or a spiking neural network; the first neural network model using a spiking neural network includes: A spiking neural network is constructed to convert the standardized multi-source detection data into a pulse sequence; Calculate the number of causal segments in the pulse sequence under different spiking neural network parameter configurations, and select the optimal initialization parameters; The spiking neural network is trained using a time-series backpropagation algorithm, and the neuron parameters are dynamically adjusted to obtain the trained spiking neural network. Analyze the pulse propagation path and activation mode in the trained spiking neural network to extract emission time-series features; The theoretical emission curve is generated based on the emission time-series characteristics.
6. The method according to claim 5, characterized in that, The step of training the spiking neural network using the time-series backpropagation algorithm includes: The gradient of the pulse generation function of the spiking neural network is calculated using the surrogate gradient method. Based on the gradient of the function, a loss function is constructed that includes the pulse sequence prediction error term and the causal segment number regularization term; The parameters of the spiking neural network are optimized using the loss function while maintaining the number of causal segments; The parameters of the spiking neural network are quantized and the connections are sparsified to compress the model size of the spiking neural network. Based on preset evaluation metrics, the detection accuracy of the compressed spiking neural network is verified on a test dataset.
7. The method according to claim 5, characterized in that, The analysis of the pulse propagation path and activation mode in the trained spiking neural network, and the extraction of emission time-series features, includes: Extract the neuron activation sequences of the spiking neural network to identify sensitive regions where changes in input cause significant changes in output; Locate the boundaries of causal segments corresponding to the sensitive regions and mark potential anomalous data points; Analyze the emission parameter combinations of the potential abnormal data points to determine the abnormal triggering conditions; Using the aforementioned abnormal triggering conditions as a benchmark, emission time-series features are extracted from the potential abnormal data points.
8. The method according to claim 1, characterized in that, The process of standardizing the multi-source detection data to generate standardized multi-source detection data includes: A physical emission model is built based on vehicle type and engine parameters to generate theoretical emission baseline data; Collect real-time OBD monitoring data and roadside remote sensing data of the vehicle to establish statistical benchmark data; Extract the timestamp information from the multi-source detection data, the OBD real-time monitoring data, and the roadside remote sensing data to construct a time alignment optimization target; Based on the aforementioned time alignment optimization objective, a genetic algorithm is used to optimize the time offset parameters to obtain time-aligned multi-source detection data. Based on the time-aligned multi-source detection data, the theoretical emission baseline data, and the statistical baseline data, a unified monitoring sequence for vehicle emission behavior is constructed to generate standardized multi-source detection data.
9. The method according to claim 8, characterized in that, The process of optimizing the time offset parameters using a genetic algorithm to obtain time-aligned multi-source detection data includes: Construct a chromosome encoding scheme to represent the time offset and scaling factor of the multi-source data; Construct a fitness function based on the mutual information loss of the multi-source detection data to evaluate the time alignment effect; Chromosomes are selected according to the fitness function, and crossover and mutation operations are performed on the selected chromosomes. The selected chromosome is iteratively optimized until the time alignment accuracy reaches the preset accuracy threshold requirement. The optimal time offset parameter is then output to obtain the time-aligned multi-source detection data.
10. A system for intelligent analysis of the authenticity of motor vehicle exhaust emission test data, characterized in that, include: The data receiving module is used to receive vehicle exhaust emission detection data, combine it with vehicle maintenance data and testing equipment calibration logs to obtain multi-source detection data, and perform standardization processing on the multi-source detection data to generate standardized multi-source detection data; wherein, the vehicle exhaust emission detection data includes vehicle VIN code, detection time, pollutant concentration value and testing equipment serial number; The organization analysis module is used to construct an association graph of the detection organizations based on the standardized multi-source detection data, and to use the isolated forest model to detect abnormal subgraphs in the association graph of the detection organizations to obtain the results of organization collaboration anomaly identification. The trajectory analysis module is used to input the target vehicle information from the standardized multi-source detection data into the first neural network model based on the trained first neural network model, generate a theoretical emission curve, calculate the time warping distance between the measured data of the target vehicle and the theoretical emission curve, and obtain the vehicle trajectory anomaly analysis results; wherein, the first neural network model includes an LSTM model and / or a spiking neural network; The early warning module is used to generate early warning information on the authenticity of detection data based on the results of the collaborative anomaly identification of the institution and the results of the vehicle trajectory anomaly analysis.
Citation Information
Cited By
Whole vehicle thermal management simulation method based on pulse neural network
CN121502917A
Intelligent metallurgical process virtual simulation method and system based on digital twinning
CN121525530A
New energy vehicle power prediction monitoring method and system based on Internet of Things
CN121742332A