Hybrid electric vehicle NOx emission prediction method and system based on dictionary learning

By combining t-SNE and dictionary learning, and combining SuperLearner regression model, the problems of low dimensionality reduction efficiency and insufficient accuracy in the NOx emission prediction of hybrid vehicles are solved, and efficient and accurate NOx emission prediction is achieved.

CN120355268AActive Publication Date: 2025-07-22SHANDONG UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510837209.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

In the NOx emission forecast of hybrid vehicles, the incremental data reduction processing is low efficiency, high computational complexity, and insufficient prediction model accuracy, making it difficult to meet the requirements of real-time and accuracy.

Method used

The dimensionality reduction method combined with t-SNE and dictionary learning is adopted, and the nonlinear relationship of high-dimensional data is processed through t-SNE, local and global structures are retained, and key features are extracted by dictionary learning, and combined with SuperLearner regression model, the nonlinear association of variables is accurately learned to improve the accuracy of NOx emission factor correction coefficient.

Benefits of technology

It effectively solves the problems of computational complexity and information redundancy of high-dimensional data, improves prediction timeliness and accuracy, reduces calculation costs, and realizes efficient management and control of NOx emissions of hybrid vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355268A_ABST
    Figure CN120355268A_ABST
Patent Text Reader

Abstract

The invention provides a hybrid electric vehicle NOx emission prediction method and system based on dictionary learning, and belongs to the field of vehicle emission prediction. Obtaining emission related parameters, an initial driving emission high-dimensional data set and incremental test data; mapping the initial driving emission high-dimensional data set into initial low-dimensional embedding through t-SNE; based on initial low-dimensional embedding, constructing a high-dimensional dictionary and a low-dimensional dictionary through dictionary learning, and performing iterative optimization; calculating a sparse coding matrix of incremental test data on the high-dimensional dictionary, and coupling the sparse coding matrix with the low-dimensional dictionary to generate incremental low-dimensional embedding; the increment is embedded into a pre-trained SuperLearner regression model in a low-dimensional mode, and an increment NOx emission factor correction coefficient is obtained; the NOx emissions are calculated based on the emissions-related parameters and the incremental NOx emissions factor correction factor. According to the method, t-SNE and dictionary learning dimension reduction are utilized, and the SuperLearner regression model is combined, so that the NOx emission prediction precision of the hybrid electric vehicle is improved, and incremental data are efficiently processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle emission prediction, and in particular to a method and system for predicting NOx emission of a hybrid electric vehicle based on dictionary learning. Background Art

[0002] As traffic demand grows, the number of cars on the road rises, and mobile sources become the main source of air pollution. Among them, NOx emissions from road vehicles account for a large proportion. Hybrid vehicles have attracted much attention as a new energy solution, but due to their poor NOx control, accurately predicting their NOx emissions has become a key issue.

[0003] In the field of vehicle emission prediction, correction coefficients are crucial to improving prediction accuracy, and the selection or extraction of key parameters is also critical. At present, research is mainly focused on two directions: using advanced algorithms for modeling and using correction coefficients to calibrate emission data. However, existing technologies have many defects. On the one hand, in terms of dimensionality reduction, traditional methods cannot efficiently process incremental data. For example, some dimensionality reduction techniques require recalculation when processing new data, which has high computational costs and low efficiency, and it is difficult to quickly obtain a low-dimensional representation of incremental data. On the other hand, the accuracy of the prediction model needs to be improved. When faced with complex vehicle emission data, some algorithms cannot fully explore the potential nonlinear correlation between variables, resulting in insufficient accuracy of the prediction model, and a large deviation between the predicted and measured values of NOx emissions, which is difficult to meet actual needs. Summary of the invention

[0004] In order to solve the above problems, the present invention proposes a method and system for predicting NOx emissions of hybrid vehicles based on dictionary learning. By combining t-SNE and dictionary learning for dimensionality reduction, t-SNE processes the nonlinear relationship of high-dimensional data and retains local and global structures. Dictionary learning extracts key features and efficiently processes incremental data, thereby avoiding the calculation complexity and information redundancy problems of traditional methods. In addition, the SuperLearner regression model is adopted to integrate the advantages of multiple algorithms, accurately learn the nonlinear correlation of variables, and improve the accuracy of the NOx emission factor correction coefficient, thereby improving the prediction accuracy of NOx emissions.

[0005] In order to achieve the above object, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning, comprising: Obtain emission-related parameters, initial driving emission high-dimensional data set and incremental test data; Mapping the initial driving emission high-dimensional dataset into an initial low-dimensional embedding via t-SNE; Based on the initial low-dimensional embedding, construct a high-dimensional dictionary and a low-dimensional dictionary through dictionary learning, and perform iterative optimization; calculate the sparse coding matrix of the incremental test data on the high-dimensional dictionary, and couple it with the low-dimensional dictionary to generate an incremental low-dimensional embedding; Input the incremental low-dimensional embedding into the pre-trained SuperLearner regression model to obtain the incremental NOx emission factor correction coefficient; Calculate the NOx emissions based on the emission-related parameters and the incremental NOx emission factor correction coefficient.

[0006] In a second aspect, the present invention provides a hybrid vehicle NOx emissions prediction system based on dictionary learning, including: A data acquisition module, configured to acquire emission-related parameters, an initial driving emission high-dimensional data set, and incremental test data; A data dimensionality reduction module, configured to map the initial driving emission high-dimensional data set into an initial low-dimensional embedding through t-SNE; A dictionary construction and optimization module, configured to construct a high-dimensional dictionary and a low-dimensional dictionary through dictionary learning based on the initial low-dimensional embedding, and perform iterative optimization; calculate the sparse coding matrix of the incremental test data on the high-dimensional dictionary, and couple it with the low-dimensional dictionary to generate an incremental low-dimensional embedding; A correction coefficient acquisition module, configured to input the incremental low-dimensional embedding into the pre-trained SuperLearner regression model to obtain the incremental NOx emission factor correction coefficient; An emission calculation module, configured to calculate the NOx emissions based on the emission-related parameters and the incremental NOx emission factor correction coefficient.

[0007] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning described in the first aspect are implemented.

[0008] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning described in the first aspect are implemented.

[0009] Compared with the prior art, the beneficial effects of the present invention are: (1) After obtaining various types of test data, the present invention uses t-SNE and dictionary learning for dimensionality reduction, effectively solving the problems of computational complexity and information redundancy caused by high-dimensional data, extracting key features, and reducing the computational cost. By iteratively optimizing the mapping relationship between dictionaries, it can efficiently process incremental data and improve the prediction timeliness. The SuperLearner regression model combines the dimensionality-reduced data to accurately learn the variable relationship and obtain a high-precision NOx emission factor correction coefficient. Finally, based on the correction coefficient and emission-related parameters, the NOx emission is calculated, significantly improving the prediction accuracy. Compared with traditional methods, the present invention takes into account both data processing efficiency and prediction accuracy, providing strong support for the effective management and control of NOx emissions in hybrid vehicles.

[0010] (2) To better predict the NOx emission factor of hybrid vehicles, the present invention applies the incremental dimensionality reduction method based on dictionary learning to the problem of processing multivariate variables affecting NOx emissions. The high-dimensional data obtained from RDE tests is embedded into a low-dimensional space through t-SNE, and then based on these low-dimensional embeddings, a high-dimensional dictionary and a low-dimensional dictionary are constructed using dictionary learning. For incremental data, by calculating its sparse coding matrix on the high-dimensional dictionary and coupling it with the pre-constructed low-dimensional dictionary, the low-dimensional embedding corresponding to the incremental data can be quickly obtained, realizing efficient dimensionality reduction through simple matrix operations.

[0011] (3) After completing dimensionality reduction, the present invention uses the low-dimensional data obtained by t-SNE dimensionality reduction to train the SuperLearner regression model to further construct an accurate NOx emission correction coefficient prediction model. The incremental low-dimensional embeddings obtained by the incremental dimensionality reduction method based on dictionary learning are input into the SuperLearner prediction model, and the corresponding correction coefficients can be directly obtained, thereby further improving the accuracy of NOx emission prediction.

[0012] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute a limitation to the present invention.

[0014] Figure 1 It is the main flowchart of a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning provided by an embodiment of the present invention; Figure 2 It is the schematic flowchart of a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning provided by an embodiment of the present invention; Figure 3 Schematic diagram of RDE test driving conditions provided by an embodiment of the present invention; Figure 4 Flowchart of an incremental dimensionality reduction method based on dictionary learning provided by an embodiment of the present invention. Detailed implementation manners

[0015] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0016] Embodiment 1 As Figure 1 shown, this embodiment discloses a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning, including the following steps: S1: Obtain emission-related parameters, an initial high-dimensional dataset of driving emissions, and incremental test data; S2: Map the initial high-dimensional dataset of driving emissions to an initial low-dimensional embedding through t-SNE; S3: Based on the initial low-dimensional embedding, construct a high-dimensional dictionary and a low-dimensional dictionary through dictionary learning, and perform iterative optimization; calculate the sparse coding matrix of the incremental test data on the high-dimensional dictionary, and couple it with the low-dimensional dictionary to generate an incremental low-dimensional embedding; S4: Input the incremental low-dimensional embedding into a pre-trained SuperLearner regression model to obtain an incremental NOx emission factor correction coefficient; S5: Calculate the NOx emissions based on the emission-related parameters and the incremental NOx emission factor correction coefficient.

[0017] Next, in combination with Figure 2 , a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning disclosed in this embodiment will be described in detail.

[0018] Currently, the prior art has proposed a method for calculating the instantaneous NOx emission factor. This method calculates the instantaneous NOx emission factor based on the emission-related parameters in the actual driving emission test data, that is, the NOx volume concentration, engine exhaust rate, engine fuel rate, and fuel flow rate data collected every second, as shown in Equation (1):

[0019] In the formula, is the instantaneous NOx emission factor (g / s) of the vehicle, which is used to characterize the amount of NOx emitted by the vehicle per unit time at a certain moment; is the NOx concentration (ppm) downstream of SCR (selective catalytic reduction); is the engine exhaust rate (kg / h); is the engine fuel rate (kg / h); is the engine fuel flow rate (L / h); is the gasoline density, set to 0.73 kg / L; M is the molecular weight of NOx. Since a high percentage of NO is converted to at a relatively fast rate under ambient conditions, its molecular weight is considered to be the same as that of , which is 46 g; 1-losses represents a correction factor obtained by validating the calculation results using RDE data; is the exhaust density, assumed to be approximately the same as the air density at 29 g / mol.

[0020] Although the conventional technology involves introducing a correction factor to calibrate the NOx emission results, there are still problems in the calculation of the correction factor: on the one hand, traditional methods often calculate based on a small number of characteristic variables, ignoring a large number of potential factors related to tail gas emissions, such as the subtle operating differences of the engine under different working conditions, etc., resulting in inaccurate correction factors. On the other hand, with the continuous accumulation of test data, the data volume is constantly increasing and the characteristic dimension is relatively high. Existing calculation methods have extremely high computational complexity and low efficiency when dealing with high-dimensional data, and it is difficult to meet the requirements of real-time and accuracy.

[0021] Therefore, in the actual driving emission test data of this embodiment, a high-dimensional data set covering multi-dimensional data such as rich vehicle information and vehicle speed is extracted. However, if it is directly used for the calculation of the correction factor, not only will it consume a huge amount of computing resources, but also the reliability of the results will be affected due to the curse of dimensionality problem. The t-SNE, a non-linear dimensionality reduction technology, can better handle the non-linear characteristics of data, balance the local and global structures of data, and dictionary learning can effectively retain key information. Therefore, this embodiment proposes a dimensionality reduction method combining t-SNE and dictionary learning to provide a more efficient and reliable data processing basis for accurately calculating the correction factor.

[0022] First, the initial driving emission high-dimensional data set described in this embodiment is obtained based on the RDE test data set of 10 hybrid vehicles. First, the high-dimensional data is used for dimensionality reduction to obtain an initial low-dimensional embedding, and then the initial low-dimensional embedding is used to assist the dimensionality reduction of the later incremental data.

[0023] The data of the high-dimensional data set includes engine characteristics, driving characteristics, and external characteristic data under different test conditions. The test conditions are divided into three types: urban (v ≤ 60 km / h), suburban (60 km / h < v ≤ 90 km / h), and highway (v > 90 km / h), as Figure 3 shown.

[0024] A parameter database related to NOx emissions is established. Specifically, the engine characteristics include four items: coolant temperature, exhaust temperature, engine speed, and intake temperature; the driving characteristics cover vehicle specific power (VSP), vehicle speed, and acceleration; the external characteristics include ambient temperature, ambient pressure, and road gradient. The above variables are all collected in real time at a frequency of once per second in the RDE test, and the parameter information is summarized in Table 1.

[0025] Table 1 Parameter information;

[0026] Preliminary processing of the data reveals that some key parameters are invalid. Therefore, preprocessing and quality control are performed on the dataset to verify the validity of the real-time measurement data and ensure that all values are non-negative to guarantee the reliability and usability of the dataset.

[0027] It should be understood that the methods of the preprocessing include missing value processing, outlier processing, data transformation, non-numeric data encoding processing, and numeric data standardization processing, etc. These methods can be selected by those skilled in the art according to the actual situation and are not limited in this embodiment.

[0028] Furthermore, since the RDE test dataset used in this embodiment contains many feature variables, directly using the original data for analysis not only has high computational complexity and low efficiency, but may also affect the accuracy and reliability of the analysis results due to problems such as multicollinearity among features. Therefore, appropriate dimensionality reduction techniques must be adopted to extract representative feature information. According to the non-linear characteristics of the data, t-SNE is selected as the main dimensionality reduction method in this embodiment.

[0029] t-SNE is a typical non-linear dimensionality reduction technique that can map high-dimensional data to a low-dimensional space while better preserving the local structure and revealing the global structure of the data. t-SNE converts the Euclidean distance between data points in the high-dimensional space into a probability distribution, so that similar points are given higher probabilities. Subsequently, another set of probability distributions is constructed in the low-dimensional space, and dimensionality reduction is achieved by minimizing the Kullback-Leibler divergence between the high- and low-dimensional probability distributions. To improve the robustness to outliers, t-SNE adopts a symmetric joint probability distribution, which not only simplifies the gradient calculation but also speeds up the optimization process. At the same time, in the low-dimensional space, t-SNE no longer uses the traditional Gaussian distribution but instead uses the Student-t distribution with a degree of freedom of 1. This distribution can effectively alleviate the "crowding problem" in the mapping from high dimensions to low dimensions, thus better revealing the global structure of the data. Given that the RDE test dataset in this embodiment has non-linear characteristics, and t-SNE can effectively balance the local and global structures of the data, t-SNE is therefore adopted as the main dimensionality reduction method.

[0030] In this embodiment, t-SNE is used to map a high-dimensional data set into a low-dimensional data set to achieve a low-dimensional embedding for extracting key features. As a typical non-linear dimensionality reduction technique, t-SNE can effectively handle the complex non-linear relationship between parameters and NOx emissions, and the t-distribution it adopts can also mitigate the influence of noise in RDE test data.

[0031] In the actual research scenario of hybrid vehicles, as time goes by, test conditions change, or to obtain more comprehensive data to improve the accuracy and generalization ability of the model, new real driving emission (RDE) test data will continuously be generated. For example, driving tests at different time periods and in different regions, tests after the vehicle has traveled different mileage, etc. may all generate new data. In this embodiment, the newly input test data is called incremental data.

[0032] Incremental data also needs to be dimensionally reduced, but conventional techniques have some limitations. Traditional dimensionality reduction methods, such as linear dimensionality reduction techniques like principal component analysis (PCA), do not perform well in dealing with complex non-linear data structures, are difficult to fully retain the key features of the data, and the calculation process is relatively complex and inefficient. When faced with a large amount of incremental data, these methods may cause information loss, thereby affecting the accuracy of subsequent model predictions. Therefore, this embodiment introduces dictionary learning, which can extract a set of representative basic elements or features from the original data to encapsulate the data features and achieve a low-dimensional representation of the entire data set. For incremental data, only the pre-learned dictionary needs to be used to calculate its sparse coding matrix, without reconstructing the complete dictionary each time, significantly shortening the calculation time and reducing the overall calculation cost, providing an efficient solution for the dimensionality reduction of incremental data.

[0033] Specifically, the goal of dictionary learning is to reconstruct the original data as a linear combination of a small number of dictionary atoms, effectively eliminating redundant information in the data. The original data can be imagined as a complex jigsaw puzzle, and each data point is a piece of the puzzle. Dictionary atoms are like some basic jigsaw puzzle modules. By combining these basic modules in different proportions (corresponding to the coefficients in the sparse coefficient matrix), an approximate representation of the original data can be reconstructed. The purpose of doing this is to use a small number of basic modules (dictionary atoms) to replace a large number of complex data points, and reconstruct the image through the linear combination of these feature blocks to extract key features, thereby achieving the effect of simplifying the data and eliminating redundant information.

[0034] The incremental dimensionality reduction method based on dictionary learning used in this embodiment has its core in combining an efficient iterative algorithm and a dynamic dictionary atom update strategy to accelerate the training process while maintaining accuracy, and is mainly divided into the following parts: (1) Initialization phase The general model of dictionary learning can be expressed as follows:

[0035] In the formula, X represents a high-dimensional data set; is the high-dimensional dictionary; C represents the sparse coefficient matrix, which is used to describe how each dictionary atom in the high-dimensional dictionary is combined with what weight (coefficient) to reconstruct the high-dimensional data set X; ε is the relative error.

[0036] Suppose is the low-dimensional data set obtained by dimensionality reduction of X, is the corresponding low-dimensional data dictionary. For a given data point , let represent its encoding in the high-dimensional dictionary , corresponding to the i-th column of the encoding matrix. Applying t-SNE to X gives the low-dimensional data set , and randomly selecting K samples from X to initialize the high-dimensional dictionary .

[0037] (2) Sparse coding stage In this stage, the high-dimensional dictionary remains fixed, and the principle of the Iterative Reweighted FISTA (IR-FISTA) algorithm is applied to solve the sparse coding matrix C. The core of the IR-FISTA algorithm is to introduce the concept of a working set and only perform accelerated iterative thresholding shrinkage (FISTA) operations on the sparse coefficients corresponding to two types of indices (in the matrix C, each element corresponds to a unique index ): ① The current non-zero sparse indices; ② The current zero coefficients whose gradients are large and may become non-zero in the next iteration.

[0038] This process avoids unnecessary calculations for the large number of parts that are stably zero in the sparse coding.

[0039] Specifically, the variables in the working set are solved as follows:

[0040] Among them, represents the intermediate value of the sparse coding matrix C at the k-th iteration; represents the value of the sparse coding matrix C at the k-th iteration; , represent the step size parameters at the k-th and k-1-th iterations respectively, , ; is the soft threshold operator, represents the step size; represents the regularization parameter; Denotes the independent variable parameter in the soft threshold operator; Denotes the sign function, Denotes the transpose.

[0041] In the sparse coding stage, the IR-FISTA algorithm is adopted in this embodiment to efficiently solve the sparse coding matrix C. Only the current non-zero coefficient indices and the zero coefficient indices with larger gradients are subjected to iterative operations, avoiding redundant calculations for a large number of stable zero-valued coefficients. When dealing with high-dimensional incremental data, the sparse coding matrix can be quickly generated, shortening the time-consuming of dimensionality reduction.

[0042] (3) Dictionary update stage In this stage, the sparse coding matrix C is fixed, and the high-dimensional dictionary is calculated by combining the learning rate decay and the dynamic update strategy based on atomic importance . This dynamic update strategy calculates the sum of the absolute values corresponding to a certain dictionary atom in the sparse matrix C as the "importance" of this dictionary atom. For the j-th atom of the dictionary , its importance

[0043] is defined as: where N is the total number of data samples, is the sparse coefficient corresponding to the j-th dictionary atom when representing the i-th data sample.

[0044] Meanwhile, on this basis, a learning rate decay mechanism is introduced, and the exponential decay form is adopted to make the base learning rate

[0045] decay as the number of dictionary learning training iterations increases, avoiding significant oscillations in the optimization process due to overly large parameter update amplitudes in the later stage of iterative training, which thus hinders the model from converging stably to the optimal solution. The mathematical expression of exponential decay is shown as follows: where is the initial base learning rate, is the decay coefficient,

[0046] represents the exponential decay factor. In each dictionary update iteration, based on the learning rate decay mechanism and the atomic importance of the dictionary, the learning rate allocation strategy is optimized, and a specific learning rate

[0047] is assigned to each atom: where is the maximum value of the atomic importance,

[0048] With the above learning rate allocation strategy, atoms with higher "importance" obtain larger learning rates to adjust to the optimal state faster; atoms with lower importance use smaller learning rates to avoid unnecessary calculations. This dynamic update strategy enables the dictionary update process to balance the needs of rapid exploration and stable and accurate convergence, thus optimizing the learning efficiency of the high-dimensional dictionary. of the learning efficiency.

[0049] Repeat the sparse coding stage and the dictionary update stage until the reconstruction error is lower than the preset threshold or the maximum number of iterations is reached. The detailed calculation process is as Figure 4 shown.

[0050] In this embodiment, through the learning rate decay and the dynamic strategy of atomic importance, the learning efficiency and adaptability of the high-dimensional dictionary are optimized. The atomic importance statistic is the sum of the absolute values of the corresponding coefficients of each atom in the sparse matrix, enabling key atoms that contribute greatly to data reconstruction to preferentially obtain larger learning rates and accelerate convergence to the optimal state; while the learning rate decay mechanism reduces the update step size in the later stage of iteration in an exponential form, avoiding oscillations in the optimization process and ensuring the stable convergence of the dictionary. This dynamic update strategy enables the dictionary to adaptively capture the key features in the RDE test data, and the generated dictionary atoms are more representative. Furthermore, when performing incremental dimensionality reduction, there is no need to reconstruct the dictionary. Only through the coupling of the pre-learned dictionary and sparse coding can new data be quickly processed, reducing the computational cost while enhancing the model's ability to represent complex data.

[0051] (4) Incremental dimensionality reduction stage Based on the iterative optimization process of dictionary learning, the solution formula for the low-dimensional data dictionary can be expressed as:

[0052] In the formula, the sparse coefficient matrix C plays a role in connecting the high-dimensional data and the low-dimensional data dictionary. Through it and the result of the high-dimensional data after dimensionality reduction to calculate the low-dimensional data dictionary.

[0053] When new data is obtained, first calculate its corresponding coding matrix on the high-dimensional dictionary; then couple this coding matrix with the low-dimensional dictionary to achieve rapid dimensionality reduction of the new data and obtain the low-dimensional embedding of the incremental data. The incremental dimensionality reduction calculation formula is shown in Equation (4):

[0054] where, represents the coding matrix of the incremental test data on its high-dimensional dictionary .

[0055] Dictionary learning can extract a set of representative "atoms" from the original data to encapsulate data features, thereby achieving a low-dimensional representation of the entire dataset. For incremental data, only the pre-learned dictionary needs to be used to calculate its sparse coding matrix, without reconstructing the complete dictionary each time. This method significantly shortens the calculation time and reduces the overall calculation cost, providing an efficient solution for the dimensionality reduction of incremental data.

[0056] As an implementation, dictionary learning can learn each column of dictionary atoms from the initial dataset, and these learned dictionary atoms can fully extract the feature information in the data. Therefore, the method based on dictionary learning can generate a dictionary with strong adaptability and data feature representation ability. The dimensionality reduction method based on dictionary learning can be simply described as: using high-dimensional data and its corresponding low-dimensional data as training samples, establishing the mapping relationship between the two, and applying this mapping to subsequent new data to quickly obtain its corresponding low-dimensional representation.

[0057] t-SNE is good at revealing the non-linear structure of data, while dictionary learning can extract sparse representations from data. Combining the two in this embodiment can more comprehensively capture the emission pattern, remove noise and redundant information, and retain the features most relevant to NOx emissions.

[0058] In this embodiment, t-SNE is used to reduce the dimensionality of the original high-dimensional dataset. The obtained low-dimensional embedding is used not only to construct a low-dimensional dictionary through dictionary learning, but also as the input of the regression prediction model to predict the correction coefficient of the NOx emission factor.

[0059] In predictive modeling problems, it is crucial to select the appropriate machine learning algorithm and the optimal parameters for a specific dataset. The SuperLearner regression algorithm avoids the problem of only being able to choose a single best model by integrating multiple basic machine learning algorithms. First, SuperLearner performs k-fold cross-validation on the data, evaluates the performance of each algorithm on the same data partition, and retains the prediction results of all basic models. Subsequently, these prediction results are used to train the meta-learner, and the meta-learner synthesizes the outputs of each model to determine the final optimal model.

[0060] To evaluate the performance of the SuperLearner regression method, this embodiment randomly generates a test dataset containing 1000 samples (rows) and 100 features (columns) to compare multiple models including the SuperLearner algorithm. The results in Table 2 show that the performance of SuperLearner can be comparable to that of the best basic model, and may even be better than all single basic models. Therefore, this embodiment selects SuperLearner as the method for constructing the NOx correction coefficient model.

[0061] Table 2 Evaluation of SuperLearner regression results;

[0062] Therefore, in this embodiment, the SuperLearner regression model is trained using the initial low-dimensional embedding based on t-SNE dimensionality reduction. As Figure 2 shown, where represents the low-dimensional dataset obtained by applying t-SNE to the high-dimensional dataset, and represent the elements corresponding to different dimensions in this low-dimensional dataset; represents the SuperLearner regression model, corresponding to the model training stage, indicating learning the prediction relationship between low-dimensional features and the correction coefficient of the NOx emission factor.

[0063] The training process of the SuperLearner regression model specifically includes: 1. Data preparation: Collect historical actual driving emission (RDE) test data of hybrid vehicles, including emission-related parameters, the historical high-dimensional dataset of driving emissions, and the measured NOx emissions.

[0064] Based on the emission-related parameters in the actual driving emission test data, calculate the theoretical NOx emission factor according to formula (1); Then, based on the historical measured NOx emissions and the theoretical NOx emission factor, derive the correction coefficient label based on the ratio.

[0065] Take the correction coefficient of the theoretical NOx emission factor as the label, and divide the low-dimensional embedding of the historical high-dimensional dataset after t-SNE dimensionality reduction into a training set and a test set according to a certain ratio. The training set is used for model training, and the test set is used for evaluating the model performance.

[0066] 2. Selection and training of basic models: Select appropriate algorithms from various basic machine learning algorithms, such as linear regression, decision tree regression, support vector regression, etc. Use the training set data to train each basic model respectively to obtain the trained basic models.

[0067] 3. k-fold cross-validation: Adopt the k-fold cross-validation method to divide the training set data into k non-overlapping subsets. For each basic model, sequentially use k - 1 of these subsets as training data, and the remaining 1 subset as validation data for training and validation, repeat k times, and record the prediction results of each validation.

[0068] 4. Meta - learner training: Integrate the prediction results of all base models in k - fold cross - validation as the input data for the meta - learner. Use this data to train the meta - learner and adjust the parameters of the meta - learner through an optimization algorithm so that it can synthesize the prediction results of each base model to obtain the final SuperLearner regression model.

[0069] 5. Model evaluation and optimization: Evaluate the performance of the trained SuperLearner regression model using the test - set data, and judge the accuracy and goodness of fit of the model by calculating metrics such as RMSE and so on. If the model performance does not meet the expectations, the selection of base models, the number of folds in k - fold cross - validation, or the parameters of the meta - learner can be adjusted, and re - training and evaluation can be carried out until the model performance meets the requirements.

[0070] Input the low - dimensional embedding of the incremental data into the trained SuperLearner regression model to obtain the corresponding incremental NOx emission factor correction coefficient. As Figure 2 shown, represents the low - dimensional data set obtained from the dictionary learning of the incremental high - dimensional data set, and represent the elements corresponding to different dimensions in this low - dimensional data set. Corresponding to the application stage after the model training is completed, the incremental NOx emission factor correction coefficient is obtained based on the low - dimensional embedding of the incremental data.

[0071] Furthermore, to verify the accuracy and effectiveness of the NOx emission calculation model, in this embodiment, the EF NOx (g / s) calculated using formula (1) will be compared with the instantaneous measured EF NOx (g / s) collected from ten HEVs during RDE tests. At this time, the coefficient (1 - losses) in formula (1) takes the value of 1, that is, no correction is performed. Through the least - squares fitting, and the coefficient of determination and the slope of the fitting curve equation are used to evaluate the overall correlation between the calculated NOx emissions and the measured values.

[0072] Table 3 lists the correlation between the calculated NOx emissions and the measured values without considering the correction coefficient, and the fitting degree between the (g / s) calculated by the formula and the measured (g / s) is quantified by parameters K, and RMSE. The parameter K is defined as the slope of the fitting curve obtained by the least - squares fitting of the calculated NOx emission rate and the measured value for a single HEV during RDE tests, and is specifically expressed as follows:

[0073] As can be seen from Table 3, there are obvious deviations between the calculated values and the measured values of NOx emissions of ten HEVs. The lowest is only 0.879.

[0074] Table 3 Evaluation indicators of ten hybrid vehicles;

[0075] Therefore, it is necessary to further calibrate by optimizing the correction coefficient in formula (1), and determine its specific value based on the multivariate variables in the RDE test data, so that the parameters K and approach the ideal value of 1. The advantages of the dimensionality reduction and correction process proposed in this embodiment are reflected in two aspects: First, it avoids the complex processing process in traditional dimensionality reduction methods, and the corresponding low-dimensional data can be obtained only through simple matrix calculations, significantly reducing the calculation time; Second, the excellent modeling ability of the SuperLearner regression method makes the correction coefficient prediction model more accurate, effectively correcting the predicted NOx emissions corresponding to the incremental data.

[0076] Based on formula (1), replacing "1-losses" with the incremental NOx emission factor correction coefficient, a more accurate NOx emission is calculated.

[0077] As an implementation, this embodiment uses three independent RDE test data sets to evaluate the performance of the proposed method that combines incremental dimensionality reduction based on dictionary learning and a prediction model based on SuperLearner. The evaluation indicators include the parameter K, and RMSE, to comprehensively evaluate the method performance from multiple dimensions.

[0078] Table 4 shows the correlation between the calculated values and the measured values of NOx emissions calculated from three HEV test data sets without introducing any correction coefficients; while Table 5 shows the accuracy of the corrected NOx emission values after adopting the proposed systematic method. The results show that applying this systematic method can significantly improve the evaluation indicators of the correlation between the predicted NOx emissions and the measured values. Specifically, the parameter K generally increases from about 0.8 - 0.9 to the range of 1 ± 0.05, increases from the lowest 0.956 to above 0.995. In addition, the RMSE also shows significant improvement.

[0079] Table 4 Evaluation indicator values of the test data set before correction;

[0080] Table 5 Evaluation indicator values of the test data set after correction;

[0081] These results show that the method described in this embodiment combines the incremental dimensionality reduction technology based on dictionary learning with the SuperLearner prediction model. It not only realizes the rapid dimensionality reduction of incremental data through simple matrix operations, simplifies the calculation process, but also can accurately predict the correction coefficient, obtaining a more accurate NOx emission estimation value compared with the uncorrected method.

[0082] In view of the deviation between the calculated value and the measured value of NOx emissions in this specific embodiment, an optimization method based on the correction coefficient is proposed. Since the number of characteristic variables obtained from the RDE test is large, this embodiment uses t-SNE to reduce the dimensionality of high-dimensional features, constructs a high-low dimensional dictionary mapping through dictionary learning, and uses simple matrix operations to quickly map the incremental data to the low-dimensional space. At the same time, a SuperLearner regression model is trained based on the t-SNE low-dimensional embedding to achieve accurate prediction of the correction coefficient from low-dimensional parameters. For incremental data, the low-dimensional embedding obtained by the incremental dimensionality reduction method based on dictionary learning is directly input into the correction coefficient model, and the corresponding NOx emission factor correction coefficient can be quickly generated, thus effectively ensuring the speed and accuracy of emission prediction.

[0083] Embodiment II This embodiment provides a NOx emission prediction system for hybrid electric vehicles based on dictionary learning, including: A data acquisition module, configured to acquire emission-related parameters, an initial high-dimensional driving emission dataset, and incremental test data; A data dimensionality reduction module, configured to map the initial high-dimensional driving emission dataset to an initial low-dimensional embedding through t-SNE; A dictionary construction and optimization module, configured to construct a high-dimensional dictionary and a low-dimensional dictionary through dictionary learning based on the initial low-dimensional embedding, and perform iterative optimization; calculate the sparse coding matrix of the incremental test data on the high-dimensional dictionary, and couple it with the low-dimensional dictionary to generate an incremental low-dimensional embedding; A correction coefficient acquisition module, configured to input the incremental low-dimensional embedding into a pre-trained SuperLearner regression model to obtain an incremental NOx emission factor correction coefficient; An emission calculation module, configured to calculate the NOx emissions based on the emission-related parameters and the incremental NOx emission factor correction coefficient.

[0084] Embodiment III This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a method for predicting NOx emissions of hybrid electric vehicles based on dictionary learning as described in Embodiment I above.

[0085] Embodiment IV This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning as described in Embodiment 1 above.

[0086] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation details, refer to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to execute any method in the present invention.

[0087] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A prediction method for NOx emissions of a hybrid vehicle based on dictionary learning, characterized in that, Including: Obtain emission-related parameters, an initial high-dimensional driving emission dataset, and incremental test data; Map the initial high-dimensional driving emission dataset into an initial low-dimensional embedding through t-SNE; Based on the initial low-dimensional embedding, construct a high-dimensional dictionary and a low-dimensional dictionary through dictionary learning, and perform iterative optimization; Calculate the sparse coding matrix of the incremental test data on the high-dimensional dictionary, and couple it with the low-dimensional dictionary to generate an incremental low-dimensional embedding; Input the incremental low-dimensional embedding into a pre-trained SuperLearner regression model to obtain an incremental NOx emission factor correction coefficient; Calculate the NOx emission based on the emission-related parameters and the incremental NOx emission factor correction coefficient.

2. The NOx emission prediction method for a hybrid vehicle based on dictionary learning according to claim 1, wherein The initial high-dimensional driving emission dataset includes engine characteristic data, driving characteristic data, and external characteristic data collected under different test conditions, and the test conditions include urban, suburban, and highway.

3. A prediction method for NOx emissions of a hybrid vehicle based on dictionary learning according to claim 1, characterized in that, The mapping of the initial high-dimensional driving emission dataset into an initial low-dimensional embedding through t-SNE specifically includes: Convert the Euclidean distance between data points in the initial high-dimensional driving emission dataset into a joint probability distribution in the high-dimensional space, so that similar data points are assigned a higher joint probability; Construct another set of joint probability distributions in the low-dimensional space, where the distance between points in the low-dimensional space is modeled by the Student-t distribution with a degree of freedom of 1; Use a symmetric joint probability distribution to represent the relationship between data points in the high-dimensional space and the low-dimensional space; Minimize the KL divergence between the joint probability distributions in the high-dimensional space and the low-dimensional space, and iteratively update the point positions in the low-dimensional space through a gradient optimization algorithm; When the divergence converges or reaches a preset number of iterations, output the final low-dimensional embedding representation, and the low-dimensional embedding representation retains the local structure and global structure features of the high-dimensional data.

4. A method for predicting the NOx emissions of a hybrid vehicle based on dictionary learning according to claim 1, characterized in that, The construction of a high-dimensional dictionary and a low-dimensional dictionary through dictionary learning based on the initial low-dimensional embedding and the iterative optimization specifically include: Based on the general model of dictionary learning, randomly sample from the high-dimensional dataset to initialize the high-dimensional dictionary; Based on the iterative reduction to optimize the sparse coding matrix, and adopt a learning rate decay and atomic importance dynamic update strategy to optimize the high-dimensional dictionary; the iterative reduction is used to perform iterative calculations only on the non-zero coefficient indices and zero coefficient indices in the sparse coding matrix whose gradients exceed a preset threshold; Based on the optimized sparse coding matrix and the initial low-dimensional embedding, obtain the low-dimensional dictionary.

5. The method for predicting the NOx emission of a hybrid vehicle based on dictionary learning according to claim 4, wherein, The adoption of a learning rate decay and atomic importance dynamic update strategy to optimize the high-dimensional dictionary specifically includes: Measure the atomic importance by statistically summing the absolute values of the corresponding atomic coefficients in the sparse coding matrix; Introduce a learning rate decay mechanism, and let the base learning rate decay exponentially with the increase of the number of iterations to avoid oscillations in the later stage of iteration; At each iteration, assign a learning rate to each atom according to the learning rate decay mechanism and atomic importance to achieve the optimization of the high-dimensional dictionary.

6. The method for predicting NOx emissions of a hybrid vehicle based on dictionary learning according to claim 1, wherein, The training process of the SuperLearner regression model includes: Calculate the initial NOx emission factor based on the emission-related parameters; obtain the initial NOx emission factor correction coefficient according to the ratio of the measured emission data and the initial NOx emission factor; The SuperLearner regression model is trained based on the initial low-dimensional embedding and the initial NOx emission factor correction coefficient.

7. The NOx emission prediction method for a hybrid vehicle based on dictionary learning according to claim 1, wherein, The NOx emissions are calculated based on the emission-related parameters and the incremental NOx emission factor correction coefficient, specifically as follows: ; Among them, is the instantaneous NOx emission factor of the vehicle, which is used to characterize the amount of NOx emitted by the vehicle per unit time at a certain moment; is the NOx concentration downstream of the selective catalytic reduction; is the engine exhaust rate; is the engine fuel rate; is the engine fuel flow rate; is the gasoline density; M is the molecular weight of NOx; 1-losses is the correction coefficient; is the exhaust density.

8. A prediction system for NOx emissions of a hybrid vehicle based on dictionary learning, characterized in that, Including: A data acquisition module, configured to acquire emission-related parameters, an initial high-dimensional driving emission dataset, and incremental test data; A data dimensionality reduction module, configured to map the initial high-dimensional driving emission dataset into an initial low-dimensional embedding by t-SNE; A dictionary construction and optimization module, configured to construct a high-dimensional dictionary and a low-dimensional dictionary through dictionary learning based on the initial low-dimensional embedding, and perform iterative optimization; Calculate the sparse coding matrix of the incremental test data on the high-dimensional dictionary, and couple it with the low-dimensional dictionary to generate an incremental low-dimensional embedding; A correction coefficient acquisition module, configured to input the incremental low-dimensional embedding into a pre-trained SuperLearner regression model to obtain an incremental NOx emission factor correction coefficient; An emission calculation module, configured to calculate the NOx emissions based on the emission-related parameters and the incremental NOx emission factor correction coefficient.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning as described in any one of claims 1-7.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a method for predicting NOx emissions of a hybrid vehicle based on dictionary learning as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Intrusion detection method based on improved dictionary learning

    CN106991435A

  • A remote sensing image retrieval method based on nonlinear dimension reduction and sparse representation

    CN109815357A

  • High-dimensional image data dimension reduction method based on manifold mapping and dictionary learning

    CN110648276A

  • Multi-modal machining center signal reconstruction method based on online dictionary learning

    CN116070091A

  • Carbon emission prediction method and system based on ISWO-LSTM model

    CN118350513A