Vehicle test data feature extraction method and system
By dynamically adjusting the sliding window and adaptive weight feature extraction method, the problems of low efficiency and redundant features in vehicle test data processing are solved, and more efficient and accurate feature extraction and data analysis are achieved to adapt to complex vehicle test environments.
Patent Information
- Application Number
- CN202510843554.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-03
AI Technical Summary
Existing vehicle test data processing relies on manual feature analysis with low efficiency; traditional feature selection methods are difficult to adapt to dynamically changing detection scenarios; fixed feature extraction processes lead to excessive redundant features; there is a lack of dynamic response mechanisms to data statistical characteristics and time series changes; and feature dimensions explode when multi-source heterogeneous data are fused and processed.
A feature extraction method that dynamically adjusts the sliding window and adaptively adjusts the weights based on scene information is adopted. Through preprocessing, time-frequency feature extraction, cross-feature construction, scoring using multiple evaluation methods, weighted summation and dimensionality reduction, low-scoring features are eliminated to form a closed-loop optimization.
It improves the intelligence and efficiency of the feature extraction process, adapts to the complex and changing vehicle testing environment, improves the accuracy and reliability of data-driven decision-making, reduces redundant computing and storage costs, and improves the speed and accuracy of model training.
Smart Images

Figure CN120744448A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle test data processing, and in particular to a method and system for extracting vehicle test data features. Background Art
[0002] With the rapid development of intelligent, connected, and electrified vehicles in the automotive industry, vehicle driving is gradually evolving towards autonomous driving. This requires that vehicles be pre-installed with autonomous driving software and hardware before leaving the factory, such as lidar, augmented reality head-up display (AR-HUD), millimeter-wave radar, and domain controllers. To ensure that vehicles have complete functionality and meet quality standards, full-function automated testing must be performed at the OEM and massive amounts of test data must be collected for automated analysis. However, the current collection and analysis of vehicle test data presents the following problems: (1) Existing vehicle test data processing relies on manual feature analysis, which is inefficient and subject to subjective bias; (2) Traditional feature selection methods (such as principal component analysis (PCA) and linear discriminant analysis (LDA)) are difficult to adapt to dynamically changing detection scenarios; (3) The fixed feature extraction process leads to excessive redundant features, which affects the efficiency of model training; (4) Lack of dynamic response mechanism to data statistical characteristics and time series changes; (5) The problem of feature dimension explosion is prominent when fusion processing multi-source heterogeneous data. Summary of the Invention
[0003] The present application provides a vehicle test data feature extraction method and system, which can solve the technical problem in the prior art of poor test data processing effect obtained by functional testing of vehicles off the production line of vehicle manufacturers.
[0004] In a first aspect, an embodiment of the present application provides a method for extracting features from vehicle test data, the method comprising: Preprocess the original test data to obtain intermediate test data that meets the preset format standards; A preset feature extraction strategy is used to extract time-frequency features from the intermediate test data, and a cross-feature is constructed based on the time-frequency features. The preset feature extraction strategy includes extracting time-domain features from the intermediate test data using a first sliding window. The size of the first sliding window is dynamically adjusted based on the data fluctuation variance, and the window is expanded when the data fluctuation variance is greater than the variance threshold, and is reduced when it is less than the variance threshold. After scoring the time-frequency features and cross-features respectively using a variety of preset evaluation methods, all scoring results are weighted and summed to obtain a comprehensive score for each feature; based on the scene information, the weights in the weighted summation are adaptively adjusted; Low-scoring features whose comprehensive scores are lower than the preset scoring threshold are eliminated to obtain high-scoring features, which are then subjected to dimensionality reduction to obtain optimized features.
[0005] In conjunction with the first aspect, in one embodiment, the format preprocessing includes multi-source data cleaning, adaptive normalization, and time series feature alignment; The multi-source data cleaning includes outlier detection through the isolation forest algorithm, missing value filling of time series data through the bidirectional long short-term memory (LSTM) network, missing value filling of non-time series data through the neighbor-based KNN interpolation method, and data denoising through the wavelet threshold denoising method; The adaptive normalization includes normalizing Gaussian distribution data by Z-Score normalization and normalizing non-Gaussian distribution data by Box-Cox transformation; The timing feature alignment includes timing feature alignment of millisecond-level data through PTP protocol hardware synchronization, timing feature alignment of data greater than millisecond level through dynamic time warping DTW algorithm software alignment, and timing feature alignment of multi-source data through cubic spline difference.
[0006] In combination with the first aspect, in one embodiment, the time domain features extracted using the first sliding window include mean, variance, skewness, kurtosis, and Hjorth parameter; The time domain features also include high-order features obtained based on mean, variance, skewness, kurtosis, and Hjorth parameter; The high-order features include approximate entropy and peak value.
[0007] In combination with the first aspect, in one embodiment, the preset feature extraction strategy includes extracting frequency domain features from the intermediate test data through fast Fourier transform FFT and power spectral density PSD.
[0008] In combination with the first aspect, in one embodiment, the preset feature extraction strategy includes generating a time-frequency graph using a continuous wavelet transform (CWT), and extracting deep time-frequency features from the time-frequency graph using a pre-trained lightweight convolutional neural model (CNN).
[0009] In conjunction with the first aspect, in one embodiment, constructing cross features based on time-frequency features specifically includes the following steps: The time-frequency features are combined according to a preset combination rule by using a genetic algorithm to obtain the cross feature.
[0010] In conjunction with the first aspect, in one embodiment, the multiple preset evaluation methods include mutual information MI, χ² test, and Embedded_Score; The weights in the weighted summation are adjusted according to the total data dimensions and total sample size of the time-frequency features and cross-features.
[0011] In conjunction with the first aspect, in one embodiment, the dimensionality reduction of high-scoring features specifically includes the following steps: Dimensionality reduction of high-scoring features is performed by improving t-SNE, principal component analysis (PCA), or autoencoders.
[0012] In conjunction with the first aspect, in one embodiment, the adaptively adjusting the weights in the weighted summation based on the scene information specifically includes the following steps: Combining the current scene information and historical feature scores, the importance of each feature is dynamically calculated through a reinforcement learning strategy to update the weights in the weighted summation. The scene information includes weather, vehicle speed, and traffic density.
[0013] In a second aspect, an embodiment of the present application provides a vehicle test data feature extraction system, the vehicle test data feature extraction system comprising: An intelligent preprocessing module is used to preprocess the original test data to obtain intermediate test data that meets the preset format standards; A dynamic feature generation module is configured to extract time-frequency features from the intermediate test data using a preset feature extraction strategy and construct cross-features based on the time-frequency features; the preset feature extraction strategy includes extracting time-domain features from the intermediate test data using a first sliding window, the size of the first sliding window being adaptively adjusted based on the extracted data; A feature evaluation and selection module is used to score the time-frequency features and cross-features respectively using multiple preset evaluation methods, and then perform weighted summation of all scoring results to obtain a comprehensive score for each feature; eliminate low-scoring features whose comprehensive scores are lower than a preset scoring threshold to obtain high-scoring features, and perform dimensionality reduction on them to obtain optimized features; The dynamic feature adjustment module is used to adaptively adjust the weights in the weighted summation based on the scene to which the features belong.
[0014] The beneficial effects of the technical solutions provided in the embodiments of the present application include: By dynamically adjusting the first sliding window, dynamically adjusting the weight of the comprehensive score based on scenario information, and implementing closed-loop optimization through feature dimensionality reduction, this solution makes the feature extraction process more intelligent and efficient, making it suitable for complex and changing vehicle testing environments (such as real-time road conditions in autonomous driving). Compared to traditional fixed feature set methods, this solution can dynamically respond to data characteristics (such as sudden fault signals) and scenario requirements (such as inclement weather), improving the accuracy and reliability of data-driven decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flow chart of an embodiment of the vehicle test data feature extraction method of the present application; Figure 2 This is a flow chart of an embodiment of a method for extracting features from vehicle test data of the present application; Figure 3 This is a schematic diagram of the functional modules of an embodiment of the vehicle test data feature extraction system of the present application. DETAILED DESCRIPTION
[0016] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0017] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0018] In a first aspect, an embodiment of the present application provides a method for extracting features from vehicle test data.
[0019] In one embodiment, referring to Figure 1 and Figure 2 , Figure 1 This is a flowchart of the first embodiment of the vehicle test data feature extraction method of this application. Figure 2 This is a flow chart of an embodiment of the vehicle test data feature extraction method of this application. Figure 1 and Figure 2 As shown, the vehicle test data feature extraction method includes: Step S1: pre-process the original test data to obtain intermediate test data that meets the preset format standard.
[0020] Step S2: Sampling a preset feature extraction strategy to extract time-frequency features from the intermediate test data, and constructing cross-features based on the time-frequency features. The preset feature extraction strategy includes extracting time-domain features from the intermediate test data using a first sliding window. The window size of the first sliding window is dynamically adjusted based on the data fluctuation variance. When the data fluctuation variance is greater than the variance threshold, the window is expanded, and when it is less than the variance threshold, the window is reduced.
[0021] Step S3: After scoring the time-frequency features and cross-features using multiple preset evaluation methods, all scoring results are weighted and summed to obtain a comprehensive score for each feature. The weights in the weighted summation are adaptively adjusted based on the scene information.
[0022] Step S4: Eliminate low-scoring features whose comprehensive scores are lower than a preset scoring threshold to obtain high-scoring features, and perform dimensionality reduction on them to obtain optimized features.
[0023] In this embodiment, the size of the first sliding window is adaptively adjusted based on data fluctuations (e.g., rapid signal changes during high-speed driving) to avoid feature loss or redundancy caused by a fixed window (e.g., expanding the window to capture trends when data fluctuates greatly on rainy days, and reducing the window to preserve details when data is stable on sunny days). This reduces invalid data processing, extracting features only during critical time periods, and reducing redundant calculations.
[0024] Automatically adjust feature evaluation weights for different driving environments (e.g., nighttime, rainy, and highway conditions) (e.g., increasing the weight of radar-related features in rainy conditions and decreasing the weight of illumination-dependent features at night) to better focus the model on key information. This avoids the limitations of a single evaluation method (e.g., mutual information ignores nonlinear relationships) and improves model generalization by adapting to complex scenarios through multi-criteria weighting.
[0025] By improving algorithms such as t-SNE or autoencoders, we can further reduce dimensionality by removing low-scoring, redundant features (e.g., from 1,000 dimensions to 50), reducing storage and computational costs while preserving the local structure and key patterns of high-scoring features. This dimensionality reduction results in more concentrated data, reduces the risk of overfitting, and accelerates model training (e.g., in classification or regression tasks, optimizing the feature set can improve accuracy by 10%-15%).
[0026] In summary, by dynamically adjusting the first sliding window, dynamically adjusting the weight of the comprehensive score based on scenario information, and forming a closed-loop optimization loop through feature dimensionality reduction, the feature extraction process becomes more intelligent and efficient, making it suitable for complex and changing vehicle testing environments (such as real-time road conditions in autonomous driving). Compared to traditional fixed feature set methods, this solution can dynamically respond to data characteristics (such as sudden fault signals) and scenario requirements (such as inclement weather), improving the accuracy and reliability of data-driven decision-making.
[0027] Furthermore, in one embodiment, the format preprocessing includes multi-source data cleaning, adaptive normalization, and time series feature alignment.
[0028] The above-mentioned multi-source data cleaning includes outlier detection through the isolation forest algorithm, missing value filling of time series data through the bidirectional long short-term memory (LSTM) network, missing value filling of non-time series data through the neighbor-based KNN interpolation method, and data denoising through the wavelet threshold denoising method.
[0029] The adaptive normalization includes normalizing Gaussian distribution data through Z-Score normalization and normalizing non-Gaussian distribution data through Box-Cox transformation.
[0030] The above-mentioned timing feature alignment includes timing feature alignment of millisecond-level data through PTP protocol hardware synchronization, timing feature alignment of data greater than millisecond level through dynamic time warping DTW algorithm software alignment, and timing feature alignment of multi-source data through cubic spline difference.
[0031] In this embodiment, an improved isolation forest algorithm is used for outlier detection. A dynamic anomaly threshold θ = μ ± 3σ × (1 + log(N)) is set, where μ and σ are the mean and standard deviation within the data sliding window, and N is the data magnitude coefficient. Anomalous data is identified by constructing an isolation tree, and outliers are isolated by randomly partitioning the data space.
[0032] For time series data (such as window lift speed, seat movement speed, and windshield wiper speed), a bidirectional LSTM (Long Short-Term Memory) network is used to predict missing values. The network structure consists of two LSTM layers (128 units per layer) plus a fully connected layer, with an input window length of T = 10 seconds. For other time series data, KNN (K-NearestNeighbor) interpolation (k = 5, based on feature similarity) is used. In many real-world datasets, some data may be missing due to various reasons (such as sensor failures and errors in the data collection process). For example, in a time series tire dataset, data such as speed and temperature may be missing at certain times. The characteristics of a bidirectional LSTM network are used to predict these missing data points. Taking time series data with some missing values as input, the bidirectional LSTM network infers the missing values by learning the relationships between known data points (including forward and backward relationships).
[0033] Wavelet threshold denoising is used for sensor signals (such as radar reflection intensity), the sym4 wavelet basis function is selected, the decomposition level L=5, and the threshold formula is: , where σ is calculated by the median estimation method ( , where D1 is the first-layer detail coefficient), and N is the signal strength. The wavelet transform decomposes the signal into different frequency subbands, analyzing the local characteristics of the signal at different scales. For noisy sensor signals (such as radar reflection intensity signals), the noise is often distributed in the high-frequency portion. By setting an appropriate threshold, wavelet coefficients below the threshold can be treated as noise and removed (or attenuated), while the signal portion corresponding to wavelet coefficients above the threshold is retained. The denoised signal is then reconstructed through an inverse wavelet transform, effectively removing noise from the signal while preserving its key characteristics.
[0034] For Gaussian distribution data, Z-Score normalization is used , where x is the original data point, μ is the mean of the data column, and σ is the standard deviation of the data column. Through this standardization, the data is transformed into a standard normal distribution with a mean of 0 and a standard deviation of 1, thus eliminating the dimension effect and making different features or variables comparable.
[0035] For non-Gaussian distributed data, use the Box-Cox transformation (Automatically search for the optimal λ∈[-5,5]) to improve the distribution of non-Gaussian distributed data, making it closer to the normal distribution, making it more consistent with the assumptions of statistical analysis, thereby improving the accuracy and reliability of statistical analysis.
[0036] The Kolmogorov-Smirnov test (significance level α = 0.05) was used to determine the data distribution type. The Kolmogorov-Smirnov test is primarily used to compare the empirical distribution function of sample data with the assumed theoretical distribution function. It calculates the maximum distance between the empirical distribution function of the sample and the cumulative distribution function of a given theoretical distribution (such as the normal distribution or uniform distribution).
[0037] Hardware synchronization based on the PTP (Precise Time Protocol) achieves microsecond-level synchronization of multi-sensor data. Software-level alignment is achieved using the DTW (Dynamic Time Warping) algorithm, which calculates timing offset compensation within a 1s window, where W represents the window size.
[0038] For camera (30fps), radar (10Hz), and sensor (5Hz) data, multi-source data alignment is achieved through cubic spline interpolation, taking the highest frequency (30Hz) as the benchmark.
[0039] Furthermore, in one embodiment, the time domain features extracted using the first sliding window include mean, variance, skewness, kurtosis, and Hjorth parameter.
[0040] The above-mentioned time domain features also include high-order features obtained based on mean, variance, skewness, kurtosis, and Hjorth parameter.
[0041] The above high-order features include approximate entropy and peak value.
[0042] In this embodiment, the mean, variance, skewness, kurtosis, and Hjorth parameter (activity, mobility, and complexity) are calculated within the first sliding window (default W = 5s). For example, the window size W′ = max(2s, v20s) is adaptively adjusted based on the vehicle speed v (v in meters per second).
[0043] Furthermore, in one embodiment, the preset feature extraction strategy includes extracting frequency domain features from the intermediate test data through FFT (fast Fourier transform) and PSD (Power Spectral Density).
[0044] In this embodiment, in signal processing, after the time domain signal is converted to the frequency domain through methods such as Fourier transform, the frequency domain representation of the signal will be obtained. For example, the time domain waveform of an audio signal reflects the change in sound pressure over time. In the frequency domain, different frequency components and their corresponding amplitude (energy) information are displayed. A variety of features can be extracted from the frequency domain representation, such as spectrum peaks, frequency band energy (such as the energy ratio of low-frequency and high-frequency parts), spectrum centroid, etc. These features can describe the characteristics of the signal at different frequencies. For example, the frequency domain energy ratio is calculated for the acceleration signal: . Introducing MFCC (Mel-Frequency Cepstrum): Extracting 13-dimensional MFCC features from audio data (such as engine noise).
[0045] Furthermore, in one embodiment, the above-mentioned preset feature extraction strategy includes generating a time-frequency graph using CWT (Continuous Wavelet Transform), and extracting deep time-frequency features from the time-frequency graph using a pre-trained lightweight convolutional neural model CNN.
[0046] In this example, a time-frequency plot is generated using the CWT, and a pre-trained lightweight CNN (MobileNetV2 fine-tuned) is used to extract deep time-frequency features. This method is used to analyze time-varying and non-stationary signals. It combines time and frequency information to provide a description of the signal's energy density or intensity at different times and frequencies. By designing a joint function of time and frequency, joint time-frequency analysis can describe the signal's energy density or intensity at different times and frequencies, providing a more comprehensive signal characterization. Joint time-frequency analysis enables the study of time-frequency filtering and time-varying signals, which is crucial for understanding and processing complex signal systems.
[0047] Furthermore, in one embodiment, when constructing the cross-features based on the time-frequency features, the following steps are specifically included: The time-frequency features are combined according to a preset combination rule through a genetic algorithm to obtain the above-mentioned cross-features.
[0048] In this embodiment, when optimizing the feature combination based on the genetic algorithm, the population size is first determined, for example, 50 feature combination individuals.
[0049] Then determine the fitness function , crossover probability pc=0.8, mutation probability pm=0.05.
[0050] Finally, a genetic algorithm is used to operate based on the existing features, combining different original features according to certain rules, and appropriately considering the importance and redundancy of function features.
[0051] MI typically stands for mutual information. MI(Xi,Y) calculates the mutual information between Xi and Y. Mutual information reflects the degree of interdependence between two variables. For example, in data collection, Xi represents the actual test value of seat massage intensity, and Y represents the passenger or driver comfort level. The value of MI(Xi,Y) can indicate the strength of the association between seat massage intensity and comfort level. In the χ²(Xi,Y) component, χ² is the chi-square test statistic. χ²(Xi,Y) calculates the chi-square value between Xi and Y. The chi-square test is often used to test whether there is an association between two categorical variables. For example, a chi-square test is used to analyze the correlation between vehicle battery consumption (Xi) and climate temperature (Y). The result can reflect the significance of the relationship. The embedded_score(Xi) component is the result of an embedded scoring function applied to Xi individually. For example, in image recognition, if Xi is a feature vector of an image, embedded_score(Xi) may be a quantitative score of this image feature based on a pre-trained model, indicating the degree to which the image meets a certain standard or target.
[0052] In particular, traditional models (such as linear regression) can only process linear relationships between features, but in real-world data, features may influence the results in nonlinear ways (for example, the product of temperature x1 and humidity x2 may jointly determine sensor performance). Combining the original features x1 and x2 into new features using mathematical operations (such as multiplication x1×log(x2) and addition sin(x3)+x42) allows the model to capture complex nonlinear relationships in the data. Algorithms automatically explore mathematical combinations between features and generate new features through nonlinear transformations (such as logarithmic, trigonometric, and square functions). This helps the model better understand complex relationships in the data and ultimately improves prediction performance. For example, in vehicle testing, x1×log(x2) may represent the nonlinear relationship between engine speed x1 and fuel pressure x2. Sin(x3)+x42 may combine the periodic fluctuations of vehicle speed x3 (such as when driving on a hilly road) with the squared effect of acceleration x42 (such as in a sudden acceleration scenario).
[0053] Furthermore, in one embodiment, the plurality of preset evaluation methods include mutual information (MI), chi-squared test, and embedded_score.
[0054] The weights in the weighted summation are adjusted according to the total data dimensions and total sample size of the above time-frequency features and cross-features.
[0055] In this embodiment, the comprehensive score calculation adopts a multi-criteria fusion score model, and its corresponding formula is: .
[0056] The dynamic weight adjustment rules are as follows: When the data dimension D>1000, (w1,w2,w3)=(0.4,0.3,0.3).
[0057] When the sample size N<104, (w1,w2,w3)=(0.5,0.5,0).
[0058] The weights w1, w2, and w3 correspond to the relative importance of MI(Xi,Y), χ²(Xi,Y), and Embedded_Score(Xi) in the final comprehensive score. These weights are usually determined based on specific application scenarios and expert experience or through some optimization algorithm.
[0059] Furthermore, in one embodiment, the dimensionality reduction of the high-scoring features specifically includes the following steps: Dimensionality reduction of high-scoring features is performed by improving t-SNE, principal component analysis (PCA), or autoencoders.
[0060] In this embodiment, dynamic dimensionality reduction is a data processing method whose primary function is to dynamically reduce the dimensionality of data based on its characteristics and task requirements during data mining, machine learning, or data analysis. For example, when processing high-dimensional image data, the original image may contain a large amount of pixel information (high dimensionality), but much of this information may be redundant or have no substantial contribution to the object recognition task (such as distinguishing between images of cats and dogs). When processing massive amounts of sensor data (such as meteorological sensors and industrial equipment sensors), dynamic dimensionality reduction can reduce data storage requirements and computational complexity, while improving the efficiency and accuracy of subsequent data analysis and model building.
[0061] For example, when using the improved t-SNE algorithm for dynamic dimensionality reduction, the initial perplexity perp=30, the learning rate η=200, and the spatial constraint term are added: (Aij is the original space similarity matrix), where This is the loss function in the original t-SNE algorithm, which typically measures the difference in similarity between data points in high-dimensional and low-dimensional space. λ is a regularization parameter that controls the influence of the spatial constraint on the overall objective function. When λ = 0, the algorithm degenerates to the original t-SNE algorithm. When λ is larger, the spatial constraint becomes more effective. zi and zj are the representations of the data points in the low-dimensional space. This component calculates the squared distance between any two data points i and j in the low-dimensional space. By minimizing the sum of the squared distances, the layout of the data points in the low-dimensional space can be constrained. Aij is the original spatial similarity matrix. It reflects the degree of similarity between data points i and j in the original high-dimensional space. If Aij is large, the two data points are relatively similar in the original space. Otherwise, they are dissimilar. In the objective function, by combining the distances in the low-dimensional space with the similarity matrix in the original space, similar data points in the original space tend to be closer in the low-dimensional space during the dimensionality reduction process.
[0062] An autoencoder is a neural network structure whose purpose is to reproduce the input data as closely as possible. It consists of two parts: an encoder and a decoder. The encoder maps the input data to a low-dimensional representation space (called a latent space), and the decoder reconstructs the original input data from this low-dimensional representation. For example, for an image, the encoder may compress it into a low-dimensional vector that contains the key feature information of the image. The decoder then regenerates an image similar to the original image based on this vector. During the training phase, the input of the autoencoder is the original data, and the output is the reconstructed data. By minimizing the difference between the input and output (for example, using the mean squared error loss function), the network continuously adjusts the weights to make the reconstruction as accurate as possible.
[0063] The encoder structure is: input layer → FC(256, ReLU) → FC(64, ReLU) → embedding layer (32 dimensions).
[0064] The loss function is: .
[0065] Where X is the original input data, and X^ is the reconstructed output data from the autoencoder. This calculation calculates the squared Euclidean distance between the original and reconstructed data. For example, in image data, if a pixel in the original image X has a value of x, and the corresponding pixel in the reconstructed image X^ has a value of x^, then this loss measures the difference between each pixel and the overall image structure during the reconstruction process. The smaller this difference, the better the autoencoder's ability to reconstruct the data. In the regularization part, q(z|X) is the posterior probability distribution of the latent variable z given the input X, and p(z) is the prior probability distribution of the latent variable z. KL divergence (Kullback-Leibler Divergence, relative entropy) measures the difference between these two distributions. In the context of autoencoders, it acts as a regularizer. For example, if we assume that the latent variable z follows a simple distribution (such as a standard Gaussian distribution as the prior p(z)), when the autoencoder learns the data features, q(z|X) should be as close to this prior distribution as possible. If the two differ significantly, it may mean that the autoencoder is overfitting the noise in the data or learning an unreasonable feature representation.
[0066] The 0.1 coefficient is used to balance the relative importance of the reconstruction error and the KL divergence loss. If the coefficient is set too large, the KL divergence will have a strong impact on the loss, potentially causing the autoencoder to focus too much on making q(z|X) close to p(z) and ignore the reconstruction error, resulting in poor reconstruction results. Conversely, if the coefficient is too small, the regularization effect of the KL divergence will be weakened, potentially leading to undesirable situations such as overfitting the model to noise in the data.
[0067] Autoencoders are designed to effectively reduce data dimensionality in dynamic dimensionality reduction. In high-dimensional data spaces, data can contain a large amount of redundant information and be computationally complex. Through the encoding process of autoencoders, high-dimensional data can be converted into a low-dimensional representation, thereby reducing data storage requirements, accelerating computation, and, to a certain extent, removing irrelevant information such as noise.
[0068] The purpose of designing this loss function is to encourage the autoencoder to not only reconstruct the input data better (by minimizing the reconstruction error) but also learn a reasonable latent variable representation (by minimizing the KL divergence to constrain the latent variable distribution).
[0069] Autoencoders can adapt to changing data. For example, when processing time-varying streaming data, the data distribution may be constantly changing. The autoencoder in the dynamic dimensionality reduction unit can dynamically adjust its parameters based on new data to maintain good dimensionality reduction results. This can involve specialized mechanisms such as adaptive learning rate adjustments and dynamic network structure adjustments. For example, if it detects that certain data features suddenly become more or less important, the autoencoder can adapt its encoding and decoding methods accordingly, effectively reducing the dimensionality of these dynamically changing data.
[0070] Furthermore, in one embodiment, the vehicle test data feature extraction method also includes feature screening based on a feature importance feedback mechanism. This mechanism uses various methods to evaluate and determine the contribution of each feature to the model's prediction results during analysis of large amounts of test data. This mechanism helps determine which features play a key role in model prediction and which ones can be ignored.
[0071] The feature importance feedback mechanism includes: Constructing a dynamic correlation map, based on the correlation dynamic map, can more accurately determine which features have the most critical impact on the target variable. If a feature has a close and stable correlation with multiple other important features, and this correlation has a significant impact on the target variable, then this feature is likely to have a high importance. and detection target , the edge weight is (calculated by SHAP value), from the feature node To the detection target node The SHAP value is a metric used to explain the output of machine learning models. Using the SHAP value when calculating edge weights means determining the features through this advanced interpretability method. For detection targets For example, in a prediction of vehicle headlight quality (detection target ) in the machine learning model, the change in headlight intensity (feature ) The impact on quality can be calculated by SHAP value to get edge weight This weight can reflect the relative importance of area in determining housing prices.
[0072] Real-time update strategy, after processing N=1000 samples, recalculate the SHAP value and update the map.
[0073] Furthermore, in one embodiment, the adaptive adjustment of the weights in the weighted summation based on the scene information specifically includes the following steps: Combining the current scene information and historical feature scores, the importance of each feature is dynamically calculated through reinforcement learning strategy to update the weights in the weighted summation.
[0074] The above scene information includes weather, vehicle speed, and traffic density.
[0075] In this embodiment, a multimodal scene classification model is first used to acquire scene information, or scene recognition. The inputs to this model are weather data (rainfall, visibility), traffic density, and time of day (day / night). The model architecture is a random forest (100 trees, maximum depth = 10) and a Bayesian network. The output is a scene label (e.g., rainy day, rush hour, low visibility), or scene information.
[0076] Then, a dynamic parameter optimizer is designed. The state space (S) contains the current feature dimension, data distribution deviation (such as data volatility), and model accuracy, which is used to characterize the system's operating state. The action space (A) includes adjusting feature thresholds (in the range δ∈[-0.1, 0.1]) and selecting dimensionality reduction dimensions (d∈{32,64,128}), and optimizing parameters through dynamic strategies. The reward function (TR) combines model performance and efficiency, and is calculated as follows: , which means giving priority to improving accuracy while reducing time cost. The optimization algorithm uses the deep Q learning network (DQN), and the exploration rate Set to 0.1 to ensure a balance between trying new strategies (exploration) and exploiting existing strategies.
[0077] By perceiving the feature processing status in real time and dynamically adjusting the threshold and dimensionality reduction parameters, the optimal configuration of feature extraction is ultimately achieved in complex scenarios.
[0078] When the cumulative amount of new data reaches the preset N = 500 samples, the incremental training process is triggered to ensure that the model can promptly incorporate the latest data information. Using an online random forest algorithm, the model structure can be dynamically adjusted (such as adding or removing decision trees). The maximum number of trees is limited to 200, balancing model complexity and computational efficiency. Features that consistently rank in the bottom 10% of importance for K = 3 consecutive training cycles are automatically frozen or removed to prevent redundant features from interfering with model learning and improve the timeliness and simplicity of the feature set.
[0079] By triggering incremental training in batches, the model can quickly adapt to changes in data distribution (such as unexpected operating conditions during vehicle testing). The dynamic structural adjustment capabilities of online random forests ensure the model's high adaptability to the data stream while limiting the maximum number of trees to prevent overfitting. A feature elimination mechanism automatically removes inefficient features, reducing computational burden and maintaining high feature set quality, making it particularly suitable for long-running vehicle testing systems.
[0080] In summary, this method provides a feature data extraction method that significantly improves processing efficiency, accuracy, and response time, ensuring the processing and analysis of existing vehicle data volumes and contributing to improved vehicle quality for OEMs. Feature extraction efficiency is increased by over 40% (measured data). Key feature recognition accuracy is increased to 92.3%. Real-time feature extraction is supported, processing over 1,000 vehicle data points per second. Feature dimension compression reaches 85%, maintaining over 98% of information content. System adaptive adjustment response time is <200ms.
[0081] In a second aspect, an embodiment of the present application also provides a vehicle test data feature extraction system.
[0082] In one embodiment, referring to Figure 3 , Figure 3 This is a functional module diagram of an embodiment of the vehicle test data feature extraction system of this application. Figure 3 As shown, the vehicle test data feature extraction system includes: The intelligent preprocessing module 1 is used to preprocess the original test data to obtain intermediate test data that meets the preset format standard.
[0083] Dynamic feature generation module 2 is configured to extract time-frequency features from the intermediate test data using a preset feature extraction strategy and construct cross-features based on the time-frequency features. The preset feature extraction strategy includes extracting time-domain features from the intermediate test data using a first sliding window, the size of which is adaptively adjusted based on the extracted data.
[0084] Feature evaluation and selection module 3 is used to score the time-frequency features and cross-features using multiple preset evaluation methods, and then perform a weighted summation of all the scoring results to obtain a comprehensive score for each feature. Low-scoring features with a comprehensive score below a preset scoring threshold are eliminated to obtain high-scoring features, which are then subjected to dimensionality reduction to obtain optimized features.
[0085] The dynamic feature adjustment module 4 is used to adaptively adjust the weights in the weighted summation based on the scene to which the features belong.
[0086] Among them, the functional implementation of each module in the above-mentioned vehicle test data feature extraction system corresponds to the various steps in the above-mentioned vehicle test data feature extraction method embodiment, and its functions and implementation processes will not be repeated here one by one.
[0087] It should be noted that the serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0088] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices. The terms "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit the "first", "second" and "third" to different types.
[0089] In the description of the embodiments of this application, the words "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.
[0090] In the description of the embodiments of the present application, unless otherwise specified, " / " means or. For example, A / B can mean A or B. The "and / or" in the text is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "plurality" means two or more than two.
[0091] In some processes described in the embodiments of the present application, multiple operations or steps are included that appear in a specific order. However, it should be understood that these operations or steps may not be performed in the order in which they appear in the embodiments of the present application or may be performed in parallel. The sequence numbers of the operations are only used to distinguish between different operations, and the sequence numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations or steps may be performed in sequence or in parallel, and these operations or steps may be combined.
[0092] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device to execute the methods described in each embodiment of this application.
[0093] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A vehicle test data feature extraction method, characterized in that: The vehicle test data feature extraction method comprises: Preprocess the original test data to obtain intermediate test data that meets the preset format standards; A preset feature extraction strategy is used to extract time-frequency features from the intermediate test data, and a cross-feature is constructed based on the time-frequency features. The preset feature extraction strategy includes extracting time-domain features from the intermediate test data using a first sliding window. The size of the first sliding window is dynamically adjusted based on the data fluctuation variance, and the window is expanded when the data fluctuation variance is greater than the variance threshold, and is reduced when it is less than the variance threshold. After scoring the time-frequency features and cross-features respectively using a variety of preset evaluation methods, all scoring results are weighted and summed to obtain a comprehensive score for each feature; based on the scene information, the weights in the weighted summation are adaptively adjusted; Low-scoring features whose comprehensive scores are lower than the preset scoring threshold are eliminated to obtain high-scoring features, which are then subjected to dimensionality reduction to obtain optimized features.
2. The vehicle test data feature extraction method according to claim 1, wherein: The format preprocessing includes multi-source data cleaning, adaptive normalization, and time series feature alignment; The multi-source data cleaning includes outlier detection through the isolation forest algorithm, missing value filling of time series data through the bidirectional long short-term memory (LSTM) network, missing value filling of non-time series data through the neighbor-based KNN interpolation method, and data denoising through the wavelet threshold denoising method; The adaptive normalization includes normalizing Gaussian distribution data by Z-Score normalization and normalizing non-Gaussian distribution data by Box-Cox transformation; The timing feature alignment includes timing feature alignment of millisecond-level data through PTP protocol hardware synchronization, timing feature alignment of data greater than millisecond level through dynamic time warping DTW algorithm software alignment, and timing feature alignment of multi-source data through cubic spline difference.
3. The vehicle test data feature extraction method according to claim 1, wherein: The time domain features extracted by the first sliding window include mean, variance, skewness, kurtosis, and Hjorth parameter; The time domain features also include high-order features obtained based on mean, variance, skewness, kurtosis, and Hjorth parameter; The high-order features include approximate entropy and peak value.
4. The vehicle test data feature extraction method according to claim 1, wherein: The preset feature extraction strategy includes extracting frequency domain features from intermediate test data through fast Fourier transform (FFT) and power spectral density (PSD).
5. The vehicle test data feature extraction method according to claim 1, wherein: The preset feature extraction strategy includes generating a time-frequency graph using a continuous wavelet transform (CWT), and extracting deep time-frequency features from the time-frequency graph using a pre-trained lightweight convolutional neural model (CNN).
6. The vehicle test data feature extraction method according to claim 1, wherein: The construction of cross features based on time-frequency features specifically includes the following steps: The time-frequency features are combined according to a preset combination rule by using a genetic algorithm to obtain the cross feature.
7. The vehicle test data feature extraction method according to claim 1, wherein: The multiple preset evaluation methods include mutual information MI, χ² test, and Embedded_Score; The weights in the weighted summation are adjusted according to the total data dimensions and total sample size of the time-frequency features and cross-features.
8. The vehicle test data feature extraction method according to claim 1, wherein: The dimensionality reduction of high-scoring features specifically includes the following steps: Dimensionality reduction of high-scoring features is performed by improving t-SNE, principal component analysis (PCA), or autoencoders.
9. The vehicle test data feature extraction method according to claim 1, wherein: Adaptively adjusting the weights in the weighted summation based on the scene information specifically includes the following steps: Combining the current scene information and historical feature scores, the importance of each feature is dynamically calculated through a reinforcement learning strategy to update the weights in the weighted summation. The scene information includes weather, vehicle speed, and traffic density.
10. A vehicle test data feature extraction system, characterized in that: The vehicle test data feature extraction system includes: An intelligent preprocessing module is used to preprocess the original test data to obtain intermediate test data that meets the preset format standards; A dynamic feature generation module is configured to extract time-frequency features from the intermediate test data using a preset feature extraction strategy and construct cross-features based on the time-frequency features; the preset feature extraction strategy includes extracting time-domain features from the intermediate test data using a first sliding window, the size of the first sliding window being adaptively adjusted based on the extracted data; A feature evaluation and selection module is used to score the time-frequency features and cross-features respectively using multiple preset evaluation methods, and then perform weighted summation of all scoring results to obtain a comprehensive score for each feature; eliminate low-scoring features whose comprehensive scores are lower than a preset scoring threshold to obtain high-scoring features, and perform dimensionality reduction on them to obtain optimized features; The dynamic feature adjustment module is used to adaptively adjust the weights in the weighted summation based on the scene to which the features belong.
Citation Information
Cited By
Power data feature extraction method based on adaptive sliding window and time-frequency fusion
CN121347986A
Radar tide level data non-stationary noise adaptive filtering method
CN121388649A
Vehicle test data processing method and device, electronic equipment and storage medium
CN121502210A
Ear-nose-throat department symptom monitoring method and system based on multi-modal data fusion
CN121512476A