A vehicle sensor data collection and collation method and system for fault detection
By constructing a vehicle health state space model and performing real-time causal analysis, the problem of insufficient quality assessment of multi-source asynchronous data streams was solved, enabling efficient fault detection and prediction, and improving the reliability of data analysis and resource utilization efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN FOXWELL TECHNOLOGY CO LTD
- Filing Date
- 2026-02-13
- Publication Date
- 2026-04-28
AI Technical Summary
Existing vehicle sensor data collection and preprocessing methods lack real-time online quality assessment and reliable quantification at the source processing level of multi-source asynchronous data streams, resulting in low accuracy and interpretability of subsequent fault prediction.
By constructing a vehicle health state space model, real-time acquisition of multi-sensor data streams and real-time operating status parameters is achieved, generating state-related quality vectors and modal cooperative offset parameters, performing causal analysis and judgment, and dynamically adjusting data collection strategies and transmission priorities.
It improves the reliability and accuracy of data analysis, enhances the sensitivity of early fault detection and the interpretability of diagnostic conclusions, and optimizes resource utilization efficiency.
Smart Images

Figure CN121725533B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle fault prediction and health management technology, and in particular to a method and system for collecting and organizing vehicle sensor data for fault detection. Background Technology
[0002] In the field of vehicle fault prediction and health management, in order to achieve health status assessment and early fault warning of key vehicle components, such as power systems, batteries and transmission mechanisms, it relies on the continuous collection and collaborative analysis of data streams generated by different types of sensors from the vehicle network. These data streams come from sources such as vibration sensors, temperature sensors, current sensors and controller status signals. They have inherent differences in physical characteristics, sampling frequency and transmission timing, constituting typical multi-source asynchronous data. Current technical practices generally adopt a centralized processing architecture, that is, at the edge computing node or data acquisition unit on the vehicle, the main tasks are to aggregate, timestamp and possibly compress or package the raw data from multiple sensors, and then transmit the data packets to the cloud or central computing platform for subsequent deep fusion analysis and model inference.
[0003] However, existing data collection and preprocessing methods have limitations at the source processing level of multi-source asynchronous data streams: due to the lack of real-time online quality assessment and reliable quantification of each data stream during the data acquisition and packaging stages, and the failure to initially associate and label data quality anomalies with the real-time operating status of vehicles or potential fault symptoms, the data packets transmitted to the back-end analysis system are essentially a quality-blind raw data set. This makes it difficult for the back-end to effectively distinguish whether data deviations are due to the instantaneous failure of the sensor itself, external interference, or early manifestations of actual component degradation or failure when performing critical multi-sensor data fusion and causal analysis. This severely restricts the accuracy, interpretability, and timeliness of fault prediction. Summary of the Invention
[0004] This invention addresses the technical problems existing in the prior art by providing a method and system for collecting and organizing vehicle sensor data for fault detection.
[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0006] A method for collecting and organizing vehicle sensor data for fault detection, comprising:
[0007] S1. Acquire multi-sensor data streams generated during vehicle operation and the vehicle's current real-time operating status parameters;
[0008] S2. Obtain a pre-established vehicle health state space model that represents the normal cluster distribution of multi-sensor data vectors under different real-time operating state parameters.
[0009] S3. Determine the corresponding normal cluster distribution from the vehicle health state space model based on the real-time operating status parameters, and compare the multi-sensor data vector with the normal cluster distribution to generate a state-related quality vector.
[0010] S4. For sensor data streams in the state-correlated mass vector that indicate the presence of quality anomalies, perform a synergy analysis between their time-frequency domain energy distribution characteristics and the current modal parameters of the vehicle system to generate modal synergy offset parameters.
[0011] S5. Combine the state-related mass vector and modal cooperative offset parameters to perform causal analysis on the sensor data stream with quality anomalies and obtain the determination result of the source of the quality anomaly.
[0012] S6. Based on the judgment result, dynamically adjust the collection strategy, preprocessing method or transmission priority of the corresponding sensor data stream.
[0013] Furthermore, S1 includes:
[0014] Multiple sensor data streams are synchronously acquired via the vehicle controller local area network and the vehicle Ethernet, including vibration sensor data stream, temperature sensor data stream and current sensor data stream.
[0015] The vehicle controller obtains the vehicle's current real-time operating status parameters, which include vehicle speed, motor speed, and battery voltage.
[0016] Furthermore, S2 includes:
[0017] Acquire historical vehicle health operation data, which includes historical multi-sensor data vectors collected under various real-time operating state parameters.
[0018] Feature extraction is performed on historical multi-channel sensor data vectors to construct a feature space;
[0019] In the feature space, conditional clustering analysis is performed on the historical multi-channel sensor data vectors based on the real-time operating status parameters associated with them, so as to form a normal cluster distribution under different real-time operating status parameters.
[0020] The normal cluster distribution and its corresponding real-time operating status parameters are associated and stored to construct a vehicle health state space model.
[0021] Furthermore, S3 includes:
[0022] Based on the current real-time operating status parameters, match the normal cluster distribution corresponding to the operating condition interval that is closest to the current real-time operating status parameters in the vehicle health state space model.
[0023] Extract the multi-sensor data vector of the current time slice aligned with the current timestamp from the multi-sensor data stream;
[0024] Calculate the multidimensional distance metric between the multi-sensor data vector of the current time slice and the center of the matched normal cluster distribution;
[0025] A state-related quality vector containing the deviation components of each sensor data stream is generated based on a multidimensional distance metric.
[0026] Furthermore, calculating the multidimensional distance metric between the multi-sensor data vector of the current time slice and the center of the matched normal cluster distribution includes: calculating the Mahalanobis distance between the multi-sensor data vector of the current time slice and the center of the matched normal cluster distribution, wherein the Mahalanobis distance is calculated based on the inversion of the covariance matrix of the matched normal cluster distribution and taking into account the statistical correlation between the dimensions of each sensor data.
[0027] Furthermore, S4 includes:
[0028] Identify sensor data streams that indicate quality anomalies from state-associated quality vectors;
[0029] The time-frequency transformation of the identified sensor data stream is performed within the time window of the continuous quality anomaly to extract its time-frequency domain energy distribution characteristics.
[0030] Based on the multi-sensor data stream of the vehicle within a time window, the current modal parameters of the vehicle system are identified by running modal analysis methods. The current modal parameters of the vehicle system include natural frequency and damping ratio.
[0031] The extracted time-frequency domain energy distribution features are correlated with the identified current modal parameters of the vehicle system, and frequency band matching analysis is performed to generate modal cooperative migration parameters that characterize the degree of cooperation between the two.
[0032] Furthermore, the correlation calculation and frequency band matching analysis of the extracted time-frequency domain energy distribution features and the identified current modal parameters of the vehicle system include: locating the frequency band corresponding to the natural frequency in the current modal parameters of the vehicle system in the time-frequency domain energy distribution; calculating the energy concentration in the corresponding location frequency band; and calculating the relative rate of change of the corresponding energy concentration with the pre-stored reference energy concentration under normal operating conditions to generate modal cooperative offset parameters.
[0033] Furthermore, S5 includes:
[0034] Based on the state-correlation quality vector, the deviation component of the sensor data stream corresponding to the quality anomaly is extracted;
[0035] Based on modal cooperative migration parameters, the correlation strength components between corresponding abnormal events and vehicle structural dynamic characteristics are extracted;
[0036] Based on the numerical combination of the deviation degree component and the correlation strength component, a matching is performed in the predefined fault mode classification logic to output a judgment result that classifies the source of quality anomalies as sensor self-interference, vehicle active control effect, or potential vehicle component failure.
[0037] Furthermore, S6 includes:
[0038] Based on the classification of the sources of quality anomalies in the judgment results, if the classification is sensor interference itself, the sampling frequency of the corresponding sensor data stream will be reduced and anti-interference filtering preprocessing will be enabled.
[0039] If classified as a vehicle active control effect, the current collection strategy is maintained, but the priority of the corresponding data stream in the transmission queue is increased;
[0040] If the fault is classified as a potential vehicle component failure, the sampling frequency of the corresponding sensor data stream is increased, data compression is disabled, and it is marked as real-time transmission.
[0041] On the other hand, the present invention provides a vehicle sensor data collection and processing system for fault detection, comprising:
[0042] The parameter acquisition module is used to acquire multi-sensor data streams generated by the vehicle during operation and the vehicle's current real-time operating status parameters;
[0043] The model building module is used to obtain a pre-established vehicle health state space model that represents the normal cluster distribution of multi-sensor data vectors under different real-time operating state parameters.
[0044] The vector generation module is used to determine the corresponding normal cluster distribution from the vehicle health state space model based on real-time operating status parameters, and to compare the multi-sensor data vectors with the normal cluster distribution to generate a state-related quality vector.
[0045] The parameter generation module is used to perform a correlation analysis between the time-frequency domain energy distribution characteristics of sensor data streams indicating quality anomalies in the state-correlated mass vector and the current modal parameters of the vehicle system, and generate modal correlation offset parameters.
[0046] The result determination module is used to perform causal analysis on sensor data streams with quality anomalies by combining the state-related quality vector and the modal cooperative offset parameter, and obtain the determination result about the source of the quality anomaly.
[0047] The judgment and adjustment module is used to dynamically adjust the collection strategy, preprocessing method, or transmission priority of the corresponding sensor data stream based on the judgment result.
[0048] The beneficial effects of this invention are:
[0049] 1. By embedding real-time online quality assessment and causal correlation analysis capabilities at the source stage of data collection and organization, the traditional quality-blind data acquisition mode has been changed. The vehicle health state space model is constructed and a state-related quality vector is generated, so that each sensor data stream is given a health metric label associated with its current operating condition at the time of acquisition. This not only moves the complex quality assessment work that was originally completed in the back end to the starting point of data generation, but more importantly, it provides input with clear quality annotations and operating condition context for all subsequent data analysis steps. This enables multi-sensor data fusion and fault diagnosis algorithms to be built on a reliable and semantically rich data foundation, significantly improving the starting point quality and reliability of subsequent analysis.
[0050] 2. By introducing time-frequency domain collaborative analysis to generate modal collaborative offset parameters, and using these parameters to make causal judgments to drive the dynamic adaptive adjustment of the collection strategy, a complete intelligent closed loop from perception, analysis to execution is achieved. This enables the system to proactively identify the physical root causes of abnormal data and accordingly configure sampling, processing, and transmission resources in a differentiated manner. This not only optimizes the utilization efficiency of limited in-vehicle communication and computing resources, but more importantly, ensures that high-value data representing real potential faults can be captured and transmitted with the highest fidelity and the highest priority. This provides a more accurate and timely data foundation for the fault prediction and health management system, enhancing the sensitivity of early fault detection and the interpretability of diagnostic conclusions. Attached Figure Description
[0051] Figure 1 This is a flowchart of a vehicle sensor data collection and processing method for fault detection according to the present invention;
[0052] Figure 2 This is a schematic diagram of the structure of a vehicle sensor data collection and processing system for fault detection according to the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Example 1: Figure 1This invention provides a method for collecting and organizing vehicle sensor data for fault detection, comprising:
[0055] S1. Acquire multi-sensor data streams generated during vehicle operation and the vehicle's current real-time operating status parameters;
[0056] S2. Obtain a pre-established vehicle health state space model that represents the normal cluster distribution of multi-sensor data vectors under different real-time operating state parameters.
[0057] S3. Determine the corresponding normal cluster distribution from the vehicle health state space model based on the real-time operating status parameters, and compare the multi-sensor data vector with the normal cluster distribution to generate a state-related quality vector.
[0058] S4. For sensor data streams in the state-correlated mass vector that indicate the presence of quality anomalies, perform a synergy analysis between their time-frequency domain energy distribution characteristics and the current modal parameters of the vehicle system to generate modal synergy offset parameters.
[0059] S5. Combine the state-related mass vector and modal cooperative offset parameters to perform causal analysis on the sensor data stream with quality anomalies and obtain the determination result of the source of the quality anomaly.
[0060] S6. Based on the judgment result, dynamically adjust the collection strategy, preprocessing method or transmission priority of the corresponding sensor data stream.
[0061] S1. Acquire the multi-sensor data streams generated by the vehicle during operation and the vehicle's current real-time operating status parameters. Specifically, this is implemented as follows:
[0062] The core of this synchronous acquisition—synchronously acquiring data from multiple sensors via the vehicle's Controller Area Network (CAN) and the in-vehicle Ethernet—lies in establishing a unified and accurate time reference for data from all heterogeneous network sources. Specifically, the CAN uses the CAN bus protocol for communication, and the data acquisition unit acts as a listening node on the bus, receiving periodic broadcast messages from different sensor control nodes. For example, a vibration sensor control node might encapsulate its measured vibration values into CAN data frames with specific identifiers and send them to the bus at fixed time intervals, such as every 10 milliseconds. Upon receiving each complete CAN message frame at the physical layer, the data acquisition unit immediately retrieves the current time from its internal high-precision clock module and uses this time value as the timestamp of that message frame. This high-precision clock module initializes and performs periodic calibration upon startup by receiving timing signals from the Global Positioning System (GPS) or a precise time protocol synchronization signal provided by the vehicle gateway, ensuring its time reference is consistent with other time-sensitive systems in the vehicle, achieving a timestamp accuracy on the order of 1 microsecond. For data transmitted via in-vehicle Ethernet, such as raw vibration waveform data streams from high-bandwidth analog-to-digital converters, the network interface controller of the data acquisition unit obtains a timestamp from the same high-precision clock module and associates it with the data packet when it arrives at its Ethernet media access control layer. In this way, regardless of whether the data comes from the controller area network or the in-vehicle Ethernet, it is marked under the same time base, thereby achieving synchronous acquisition of multiple sensor data streams.
[0063] The multi-sensor data stream includes vibration sensor data stream, temperature sensor data stream, and current sensor data stream. The vibration sensor data stream originates from an accelerometer mounted on the vehicle's powertrain housing. Its output analog voltage signal is first filtered by anti-aliasing and then digitized by an analog-to-digital converter at a set sampling frequency. This sampling frequency is set based on the highest frequency of the fault characteristic frequency component to be analyzed and follows the Nyquist sampling theorem. For example, if the highest fault characteristic frequency to be analyzed is 2000 Hz, the sampling frequency must be at least 4000 Hz; in practical applications, to provide a certain frequency analysis margin, the sampling frequency can be set to 5000 Hz or 10000 Hz. The digitized series of sampling points arranged in chronological order constitutes the vibration sensor data stream. The temperature sensor data stream originates from a thermistor embedded in the drive motor windings. Its resistance value is converted into a voltage signal by a measurement circuit and then acquired by an analog-to-digital converter at a lower sampling frequency, such as 1 Hz, because temperature changes are usually a slow process. The current sensor data stream originates from the measurement of the three-phase AC current of the motor. A closed-loop Hall effect current sensor is typically used. Its analog output is conditioned and then acquired by an analog-to-digital converter at a sampling frequency that matches the motor control frequency. For example, when the motor control frequency is 10 kHz, the current sampling frequency can be set to 20 kHz.
[0064] Simultaneously, the vehicle's current real-time operating status parameters are acquired through the vehicle controller. Here, the vehicle controller refers to the electronic control unit within the vehicle responsible for specific function control, such as the vehicle controller, motor controller, and battery management system controller. Real-time acquisition means that the data acquisition unit reads parameter values from these controllers via the vehicle controller area network (LAN) through periodic requests or passive listening. For example, the data acquisition unit can send a standard diagnostic message requesting vehicle speed parameters to the vehicle controller every 20 milliseconds. Upon receiving the request, the vehicle controller returns a response message containing the current vehicle speed data, calculated from the pulse signals of the wheel speed sensors, in kilometers per hour. The data acquisition unit also associates a timestamp from the same high-precision clock module with the moment it receives the response message. Motor speed parameters are acquired from the motor controller in a similar manner, for example, every 10 milliseconds. This parameter is obtained by analyzing the signal from the motor's resolver, in revolutions per minute (rpm). Battery voltage parameters are acquired from the battery management system controller, for example, every 100 milliseconds. This parameter is the measured total terminal voltage of the battery pack, in volts. These acquisition cycles are set based on the rate of physical change of each parameter and the needs of subsequent analysis. For example, vehicle speed and motor speed change rapidly, so shorter acquisition cycles are used; the total battery voltage is relatively stable, so longer acquisition cycles are used. All successfully acquired parameter values and their corresponding timestamps together constitute the vehicle's current real-time operating status parameters. These parameter values, along with the aforementioned timestamped multi-channel sensor data streams, will be transmitted to subsequent processing steps and aligned and correlated based on the timestamps.
[0065] S2. Obtain a pre-established vehicle health state space model representing the normal cluster distribution of multi-sensor data vectors under different real-time operating state parameters. Specifically, this is implemented as follows:
[0066] The pre-establishment process of the vehicle health state space model is completed on a cloud server or high-performance computing platform before the vehicle leaves the factory or in the early stages of its use. This process first requires acquiring historical vehicle health operation data. Here, historical vehicle health operation data refers to the raw dataset collected and filtered and cleaned under various operating conditions when the vehicle is known to be in a fault-free and healthy operating state. The data acquisition method is the same as the real-time acquisition method described in step S1, that is, multi-channel sensor data streams are synchronously acquired through the vehicle controller area network and the vehicle Ethernet, and the corresponding real-time vehicle operating status parameters are acquired synchronously. For example, in vehicle bench tests or actual road tests, the vehicle is operated under various operating conditions covering its design operating range, such as idling, constant speed cruising, acceleration, deceleration, and driving on different slopes, and all sensor data and corresponding status parameters are continuously recorded. The duration of each recording session must be set to ensure the capture of a stable operating state under that condition. This duration is based on a factor of several times the time required for the vehicle system's main dynamic processes to reach steady state, such as 3 to 5 times the time constant of the main thermodynamic processes or the stabilization time of the control loop. Therefore, a single recording session might be set to 3 minutes or 5 minutes. The recorded raw time-series data undergoes preprocessing, including removing abnormal null values caused by momentary communication interruptions, correcting outliers that significantly exceed the physical range, and strictly aligning all sensor data streams according to the high-precision timestamps in step S1 to ensure that the measured values of each sensor correspond one-to-one with the vehicle's state parameters at the same moment. Subsequently, the continuous data stream is cut into data segments of fixed duration. The duration of these segments is set to include sufficient dynamic information while being shorter than the typical time for significant changes in the vehicle's operating state. For example, for cases where mechanical vibration response is of interest, this duration can be set to 1 second to include tens to hundreds of vibration cycles. The real-time vehicle operating state parameters at the start of each segment are used as the operating condition label for that segment. From each data segment, a combination containing the current measurements of all specified sensors is extracted according to a predefined time window, such as every 10 milliseconds. This combination constitutes a historical multi-sensor data vector. Ultimately, the vehicle's historical health operation data contains tens of thousands of such historical multi-sensor data vectors collected under various real-time operating state parameters. Each vector is precisely associated with a specific set of real-time operating state parameters, such as vehicle speed of 50 kilometers per hour, motor speed of 3000 revolutions per minute, and battery voltage of 400 volts.
[0067] Feature extraction is performed on historical multi-channel sensor data vectors to construct a feature space. This feature extraction aims to extract, from the raw, high-dimensional sensor data vectors that may contain redundant information, features that effectively characterize the vehicle's health status and are more discriminative and have lower dimensionality. In practice, this is not a simple stacking of all raw sensor measurements, but rather the extraction of physically meaningful statistical or frequency domain features from different types of sensor data streams. For example, for vibration acceleration sequences in vibration sensor data vectors, extracted features may include the root mean square value, peak factor, kurtosis index in the time domain, and the energy proportion within a specific frequency band after conversion to the frequency domain via Fast Fourier Transform, such as the energy in the frequency band near the fundamental frequency harmonics of the motor. For temperature sensor data vectors, due to their slow changes, features can be selected as the average value and slope of change over a data segment. For current sensor data vectors, features may include the effective value of the current, total harmonic distortion (THD), etc. After the above feature extraction, each historical multi-channel sensor data vector is transformed into a feature vector composed of multiple feature components. The set of feature vectors corresponding to all historical data constitutes a multi-dimensional feature space. The dimensions and physical meanings of each feature component in the feature space may differ. Therefore, standardization is necessary before proceeding with further analysis. One method for standardization is Z-score standardization, which involves: for each feature component dimension in the feature space, first calculating the arithmetic mean and standard deviation of all historical feature vectors along that dimension; then, for each historical feature vector, subtracting the arithmetic mean of that dimension from its original value, and dividing by the standard deviation, to obtain the standardized value. After this process, the numerical distribution of each feature component dimension is adjusted to have a mean of 0 and a standard deviation of 1.
[0068] In the feature space, conditional clustering analysis is performed on the historical multi-channel sensor data vectors based on the real-time operating status parameters associated with them, to form normal cluster distributions under different real-time operating status parameters. Here, conditional clustering analysis means that the clustering process is constrained by the real-time operating status parameters. The implementation consists of two steps. The first step is the division of operating condition intervals. Based on the distribution range of real-time operating status parameters in the vehicle's historical health operation data, each status parameter dimension is independently divided into several continuous intervals. The division is based on ensuring that each interval contains a sufficient number of historical data samples, and the interval boundaries are usually chosen in areas with low parameter distribution density or at thresholds with engineering significance. For example, the vehicle speed dimension can be divided based on the cumulative distribution function of vehicle speed values in historical data, taking the vehicle speed values corresponding to cumulative probabilities of 30% and 70% as boundary values, thereby dividing the vehicle speed into low, medium, and high intervals. Specifically, if the 30th percentile corresponds to a vehicle speed of 30 km / h and the 70th percentile corresponds to a vehicle speed of 80 km / h, then the vehicle speed range is divided into 0-30 km / h, 31-80 km / h, and above 81 km / h. A specific operating condition range is defined by a combination of specific ranges for each state parameter dimension. For example, a vehicle speed of 31-80 km / h and a motor speed of 2001-5000 rpm constitute an operating condition range. Each historical feature vector is uniquely assigned to a specific operating condition range based on its associated real-time operating state parameter value. The second step is to cluster the historical feature vectors within each operating condition range. The K-means clustering algorithm is used, which first requires a preset number of clusters K for each operating condition range. The K value can be automatically determined based on the number and distribution complexity of historical feature vectors within the range using the elbow rule: calculate the total within-class variance of the clustering results under different K values, plot its curve as a function of K, and select the K value corresponding to the inflection point of the curve, i.e., the elbow point, as the final value. After determining the value of K, the algorithm executes as follows: K cluster centers are randomly initialized; iterative calculations are performed. In each iteration, the Euclidean distance from each feature vector within the interval to all cluster centers is calculated, and the distance is assigned to the cluster represented by the nearest cluster center; then, based on all feature vectors assigned to each cluster, the center point position of that cluster is recalculated, i.e., the average value of all feature vectors in each dimension within that cluster is taken as the new center point; the iteration stops when the movement distance of all cluster centers calculated in two consecutive iterations is less than a preset cluster center point movement distance threshold. This threshold is used to determine whether the center point has stabilized. The cluster center point movement distance threshold is set based on the scale after feature space standardization. Since each dimension is standardized to a mean of 0 and a standard deviation of 1, the center point movement distance is usually set to a positive number much less than 1, such as 0.001.The specific comparison method is as follows: Calculate the Euclidean distance between each new center point and the corresponding old center point in the previous iteration after each iteration. If the moving distance of all K center points is less than the moving distance threshold of the cluster center point, it is determined to be converged and the iteration stops; if the number of iterations reaches the preset maximum number of iterations, such as 100, it is forcibly stopped. After the algorithm converges, the historical feature vectors within each operating condition interval are divided into K clusters. Each cluster is characterized by its cluster center vector and a covariance matrix describing its distribution shape and range, constituting one or more normal cluster distributions under the operating condition interval defined by the specific combination of real-time operating state parameters.
[0069] A vehicle health state space model is constructed by associating and storing the normal cluster distributions and their corresponding real-time operating status parameters. After completing conditional clustering analysis for all operating condition intervals, the upper and lower boundary values of the defined parameters (i.e., the combinations of real-time operating status parameters) for each operating condition interval are used as index keys. The specific parameters of each normal cluster distribution obtained through clustering within that interval are stored as the content, and this is done in a structured, associative manner. The stored parameters for each cluster include: the coordinate values of each dimension of the cluster center vector, the values of each element of the covariance matrix representing the cluster distribution, and the number of historical feature vectors contained in the cluster. All different operating condition intervals and their corresponding sets of all cluster distribution parameters together constitute a complete and queryable vehicle health state space model. In subsequent online use, when a set of real-time operating status parameters is input, the model can find the matching operating condition interval by comparing whether each parameter value falls within the boundary range of each operating condition interval, and quickly retrieve the parameters of one or more normal cluster distributions corresponding to that interval, thus providing a benchmark for online health assessment. The model can be stored in a multidimensional lookup table or a database table. The final vehicle health state space model is deployed to the edge computing unit or data acquisition unit on the vehicle so that it can be invoked in real time in step S3.
[0070] S3. Determine the corresponding normal cluster distribution from the vehicle health state space model based on real-time operating status parameters, and compare the multi-sensor data vector with the normal cluster distribution to generate a state-related quality vector. Specifically, this is implemented as follows:
[0071] Based on the current real-time operating status parameters, the system matches the normal cluster distribution corresponding to the operating condition interval closest to the current real-time operating status parameters in the vehicle health state space model. This process is executed online in real-time in the vehicle's computing unit. The computing unit receives the vehicle's current real-time operating status parameters from step S1, such as a set of specific values including vehicle speed, motor speed, and battery voltage. The goal of the matching is to find the operating condition interval that best matches the set of real-time parameters among the many operating condition intervals stored in the pre-established vehicle health state space model. The matching is specifically implemented by calculating the weighted Euclidean distance between the real-time operating status parameter vector and the center parameter vector of each operating condition interval in the model. The center parameter vector of each operating condition interval is obtained during the model building phase by calculating the average of the real-time operating status parameters of all historical data segments falling into that interval, and is stored together with the interval. The calculation of the weighted Euclidean distance requires setting a weight coefficient for each state parameter dimension. This weight coefficient is used to reflect the importance of that dimension parameter in distinguishing different operating conditions. The weight coefficient can be set based on the range of variation of that dimension parameter in historical data. Dimensions with a larger range of variation are usually given a lower weight to prevent them from dominating the distance calculation. For example, if the historical variation range of vehicle speed is 0 to 120 kilometers per hour, and the historical variation range of motor speed is 0 to 6000 revolutions per minute, then a weighting coefficient of 1 / 120 = 0.0083 can be set for the vehicle speed dimension, and a weighting coefficient of 1 / 6000 = 0.000167 can be set for the motor speed dimension. The calculation steps for the weighted Euclidean distance are as follows: First, subtract the value of the center parameter vector of the working condition interval to be matched in the corresponding dimension from the value of the current real-time operating status parameter in each dimension to obtain the difference in each dimension; second, multiply the difference in each dimension by the pre-set weighting coefficient of that dimension to obtain the weighted difference; then, sum the squares of the weighted differences in all dimensions; finally, take the square root of the sum of squares, and the resulting value is the weighted Euclidean distance between the current parameter and the center of the working condition interval. The calculation unit traverses all predefined working condition intervals in the model, calculates the weighted Euclidean distance between the current parameter and each of them, and selects the working condition interval with the smallest calculated distance value as the preliminary matching result. If all calculated weighted Euclidean distance values are greater than a preset operating condition matching distance threshold, the current operating condition is determined to be outside the model's coverage. This operating condition matching distance threshold is set based on the weighted Euclidean distance distribution between all parameter vectors and the center vector of their respective operating condition intervals in historical health data; for example, the 95th percentile of this distance distribution is used as the threshold. If no valid operating condition is found, a default processing procedure is triggered, such as using a general cluster distribution covering all operating conditions for subsequent calculations. After successfully matching an operating condition interval, the specific parameters of one or more normal cluster distributions associated with that interval are retrieved from the vehicle health state space model.
[0072] The multi-sensor data vector of the current time slice, aligned with the current timestamp, is extracted from the multi-sensor data stream. The current time slice refers to a data window of fixed duration that ends at the current processing moment and extends backward. The duration of this window is consistent with the duration of the data segment used in model building in step S2, for example, 1 second. The computing unit extracts all original sensor sampling points within this time window from the buffered data area with precise timestamps provided in step S1. Since the sampling frequencies of different sensor data streams may differ, low-frequency data needs to be interpolated first to obtain synchronized values for all sensors at equally spaced time points within the same series. Linear interpolation can be used, i.e., calculating the estimated value at each high-frequency data sampling point based on the timestamps and values of adjacent low-frequency data sampling points. After time alignment and interpolation, a set containing the measurements of all specified sensors at each equally spaced time point is obtained; this set constitutes the original multi-sensor data vector of the current time slice. Next, the same feature extraction operation as in step S2 must be performed on this original vector. Following the same feature type and calculation method defined in step S2, the original multi-channel sensor data vector for the current time slice is transformed into a feature vector. This feature vector also needs to undergo the same standardization process as in step S2, that is, using the historical arithmetic mean and historical standard deviation of each feature component dimension stored in step S2, the mean is subtracted from each component of the feature vector and then divided by the standard deviation, thereby obtaining a standardized multi-channel sensor data feature vector for the current time slice.
[0073] The multidimensional distance metric between the multi-sensor data vector of the current time slice and the center of the matched normal cluster distribution is calculated. Specifically, it involves calculating the Mahalanobis distance between the standardized multi-sensor data feature vector of the current time slice and the cluster center vector of each normal cluster distribution within the matched operating range. For a specific normal cluster distribution, its cluster center vector and covariance matrix have been retrieved from the model. The Mahalanobis distance calculation process is as follows: First, calculate the difference vectors between the multi-sensor data feature vector of the current time slice and the cluster center vector of the cluster in each dimension; second, obtain the covariance matrix of the cluster; then, calculate the inverse of the covariance matrix; next, perform matrix operations, which mathematically mean transposing the difference vector to obtain a row vector, multiplying this row vector by the inverse of the covariance matrix, and then performing a dot product operation between the result of the multiplication and the original difference vector to obtain a scalar value; finally, take the arithmetic square root of this scalar value to obtain the Mahalanobis distance of the multi-sensor data feature vector of the current time slice relative to the specific cluster. The calculation unit performs the above calculation on all clusters within the matched operating condition interval, obtaining a set of Mahalanobis distances. Finally, the smallest Mahalanobis distance value in this set is taken as a multidimensional distance metric representing the degree of deviation of the current data from the overall normal state under that operating condition. If this smallest Mahalanobis distance value exceeds a preset Mahalanobis distance anomaly threshold, the current state is directly determined to have a significant anomaly. The Mahalanobis distance anomaly threshold is set based on the following: since the squared value of the Mahalanobis distance approximately follows a chi-square distribution with degrees of freedom equal to the dimension of the feature space, the corresponding quantile value can be found in the chi-square distribution table according to the selected significance level, and the arithmetic square root of this quantile value is set as the Mahalanobis distance anomaly threshold. For example, if the feature space dimension is 10, the selected significance level is 0.001, and the quantile with 10 degrees of freedom and a cumulative probability of 0.999 found in the chi-square distribution table is approximately 29.588, then the Mahalanobis distance anomaly threshold can be set as the arithmetic square root of 29.588, approximately equal to 5.44. This threshold is used for initial filtering of significant deviations.
[0074] A state-related quality vector containing deviation components of each sensor data stream is generated based on a multi-dimensional distance metric. The state-related quality vector is a vector with a dimension equal to the number of sensor data stream types of interest. One generation method is to decompose the contribution based on the best-matching cluster corresponding to the calculated minimum Mahalanobis distance. Specifically, the absolute values of each component of the intermediate result obtained during the Mahalanobis distance calculation—that is, the difference vector after linear transformation of the covariance matrix inverse matrix—are summed. Then, the absolute value of each component is divided by the sum to obtain the normalized contribution of each feature dimension to the overall Mahalanobis distance. Then, according to the mapping relationship extracted in step S2, the contributions of multiple feature dimensions belonging to the same original sensor data stream are added together as the initial deviation components of that sensor data stream. Another more direct method is to perform calculations at the feature subspace level: for each feature subset corresponding to a sensor data stream, the Euclidean distance between that subset and the subset corresponding to the center vector of the best-matching cluster in the feature vector of the multi-sensor data stream in the current time slice is calculated, resulting in a set of sensor-granular sub-distances. Then, these sub-distance values are mapped to deviation scores between 0 and 1 using a preset normalization function. The parameters of this normalization function are determined by analyzing the statistical distribution of each sensor sub-distance in historical health data. For example, the 95th percentile of a sensor sub-distance value in historical data is set as a reference point mapped to a deviation score of 0.8. Distance values below this reference point are mapped to between 0 and 0.8 using linear interpolation, while distance values above this reference point are mapped to between 0.8 and 1.0. The resulting state-related quality vector has each component with a value between 0 and 1; a larger value indicates a greater deviation of the corresponding sensor data stream from its normal pattern. This vector will be passed to subsequent step S4 for further analysis.
[0075] S4. For sensor data streams indicating quality anomalies in the state-correlated mass vector, perform a correlation analysis between their time-frequency domain energy distribution characteristics and the current modal parameters of the vehicle system to generate modal correlation offset parameters. Specifically, this is implemented as follows:
[0076] The sensor data streams indicating quality anomalies are identified from the state-associated quality vector. The state-associated quality vector, generated and input in step S3, consists of components corresponding to a single sensor data stream, each value representing the deviation of that data stream from its normal mode. The identification process requires setting an anomaly threshold to determine which component indicates a quality anomaly. This threshold is set based on statistical analysis of the historical state-associated quality vector values for each component under healthy vehicle conditions. Specifically, during the vehicle health state space model construction phase, a large number of state-associated quality vector samples under healthy conditions are collected. For each component corresponding to a sensor data stream, the statistical distribution of all historical sample values is calculated, and the 95th percentile of this distribution is taken as the anomaly threshold for that sensor data stream. For example, for the component corresponding to the front axle vibration sensor data stream, if the 95th percentile of its historical healthy sample values is 0.85, then its anomaly threshold is set to 0.85. During online identification, each component of the current state-associated quality vector is compared with its corresponding anomaly detection threshold. If the value of a component exceeds its anomaly detection threshold, the sensor data stream corresponding to that component is determined to be a sensor data stream indicating a quality anomaly. All identified sensor data streams will then proceed to the subsequent collaborative analysis process.
[0077] The identified sensor data stream undergoes time-frequency transformation within the time window of sustained quality anomalies to extract its time-frequency domain energy distribution characteristics. Here, the time window of sustained quality anomalies refers to the period from the starting point when a significant deviation in the state is determined in step S3 until the current processing time. The length of this time window needs to balance frequency resolution and time change capture capability. It is based on the period length of the lowest frequency component of interest in the signal being analyzed, typically requiring the time window length to include at least 10 complete periods of that lowest frequency. For example, if the focus is on vehicle chassis vibration, its lowest natural frequency might be 1 Hz with a period of 1 second; in this case, the time window length could be set to 10 seconds. The time-frequency transformation uses a short-time Fourier transform. During execution, the window function and window length need to be set. A Hanning window is chosen to reduce spectral leakage. The window length setting requires a trade-off between time resolution and frequency resolution, based on the rate of change of the frequency components of the signal being analyzed; a longer window length results in higher frequency resolution but more ambiguous time positioning. The window length can be set to one-tenth to one-fifth of the time window length. For example, for a 10-second time window, the window length can be set to 1 second. After performing a short-time Fourier transform on the sensor data stream within the time window, a two-dimensional complex matrix is obtained, where the rows correspond to frequency points and the columns correspond to time points. The square of the modulus of each element represents the signal power at that frequency point at that moment. This two-dimensional power matrix is the extracted time-frequency domain energy distribution feature.
[0078] Based on multi-sensor data streams from the vehicle within a time window, operational modal analysis (EMA) is used to identify the current modal parameters of the vehicle system. Here, the current modal parameters specifically refer to the structural dynamic characteristic parameters of the vehicle under actual operating conditions within the time window. The EMA employs frequency domain decomposition. Specifically, multiple sensor data streams within the time window that are spatially related to the identified abnormal sensor data streams are selected as input; for example, vertical vibration acceleration data streams from four suspension mounting points of the vehicle are selected simultaneously. First, these data streams are preprocessed, including removing the mean to eliminate the DC component and performing bandpass filtering, with the filtering frequency band covering the main vibration frequency range of the vehicle structure, for example, 0.5 Hz to 50 Hz. Subsequently, the cross-power spectral density matrix between these multiple data streams is calculated. The calculation employs the Welch average periodogram method: each data stream is divided into multiple overlapping segments, each segment having a length equal to the aforementioned window length, with an overlap rate of 50%. After applying a Hanning window to each data segment, a Fast Fourier Transform (FFT) is performed. For each pair of data streams from different sensors, the conjugate product of the Fourier transform results of their corresponding data segments is calculated and averaged across all data segments to obtain a cross-power spectral density (CPS) estimate. This process is repeated for all sensor pairs to form a CPS matrix. Next, within the band of interest, singular value decomposition (SVD) is performed on this CPS matrix at each frequency point. The decomposition yields a set of singular value spectra arranged in order of magnitude. The criterion for identifying significant modes is: if a local peak appears in the first singular value at a certain frequency in the singular value spectrum, and this peak exceeds a preset noise level threshold, then a mode is considered to exist at that location. This noise level threshold is estimated by adding three standard deviations to the average of the singular value spectra across the entire band of interest. For each identified peak, its corresponding frequency is the estimated intrinsic frequency of that mode. The damping ratio of this mode is estimated by calculating the half-power bandwidth at the peak value: within the half-power bandwidth, two frequency points where the singular value drops by 3 dB are found; the difference between these two frequency points is divided by twice the estimated natural frequency value to obtain the estimated damping ratio. The finally identified current modal parameters of the vehicle system include one or more natural frequency values and their corresponding damping ratio values.
[0079] The extracted time-frequency domain energy distribution features are correlated with the identified current modal parameters of the vehicle system, and frequency band matching analysis is performed to generate modal cooperative migration parameters characterizing the degree of cooperation between the two. The specific implementation involves the following steps: The first step is frequency band localization. In the time-frequency domain energy distribution features, the frequency band corresponding to each natural frequency in the identified current modal parameters of the vehicle system is located. The range of each frequency band is centered on the natural frequency value, and the width extending to both sides is determined by the half-power bandwidth of that mode. Specifically, the lower limit of the frequency band is the natural frequency minus half of the half-power bandwidth, and the upper limit of the frequency band is the natural frequency plus half of the half-power bandwidth. For example, for a mode with a natural frequency of 10 Hz and a damping ratio of 0.02, its half-power bandwidth is 10 Hz × 2 × 0.02 = 0.4 Hz, then the corresponding frequency band is 9.8 Hz to 10.2 Hz. The second step is calculating the energy concentration. For each frequency band being located, the power values of all frequency points within that band over the entire time window are summed in the time-frequency domain energy distribution characteristics to obtain the total energy of that band. Simultaneously, the sum of the power values of all frequency points within the entire analysis band of interest, for example, from 0.5 Hz to 50 Hz, over the entire time window is calculated to obtain the total energy. The total energy of this band is divided by the total energy of the entire analysis band; the resulting ratio is the energy concentration within the corresponding band for that mode. The third step is to calculate the modal cooperative migration parameters. The current energy concentration calculated in the second step is compared with a pre-stored reference energy concentration when the vehicle is in a healthy state under the same or similar operating conditions. This reference energy concentration is derived from the vehicle health state space model construction phase, calculated and stored using the same steps on historical health data under the same operating condition range. The comparison method is to calculate the relative rate of change, i.e., the current energy concentration minus the reference energy concentration, and then divided by the reference energy concentration.
[0080] For example, if the current energy concentration is 0.18 and the reference energy concentration is 0.15, the relative rate of change is calculated as (0.18-0.15) / 0.15=0.2. This relative rate of change is used as a modal cooperative migration parameter. If multiple modes are identified, the above process is repeated for each mode to generate a vector consisting of multiple relative rate of change values, which serves as the comprehensive modal cooperative migration parameter.
[0081] S5. Combining the state-related quality vector and modal cooperative offset parameters, perform causal analysis on the sensor data stream with quality anomalies to determine the source of the quality anomalies. Specifically, this is implemented as follows:
[0082] Based on the state-related quality vector, the deviation degree components of the corresponding sensor data streams with quality anomalies are extracted. The state-related quality vector is generated and input in step S3, and each component value represents the degree to which the corresponding sensor data stream deviates from the normal mode. This value is a scalar between 0 and 1. The extraction process is as follows: directly read the values of the state-related quality vector components corresponding to the sensor data streams identified in step S4 as indicating quality anomalies. These read values are the deviation degree components corresponding to each abnormal data stream. For example, if step S4 identifies quality anomalies in the front axle vibration sensor and the motor current sensor, then the component values corresponding to these two sensors are read from the current state-related quality vector. For example, if the component value of the front axle vibration sensor is 0.92 and the component value of the motor current sensor is 0.78, then 0.92 and 0.78 are the extracted deviation degree components. If there are multiple abnormal data streams, a set of deviation degree components is obtained.
[0083] Next, the correlation strength component between the corresponding anomalous event and the dynamic characteristics of the vehicle structure is extracted based on the modal cooperative migration parameters. The modal cooperative migration parameters are generated and input in step S4, and may be in the form of a scalar, such as the relative rate of change of a single mode, or a vector, such as the relative rates of change of multiple modes. The purpose of extracting the correlation strength component is to convert it into a single scalar value that can quantify the overall correlation between the anomalous event and the dynamic characteristics of the vehicle structure. The specific extraction rules are as follows: If the modal cooperative migration parameter is a scalar, its absolute value is directly taken as the correlation strength component. If the modal cooperative migration parameter is a vector, the vector norm calculation method is used. Common norm calculations include taking the maximum value of the absolute values of all elements in the vector, or calculating the arithmetic mean of the absolute values of all elements.
[0084] For example, if the modal cooperative offset parameter vector is [0.2, -0.05, 0.15], then the maximum value after taking the absolute value is 0.2, and the arithmetic mean is (0.2+0.05+0.15) / 3≈0.133. The maximum value of the absolute value is taken as the correlation strength component. This correlation strength component is also a dimensionless scalar value. The larger its value, the stronger the correlation between the observed anomaly and the dynamic characteristics of the vehicle structure.
[0085] Then, according to the numerical combination of the deviation degree component and the correlation strength component, it is matched in the predefined fault mode classification logic to output a determination result that classifies the source of quality abnormality as sensor self-interference, vehicle active control effect, or potential fault of vehicle components. The predefined fault mode classification logic is a decision rule-based classifier, and its establishment process is as follows: In the model development stage, a training data set containing multiple groups of samples is collected. Each sample contains a set of input features and a true fault category label confirmed by experts or subsequent diagnosis. The input features are two key indicators corresponding to each sample: The first indicator is the maximum value of the deviation degree components of all abnormal sensors in the sample, denoted as MaxD; the second indicator is the correlation strength component of the sample, denoted as CStrength. The true fault category label is one of the three categories: sensor self-interference, vehicle active control effect, potential fault of vehicle components. In the feature space (MaxD, CStrength), scatter plots of all training samples are drawn and colored according to their true labels. By observing the aggregation regions of samples of different categories in the two-dimensional feature space, boundary rules for dividing these three categories are manually defined or automatically induced by a simple classifier, such as a decision tree. A typical decision rule induced by observation is as follows: First, set two thresholds, namely the deviation degree threshold Td and the correlation strength threshold Tc. The deviation degree threshold Td is used to distinguish the abnormal significance degree of the sensor readings itself, and its value is determined by analyzing the distribution of MaxD values of all sensor self-interference and potential fault of vehicle components class samples in the training data set. For example, take the 70th percentile of the overall MaxD values of these two types of samples. The correlation strength threshold Tc is used to distinguish the correlation strength between the abnormality and the structural dynamics, and its value is determined by analyzing the distribution of CStrength values of all vehicle active control effect and potential fault of vehicle components class samples in the training data set. For example, take the 70th percentile of the overall CStrength values of these two types of samples. Based on the thresholds Td and Tc, the matching rules of the classification logic are defined as: Rule 1, if CStrength < Tc and MaxD >= Td for the current sample, it is determined as sensor self-interference. Rule 2, if CStrength >= Tc and MaxD < Td for the current sample, it is determined as vehicle active control effect. Rule 3, if CStrength >= Tc and MaxD >= Td for the current sample, it is determined as potential fault of vehicle components. For the current online analysis, first select the maximum value from the set of extracted deviation degree components as the current MaxD, and use the extracted correlation strength component as the current CStrength. Then, match the current numerical pair (MaxD, CStrength) with the above predefined rules. For example, assume the predefined deviation degree threshold Td = 0.85 and the correlation strength threshold Tc = 0.15.If the current MaxD=0.92 and the current CStrength=0.1, then since CStrength=0.1<0.15(Tc) and MaxD=0.92>=0.85(Td), according to Rule 1, the output judgment result is sensor self-interference. This judgment result will be passed to step S6 to guide the subsequent adjustment of the data collection strategy.
[0086] S6. Based on the judgment result, dynamically adjust the collection strategy, preprocessing method, or transmission priority of the corresponding sensor data stream, specifically as follows:
[0087] This step receives the judgment result regarding the source of the quality anomaly output from step S5 as input. This judgment result classifies the source of the quality anomaly as one of sensor self-interference, vehicle active control effects, or potential vehicle component failures. Dynamic adjustment is achieved by executing control logic within the vehicle data acquisition and communication management function. This logic differentiates the resource parameters for data acquisition, processing, and transmission based on the physical nature and urgency of the anomaly implied by the judgment result.
[0088] If the classification result indicates sensor-specific interference, the sampling frequency of the corresponding sensor data stream is reduced and anti-interference filtering preprocessing is enabled. The specific steps for reducing the sampling frequency are as follows: First, obtain the nominal sampling frequency of the sensor data stream defined during the vehicle design phase; for example, the nominal sampling frequency of a temperature sensor is 10 Hz. Second, determine a minimum effective sampling frequency based on the maximum effective change frequency of the physical quantity monitored by the sensor. The maximum effective change frequency is obtained by analyzing the theoretical rate of change of the physical quantity under all possible operating conditions of the vehicle; for example, for battery temperature, its maximum effective change frequency can be assessed as 0.1 Hz. According to the Nyquist sampling theorem, the minimum effective sampling frequency should be greater than twice the maximum effective change frequency. Considering a certain engineering margin, it can be set to 5 times the maximum effective change frequency, i.e., 0.5 Hz. When sensor-specific interference is determined, the sampling frequency is adjusted from the nominal 10 Hz to an intermediate value between the minimum effective sampling frequency of 0.5 Hz and the nominal sampling frequency of 10 Hz, such as 2 Hz. The selection principle for this intermediate value is to minimize the amount of data while ensuring effective capture of the true changes in physical quantities. Specifically, it can be four times the lowest effective sampling frequency, i.e., 0.5 Hz multiplied by 4 equals 2 Hz. Enabling anti-interference filtering preprocessing means applying a digital low-pass filter to the data stream after the data acquisition stage. The cutoff frequency of this filter is set to one-quarter of the adjusted sampling frequency of 2 Hz, i.e., 0.5 Hz, to filter out interference noise above this frequency. The filter can be designed using the window function method to design a finite impulse response filter, and its order is determined according to the required stopband attenuation, for example, set to order 64.
[0089] If the determination result classifies it as a vehicle active control effect, the current collection strategy is maintained, but the priority of the corresponding data stream in the transmission queue is increased. Maintaining the current collection strategy means not modifying the sampling frequency or any preprocessing parameters of the sensor data stream. Increasing the transmission priority is achieved by assigning a higher transmission priority level to the data stream. In vehicle network communication protocols, several transmission priority levels are predefined for different types of data streams, such as from level 1 to level 5, with level 5 being the highest. Under normal circumstances, sensor data streams use level 3 by default. When a vehicle active control effect is determined, the transmission priority level of subsequent data packets generated by that data stream is increased to level 4. The packet scheduler sends data in descending order of priority. For example, when there are data packets to be sent in the transmission buffer, the scheduler first checks for a data packet with priority level 5 and sends it, then checks for level 4, and so on. By increasing the data stream from level 3 to level 4, its queuing time in the communication link can be significantly reduced, thus achieving priority transmission.
[0090] If the assessment classifies the data stream as a potential fault in a vehicle component, the sampling frequency of the corresponding sensor data stream is increased, data compression is disabled, and it is marked for real-time transmission. The specific operation of increasing the sampling frequency follows a predefined boosting strategy table, which is developed during vehicle development based on fault diagnosis requirements. The table associates different potential fault types with recommended sampling frequency boosting factors. For example, for a vibration sensor, if the fault is suspected to be related to a rolling bearing, the sampling frequency boosting factor is set to 2 times; if it is suspected to be related to gear meshing, the boosting factor is set to 4 times. Assuming the nominal sampling frequency of the vibration sensor is 5 kHz, and the current assessment points to a potential bearing fault, the sampling frequency is increased to twice the nominal value, i.e., 10 kHz, according to the boosting strategy table. Disabling data compression means bypassing any form of data compression algorithm in the data stream processing pipeline, directly sending the raw, uncompressed sampled data to the network packet module. Marking it for real-time transmission means, at the network layer, tagging the data packets of this data stream with a service type label identifying real-time streaming media, for example, setting its Differential Service Code Point field in the IP protocol to an accelerated forwarding value. Meanwhile, in the vehicle's internal gateway, a logical channel with guaranteed bandwidth and the highest scheduling priority is configured for this data stream to ensure that its data packets can be forwarded to the target endpoint with the lowest possible latency and minimal jitter.
[0091] After all the above dynamic adjustments take effect, they will continue for a predefined monitoring period. The length of this period is set to match the time window length used for time-frequency analysis in step S4 to ensure that at least one complete analysis cycle is covered, for example, it is set to 5 seconds. After this period ends, the system will automatically restore the parameters of the corresponding sensor data stream to their default configurations before adjustment. If, during the monitoring period, step S5 produces a different judgment result based on new data, the dynamic adjustment logic of this step will be re-executed immediately according to the new judgment result, overriding the current adjustment strategy. This design enables the entire data collection and processing method to form a closed loop of perception, analysis, decision-making, and execution, adaptively optimizing resource allocation according to the real-time needs of fault diagnosis.
[0092] Example 2: Figure 2 A schematic diagram of a vehicle sensor data collection and processing system for fault detection according to the present invention is provided. The system includes:
[0093] The parameter acquisition module is used to acquire multi-sensor data streams generated by the vehicle during operation and the vehicle's current real-time operating status parameters;
[0094] The model building module is used to obtain a pre-established vehicle health state space model that represents the normal cluster distribution of multi-sensor data vectors under different real-time operating state parameters.
[0095] The vector generation module is used to determine the corresponding normal cluster distribution from the vehicle health state space model based on real-time operating status parameters, and to compare the multi-sensor data vectors with the normal cluster distribution to generate a state-related quality vector.
[0096] The parameter generation module is used to perform a correlation analysis between the time-frequency domain energy distribution characteristics of sensor data streams indicating quality anomalies in the state-correlated mass vector and the current modal parameters of the vehicle system, and generate modal correlation offset parameters.
[0097] The result determination module is used to perform causal analysis on sensor data streams with quality anomalies by combining the state-related quality vector and the modal cooperative offset parameter, and obtain the determination result about the source of the quality anomaly.
[0098] The judgment and adjustment module is used to dynamically adjust the collection strategy, preprocessing method, or transmission priority of the corresponding sensor data stream based on the judgment result.
[0099] All calculations involved in the embodiments are dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to the actual situation.
[0100] It should be noted that this invention can be deployed on the device itself to realize embedded applications, or it can run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.
[0101] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions according to the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wireless or wired transmission; wired transmission methods include optical fiber, twisted pair, coaxial cable, etc.; wireless transmission includes infrared, microwave, etc. Computer-readable storage media can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more sets of available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.
[0102] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0103] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0104] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0105] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0106] If a function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0107] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0108] In conclusion, the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for collecting and organizing vehicle sensor data for fault detection, characterized in that, include: S1. Acquire multi-sensor data streams generated during vehicle operation and the vehicle's current real-time operating status parameters; S2. Obtain a pre-established vehicle health state space model representing the normal cluster distribution of multi-sensor data vectors under different real-time operating state parameters, including: Acquire historical vehicle health operation data, which includes historical multi-sensor data vectors collected under various real-time operating state parameters. Feature extraction is performed on historical multi-channel sensor data vectors to construct a feature space; In the feature space, conditional clustering analysis is performed on the historical multi-channel sensor data vectors based on the real-time operating status parameters associated with them, so as to form a normal cluster distribution under different real-time operating status parameters. The normal cluster distribution and its corresponding real-time operating status parameters are associated and stored to construct a vehicle health state space model. S3. Determine the corresponding normal cluster distribution from the vehicle health state space model based on real-time operating status parameters, and compare the multi-sensor data vector with the normal cluster distribution to generate a state-related quality vector, including: Based on the current real-time operating status parameters, match the normal cluster distribution corresponding to the operating condition interval that is closest to the current real-time operating status parameters in the vehicle health state space model. Extract the multi-sensor data vector of the current time slice aligned with the current timestamp from the multi-sensor data stream; Calculate the multidimensional distance metric between the multi-sensor data vector of the current time slice and the center of the matched normal cluster distribution; A state-related quality vector containing the deviation components of each sensor's data stream is generated based on a multidimensional distance metric. S4. For sensor data streams indicating quality anomalies in the state-correlated mass vector, perform a co-operation analysis of their time-frequency domain energy distribution characteristics with the current modal parameters of the vehicle system to generate modal co-operational offset parameters, including: Identify sensor data streams that indicate quality anomalies from state-associated quality vectors; The time-frequency transformation of the identified sensor data stream is performed within the time window of the continuous quality anomaly to extract its time-frequency domain energy distribution characteristics. Based on the multi-sensor data stream of the vehicle within a time window, the current modal parameters of the vehicle system are identified by running modal analysis methods. The current modal parameters of the vehicle system include natural frequency and damping ratio. The extracted time-frequency domain energy distribution features are correlated with the identified current modal parameters of the vehicle system, and frequency band matching analysis is performed to generate modal cooperative migration parameters that characterize the degree of cooperation between the two. This includes: locating the frequency band corresponding to the natural frequency in the current modal parameters of the vehicle system in the time-frequency domain energy distribution; calculating the energy concentration in the corresponding location frequency band; and calculating the relative rate of change of the corresponding energy concentration with the pre-stored reference energy concentration under normal operating conditions to generate modal cooperative migration parameters. S5. Combine the state-related mass vector and modal cooperative offset parameters to perform causal analysis on the sensor data stream with quality anomalies and obtain the determination result of the source of the quality anomaly. S6. Based on the judgment result, dynamically adjust the collection strategy, preprocessing method or transmission priority of the corresponding sensor data stream.
2. The method for collecting and organizing vehicle sensor data for fault detection according to claim 1, characterized in that, S1 includes: Multiple sensor data streams are synchronously acquired via the vehicle controller local area network and the vehicle Ethernet, including vibration sensor data stream, temperature sensor data stream and current sensor data stream. The vehicle controller obtains the vehicle's current real-time operating status parameters, which include vehicle speed, motor speed, and battery voltage.
3. The method for collecting and organizing vehicle sensor data for fault detection according to claim 1, characterized in that, Calculating the multidimensional distance metric between the multi-sensor data vector of the current time slice and the center of the matched normal cluster distribution includes: calculating the Mahalanobis distance between the multi-sensor data vector of the current time slice and the center of the matched normal cluster distribution, wherein the Mahalanobis distance is calculated based on the inversion of the covariance matrix of the matched normal cluster distribution and taking into account the statistical correlation between the dimensions of each sensor data.
4. The method for collecting and organizing vehicle sensor data for fault detection according to claim 1, characterized in that, S5 include: Based on the state-correlation quality vector, the deviation component of the sensor data stream corresponding to the quality anomaly is extracted; Based on modal cooperative migration parameters, the correlation strength components between corresponding abnormal events and vehicle structural dynamic characteristics are extracted; Based on the numerical combination of the deviation degree component and the correlation strength component, a matching is performed in the predefined fault mode classification logic to output a judgment result that classifies the source of quality anomalies as sensor self-interference, vehicle active control effect, or potential vehicle component failure.
5. The method for collecting and organizing vehicle sensor data for fault detection according to claim 1, characterized in that, S6 include: Based on the classification of the sources of quality anomalies in the judgment results, if the classification is sensor interference itself, the sampling frequency of the corresponding sensor data stream will be reduced and anti-interference filtering preprocessing will be enabled. If classified as a vehicle active control effect, the current collection strategy is maintained, but the priority of the corresponding data stream in the transmission queue is increased; If the fault is classified as a potential vehicle component failure, the sampling frequency of the corresponding sensor data stream is increased, data compression is disabled, and it is marked as real-time transmission.
6. A vehicle sensor data collection and processing system for fault detection, used to implement the vehicle sensor data collection and processing method for fault detection as described in any one of claims 1-5, characterized in that, include: The parameter acquisition module is used to acquire multi-sensor data streams generated by the vehicle during operation and the vehicle's current real-time operating status parameters; The model building module is used to obtain a pre-established vehicle health state space model that represents the normal cluster distribution of multi-sensor data vectors under different real-time operating state parameters. The vector generation module is used to determine the corresponding normal cluster distribution from the vehicle health state space model based on real-time operating status parameters, and to compare the multi-sensor data vectors with the normal cluster distribution to generate a state-related quality vector. The parameter generation module is used to perform a correlation analysis between the time-frequency domain energy distribution characteristics of sensor data streams indicating quality anomalies in the state-correlated mass vector and the current modal parameters of the vehicle system, and generate modal correlation offset parameters. The result determination module is used to perform causal analysis on sensor data streams with quality anomalies by combining the state-related quality vector and the modal cooperative offset parameter, and obtain the determination result about the source of the quality anomaly. The judgment and adjustment module is used to dynamically adjust the collection strategy, preprocessing method, or transmission priority of the corresponding sensor data stream based on the judgment result.
Citation Information
Patent Citations
Tire pressure sensor activation method and device, equipment and storage medium
CN119459192A
Automobile data analysis method and system based on intelligent diagnostic instrument
CN119541080A