A vehicle fault analysis method, system, device and train

By performing correlation grouping and feature variable analysis on vehicle sensor data, and using intra-group and inter-group feature variables to determine vehicle faults, the problem of high detection difficulty caused by sensor data correlation is solved, and more efficient fault identification is achieved.

CN116625428BActive Publication Date: 2026-03-31CRRC QINGDAO SIFANG CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

The difficulty in detecting vehicle faults in existing technologies is mainly due to the strong correlation between multiple data detected by sensors, which makes data analysis difficult and makes it impossible to accurately determine vehicle faults.

Method used

The first dataset is generated by performing correlation grouping processing on the raw feature variables collected by the sensors on the vehicle. Then, the second dataset is generated by performing intra-group feature variable analysis based on the first dataset. Finally, outlier detection algorithms and clustering algorithms are used to identify abnormal carriages.

Benefits of technology

It improves the accuracy and effectiveness of vehicle fault diagnosis, enabling more accurate identification of abnormal carriages and reducing misjudgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116625428B_ABST
    Figure CN116625428B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle fault analysis method, system and device and a train, and relates to the field of vehicle control. Raw feature variables collected by each sensor on the vehicle are subjected to correlation grouping processing to generate a first data set, in-group analysis is performed based on each data point in the first data set, in-group data abnormal correlation groups are determined, and feature values of each correlation group in a second data set are set as unique feature variables of the correlation group, and an abnormal car is determined based on the unique feature variables in each data point in the second data set and / or based on raw feature variables in the in-group data abnormal correlation group, so that the abnormal car is judged based on in-group and / or inter-group feature variables. The correlation between each raw feature variable is utilized, in-group feature variable analysis is performed based on the first data set, inter-group feature variable analysis is performed based on the second data set, the vehicle fault is judged, and the accuracy and effectiveness of the judgment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle control, and in particular to a vehicle fault analysis method, system, device, and train. Background Technology

[0002] In existing technologies, vehicle fault detection typically involves deploying numerous sensors throughout the vehicle to collect various data points and perform fault detection and diagnosis. Based on the results, control, inspection, and maintenance operations are then performed. However, a vehicle is a complex system. Due to the interactions between its subsystems and the deep coupling of multiple physical quantities in a fault, there are strong correlations between the multiple data points detected by sensors and the various variables. This makes data analysis difficult, and fault detection cannot be performed based on a single data point, resulting in significant challenges in vehicle fault detection. Summary of the Invention

[0003] The purpose of this invention is to provide a vehicle fault analysis method, system, device, and train. By utilizing the correlation between various original feature variables, intra-group feature variable analysis is performed based on a first dataset, and inter-group feature variable analysis is performed based on a second dataset to determine vehicle faults, thereby improving the accuracy and effectiveness of the determination.

[0004] To address the aforementioned technical problems, this invention provides a vehicle fault analysis method, comprising:

[0005] The raw feature variables collected by various sensors on the vehicle are subjected to correlation grouping processing to generate a first dataset, which includes multiple data points, and each data point includes multiple correlation groups.

[0006] The feature values ​​of each correlation group in each data point in the first dataset are respectively set as unique feature variables in each correlation group of each data point to generate the second dataset;

[0007] Based on the data points in the first dataset, perform intra-group analysis to determine the grouping of abnormal data correlations within the group;

[0008] The abnormal carriages are determined based on the unique feature variables in each of the data points in the second dataset and / or based on the original feature variables in the grouping of data anomalies within the group.

[0009] Preferably, before generating the first dataset by performing correlation grouping processing on the raw feature variables collected by each sensor on the vehicle, the following steps are also included:

[0010] Remove invalid feature variables from the original feature variables collected by each of the sensors on the vehicle, and determine the remaining original feature variables.

[0011] Preferably, invalid feature variables are removed from the original feature variables collected by each of the sensors on the vehicle, and the remaining original feature variables are determined, including:

[0012] The invalid feature variables in the original feature variables collected by each of the sensors on the vehicle are cleared, and the original feature variables after clearing the invalid feature variables are resampled to determine the original feature variables.

[0013] Preferably, the feature values ​​of each correlation group in each data point of the first dataset are respectively set as unique feature variables in each correlation group of each data point to generate a second dataset, including:

[0014] The median of each correlation group in each data point of the first dataset is set as the unique feature variable in each correlation group of each data point to generate the second dataset.

[0015] Preferably, the feature values ​​of each correlation group in each data point of the first dataset are respectively set as unique feature variables in each correlation group of each data point to generate a second dataset, including:

[0016] Set the feature values ​​of each correlation group in each data point in the first dataset as unique feature variables in each correlation group in each data point;

[0017] The second dataset is generated by setting the mean and standard deviation of the unique feature variables of the same correlation group of a predetermined number of neighboring data points in the first dataset as additional feature variables of the data points according to the time order of the data points.

[0018] Preferably, determining the abnormal carriages based on the unique feature variables in each data point of the second dataset and / or based on the original feature variables in the grouping of data anomalies within the group includes:

[0019] Outlier detection algorithms are used to detect outliers in each unique feature variable and additional feature variable of each data point in the second dataset to identify the abnormal carriages.

[0020] And / or, based on the outlier detection algorithm, outlier detection is performed on each original and feature variable in the group of abnormal correlation data within the group to identify the abnormal carriage.

[0021] Preferably, outlier detection is performed on each unique feature variable and the additional feature variable of each data point in the second dataset based on an outlier detection algorithm to determine the abnormal carriage, including:

[0022] The joint probability density of each data point in the second dataset is calculated based on a multidimensional Gaussian distribution.

[0023] The anomaly score of each data point is determined based on the joint probability density of each data point.

[0024] Outlier detection algorithms are used to detect outliers in each unique feature variable and additional feature variable of each data point in the second dataset, and the abnormal carriages are determined based on the abnormal scores of each data point.

[0025] Preferably, after performing intra-group analysis based on each data point in the first dataset to determine the intra-group data abnormal correlation grouping, the method further includes:

[0026] Based on the clustering algorithm, the original feature variables in the abnormal correlation grouping of the data within the group are visualized and clustered to identify the abnormal carriages.

[0027] Preferably, based on each data point in the first dataset, an in-group analysis is performed to determine in-group data anomaly correlation grouping, including:

[0028] For each original feature variable in each of the correlation groups in the first dataset, perform pairwise correlation judgment, and set the correlation group where the original feature variable with abnormal correlation is located as the abnormal correlation group of the data in the group.

[0029] To address the aforementioned technical problems, the present invention provides a vehicle fault analysis system, comprising:

[0030] A grouping unit is used to perform correlation grouping processing on the raw feature variables collected by various sensors on the vehicle to generate a first dataset. The first dataset includes multiple data points, and each data point includes multiple correlation groups.

[0031] The variable setting unit is used to set the feature values ​​of each correlation group in each data point in the first dataset as unique feature variables in each correlation group of each data point, thereby generating the second dataset;

[0032] The analysis unit is used to perform intra-group analysis based on each data point in the first dataset to determine the intra-group data abnormal correlation grouping;

[0033] The determining unit is used to determine the abnormal carriage based on the unique feature variables in each of the data points in the second dataset and / or based on the original feature variables in the grouping of data anomalies within the group.

[0034] To solve the above-mentioned technical problems, the present invention provides a vehicle fault analysis device, comprising:

[0035] Memory, used to store computer programs;

[0036] A processor is used to execute the computer program to implement the steps of the vehicle fault analysis method as described above.

[0037] To address the aforementioned technical problems, the present invention provides a train, including the vehicle fault analysis device as described above.

[0038] This application provides a vehicle fault analysis method, system, device, and train, relating to the field of vehicle control. The method involves performing correlation grouping processing on the raw feature variables collected by various sensors on the vehicle to generate a first dataset. Based on each data point in the first dataset, intra-group analysis is performed to determine abnormal correlation groups within each group. The feature values ​​of each correlation group in a second dataset are then set as unique feature variables for each correlation group. Abnormal carriages are identified based on the unique feature variables in each data point of the second dataset and / or based on the raw feature variables in the abnormal correlation groups within each group, thus achieving the judgment of abnormal carriages based on intra-group and / or inter-group feature variables. By utilizing the correlation between various raw feature variables, intra-group feature variable analysis is performed based on the first dataset, and inter-group feature variable analysis is performed based on the second dataset to determine vehicle faults, improving the accuracy and effectiveness of the judgment. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the prior art and embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 A flowchart illustrating a vehicle fault analysis method provided by the present invention;

[0041] Figure 2 A diagram illustrating the correlation between variables within a group, provided by this invention;

[0042] Figure 3 A schematic diagram provided in this application showing data point visualization processing using a clustering algorithm;

[0043] Figure 4This invention provides a schematic diagram of the structure of a vehicle fault analysis system.

[0044] Figure 5 This is a schematic diagram of the structure of a vehicle fault analysis device provided by the present invention. Detailed Implementation

[0045] The core of this invention is to provide a vehicle fault analysis method, system, device, and train. By utilizing the correlation between various original feature variables, intra-group feature variable analysis is performed based on a first dataset, and inter-group feature variable analysis is performed based on a second dataset to determine vehicle faults, thereby improving the accuracy and effectiveness of the determination.

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] Please refer to Figure 1 , Figure 1 A flowchart illustrating a vehicle fault analysis method provided by the present invention includes:

[0048] S11: Perform correlation grouping processing on the raw feature variables collected by various sensors on the vehicle to generate the first dataset. The first dataset includes multiple data points, and each data point includes multiple correlation groups.

[0049] Existing vehicles already have a large number of sensors deployed, facilitating the collection of various vehicle data, such as temperature, current, force, spring pressure, vertical indices, and lateral indices. Current technologies directly analyze and diagnose vehicle faults based on this data for control, detection, and maintenance tasks. However, due to the interactions between subsystems in complex systems and the deep coupling relationships between various physical quantities in faults, strong correlations can exist between different data points, increasing the difficulty of analysis. For example, vehicle temperature includes temperatures from multiple locations or devices, and these temperatures exhibit certain correlations. There is a correlation between the stator temperatures of motors on different axes (1-4) and motors on different axes (2), and between the bearing temperatures at the drive ends of motors on different axes (1-4). Relying solely on the temperature of one device or location is insufficient to accurately determine whether a vehicle fault has occurred.

[0050] To address the aforementioned technical issues, this application does not directly determine vehicle malfunction based on individual original feature variables. Instead, it performs correlation grouping processing on the original feature variables. For example, the stator temperatures of motors on axes 1-4 and 2 are grouped into a "motor stator temperature" group, and the drive end bearing temperatures of motors on axes 1-4 and 2 are grouped into a "motor drive end bearing temperature" group. In other words, each original feature variable in the first dataset is distributed according to its corresponding correlation group. Since the sensors collect data sequentially, one data point is collected at each time point. Therefore, the first dataset includes multiple data points. The data collected by each sensor at each time point becomes the data within each data point. After processing, each data point includes multiple correlation groups, and each correlation group includes one or more correlated original feature variables.

[0051] For example, the first data point corresponds to the first time point, and the original feature variables of the data collected by each sensor at the first time point after processing are the original feature variables of the first data point. Moreover, the original feature variables of the first data point are distributed according to the grouping after correlation grouping, that is, the original feature variables with correlation are the original feature variables in the same correlation group. The second data point corresponds to the second time point, and the original feature variables of the data collected by each sensor at the second time point after processing are the original feature variables of the second data point. Moreover, the original feature variables of the second data point are distributed according to the grouping after correlation grouping, that is, the original feature variables with correlation are the original feature variables in the same correlation group. And so on, the first dataset includes multiple data points, and the original feature variables in each data point are all processed by correlation grouping.

[0052] To further describe this in tabular form, we can assume that each row of original feature variables represents the original feature variables obtained from data collected at the same time point after processing, and each column of original feature variables represents the original feature variables obtained from data collected by the same sensor at different time points after processing. Moreover, each original feature variable in each row corresponds to its own vehicle model, train number, and carriage number. Thus, we can determine the original feature variables in a certain carriage within a certain data point and in a certain correlation group.

[0053] As a preferred embodiment, before generating the first dataset by performing correlation grouping processing on the raw feature variables collected by various sensors on the vehicle, the following steps are also included:

[0054] Remove invalid feature variables from the original feature variables collected by each sensor on the vehicle, and determine the remaining original feature variables.

[0055] Specifically, when processing the data collected by the sensor and converting it into original feature variables, invalid feature variables in each original feature variable collected by the sensor can be removed, such as invalid data points, or invalid original feature variables in each data point, in order to reduce the amount of data processing.

[0056] Accordingly, since the data points are sorted according to time data and each data point corresponds to a specific carriage, a smooth window can be used to fill in missing values ​​for each original collected feature variable after removing invalid feature variables. For example, the average value of the original collected feature variables of three data points with adjacent time points can be used to fill in the missing value. Specifically, if the time difference between two adjacent data points is normally 1 second, but the time difference between the second and third data points is 2 seconds, then there is a missing data point between the second and third data points. In this case, the average value of the same original collected feature variable in the first, second, and third data points can be used to fill in the missing value between the second and third data points, or the average value of the same original collected feature variable in the second, third, and fourth data points can be used to fill in the missing value between the second and third data points.

[0057] As a preferred embodiment, invalid feature variables are removed from the original feature variables collected by each sensor on the vehicle, and the remaining original feature variables are determined, including:

[0058] Remove invalid feature variables from the original feature variables collected by each sensor on the vehicle, and then resample each original feature variable after removing invalid feature variables to determine each original feature variable.

[0059] After removing all invalid feature variables, considering that there may still be a large number of data points, we can resample each of the original collected feature variables after removing invalid feature variables, sampling only a small number of data points to determine each original feature variable.

[0060] For example, if the time difference between two adjacent data points in each of the original collected feature variables after removing invalid feature variables is 1 second, then resampling will only collect data points in each of the original collected feature variables where the time difference between two adjacent data points is 3 seconds (for example), such as collecting the first data point, the fourth data point, the seventh data point, and so on in each of the original collected feature variables, in order to reduce the amount of data processing.

[0061] It should be noted that, for example, in this application, the eight original feature variables collected by the four stator temperature sensors for the 1-channel 1-4 axis motors and the four stator temperature sensors for the 2-channel 1-4 axis motors are grouped into groups based on motor stator temperature correlation; the eight original feature variables collected by the four drive-end bearing temperature sensors for the 1-channel 1-4 axis motors and the four drive-end bearing temperature sensors for the 2-channel 1-4 axis motors are grouped into groups based on motor drive-end bearing temperature correlation; the eight original feature variables collected by the four non-drive-end bearing temperature sensors for the 1-channel 1-4 axis motors and the four non-drive-end bearing temperature sensors for the 2-channel 1-4 axis motors are grouped into groups based on motor drive-end bearing temperature correlation; The eight raw characteristic variables collected by the four transmission end bearing temperature sensors are grouped into groups based on the temperature correlation of the bearings at the non-transmission end of the motor; the sixteen raw characteristic variables collected by the eight axle box temperature sensors (one channel for axes 1-8) and the eight axle box temperature sensors (two channels for axes 1-8) are grouped into groups based on the temperature correlation of the axle box; the four motor-side bearing temperature sensors (one channel for axes 1-4), the four wheel-side bearing temperature sensors (one channel for axes 1-4), the four motor-side bearing temperature sensors (one channel for axes 1-4), and the four large gearbox temperature sensors (one channel for axes 1-4) are also grouped into groups based on the temperature correlation of the large gearbox. Thirty-two original characteristic variables collected by four wheel box wheel-side bearing temperature sensors, four 1-4 axis small gearbox motor-side bearing temperature sensors, four 1-4 axis small gearbox wheel-side bearing temperature sensors, four 1-4 axis large gearbox motor-side bearing temperature sensors, and four 1-4 axis large gearbox wheel-side bearing temperature sensors are classified into gear temperature correlation groups; two original characteristic variables collected by one motor current sensor and two motor current sensors are classified into current correlation groups; four original characteristic variables collected by one motor actual force sensor, two motor actual force sensors, one motor set force sensor, and two motor set force sensors are classified into force correlation groups; two original characteristic variables collected by the air spring pressure press_1 sensor and air spring pressure press_2 sensor are classified into air spring pressure correlation groups; one vertical index v1 sensor and two vertical index v2 sensors are classified into vertical index correlation groups; one lateral index t1 sensor and two lateral index t2 sensors are classified into lateral index correlation groups. Of course, in the case of multiple data points, each data point is a set of raw feature variables collected by each sensor at different time points. After the raw feature variables are grouped by correlation, the raw feature variables in each data point are distributed according to the correlation group.

[0062] S12: Set the feature values ​​of each correlation group in each data point of the first dataset as unique feature variables in each correlation group of each data point to generate the second dataset;

[0063] After grouping the original feature variables of each data point in the first dataset, since the original feature variables in the same correlation group are correlated, the feature values ​​of the original feature variables in each correlation group of each data point in the first dataset are taken. That is, each correlation group of each data point in the second dataset includes only one unique feature variable. The unique feature variable in the correlation group represents each original feature variable in the correlation group.

[0064] Based on this, it can be seen that the difference between the first dataset and the second dataset is that each correlation group of each data point in the first dataset includes at least one original feature variable, while each correlation group of each data point in the second dataset includes only one unique feature variable, that is, each data point in the second dataset includes multiple unique feature variables.

[0065] S13: Based on each data point in the first dataset, perform intra-group analysis to determine the grouping of abnormal correlations in the intra-group data;

[0066] Within-group analysis is performed on each data point in the first dataset. That is, since the original feature variables in each correlation group are correlated, within-group analysis can be performed only on the feature variables in the same correlation group to identify abnormal correlation groups within the group data.

[0067] As a preferred embodiment, within-group analysis is performed based on each data point in the first dataset to determine the grouping of abnormal correlations within the data, including:

[0068] For each original feature variable in each correlation group in the first dataset, perform pairwise correlation judgment and set the correlation group where the original feature variable with abnormal correlation is located as the in-group data abnormal correlation group.

[0069] When performing within-group analysis based on the first dataset, the anomalies within the group can be determined by assessing the pairwise correlations of the original feature variables within the same correlation group. Specifically, grouping can be done according to different train models, train numbers, and carriage numbers; that is, data points from the same carriage can be grouped together, and a variable relationship diagram can be drawn, such as... Figure 2 As shown, Figure 2 The diagram illustrates the correlation between variables within a group, as provided by this invention. As can be seen, taking V1 and V2 in the same correlation group as an example, there are one or more carriages where the correlation between V1 and V2 is different from that between other carriages. Therefore, the correlation group where V1 and V2 are located is an abnormal correlation group within the group.

[0070] S14: Identify abnormal carriages based on unique feature variables in each data point of the second dataset and / or original feature variables in the grouping of data anomalies within the group.

[0071] After determining the second dataset, the faulty carriages can be identified directly based on the unique feature variables of each data point in the second dataset, or the faulty carriages can be identified by grouping the data anomalies within the group determined by the first dataset. In other words, the faulty carriages identified by the feature variables within and / or between groups can be used as the final fault analysis results.

[0072] In summary, this application utilizes the correlation between various original feature variables to perform intra-group feature variable analysis based on the first dataset and inter-group feature variable analysis based on the second dataset to determine vehicle faults, thereby improving the accuracy and effectiveness of the determination.

[0073] Based on the above embodiments:

[0074] In a preferred embodiment, the feature values ​​of each correlation group in each data point of the first dataset are respectively set as unique feature variables in each correlation group of each data point to generate a second dataset, including:

[0075] The median of each correlation group in each data point of the first dataset is set as the unique feature variable in each correlation group of each data point to generate the second dataset.

[0076] In this embodiment, when determining the second dataset based on the first dataset, the median of each original feature variable in each correlation group is specifically set as the unique feature variable in that correlation group. For example, the median of each motor stator temperature in the motor stator temperature correlation group among the data points is set as the unique feature variable in the motor stator temperature correlation group. Of course, this application considers that the median can avoid the adverse effects of excessively large or small values ​​on the unique feature variable, so the median is used as the unique feature variable. However, the mean can also be chosen as the unique feature variable according to the actual situation, and this application does not limit this.

[0077] In a preferred embodiment, the feature values ​​of each correlation group in each data point of the first dataset are respectively set as unique feature variables in each correlation group of each data point to generate a second dataset, including:

[0078] Set the feature values ​​of each correlation group in each data point in the first dataset as unique feature variables in each correlation group of each data point;

[0079] The second dataset is generated by setting the mean and standard deviation of the unique feature variables of the same correlation group of a preset number of neighboring data points in the first dataset as additional feature variables of the data points according to the time order of the data points.

[0080] In this embodiment, after determining the unique feature variables in each correlation group, the mean and standard deviation of the unique feature variables of the same correlation group of neighboring data points of a preset number are also set as additional feature variables of the data points. For example, the mean and standard deviation of the unique feature variables of the same correlation group in the first to thirtieth data points are set as additional feature variables of the first to thirtieth data points, and the mean and standard deviation of the unique feature variables of the same correlation group in the thirty-first to sixtieth data points are set as additional feature variables of the thirty-first to sixtieth data points. Of course, since each data point includes multiple correlation groups, multiple additional feature variables can be added to each data point.

[0081] As a preferred embodiment, identifying abnormal carriages based on unique feature variables in each data point of the second dataset and / or based on original feature variables in the grouping of data anomaly correlations within the group includes:

[0082] Outlier detection algorithms are used to detect outliers in each unique feature variable and additional feature variable of each data point in the second dataset in order to identify abnormal carriages.

[0083] And / or, based on the outlier detection algorithm, outlier detection is performed on each original and feature variable in the group of abnormal correlation data within the group to identify abnormal carriages.

[0084] When identifying abnormal carriages based on unique feature variables of each data point in the second dataset, outlier detection can be performed on each unique feature variable of each data point in the second dataset. First, the mean and variance of the unique feature variables of each data point in the second dataset are standardized. Then, principal component analysis is used to reduce the dimensionality of the second dataset and extract principal components. An outlier detection model is established, and the extracted principal component data is used to fit the model. The outlier detection model here can employ techniques such as Isolation Forest, Elliptic Envelope, Local Outlier Factor, or One-Class Support Vector Machine. The system uses any one of four models (SVM, SVM, etc.) to perform outlier detection. Based on a preset outlier threshold, each data point in the second dataset is categorized as either outlier or non-outlier. The system outputs the anomaly score and type (outlier status) for each data point. It calculates the proportion of outliers identified in each carriage based on outlier detection, as well as the average anomaly score for all data points in each carriage. The average anomaly score for all data points in each carriage determines whether a carriage is faulty, i.e., whether it is an abnormal carriage. The proportion of outliers in each carriage determines whether the calculation results are abnormal. If the proportion is low, it indicates possible interference or other factors, and the carriage may not actually be faulty. Therefore, if a carriage has a high proportion of outliers and an average anomaly score higher than the preset anomaly score, it can be determined that the carriage is an abnormal carriage.

[0085] Accordingly, when identifying abnormal carriages based on the original feature variables in the intra-group data anomaly correlation grouping of each data point in the first dataset, outlier detection can be performed on each original feature variable in the intra-group data anomaly correlation grouping of each data point in the first dataset. First, the mean and variance of the original feature variables of each data point in the first dataset are standardized. Then, principal component analysis is used to reduce the dimensionality of the first dataset and extract principal components. An outlier detection model is established, and the extracted principal component data is used to fit the model. The outlier detection model here can be an Isolation Forest, an Elliptic Envelope, a Local Outlier Factor, or a One-Class Support Vector Machine. The system uses any one of four models (SVM, SVM, etc.) to perform outlier detection. Based on a preset outlier threshold, each data point in the first dataset is categorized as either outlier or non-outlier. The system outputs the anomaly score and type (outlier status) for each data point. It calculates the proportion of outliers identified in each carriage based on outlier detection, as well as the average anomaly score for all data points in each carriage. The average anomaly score for all data points in each carriage determines whether the carriage is faulty, i.e., whether it is an abnormal carriage. The proportion of outliers in each carriage determines whether the calculation results are abnormal. If the proportion is low, it indicates possible interference or other factors, and the carriage may not actually be faulty. Therefore, if the proportion of outliers in a carriage is high, and the average anomaly score is higher than the preset anomaly score, then that carriage can be identified as an abnormal carriage.

[0086] In addition, an interval range for abnormal scores can be set. When the average abnormal score falls within different interval ranges, the fault level of the carriage can be determined. For example, the higher the average abnormal score, the higher the fault level of the corresponding carriage.

[0087] As a preferred embodiment, outlier detection is performed on each unique feature variable and additional feature variable of each data point in the second dataset based on an outlier detection algorithm to identify abnormal carriages, including:

[0088] Calculate the joint probability density of each data point in the second dataset based on the multidimensional Gaussian distribution;

[0089] The anomaly score for each data point is determined based on the joint probability density of each data point.

[0090] Outlier detection algorithms are used to detect outliers in each unique feature variable and additional feature variable of each data point in the second dataset, and abnormal carriages are identified based on the abnormal scores of each data point.

[0091] In this embodiment, statistical analysis of inter-group feature variables is performed on the second dataset based on a multidimensional Gaussian distribution. First, a multidimensional Gaussian distribution model is established, and the mean and variance of the unique and additional feature variables for each data point in the second dataset are calculated to fit the model. The fitted multidimensional Gaussian distribution model is then used to calculate the joint probability density of each data point. Each data point is scored based on its joint probability density, and data points with a joint probability density below a preset threshold are designated as outliers, assigned higher anomaly scores, and their anomaly scores and classifications are output. The joint probability density of data point i is denoted as fi, and the anomaly score for data point i is...

[0092]

[0093] As a preferred embodiment, after performing within-group analysis on each data point in the first dataset to determine the grouping of abnormal correlations within the data, the method further includes:

[0094] Based on clustering algorithms, the original feature variables in the abnormal correlation grouping of data within the group are visualized and clustered to identify abnormal carriages.

[0095] After identifying the abnormal correlation groups within the data set, clustering algorithms can be used to analyze the characteristic variables within the groups to identify abnormal carriages. Specifically, the original characteristic variables belonging to the same correlation group in the first dataset are extracted, and the mean and variance of each original characteristic variable are calculated and standardized. Principal component analysis is used to reduce the dimensionality of each data point in the first dataset and extract principal components. The K-means method is used to cluster the original characteristic variables in the same correlation group in the first dataset, with the number of clusters being the number of different car models, train numbers, and carriage numbers in the first dataset. Correlation analysis and visualization are then performed. A heatmap is used to visualize the correlation between the clustered categories and the data points corresponding to each carriage. A distribution map of the clustered data points on a two-dimensional plane is plotted using the dimensionality-reduced principal components as coordinates, and compared with the distribution map of the data points of each carriage on a two-dimensional plane to analyze the possibility of significant differences between a single carriage and other carriages. Figure 3 As shown, Figure 3 This application provides a schematic diagram of data point visualization processing using a clustering algorithm. The diagram shows how the average anomaly score of each carriage, determined by an outlier detection algorithm, combined with the distribution of data points within each carriage as determined by a clustering algorithm, can further accurately determine whether a carriage is faulty. For example, if the average anomaly score of carriage E68_CR400AF-BZ-2249_4 is higher than a preset anomaly score, and... Figure 3The data points corresponding to carriage E68_CR400AF-BZ-2249_4 shown in the figure are concentrated in the fifth cluster, so the second carriage is identified as an abnormal carriage.

[0096] Please refer to Figure 4 , Figure 4 This invention provides a schematic diagram of a vehicle fault analysis system, which includes:

[0097] Grouping unit 41 is used to perform correlation grouping processing on the raw feature variables collected by various sensors on the vehicle to generate a first dataset. The first dataset includes multiple data points, and each data point includes multiple correlation groups.

[0098] The variable setting unit 42 is used to set the feature values ​​of each correlation group in each data point in the first dataset as unique feature variables in each correlation group of each data point, thereby generating the second dataset;

[0099] Analysis unit 43 is used to perform intra-group analysis based on each data point in the first dataset to determine the intra-group data abnormal correlation grouping;

[0100] The determination unit 44 is used to determine the abnormal carriages based on the unique feature variables in each data point of the second dataset and / or the original feature variables in the grouping of data anomalies within the group.

[0101] For an introduction to the vehicle fault analysis system provided by the present invention, please refer to the above method embodiments; the present invention will not be described in detail here.

[0102] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the structure of a vehicle fault analysis device provided by the present invention.

[0103] The vehicle fault analysis device includes:

[0104] Memory 51 is used to store computer programs;

[0105] The processor 52 is used to execute computer programs to implement the steps of the vehicle fault analysis method described above.

[0106] For a description of the vehicle fault analysis device provided by the present invention, please refer to the above method embodiments; the present invention will not be described again here.

[0107] The train of the present invention includes the aforementioned vehicle fault analysis device.

[0108] For an introduction to the train provided by this invention, please refer to the above method embodiments; the invention itself will not be described in detail here.

[0109] The computer-readable storage medium of the present invention stores a computer program, which, when executed by the processor 52, implements the steps of the vehicle fault analysis method described above.

[0110] For a description of the computer-readable storage medium provided by the present invention, please refer to the above method embodiments; the present invention will not be described again here.

[0111] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0112] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A vehicle failure analysis method characterized by, The method comprises: correlation grouping of original feature variables collected by various sensors on a vehicle to generate a first data set, the first data set comprising a plurality of data points, and each data point comprising a plurality of correlation groups; setting the feature values in each correlation group in each data point in the first data set as the unique feature variable in each correlation group in each data point, respectively, to generate a second data set; intra-group analysis based on each data point in the first data set to determine intra-group data abnormal correlation groups; determining an abnormal vehicle compartment based on the unique feature variables in each data point in the second data set and the original feature variables in the intra-group data abnormal correlation groups; setting the feature values in each correlation group in each data point in the first data set as the unique feature variable in each correlation group in each data point, respectively, to generate a second data set, comprising: setting the feature values in each correlation group in each data point in the first data set as the unique feature variable in each correlation group in each data point; setting the mean and standard deviation of the unique feature variables of the same correlation group of a predetermined number of adjacent data points in the first data set as additional feature variables of the data points in the order of time of the data points to generate the second data set; determining an abnormal vehicle compartment based on the unique feature variables in each data point in the second data set and the original feature variables in the intra-group data abnormal correlation groups, comprising: performing outlier detection on each unique feature variable and the additional feature variables of each data point in the second data set based on an outlier detection algorithm, and performing outlier detection on each original feature variable in the intra-group data abnormal correlation groups based on the outlier detection algorithm to determine the abnormal vehicle compartment; performing outlier detection on each unique feature variable and the additional feature variables of each data point in the second data set based on an outlier detection algorithm to determine the abnormal vehicle compartment, comprising: calculating the joint probability density of each data point in the second data set based on a multi-dimensional Gaussian distribution; determining the anomaly score of each data point based on the joint probability density of each data point; performing outlier detection on each unique feature variable and the additional feature variables of each data point in the second data set based on an outlier detection algorithm, and determining the abnormal vehicle compartment based on the anomaly score of each data point; calculating the joint probability density of each data point in the second data set based on a multi-dimensional Gaussian distribution, comprising: establishing a multi-dimensional Gaussian distribution model, calculating the mean and variance of the unique feature variables and the additional feature variables in each data point in the second data set to fit the multi-dimensional Gaussian distribution model; and using the fitted multi-dimensional Gaussian distribution model to calculate the joint probability density of each data point; determining the anomaly score of each data point based on the joint probability density of each data point, comprising: According to the joint probability density of each data point, each data point is scored, data points with joint probability density lower than a preset probability density threshold are recorded as abnormal data points, the abnormal data points are given higher abnormal scores, and the abnormal scores and the abnormal classification of each data point are output; the joint probability density of the data points is recorded as , and the abnormal score of the data point is: ; determining an intra-group data abnormal correlation group based on each of the data points in the first data set, comprising: judging a correlation between each original feature variable in each of the correlation groups in the first data set, and setting a correlation group in which an original feature variable with abnormal correlation is located as the intra-group data abnormal correlation group.

2. The vehicle trouble analysis method according to claim 1, characterized by, Before the original feature variables collected by each sensor on the vehicle are processed by correlation grouping to generate a first data set, the method further comprises: clearing invalid feature variables in the original feature variables collected by each sensor on the vehicle, and determining each of the remaining original feature variables.

3. The vehicle fault analysis method of claim 2, wherein, clearing invalid feature variables in the original feature variables collected by each sensor on the vehicle, and determining each of the remaining original feature variables, comprising: clearing the invalid feature variables in the original feature variables collected by each sensor on the vehicle, and resampling each of the original feature variables after the invalid feature variables are cleared, to determine each of the original feature variables.

4. The vehicle fault analysis method of claim 1, wherein, setting a feature value in each correlation group in each of the data points in the first data set as a unique feature variable in each correlation group in each of the data points, to generate a second data set, comprising: setting a median in each correlation group in each of the data points in the first data set as a unique feature variable in each correlation group in each of the data points, to generate the second data set.

5. The vehicle fault analysis method of claim 1, wherein, After determining the intra-group data abnormal correlation group based on each of the data points in the first data set, the method further comprises: performing visual clustering processing on each original feature variable in the intra-group data abnormal correlation group based on a clustering algorithm, to determine the abnormal vehicle compartment.

6. A vehicle fault analysis system characterized by, comprising: a grouping unit configured to process original feature variables collected by each sensor on a vehicle by correlation grouping, to generate a first data set, the first data set comprising a plurality of data points, and each of the data points comprising a plurality of correlation groups; a variable setting unit configured to set a feature value in each correlation group in each of the data points in the first data set as a unique feature variable in each correlation group in each of the data points, to generate a second data set; an analysis unit configured to determine an intra-group data abnormal correlation group based on each of the data points in the first data set; a determination unit configured to determine an abnormal vehicle compartment based on the unique feature variables in each of the data points in the second data set and original feature variables in the intra-group data abnormal correlation group; the variable setting unit is specifically configured to: setting a feature value in each of the correlation groups in each of the data points in the first data set as a unique feature variable in each of the correlation groups in each of the data points respectively, setting a mean value and a standard deviation of the unique feature variables in the same correlation group of a preset number of adjacent data points in the first data set as an additional feature variable of the data points in a time sequence of the data points, and generating the second data set; The determination unit is specifically configured to: detect outliers of each of the unique feature variables and the additional feature variables of each of the data points in the second data set based on an outlier detection algorithm, and detect outliers of each of the original feature variables in the in-group data abnormal correlation group based on the outlier detection algorithm, to determine the abnormal carriage; detect outliers of each of the unique feature variables and the additional feature variables of each of the data points in the second data set based on an outlier detection algorithm, to determine the abnormal carriage, including: calculating a joint probability density of each of the data points in the second data set based on a multi-dimensional Gaussian distribution; determining an abnormal score of each of the data points based on the joint probability density of each of the data points; detecting outliers of each of the unique feature variables and the additional feature variables of each of the data points in the second data set based on an outlier detection algorithm, and determining the abnormal carriage based on the abnormal score of each of the data points; calculating a joint probability density of each of the data points in the second data set based on a multi-dimensional Gaussian distribution, including: establishing a multi-dimensional Gaussian distribution model, calculating a mean value and a variance of the unique feature variables and the additional feature variables in each of the data points in the second data set to fit the multi-dimensional Gaussian distribution model, and calculating a joint probability density of each of the data points using the fitted multi-dimensional Gaussian distribution model; determining an abnormal score of each of the data points based on the joint probability density of each of the data points, including: According to the joint probability density of each data point, each data point is scored, data points with joint probability density lower than a preset probability density threshold are recorded as abnormal data points, the abnormal data points are given higher abnormal scores, and the abnormal scores and the abnormal classification of each data point are output; the joint probability density of the data points is recorded as , and the abnormal score of the data point is: ; The analysis unit is specifically configured to: judging a correlation between each of the original feature variables in each of the correlation groups in the first data set, and setting a correlation group in which the original feature variable with abnormal correlation as the in-group data abnormal correlation group.

7. A vehicle failure analysis apparatus characterized by comprising: including: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the vehicle fault analysis method according to any one of claims 1 to 5.

8. A train characterized by including the vehicle fault analysis device according to claim 7.

Citation Information

Patent Citations

  • Industrial data quality analysis method

    CN113570254A

  • Anomaly detection system and method

    CN114577252A

  • Vehicle fault early warning method and system based on high-frequency time sequence data

    CN114676782A

  • Method and system for training a big data machine to defend

    US20170169360A1