Real-time processing method of MVB bus data based on cloud computing

By dividing the operation stage of the traction motor in the MVB bus data processing and calculating the importance of data in each dimension, the problem that the difference in the sensitivity of data in different dimensions is not considered to be solved, and more accurate abnormality detection results are achieved.

CN119739092BActive Publication Date: 2025-05-20XIAN SHENXI ELECTRIC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510245313.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-20
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

In MVB bus data processing, the prior art fails to effectively consider the differences in sensitivity of data from different dimensions to faults, resulting in inaccurate abnormal detection results.

Method used

By dividing the dimensional data of the traction motor into acceleration, constant speed and deceleration stages, the importance of the dimensional data in each stage is calculated, and the weight is allocated according to the importance, and the path length is adjusted to calculate the abnormal score.

Benefits of technology

This method can more accurately reflect the behavior characteristics of the traction motor at different operating stages, improve the accuracy of abnormal detection, and reduce false alarms and missed alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119739092B_ABST
    Figure CN119739092B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing, and in particular to a real-time processing method for MVB bus data based on cloud computing, the method comprising: using an MVB bus interface device to collect dimensional data of a traction motor in real time; dividing the dimensional data of the traction motor into acceleration, constant speed and deceleration stages, and calculating the importance of each dimensional data in each stage; assigning a weight to each dimensional data based on the importance of each dimensional data in the current stage, and obtaining a corresponding dimensional weight; weighting and adjusting the path length of the dimensional data according to the dimensional weight to obtain a weighted path length, and calculating anomaly scores of the dimensional data based on the weighted path length; and performing anomaly detection according to the anomaly score. The present invention effectively solves the problem that simple average path length calculation easily ignores the difference between dimensions, and can more accurately identify anomalies in the data, thereby improving the accuracy of anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing. More specifically, the present invention relates to a real-time processing method for MVB bus data based on cloud computing. Background Art

[0002] The Multifunction Vehicle Bus (MVB) is a communication protocol applied in the transportation field. It can connect various devices in a train, such as traction control units, braking systems, door controls, air conditioning systems, etc., so as to achieve real-time communication and data exchange between devices. As an important part of the train communication network, the data collected by the MVB bus contains various key information during the train operation, and this data is crucial for the safe operation and fault diagnosis of the train.

[0003] The Isolation Forest algorithm is an unsupervised anomaly detection algorithm based on a tree model. In the scenario of MVB bus data collection, the Isolation Forest algorithm can analyze the collected data and quickly identify abnormal data that deviates significantly from the normal data pattern. These abnormal data are of great significance for reminding relevant personnel to conduct inspections and repairs in a timely manner, and contribute to improving the reliability and safety of the equipment.

[0004] However, when performing fault early warning, the sensitivities of data in different dimensions to faults are different. For example, the data of certain sensors may be more likely to reflect specific fault types, while the data of other sensors may be less sensitive. Therefore, if the anomaly score is simply calculated based on the average path length of each data point in all isolation trees, it is easy to ignore the differences between these dimensions, resulting in inaccurate calculation results and further affecting the accuracy of the early warning results. Summary of the Invention

[0005] To solve the above technical problem that the differences in the sensitivities of data in different dimensions to faults are not considered, resulting in inaccurate processing results of the collected data, the present invention provides the following technical solutions.

[0006] A real-time processing method for MVB bus data based on cloud computing, comprising:

[0007] Using the MVB bus interface device to collect the dimensional data of the traction motor in real time; calculating the anomaly score of each dimensional data using an improved Isolation Forest algorithm; performing anomaly detection based on the anomaly score;

[0008] The improvement process is as follows:

[0009] Divide the dimensional data of the traction motor into acceleration, constant speed, and deceleration phases, calculate the importance of each dimensional data under each phase; assign weights to each dimensional data based on the importance of the dimensional data in the current phase to obtain the corresponding dimensional weights; perform weighted adjustment on the path length of the dimensional data according to the dimensional weights to obtain the weighted path length, and calculate the anomaly score of the dimensional data based on the weighted path length.

[0010] The present invention first divides the dimensional data of the traction motor into phases, avoiding unified processing of all data, which helps to identify the behavioral characteristics of the motor in different operating phases, and further accurately understand the importance of each dimensional data in different phases. Then, through the behavioral characteristics of the motor in each phase, calculate the importance of each dimensional data under each phase, which can clarify which dimensions have a significant impact on the operating state of the motor in a specific phase, so that important dimensions will be given higher weights when calculating the path length, and further accurately reflect the true anomaly degree of the data points.

[0011] Through the above operations, the algorithm can more accurately reflect the behavioral characteristics of the traction motor in different operating phases, more sensitively capture abnormal data points, and reduce false alarms and missed alarms caused by ignoring differences between dimensions.

[0012] Preferably, the dimensional data includes at least one of electrical parameters, rotational speed, and temperature.

[0013] Preferably, the process of dividing the dimensional data of the traction motor into acceleration, constant speed, and deceleration phases includes:

[0014] Synchronously collect the speed data of the train based on the collected dimensional data of the traction motor, sort them in chronological order, and construct a speed sequence;

[0015] Starting from the second item of the speed sequence, calculate the difference value with its previous item in turn. If the difference value is positive, the train at the moment corresponding to the speed data is in the acceleration phase. If the difference value is zero, the train at the moment corresponding to the speed data is in the constant speed phase. If the difference value is negative, the train at the moment corresponding to the speed data is in the deceleration phase.

[0016] Through the difference value of the speed data, it can accurately judge whether the train is in the acceleration, constant speed, or deceleration phase at a certain moment, which helps to more precisely understand the operating state of the train and provides a basis for subsequent performance analysis of the traction motor under different working conditions.

[0017] Preferably, the calculation process of the importance includes:

[0018] Calculate the degree of change of all sub-data under any one dimensional data, and take the average value of the degrees of change of all sub-data as the importance of the dimensional data.

[0019] By calculating the degree of change of sub - data and taking the mean value to represent the importance of dimensional data, it can intuitively show the contribution of each dimension to the overall data change. In the dimensional data of a traction motor, for example, data of different dimensions such as current, voltage, and temperature reflect the operating state of the motor to different degrees. For those dimensions that change more violently during operation and are crucial to the motor performance, such as the change of current during the acceleration stage, their importance will be highlighted, and then it can quickly identify which dimensions are the key factors affecting the operation of the traction motor and which are relatively less important factors.

[0020] Preferably, the process of obtaining the dimension weight includes:

[0021] Select any one dimension, and take the ratio of the importance of this dimension to the sum of the importance of all dimensions as the dimension weight of the data of this dimension.

[0022] In the traditional isolation forest algorithm, all dimensions are given the same weight, which may cause dimensions with low importance to have an unnecessary impact on the anomaly detection result, thus increasing the possibility of misjudgment. By assigning weights to each dimension, the algorithm can reduce the influence of unimportant dimensions, thereby reducing misjudgment and improving the reliability of the detection result.

[0023] Preferably, the weighted path length satisfies the relational expression:

[0024] ; where is the weighted path length of the th dimension data, is the path length of the th dimension data in the th isolation tree, is the dimension weight of the dimension data corresponding to the th isolation tree, is the total number of isolation trees.

[0025] Preferably, the anomaly detection based on the anomaly score includes:

[0026] When the anomaly score of the dimensional data of the traction motor exceeds the preset anomaly threshold, trigger the warning mechanism.

[0027] Preferably, the process of obtaining the degree of change includes:

[0028] Under any one dimension, select a set number of sub - data as the target data;

[0029] Calculate the numerical difference and time difference between each sub - data and the target data;

[0030] Take the product of the normalized numerical difference and the time difference as the degree of change of the corresponding sub - data.

[0031] For each sub - data, calculate its numerical difference from the target data, which reflects the degree of numerical change of the sub - data; at the same time, calculate the time difference between the sub - data and the target data, taking into account the time factor of data change, which helps to understand the timeliness of data change.

[0032] Preferably, the calculation process of the importance includes:

[0033] Calculate the degree of change of all sub - data under any one - dimensional data, sort them from small to large, and select the degree of change corresponding to its quartiles as the importance of this dimensional data.

[0034] Preferably, the process of dividing the dimensional data of the traction motor into acceleration, constant - speed, and deceleration stages includes:

[0035] Synchronously collect the speed data of the train based on the collected dimensional data of the traction motor, sort them in chronological order, and construct a speed sequence.

[0036] For each moment of the speed sequence, calculate its speed change rate. According to the positive, negative, and zero values of the speed change rate, preliminarily divide the running state of the train into an acceleration stage, a constant - speed stage, and a deceleration stage; if the continuous moments of a certain stage are less than a preset threshold, merge this stage into the previous stage.

[0037] Considering the possible short - term abnormal situations or data acquisition inaccuracies in actual operation, by setting the continuous - moment threshold, those stages that may be abnormal or inaccurate can be merged into the previous more reasonable stage, thereby reducing the impact of data acquisition errors or temporary anomalies on the overall division result, which reflects a certain degree of fault tolerance.

[0038] The beneficial effects of the present invention are:

[0039] The present invention adopts an improved isolation forest algorithm. By dividing the operating stages of the traction motor (acceleration, constant - speed, deceleration), calculating the importance and weight of each dimensional data in each stage, and then adjusting the path length and calculating the anomaly score. This method effectively solves the problem that the simple average path - length calculation is prone to ignoring the differences between dimensions, and can thus more accurately identify abnormal situations in the data and improve the accuracy of anomaly detection. Description of the Drawings

[0040] Figure 1 It is a flowchart of the method of steps S1 - S3 in the real - time processing method of MVB bus data based on cloud computing in the embodiment of the present invention. Detailed Embodiments

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments.

[0042] The application scenario of the present invention is as follows: real-time collection of the operation data of the traction motor of the train through the MVB bus interface, and the use of an improved anomaly detection algorithm to analyze and process these data to identify anomalies or potential faults, and immediately send a warning signal when an anomaly is detected.

[0043] The following will describe in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings.

[0044] Referring to Figure 1 , the real-time processing method of MVB bus data based on cloud computing includes steps S1 - S3, specifically as follows:

[0045] S1: Use the MVB bus interface device to collect the dimensional data of the traction motor in real time.

[0046] The traction motor is a key component for the power output of the train. If the traction motor fails, such as overload or overheat, it may cause the power interruption of the train and even lead to serious accidents such as fires. Therefore, the real-time monitoring of the traction motor is crucial.

[0047] In one embodiment, first select an interface device that supports MVB bus communication, such as an MVB board or adapter, connect the MVB bus interface device to the control system of the traction motor and ensure that the communication link is normal.

[0048] Install electrical parameter sensors, speed sensors, and temperature sensors at the corresponding positions of the traction motor respectively, start the MVB bus interface device, and collect and transmit the multi-dimensional data of the traction motor in real time at a fixed sampling period (such as 10 milliseconds), including electrical parameters (voltage, current, power, etc.), speed, and temperature.

[0049] Among them, the collected multi-dimensional data is processed, such as filtering, amplification, etc., to improve the accuracy and stability of the data.

[0050] S2: Use the improved isolation forest algorithm to calculate the anomaly scores of each dimensional data.

[0051] In one embodiment, through the MVB bus interface, while collecting the dimensional data of the traction motor according to the collection frequency set in S1 above, collect the speed data of the train.

[0052] It should be noted that to ensure data synchronization, time stamps are marked on the train speed data and dimensional data during the collection process to ensure that each data point can accurately correspond to its collection time.

[0053] Sort the obtained train speed data according to the marked timestamps to form a speed sequence. Starting from the second item of the speed sequence, calculate the difference values with its previous item in turn.

[0054] Among them, if the difference item is positive, it means that the train speed at the current moment is increasing, that is, the train is in the acceleration stage. If the difference item is 0, it means that the train speed at the current moment remains unchanged, that is, the train is in the constant speed stage. If the difference item is negative, it means that the train speed at the current moment is decreasing, that is, the train is in the deceleration stage. Furthermore, divide the train operation state into acceleration, constant speed, and deceleration stages.

[0055] Furthermore, according to the traction motor dimension data collected at the same time as the train speed sequence, divide these dimension data into acceleration, constant speed, and deceleration stages as well.

[0056] It should be noted that classifying the traction motor data according to different operation stages can avoid mixing the data of all stages for calculation, which can reduce the calculation errors caused by the mixing of different operation states.

[0057] In one embodiment, calculate the change degree of the sub-data of each data in a specific dimension.

[0058] Exemplarily, taking the temperature dimension data as an example, for each sub-data under the temperature dimension data, calculate the absolute value of the difference between each sub-data and all other sub-data respectively, and sort these absolute values of the differences in ascending order to obtain a difference sequence. Select the sub-data corresponding to the first K smallest differences in this difference sequence as the target data.

[0059] Furthermore, calculate the sum of the numerical differences between each sub-data and all the target data, and perform normalization to obtain the numerical differences of each sub-data; calculate the square root of the mean of the squares of the differences between the collection times of each sub-data and all the target data, and perform normalization to obtain the time differences of each sub-data.

[0060] Take the product of the above numerical differences and time differences as the change degree of the corresponding sub-data. The change degree satisfies the relational expression as:

[0061]

[0062] In the formula, is the change degree of the th sub-data of the temperature dimension, is the value of the th sub-data of the temperature dimension, is the value of the th sub-data among the target data of the th sub-data, is the total number of target data, is the acquisition time corresponding to the th sub - data in the temperature dimension, is the th sub - data, and is the acquisition time corresponding to the th sub - data in the target data of the

[0063] Among them, if the numerical difference between the th sub - data and its target data is larger, and the sampling times of these target data and the sampling time of the th sub - data also have a large difference, then the th sub - data has a higher degree of change, and the th sub - data has a more obvious change characteristic.

[0064] Furthermore, according to the above calculation method of the change degree of the th sub - data, the change degrees of all sub - data in the temperature dimension are obtained in the same way.

[0065] Furthermore, the mean value of the change degrees of all sub - data in the temperature dimension is used as the importance of the temperature dimension.

[0066] Furthermore, according to the above calculation method of the importance of the temperature dimension, the importance corresponding to all dimensions can be obtained in the same way.

[0067] In different stages of train operation, the sensitivities of data in different dimensions to faults are different, and the degree of change of data points can be used as an index to measure the importance degree of dimensions. The greater the degree of change, the higher the sensitivity of the data in this dimension to faults, that is, the higher the importance degree of this dimension in fault detection and diagnosis.

[0068] In one embodiment, weights are assigned to each dimension data based on the importance of each dimension data in the current stage to obtain the corresponding dimension weights. Exemplarily, still taking the temperature dimension as an example, the dimension weight of the temperature dimension satisfies the relational expression:

[0069]

[0070] In the formula, is the dimension weight of the temperature dimension, is the importance of the temperature dimension, is the sum of the importance of all dimensions.

[0071] Furthermore, for the dimension data of each stage, the genetic algorithm constructs multiple isolation trees, and each tree is continuously divided by data of the same dimension. Among them, for each data point, the path lengths of it in all isolation trees are weighted and summed based on the corresponding dimension weights, that is, it satisfies the relational expression:

[0072]

[0073] Wherein, is the weighted path length of the data in the th dimension, is the path length of the data in the th dimension in the th isolated tree, is the dimension weight of the dimension data corresponding to the th isolated tree, is the total number of isolated trees.

[0074] Furthermore, the weighted path lengths of all dimension data are obtained. According to the path lengths of each dimension data, the genetic algorithm assigns an anomaly score to each dimension data. The higher the score, the more likely the dimension data is an outlier.

[0075] It should be noted that obtaining the anomaly score is part of the isolation forest algorithm and will not be elaborated here in detail.

[0076] In another embodiment, a fault tolerance space is considered when dividing the train running state.

[0077] Specifically, the obtained train speed data is sorted according to its marked timestamp to form a speed sequence. Calculate the speed change rate of the train at the th moment (usually the ratio of the speed difference between the th moment and the previous moment to the time difference). If the speed change rate at the th moment is positive, it means the train is accelerating. If the speed change rate at the th moment is zero, it means the train is in a constant speed stage. If the speed change rate at the th moment is negative, it means the train is decelerating.

[0078] According to the speed change rate at the th moment, the running state of the train can be divided into an acceleration stage, a constant speed stage, and a deceleration stage. However, to avoid the influence of minor adjustments of the train during operation on the judgment of the train running stage, by analyzing the duration of each stage, that is, the number of consecutive moments is less than a preset threshold (obtained through experiments and debugging), then it is considered that this stage is caused by the minor adjustment of the train, and thus this stage is merged into the previous stage.

[0079] In another embodiment, after obtaining the change degrees of all sub-data in the temperature dimension in the same way as the calculation method of the change degree of the above-mentioned rd sub-data, sort all the change degrees from small to large, and select the change degree corresponding to the quartile as the importance of this temperature dimension.

[0080] S3: Perform anomaly detection based on the anomaly score.

[0081] In one embodiment, the preset anomaly threshold is 0.85. When the anomaly score of the dimensional data of the traction motor exceeds the preset anomaly threshold, the system triggers an early warning mechanism to notify relevant staff to check for possible problems and take necessary measures to prevent potential failures or accidents.

[0082] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent shall be subject to the appended claims.

Claims

1. A real-time processing method for MVB bus data based on cloud computing, characterized in that: include: Use MVB bus interface equipment to collect traction motor dimensional data in real time; The improved isolation forest algorithm is used to calculate the anomaly score of each dimension of data; Anomaly detection based on anomaly scores; The improved process is: The dimensional data of the traction motor is divided into acceleration, constant speed and deceleration stages. The division process includes: The speed data of the train is synchronously collected based on the collected dimensional data of the traction motor, and is sorted in chronological order to construct a speed sequence; Starting from the second item of the speed sequence, the difference value is calculated with the previous item in turn. If the difference value is positive, the train at the time corresponding to the speed data is in the acceleration stage. If the difference value is zero, the train at the time corresponding to the speed data is in the constant speed stage. If the difference value is negative, the train at the time corresponding to the speed data is in the deceleration stage. Calculate the importance of each dimension of data at each stage, including: Calculate the degree of change of all sub-data under any dimension data, and take the mean of the degree of change of all sub-data as the importance of the dimension data; Assign weights to each dimension data based on its importance at the current stage to obtain the corresponding dimension weights; perform weighted adjustment on the path length of the dimension data according to the dimension weights to obtain the weighted path length, and calculate the anomaly score of the dimension data based on the weighted path length; The process of obtaining the degree of change includes: In any dimension, select a set number of sub-data as target data; Calculate the sum of the numerical differences between each sub-data and all target data, and normalize them to obtain the numerical differences of each sub-data; calculate the square root of the mean of the squares of the differences between the acquisition times of each sub-data and all target data, and normalize them to obtain the time differences of each sub-data; The product of the normalized numerical difference and the time difference is taken as the degree of change of the corresponding sub-data.

2. The method for real-time processing of MVB bus data based on cloud computing according to claim 1 is characterized in that: The dimensional data includes at least one of electrical parameters, rotation speed, and temperature.

3. The method for real-time processing of MVB bus data based on cloud computing according to claim 2 is characterized in that: The process of obtaining the dimension weight includes: Select any dimension and take the ratio of the importance of this dimension to the sum of the importance of all dimensions as the dimension weight of the data of this dimension.

4. The method for real-time processing of MVB bus data based on cloud computing according to claim 3 is characterized in that: The weighted path length satisfies the relationship: ; In the formula, For the The weighted path length of the dimension data, For the The dimension data is in The path length in an isolated tree, For the The dimension weight of the dimension data corresponding to an isolated tree, is the total number of isolated trees.

5. The method for real-time processing of MVB bus data based on cloud computing according to claim 4 is characterized in that: The anomaly detection according to the anomaly score comprises: When the abnormal score of the dimensional data of the traction motor exceeds the preset abnormal threshold, the early warning mechanism is triggered.

6. The method for real-time processing of MVB bus data based on cloud computing according to claim 1 is characterized in that: The calculation process of the importance includes: Calculate the degree of change of all sub-data under any dimension data, sort them from small to large, and select the degree of change corresponding to its quartile as the importance of the dimension data.

7. The method for real-time processing of MVB bus data based on cloud computing according to claim 1, characterized in that: The process of dividing the dimensional data of the traction motor into acceleration, constant speed and deceleration stages includes: The speed data of the train is synchronously collected based on the collected dimensional data of the traction motor, and is sorted in chronological order to construct a speed sequence; For each moment in the speed sequence, the speed change rate is calculated. According to the positive, negative and zero values ​​of the speed change rate, the train's operating state is preliminarily divided into acceleration stage, constant speed stage and deceleration stage. If the continuous moments of a certain stage are less than the preset threshold, the stage is merged into the previous stage.

Citation Information

Patent Citations

  • Intelligent train traction fault big data abnormality detection and recognition method

    CN109506963A

  • Automobile starter monitoring method based on data analysis

    CN117238058A