Data Integration Method for Orbit Train Fusion Host
By collecting, preprocessing and focusing on track train data in multi-dimensional timing data, combined with parallel coding technology, the problems of low compression efficiency and high computing resource consumption in track train data processing are solved, and efficient data processing and compression are achieved.
Patent Information
- Application Number
- CN202510157511.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-13
AI Technical Summary
In the processing of rail train data, the prior art directly performs unified compression storage, resulting in low compression efficiency, high computing resources consumption, and affecting data processing efficiency.
Multi-dimensional timing data during the track train operation is collected, target dimensions and non-standard factors are determined, attention sequences are constructed, and data in the optimal period is encoded using parallel encoding technology.
It improves data encoding efficiency, realizes segmented compression, improves data processing efficiency, and reduces computing resource consumption.
Smart Images

Figure CN119621689B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a data integration method for an integrated host for rail trains. Background Art
[0002] The data of rail trains is diverse, including detailed information in multiple aspects such as train operation status, passenger information, and environmental monitoring. The amount of data is extremely large, posing a huge challenge to data processing and storage. In the process of data integration for an integrated host for rail trains, data compression and storage are essential steps. If all the data for each train's each operation is directly and uniformly compressed and stored, the compression efficiency is low. Moreover, since the data is uniformly compressed, when performing subsequent data analysis, all the data needs to be uniformly decompressed, consuming a large amount of computer resources and taking a long time, which affects the overall data processing efficiency. Summary of the Invention
[0003] In order to solve the technical problem that the existing method of directly and uniformly compressing data for rail trains has low efficiency and is not conducive to subsequent decoding and analysis, the purpose of the present invention is to provide a data integration method for an integrated host for rail trains. The specific technical solution adopted is as follows:
[0004] Collect multi-dimensional time-series data during the operation of the rail train and perform preprocessing;
[0005] Define any shift of the rail train as the target shift, determine the target dimension between the target shift and non-target shifts based on the preprocessed multi-dimensional time-series data, obtain the corresponding specificity through the target dimension, construct a set of target dimensions, and obtain the weight of each dimension of the target shift; analyze the dissimilarity between the target shift and non-target shifts to determine the non-standard factor of the target shift.
[0006] The operation process of the rail train includes multiple initial time periods. Based on the non-standard factor and multi-dimensional time-series data, determine the first attention degree and the second attention degree of each initial time period respectively, and construct an attention sequence of the target shift from the starting point to the ending point by combining the first attention degree and the second attention degree, so as to obtain multiple optimal time periods and the corresponding attention degrees.
[0007] Parallelly encode the time-series data of all dimensions corresponding to the optimal time periods, and use the attention degree corresponding to the optimal time period as a label to complete data integration.
[0008] Preferably, collecting multi-dimensional time-series data during the operation of the rail train and performing preprocessing means that the multi-dimensional time-series data includes any data of speed, acceleration, train motor temperature, brake cylinder pressure, vibration data time series, and real-time position coordinates from the starting point to the ending point, and the preprocessing method is standardization processing.
[0009] Preferably, define any train trip in the rail train as the target trip, determine the target dimension between the target trip and non-target trips based on the preprocessed multi-dimensional time series data, obtain the corresponding specificity through the target dimension, and construct a set of target dimensions to obtain the weight of each dimension of the target trip, including:
[0010] Define two adjacent stations as Station One and Station Two respectively. Based on the target trip and any non-target trip, obtain the time series of each dimension from Station One to Station Two, calculate the dynamic time warping distance of the time series of the same dimension, and sort them from largest to smallest. The dimension corresponding to the largest dynamic time warping distance is the target dimension, and obtain the target dimension corresponding to the target trip and each non-target trip;
[0011] Calculate the specificity corresponding to each target dimension, construct a set of target dimensions, count the number of occurrences of each same target dimension, calculate the weight corresponding to each same target dimension, and assign the weight of the dimension in the non-target dimension set to obtain the weight of each dimension of the target trip.
[0012] Preferably, the calculation formula for the specificity corresponding to each target dimension is:
[0013]
[0014] Among them, represents the specificity of the target dimension; represents the dynamic time warping distance corresponding to the target dimension; represents the mean value of the dynamic time warping distances of all non-target dimensions.
[0015] Preferably, the calculation formula for the weight corresponding to each same target dimension is:
[0016]
[0017] Among them, represents the weight of the th same target dimension in the set of target dimensions; represents the total number of target dimensions in the set of target dimensions; represents the th total number of the same target dimension in the set of target dimensions; represents the maximum value of the specificities of all target dimensions in the th same target dimension in the set of target dimensions; represents the linear normalization function.
[0018] Preferably, analyze the dissimilarity between the target trip and non-target trips to determine the non-standard factors of the target trip, including:
[0019] Calculate the dissimilarity between the target shift and each non-target shift from Site 1 to Site 2, and determine the clustering distance.
[0020] Divide all shifts of the rail train into multiple clustering clusters based on the clustering distance, define the clustering cluster with the largest number of shifts as the standard clustering cluster, obtain the clustering distance between the target shift and the clustering center of the standard clustering cluster, and obtain the non-standard factor of the target shift.
[0021] Preferably, the calculation formula for dissimilarity is:
[0022]
[0023] Where, represents dissimilarity; represents the total number of different dimensions in all dimensions; represents the weight of the th dimension of the target shift from Site 1 to Site 2; represents the weight of the th dimension of the non-target shift from Site 1 to Site 2; represents the dynamic time warping distance of the th dimension time series of the target shift and the non-target shift from Site 1 to Site 2.
[0024] Preferably, the operation process of the rail train includes multiple initial time periods. Determine the first attention degree and the second attention degree of each initial time period based on the non-standard factor and multi-dimensional time series data, including:
[0025] The initial time period includes the starting acceleration stage, the constant speed operation stage, and the deceleration braking stage in the operation process of the rail train. Determine the first attention degree of the target shift in the starting acceleration stage and the deceleration braking stage through the non-standard factor, which is the non-standard factor;
[0026] Divide the constant speed operation stage into several sequence small segments according to the duration, and obtain the wave peaks and wave valleys of each sequence small segment in the constant speed operation stage. Calculate the speed instability corresponding to the constant speed operation stage, and the corresponding calculation formula is:
[0027]
[0028] Where, represents speed instability; represents the variance of all speeds in the constant speed operation stage; represents the maximum value of the speed sudden change of all wave peaks in the constant speed operation stage; represents the duration of the constant speed operation stage; represents the total number of wave peaks in the constant speed operation stage;
[0029] Calculate the initial adjustment factor of attention in the uniform running stage. The corresponding calculation formula is:
[0030]
[0031] Wherein, represents the initial adjustment factor of attention; represents the speed instability; represents the average value of all cruising speeds in the uniform running stage; represents the predetermined cruising speed in the uniform running stage;
[0032] Determine the position coordinates of the rail train at the start time and end time of the uniform running stage, construct a speed standard sequence, and determine the second attention in the uniform running stage.
[0033] Preferably, the calculation formula of the second attention is:
[0034]
[0035] Wherein, represents the second attention; represents the non-standard factor; represents the sequence composed of all sequence sub-segments in the uniform running stage; represents the speed standard sequence; represents and the dynamic time warping distance of; represents the linear normalization function.
[0036] Preferably, combine the first attention and the second attention to construct the attention sequence of the target shift from the starting point to the ending point, and obtain multiple optimal time periods and the corresponding attention, including:
[0037] Determine the attention of each initial time period based on the first attention and the second attention, and construct the attention sequence of the target shift from the starting point to the ending point with the attention of all initial time periods;
[0038] Divide the attention sequence into multiple sequence segments by clustering, merge all the initial time periods corresponding to each sequence segment to form each optimal time period, calculate the average value of the attention of all the initial time periods in the corresponding sequence segment, and determine the attention of each optimal time period.
[0039] The present invention has the following beneficial effects:
[0040] The data integration method for the integrated host of rail trains refers to a method of integrating and uniformly processing data from different systems and devices on rail trains. Through data integration, data sharing, analysis, and application can be achieved to improve the safety, efficiency, and service quality of train operation. By collecting data, it determines the weights between different shifts of rail trains, analyzes the data to determine the optimal time periods of shifts, and based on the optimal time periods, uses block Huffman coding to parallelly encode the data of each time period, improving the coding efficiency and achieving segmented compression, thereby enhancing the compression efficiency. This enables, when analyzing subsequent rail train operation problems, only decoding the data of the optimal time periods of the problems, improving the processing efficiency of the integrated data. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 FIG. is a schematic diagram of the steps of a data integration method for the integrated host of rail trains provided by an embodiment of the present invention;
[0043] Figure 2 FIG. is a schematic diagram of the acquisition of multi-dimensional time-series data of a data integration method for the integrated host of rail trains provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in combination with the drawings and preferred embodiments, details the specific implementation manner, structure, features, and effects of a data integration method for the integrated host of rail trains proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0046] The following specifically describes the specific solution of a data integration method for the integrated host of rail trains provided by the present invention in combination with the drawings.
[0047] Please refer to Figure 1, which shows a schematic diagram of the steps of a data integration method for an on - track train fusion host provided by an embodiment of the present invention. The method includes:
[0048] Step S1: Collect multi - dimensional time - series data during the operation of the on - track train and perform pre - processing;
[0049] Step S2: Define any shift in the on - track train as the target shift. Based on the pre - processed multi - dimensional time - series data, determine the target dimension between the target shift and non - target shifts, obtain the corresponding specificity through the target dimension, construct a target dimension set, and obtain the weight of each dimension of the target shift; analyze the dissimilarity between the target shift and non - target shifts to determine the non - standard factors of the target shift;
[0050] Step S3: The operation process of the on - track train includes multiple initial time periods. Based on the non - standard factors and multi - dimensional time - series data, determine the first attention degree and the second attention degree of each initial time period respectively, and construct an attention sequence of the target shift from the starting point to the ending point by combining the first attention degree and the second attention degree, so as to obtain multiple optimal time periods and the corresponding attention degrees;
[0051] Step S4: Parallelly encode all the dimension time - series data corresponding to the optimal time periods, and use the attention degree corresponding to the optimal time period as a label to complete data integration.
[0052] Explanation: The data integration method for an on - track train fusion host refers to integrating data from different sources to improve the safety, efficiency, and maintainability of the rail transit system. In this regard, it can perform fault detection and predictive maintenance to reduce downtime; integrate passenger flow data and real - time operation data to optimize train scheduling and improve transportation efficiency. Integrating this method with the host to achieve intelligent and automated analysis helps to realize more efficient management of the rail transit system.
[0053] Please refer to Figure 2 , which shows a schematic diagram of the collection of multi - dimensional time - series data of a data integration method for an on - track train fusion host.
[0054] Furthermore, Step S1 is that the multi - dimensional time - series data includes any data such as speed, acceleration, train motor temperature, brake cylinder pressure, vibration data time - series sequence, and real - time position coordinates from the starting point to the ending point, and the pre - processing method is standardization processing.
[0055] As an alternative embodiment, the most common operation modes in railway track transportation are "fixed-line operation" or "circular operation", that is, fixed starting and ending points, periodic schedules, repeated routes, and standardized operations. Since trains on the same line need to adapt to the same track conditions, signal systems, and operation requirements, the train models on the same line are usually basically similar. Therefore, in this embodiment, multi-dimensional time-series data of a rail train from the starting point to the ending point is collected based on any fixed track line.
[0056] Specifically, speed sensors, acceleration sensors, temperature sensors, pressure sensors, vibration sensors, etc. are installed at corresponding positions on the rail train to monitor the running state of the rail train in real time, and the time-series sequences of the speed, acceleration, train motor temperature, brake cylinder pressure, and vibration data of each schedule of the rail train from the starting point to the ending point are determined; then the real-time position coordinates of the rail train are tracked through GPS (Global Positioning System); then the above-mentioned multi-dimensional time-series data collected is standardized, such as by min-max standardization, z-score (Standard Score) standardization, etc., to unify the dimension and eliminate the influence brought by different dimensions, ensuring that the data is compared and analyzed on the same magnitude and improving the data processing efficiency.
[0057] Furthermore, in step S2, any schedule of the rail train is defined as the target schedule, the target dimension between the target schedule and non-target schedules is determined based on the preprocessed multi-dimensional time-series data, the specificity corresponding to the target dimension is obtained through the target dimension, and a target dimension set is constructed to obtain the weight of each dimension of the target schedule, including:
[0058] Step S211: Define two adjacent stations as Station One and Station Two respectively. Based on the target schedule and any non-target schedule, the time-series sequences of each dimension from Station One to Station Two are obtained, the dynamic time warping distance of the time-series sequences of the same dimension is calculated, and they are sorted from large to small. The dimension corresponding to the largest dynamic time warping distance is the target dimension, and the target dimension corresponding to the target schedule and each non-target schedule is obtained.
[0059] It should be noted that the dynamic time warping distance, that is, DTW (Dynamic Time Warping), is a technology used to measure and compare the similarity between two time-series sequences; by adjusting the time points in the time-series sequence, the time series is elastically stretched or compressed to find the best match, so that the two sequences are aligned in time to more accurately compare their shapes and patterns.
[0060] Specifically, calculate the dynamic time warping distance of time series sequences of the same dimension. At this time, the smaller the DTW distance between the target shift corresponding to the time series sequence of the same dimension and any non-target shift, the more similar the time series sequences between the two shifts are. Sort all the calculated dynamic time warping distances from largest to smallest. Among them, the dimension corresponding to the largest dynamic time warping distance is denoted as the target dimension. Similarly, calculate the target dimension corresponding to the target shift and each non-target shift.
[0061] It can be understood that when the environment of the rail train changes, the impact on its operation data is comprehensive. Due to environmental changes, the changes in the operation data of different shifts in each dimension are usually relatively small, which belongs to the normal range of changes. The impact caused by faults is more concentrated and specific, and the data changes of different shifts in specific dimensions are relatively significant, which belongs to important changes. Therefore, different weights should be assigned to the similarity of data in different dimensions to enhance the reliability of data analysis.
[0062] Step S212: Calculate the specificity corresponding to each target dimension, construct a set of target dimensions, count the number of occurrences of each same target dimension, calculate the weight corresponding to each same target dimension, and assign the weight of the dimension in the non-target dimension set to obtain the weight of each dimension of the target shift.
[0063] It should be noted that specificity refers to the specific changes that may occur in the rail train, that is, it indicates that the dimension currently analyzed belongs to the category of important changes, indicating that some changes in the rail train may have a significant impact on the overall function or stability.
[0064] Furthermore, in step S212, the calculation formula for the specificity corresponding to each target dimension is:
[0065]
[0066] Among them, represents the specificity of the target dimension; represents the dynamic time warping distance corresponding to the target dimension; represents the mean value of the dynamic time warping distances of all non-target dimensions.
[0067] It can be explained that the greater the specificity of the target dimension, that is, the greater the difference between the dynamic time warping distance corresponding to the target dimension and the mean value of the dynamic time warping distances of all non-target dimensions, the more significant the difference in the performance of the target dimension in the time series from other dimensions, indicating that the target dimension is more likely to be the specific dimension of the fault. At this time, a larger weight is assigned to this dimension to improve the accuracy of fault detection for subsequent detection.
[0068] For better illustration, in this embodiment, a target dimension set is constructed, that is, for the multiple dimensions in the target dimension, a series of key measurement criteria and indicators are defined and set to facilitate effective evaluation and monitoring in a project or task.
[0069] Specifically, as can be seen above, the target dimension set includes the target shift and the target dimensions of all other non-target shifts, which can be any multi-dimensional time series data such as speed, acceleration, temperature, etc.; for simplicity, based on speed and acceleration, the constructed target dimension data set is , where represents the speed dimension, represents the acceleration dimension, represents the target shift and the target dimension of the target shift and other non-target shift 1 is the speed dimension ; represents the target shift , , , similarly; that is, the target dimension set is {the target dimension of the target shift and other non-target shift 1, the target dimension of the target shift and other non-target shift 2,..., the target dimension of the target shift and other non-target shifts }.
[0070] Furthermore, in step S212, the calculation formula for the weight corresponding to each same target dimension is:
[0071]
[0072] where represents the weight of the th same target dimension in the target dimension set; represents the total number of target dimensions in the target dimension set; represents the th total number of the same target dimension in the target dimension set; represents the maximum value of the specificity of all target dimensions in the th same target dimension in the target dimension set; represents the linear normalization function.
[0073] It can be understood that in the target dimension set, the number of occurrences of each same target dimension is counted. The more times it appears, it indicates that when performing similarity analysis between the target shift and other non-target shifts, this target dimension is basically a fault dimension, so a larger weight should be assigned.
[0074] As an alternative implementation, assign weights to the dimensions in the non-target dimension set, that is, for the dimensions not in the target dimension set, assign the weight of each dimension as 1, and combine the above calculation of weights to obtain the weight of each dimension of the target shift.
[0075] Further, in step S2, analyze the dissimilarity between the target shift and non-target shifts to determine the non-standard factors of the target shift, including:
[0076] Step S221: Calculate the dissimilarity between the target shift and each non-target shift from station one to station two to determine the clustering distance.
[0077] Further, in step S221, the calculation formula for dissimilarity is:
[0078]
[0079] where, represents dissimilarity; represents the total number of different dimensions among all dimensions; represents the weight of the th dimension of the target shift from station one to station two; represents the weight of the th dimension of the non-target shift from station one to station two; represents the dynamic time warping distance of the th dimension time series of the target shift and the non-target shift from station one to station two.
[0080] For better illustration, dissimilarity refers to the dissimilarity of the operation data of the target shift and any non-target shift from station one to station two; based on this, analyze the non-standard factors, which represent the significant different features or factors that occur between the target shift and any non-target shift from station one to station two, and may include but are not limited to differences in speed, differences in shift time, inconsistencies in staff allocation, particularities in task assignment, and variations in work processes, etc.
[0081] It should be noted that the dissimilarity between the target shift and any non-target shift is determined as the clustering distance.
[0082] Step S222: Divide all shifts of the rail train into multiple clustering clusters based on the clustering distance, define the clustering cluster with the largest number of shifts as the standard clustering cluster, obtain the clustering distance between the target shift and the clustering center of the standard clustering cluster, and obtain the non-standard factors of the target shift.
[0083] Specifically, using the K-means clustering algorithm, all shifts are divided into multiple clusters, that is, the operation data of all shifts from site one to site two in each cluster are similar; among them, the K-means clustering algorithm is an unsupervised learning algorithm widely used in the fields of data mining and machine learning. The main purpose is to divide the sample points in the data set into K clusters, so that the sample points within each cluster are as similar as possible, while the sample points between different clusters are as different as possible, that is, this goal is achieved by iteratively optimizing the sum of squared errors within the cluster; in this embodiment, the value of K is obtained according to the elbow method, that is, by calculating the sum of squared errors under different numbers of clusters, and then drawing a graph with the number of clusters on the horizontal axis and the sum of squared errors on the vertical axis. By observing the curve in the graph, a point in the shape of an "elbow" is found, that is, the point where the sum of squared errors begins to decrease significantly more slowly, and the corresponding number of clusters is confirmed as the optimal value of K.
[0084] It can be explained that the number of problems occurring in the operation of the rail train is relatively small compared to the normal number, and it is an even less likely event for the same operation problem to occur multiple times. Therefore, the cluster with the largest number of shifts is used as the standard cluster, that is, the time series of each dimension corresponding to the cluster center of the standard cluster for the rail train from site one to site two is used as the time series of each dimension for the rail train from site one to site two. It covers the performance of the rail train in each dimension during the entire process from departure from site one to arrival at site two, including but not limited to multiple aspects such as speed, acceleration, energy consumption, and passenger flow. By analyzing the time series of these standard clusters, we can better understand and evaluate the performance of the rail train under normal operating conditions, thus providing strong data support for subsequent maintenance, optimization, and decision-making.
[0085] Specifically, the clustering distance between the target shift and the cluster center of the standard cluster is used as the non-standard factor for the operation data of the target shift from site one to site two; the larger the clustering distance, the greater the difference between the time series of each dimension of the target shift from site one to site two and the standard time series of each dimension, and the greater the possibility that there are operation problems for the target shift from site one to site two.
[0086] Furthermore, in step S3, the operation process of the rail train includes multiple initial time periods. Based on the non-standard factor and multi-dimensional time series data, the first attention degree and the second attention degree of each initial time period are determined respectively, including:
[0087] Step S311: The initial time period includes the starting acceleration stage, the constant speed operation stage, and the deceleration braking stage in the operation process of the rail train. The first attention degree of the target shift in the starting acceleration stage and the deceleration braking stage is determined through the non-standard factor, which is the non-standard factor.
[0088] Preferably, the speed of a known rail train can directly reflect the operating efficiency of the rail train. Therefore, in the embodiments, the analysis is carried out mainly from the dimension of speed. Among them, the start time and end time of each stage, namely the start-up acceleration stage, the constant-speed operation stage, and the deceleration braking stage, are known, and this information can be obtained through monitoring systems such as the train control system, the data recorded by on-vehicle equipment, and the ground monitoring system.
[0089] It should be noted that since the start and stop stages of a rail train are one of the stages most prone to accidents during train operation, precise control is required for both the start-up acceleration stage and the deceleration braking stage to avoid collisions, derailments, or other safety accidents. Therefore, the first attention degree of the target shift in the start-up acceleration stage and the deceleration braking stage is assigned as a non-standard factor, that is .
[0090] Step S312: Divide the constant-speed operation stage into several sequence sub-segments according to the duration, and obtain the peaks and valleys of each sequence sub-segment in the constant-speed operation stage, and calculate the speed instability corresponding to the constant-speed operation stage. The corresponding calculation formula is:
[0091]
[0092] Among them, represents the speed instability; represents the variance of all speeds in the constant-speed operation stage; represents the maximum value of the speed sudden change of all peaks in the constant-speed operation stage; represents the duration of the constant-speed operation stage; represents the total number of peaks in the constant-speed operation stage.
[0093] For better illustration, the speed of the rail train in the constant-speed operation stage usually remains relatively stable. When there are large fluctuations in the speed, it indicates that there are problems with the operation of the rail train.
[0094] Specifically, taking the target shift as an example, the target shift is divided into multiple initial time periods between station one and station two. The initial time periods include the start-up acceleration stage, the constant-speed operation stage, and the deceleration braking stage during the operation of the rail train. Then, the constant-speed operation stage is divided into several sequence sub-segments according to the duration. Preferably, in this embodiment, the constant-speed operation stage is divided into 10 sequence sub-segments according to the duration. At this time, the initial time period is {start-up acceleration stage, constant-speed operation sub-segment 1, constant-speed operation sub-segment 2,..., constant-speed operation sub-segment 10, deceleration braking stage}.
[0095] Through the peak and trough detection algorithm, the peaks and troughs of each small segment of the sequence in the uniform speed running stage are obtained; and the speed difference and time interval between each peak and the adjacent trough are counted, and the ratio of the speed difference to the time interval is used as the speed sudden change of each peak; among them, the peak and trough detection algorithm is a technology used to identify and analyze local maximum and minimum values in data sequences. A peak refers to a local maximum point in a data sequence, and the values on both sides are smaller than it; and a trough refers to a local minimum point, and the values on both sides are larger than it.
[0096] To explain, It represents the variance of all speeds during the uniform speed operation stage. The larger the variance, the more drastic the speed change and the more unstable it is. It indicates the maximum value of the speed sudden change of all wave peaks in the uniform speed running stage. The larger the value, the greater the speed change in a short time, and the more unstable it is. The larger the value, the more unstable the speed of the train will be. The larger it is, the greater the possibility of problems with the rail train, and the more attention it needs.
[0097] Step S313: Calculate the initial adjustment factor of the attention degree in the uniform speed running stage, and the corresponding calculation formula is:
[0098]
[0099] in, represents the initial adjustment factor of attention; Indicates speed instability; It represents the mean value of all cruising speeds during the uniform speed operation stage; Indicates the predetermined cruising speed during the uniform speed operation phase.
[0100] It can be explained that The larger it is, the greater the possibility that there will be problems with the rail train and the more attention it needs.
[0101] Step S314: determine the track train position coordinates at the start and end times of the uniform speed operation phase, construct a speed standard sequence, and determine the second attention level of the uniform speed operation phase.
[0102] Understandably, the determination of the above values all indicates that there may be problems with the current rail train. During the uniform running stage of the rail train, there are many reasons for speed fluctuations. Among them, the basic conditions of the track, such as the slope and curvature changes of the track, cause speed fluctuations, which is a normal phenomenon; mechanical failures, such as failures in the traction system, braking system or suspension system, may all lead to unstable speeds; changes in track conditions, such as obstacles on the track or track damage, may affect the speed. Therefore, it is necessary to continue to analyze whether the speed fluctuation is a normal phenomenon or caused by problems, so data analysis is carried out to determine the second attention level.
[0103] Furthermore, in step S314, the calculation formula for the second attention level is:
[0104]
[0105] Wherein, represents the second attention level; represents the non-standard factor; represents the sequence composed of all sequence small segments during the uniform running stage; represents the speed standard sequence; represents and the dynamic time warping distance between; represents the linear normalization function.
[0106] Specifically, continuing from the above, taking the target shift as an example, using the position coordinates of the rail train at the start and end times of any uniform running small segment from station one to station two of the target shift as the start and end target coordinates, in the standard time series sequences of each dimension of the rail train from station one to station two, obtain the speed standard sequence within the start target coordinates to the end target coordinates. Among them, the track path of the sequence small segment and its corresponding speed standard sequence is the same, that is, the basic conditions of the track are the same. When the sequence small segment is less similar to the corresponding speed standard sequence, it indicates that the speed fluctuation during the uniform running stage is probably a fault phenomenon.
[0107] It can be explained that represents and the dynamic time warping distance between. The larger this value is, the less similar the uniform running stage is to its corresponding speed standard sequence, then is more credible.
[0108] Furthermore, in step S3, combine the first attention level and the second attention level to construct the attention level sequence of the target shift from the starting point to the ending point, and obtain multiple optimal time periods and corresponding attention levels, including:
[0109] Step S321: Determine the attention degree of each initial time period based on the first attention degree and the second attention degree, and construct the attention degree sequences of all initial time periods from the starting point to the ending point of the target shift;
[0110] Step S322: Divide the attention degree sequence into multiple sequence segments through clustering, merge all the initial time periods corresponding to each sequence segment to form each optimal time period, calculate the mean value of the attention degrees of all the initial time periods in the corresponding sequence segment, and determine the attention degree of each optimal time period.
[0111] Specifically, based on the above calculations, similarly obtain all the attention degrees of each initial time period between any two stations of any shift of the rail train, and further obtain all the attention degrees between the starting station and the ending station of the target shift. Construct the obtained all attention degrees into an attention degree sequence from the starting point to the ending point; use the K-means clustering algorithm to divide the attention degree sequence into multiple sequence segments, merge all the initial time periods corresponding to each sequence segment to form each optimal time period, calculate the mean value of the attention degrees of all the initial time periods in the corresponding sequence segment, and use it as the attention degree of each optimal time period.
[0112] It can be explained that in step S4, parallel encoding is performed on all the dimension time series data corresponding to the optimal time period, and the attention degree corresponding to the optimal time period is used as a label to complete data integration.
[0113] Specifically, use block Huffman coding to perform parallel encoding on all the train operation data within each optimal time period, that is, all the dimension data, to obtain a compressed encoding set corresponding to each shift and each optimal time period, and use the attention degree of each optimal time period as the label of its compressed encoding set; then store the compressed encoding sets and their labels of all the optimal time periods of each shift in the database to complete data integration; among them, block Huffman coding is to divide the data into multiple small blocks, and then perform Huffman coding on each small block independently to achieve an efficient compression effect. It is a variable-length coding method that assigns different lengths of codes according to the frequency of each character appearing in the data. Characters with high frequencies use shorter codes, and characters with low frequencies use longer codes.
[0114] It can be understood that concentrating data with different attention degrees in one optimal time period and adding an identification label to the compressed encoding set of each optimal time period can simplify data retrieval to quickly locate the optimal time period containing the concerned data, so that when analyzing the train operation problems of each shift subsequently, each optimal time period can be decoded in descending order according to the size of the label until the cause of the problem is analyzed, and then the subsequent optimal time periods are no longer decoded, improving the effect of problem analysis.
[0115] The data integration method for the integrated host of rail trains refers to a method of integrating and uniformly processing data from different systems and devices on rail trains. Through data integration, data sharing, analysis, and application can be achieved to improve the safety, efficiency, and service quality of train operation. It determines the weights between different shifts of rail trains by collecting data, analyzes the data to determine the optimal time period for shifts, and uses block Huffman coding to parallelly encode the data for each time period based on the optimal time period, improving the coding efficiency and achieving segmented compression, thereby improving the compression efficiency. This enables, when analyzing subsequent rail train operation problems, only decoding the data for the optimal time period of the problem, improving the processing efficiency of the integrated data.
[0116] It should be noted that the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0117] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A data integration method for a rail train fusion host, characterized in that: The method comprises: Collect multi-dimensional time series data of rail train operation and perform preprocessing; Define any train in the rail train as the target train, determine the target dimension between the target train and the non-target train based on the preprocessed multi-dimensional time series data, obtain the corresponding specificity through the target dimension, and construct the target dimension set to obtain the weight of each dimension of the target train; analyze the dissimilarity between the target train and the non-target train, and determine the non-standard factor of the target train; The operation process of a rail train includes multiple initial time periods. The first attention level and the second attention level of each initial time period are determined based on non-standard factors and multi-dimensional time series data. The attention level sequence of the target train from the starting point to the end point is constructed by combining the first attention level and the second attention level, and multiple optimal time periods and corresponding attention levels are obtained. All dimensional time series data corresponding to the optimal period are encoded in parallel, and the attention level corresponding to the optimal period is used as a label to complete data integration; Define any train in the rail train as the target train, determine the target dimension between the target train and the non-target train based on the preprocessed multi-dimensional time series data, obtain the corresponding specificity through the target dimension, and construct the target dimension set to obtain the weight of each dimension of the target train, including: Define two adjacent stations as Station 1 and Station 2, obtain the time series of each dimension from Station 1 to Station 2 based on the target shift and any non-target shift, calculate the dynamic time warping distance of the time series of the same dimension, and sort them from large to small. The dimension corresponding to the largest dynamic time warping distance is the target dimension, and the target dimension corresponding to the target shift and each non-target shift is obtained; Calculate the specificity corresponding to each target dimension, build a target dimension set, count the number of occurrences of each identical target dimension, calculate the weight corresponding to each identical target dimension, assign the weight of the dimension in the non-target dimension set, and obtain the weight of each dimension of the target shift; The rail train operation process includes multiple initial periods. The first attention level and the second attention level of each initial period are determined based on non-standard factors and multi-dimensional time series data, including: The initial period includes the start-up acceleration stage, the uniform speed running stage and the deceleration braking stage during the operation of the rail train. The first attention level of the target train in the start-up acceleration stage and the deceleration braking stage is determined by the non-standard factor. Nonstandard factors; The uniform speed running stage is divided into several small sequence segments according to the duration, and the peaks and troughs of each small sequence segment in the uniform speed running stage are obtained to calculate the speed instability corresponding to the uniform speed running stage. The corresponding calculation formula is: in, Indicates speed instability; represents the variance of all speeds during the uniform speed running stage; It indicates the maximum value of the sudden change of speed of all wave crests in the uniform speed running stage; Indicates the duration of the uniform speed running phase; Indicates the total number of wave peaks in the uniform speed running stage; Calculate the initial adjustment factor of attention in the uniform speed running stage. The corresponding calculation formula is: in, represents the initial adjustment factor of attention; Indicates speed instability; It represents the mean value of all cruising speeds during the uniform speed operation stage; Indicates the cruise speed scheduled during the uniform speed operation phase; Determine the track train position coordinates at the start and end times of the uniform speed operation phase, construct a speed standard sequence, and determine the second degree of attention of the uniform speed operation phase; The calculation formula for the second attention is: in, Indicates the second level of attention; represents a nonstandard factor; Indicates the sequence composed of all sequence sub-segments in the uniform speed running stage; Indicates the speed standard sequence; express and Dynamic time warping distance of represents the linear normalization function.
2. A data integration method for a rail train fusion host as claimed in claim 1, characterized in that: The multi-dimensional time series data of the rail train during operation is collected and preprocessed, wherein the multi-dimensional time series data includes any data of the speed, acceleration, train motor temperature, brake cylinder pressure, vibration data time series sequence and real-time position coordinates from the starting point to the end point, and the preprocessing method is standardized processing.
3. A data integration method for a rail train fusion host as claimed in claim 1, characterized in that: The calculation formula for the specificity of each target dimension is: in, Indicates the specificity of the target dimension; Indicates the dynamic time warping distance corresponding to the target dimension; Represents the mean of the dynamic time warping distances of all non-target dimensions.
4. A data integration method for a rail train fusion host as claimed in claim 1, characterized in that: The calculation formula for the weight corresponding to each same target dimension is: in, Indicates the target dimension set The weights of the same target dimension; Represents the total number of target dimensions in the target dimension set; Indicates the first The total number of the same target dimensions; Indicates the first The maximum value of the specificity of all target dimensions in the same target dimension; represents the linear normalization function.
5. The data integration method for rail train fusion host according to claim 1, characterized in that: Analyze the dissimilarity between target shifts and non-target shifts and determine the non-standard factors of the target shifts, including: Calculate the dissimilarity between the target shift and each non-target shift from station one to station two to determine the cluster distance; All train schedules are divided into multiple clusters based on clustering distance, and the cluster with the largest number of schedules is defined as the standard cluster. The clustering distance between the target schedule and the cluster center of the standard cluster is obtained to obtain the non-standard factor of the target schedule.
6. A data integration method for a rail train fusion host as claimed in claim 5, characterized in that: The formula for calculating dissimilarity is: in, Indicates dissimilarity; Represents the total number of different dimensions in all dimensions; Indicates the target shift from station 1 to station 2. The weight of each dimension; Indicates the non-target shift from station 1 to station 2. The weight of each dimension; Indicates the target shift and non-target shift from station one to station two. Dynamic time warping distance of time series in one dimension.
7. A data integration method for a rail train fusion host as claimed in claim 1, characterized in that: Combining the first attention and the second attention, the attention sequence of the target shift from the starting point to the end point is constructed to obtain multiple optimal time periods and corresponding attention, including: Determine the attention level of each initial period based on the first attention level and the second attention level, and construct the attention levels of all initial periods into an attention level sequence from the starting point to the end point of the target shift; The attention sequence is divided into multiple sequence segments by clustering, all initial time periods corresponding to each sequence segment are merged to form each optimal time period, the mean of the attention of all initial time periods in the corresponding sequence segment is calculated, and the attention of each optimal time period is determined.
Citation Information
Patent Citations
Time series data processing method, device and equipment and readable storage medium
CN115422264A
Power transaction method and system based on multi-dimensional data analysis
CN117764637A