A multi-dimensional orbit data association method, device, equipment and medium
By conducting extreme and difference trend analysis on the track and walking part data, frequent item sets are generated and correlation rules are screened, the problem of multi-source data fusion is solved, and the accuracy and efficiency of track data correlation analysis is achieved, and anomalies can be discovered in a timely manner and potential faults can be predicted.
Patent Information
- Application Number
- CN202510388947.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art is difficult to effectively integrate and analyze multi-source heterogeneous rail transit data, resulting in insufficient accuracy and reliability of correlation analysis results, especially in complex and diverse operating conditions, which is difficult to accurately mine the correlation relationship of track data.
By analyzing the extreme change direction and data difference change trend of the multi-source data subset, frequent item sets are generated and correlation rules are calculated, effective correlation rules are filtered using support degree and confidence thresholds, track and walking part data are integrated, and layer-by-layer search and pruning technology are used to process the data set.
It improves the accuracy and comprehensiveness of track data correlation analysis, can promptly detect potential failures, reduce computing resource consumption, accurately locate key data items, and generate real and reliable correlation rules.
Smart Images

Figure CN119903290B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining, and particularly relates to a multi-dimensional track data association method, device, equipment and medium. Background Art
[0002] With the development of rail transit, the amount of data generated in operation is huge and the types are complex and diverse, such as line conditions, train speeds, passenger loads, climate changes, random distributions of track irregularities, uncertainties of passenger behaviors, etc., which bring great challenges to processing and analysis. Track association analysis needs to integrate multi-source data. However, data such as running gear trend data, track vibration and shock trend data, and track design parameters vary greatly in format, semantics, and spatio-temporal scales, making it difficult to fuse.
[0003] In summary, how to accurately mine the association relationships of rail transit data to obtain accurate association analysis results is a technical problem to be solved in this field. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a multi-dimensional track data association method, device, equipment and medium, which can accurately mine the association relationships of rail transit data to obtain accurate association analysis results. The specific solutions are as follows:
[0005] In a first aspect, the present application discloses a multi-dimensional track data association method, including:
[0006] Performing trend analysis on the extreme value change directions of each track data in a multi-source data subset within a preset time range to obtain a first analysis result;
[0007] Performing trend analysis on the data difference changes between each running gear data in the multi-source data subset to obtain a second analysis result;
[0008] Integrating and marking the first analysis result, the second analysis result, the track data and the running gear data in the multi-source data subset according to the common dimensions to obtain a marked data set;
[0009] Performing layer-by-layer search and pruning processing on the marked data set to obtain a frequent item set composed of marked data items that meet the preset support threshold condition;
[0010] Generating a number of association rules based on the non-empty subset information of the frequent item set, calculating the confidence information of each association rule, and taking the association rules whose confidence information meets the preset confidence threshold condition as valid track data association rules.
[0011] Optionally, before performing trend analysis on the extreme value change directions of each track data in the multi-source data subset within a preset time range, the following steps are further included:
[0012] Screen a multi-source data subset related to the track association target from the multi-source data set including track vibration and shock data, running gear data, track design parameters, and track operation and maintenance data.
[0013] Optionally, after screening the multi-source data subset related to the track association target from the multi-source data set including track vibration and shock data, running gear data, track design parameters, and track operation and maintenance data, the following steps are further included:
[0014] Delete the abnormal data in the multi-source data subset to obtain a cleaned multi-source data subset;
[0015] Perform data filling processing on the cleaned multi-source data subset according to a preset data filling rule to obtain a preprocessed multi-source data subset; wherein, the preset data filling rule includes filling the numerical missing data in the cleaned multi-source data subset by using the median method, and / or filling the discrete type missing data in the cleaned multi-source data subset by using the mode method;
[0016] Correspondingly, performing trend analysis on the extreme value change directions of each track data in the multi-source data subset within a preset time range to obtain a first analysis result; performing trend analysis on the data difference change between each running gear data in the multi-source data subset to obtain a second analysis result, includes:
[0017] Perform trend analysis on the extreme value change directions of each track data in the preprocessed multi-source data subset within a preset time range to obtain a first analysis result;
[0018] Perform trend analysis on the data difference change between each running gear data in the preprocessed multi-source data subset to obtain a second analysis result.
[0019] Optionally, performing trend analysis on the extreme value change directions of each track data in the multi-source data subset within a preset time range to obtain a first analysis result, includes:
[0020] Decompose the track vibration values of each track data to obtain a decomposed target vibration trend, target seasonal component, and target residual information that change with time as the first analysis result.
[0021] Optionally, the step of decomposing the track vibration values of each track data to obtain a decomposed target vibration trend, target seasonal component, and target residual information that change with time as the first analysis result, includes:
[0022] Set the period length of the seasonal component according to the periodic characteristics of the orbital vibration values of each orbital data;
[0023] Initialize the vibration trend and the seasonal component as empty to set the orbital vibration value as the residual information;
[0024] Perform LOESS smoothing on the residual information according to a preset smoothing window size, and use the smoothed result as the current vibration trend;
[0025] Remove the current vibration trend from the orbital vibration value to obtain detrended data;
[0026] Segment the detrended data according to the period length to obtain each detrended data segment, and perform LOESS smoothing on each detrended data segment to obtain a preliminary seasonal component;
[0027] Perform within-period averaging on the preliminary seasonal component to obtain a corrected seasonal component;
[0028] Remove the corrected seasonal component from the orbital vibration value to obtain deseasonalized data;
[0029] Perform LOESS smoothing on the deseasonalized data according to a preset smoothing window size to update the current vibration trend, obtain a new current vibration trend, and jump to execute the step of removing the current vibration trend from the orbital vibration value until the change information of the vibration trend and the seasonal component is less than a preset change threshold, and output the current vibration trend and the corrected seasonal component as the target vibration trend and the target seasonal component;
[0030] Use the orbital vibration value, the target vibration trend, and the target seasonal component to determine the target residual information to obtain a first analysis result.
[0031] Optionally, the using the orbital vibration value, the target vibration trend, and the target seasonal component to determine the target residual information includes:
[0032] By Determine the target residual information;
[0033] Wherein, Represents The target vibration trend at time Represents The target seasonal component at time Represents The target residual information at time Represents The orbital vibration value at time
[0034] Optionally, performing trend analysis on the data difference changes between the data of each running part in the multi-source data subset to obtain a second analysis result, including:
[0035] Pairing the data of each running part to obtain a number of running part data pairs;
[0036] Performing a difference operation on each of the running part data pairs to obtain the difference results of each of the running part data pairs; the difference results include positive difference results and negative difference results;
[0037] Determining the second analysis result of the running part data based on the quantity information of the positive difference results and the quantity information of the negative difference results.
[0038] Optionally, the determining the second analysis result of the running part data based on the quantity information of the positive difference results and the quantity information of the negative difference results includes:
[0039] Taking the sum of the quantity of the positive difference results and the quantity of the negative difference results as the total number of data pairs;
[0040] Performing binomial distribution processing using the total number of data pairs, the success probability of the positive difference result / negative difference result of the running part data pair, and the number of successful times to obtain a success probability value, and calculating significant level information based on the success probability value;
[0041] Analyzing the trend of the data difference change of the running part data according to the size result of the quantity information of the positive difference results and the quantity information of the negative difference results and the significant level information to obtain a second analysis result.
[0042] Optionally, the analyzing the trend of the data difference change of the running part data according to the size result of the quantity information of the positive difference results and the quantity information of the negative difference results and the significant level information to obtain a second analysis result includes:
[0043] When the quantity of the positive difference results is greater than the quantity of the negative difference results and the significant level information is less than a preset significant level threshold, it is determined that the data difference change of the running part data is an upward trend, and a second analysis result with an upward trend analysis result is obtained;
[0044] When the quantity of the positive difference results is less than the quantity of the negative difference results and the significant level information is less than a preset significant level threshold, it is determined that the data difference change of the running part data is a downward trend, and a second analysis result with a downward trend analysis result is obtained.
[0045] Optionally, integrating and labeling the first analysis result, the second analysis result, the track data, and the running gear data in the multi-source data subset according to the common dimension to obtain a labeled data set, including:
[0046] Integrating the track data and the running gear data according to the common dimension to obtain a process integration data set;
[0047] Inserting the first analysis result and the second analysis result into the process integration data set in field form to obtain an integrated data set;
[0048] Performing interval partitioning on the integrated data of numerical type in the integrated data set, and performing discrete labeling on the integrated data after partitioning the intervals according to the discrete labeling method to obtain a labeled data set.
[0049] Optionally, performing layer-by-layer search and pruning on the labeled data set to obtain a frequent item set composed of labeled data items that meet the preset support threshold condition, including:
[0050] Scanning the labeled data set to count the occurrence times information of each labeled data item in the labeled data set;
[0051] Calculating the support of the occurrence times information of each labeled data item in the data record quantity of the labeled data set;
[0052] Pruning the labeled data items with support less than the preset support threshold condition to obtain updated labeled data items;
[0053] Connecting the updated labeled data items in each data record according to the preset connection rule, using the connected labeled data items as new labeled data items, and jumping to execute the step of counting the occurrence times information of each labeled data item in the labeled data set until pruning ends, and outputting the frequent item set.
[0054] In a second aspect, the present application discloses a multi-dimensional track data association device, including:
[0055] A first trend analysis module, configured to perform trend analysis on the extreme value change direction of each track data in the multi-source data subset within a preset time range to obtain a first analysis result;
[0056] A second trend analysis module, configured to perform trend analysis on the data difference change between each running gear data in the multi-source data subset to obtain a second analysis result;
[0057] An integration and marking module, configured to integrate and mark the first analysis result, the second analysis result, the track data in the multi-source data subset, and the running gear data according to a common dimension to obtain a marked data set;
[0058] A search processing module, configured to perform a layer-by-layer search and pruning process on the marked data set to obtain a frequent item set composed of marked data items that meet a preset support threshold condition;
[0059] An association generation module, configured to generate a number of association rules based on the information of each non-empty subset of the frequent item set, calculate the confidence information of each association rule, and use the association rules whose confidence information meets the preset confidence threshold condition as valid track data association rules.
[0060] In a third aspect, the present application discloses an electronic device, including:
[0061] A memory, configured to store a computer program;
[0062] A processor, configured to execute the computer program to implement the steps of the multi-dimensional track data association method disclosed above.
[0063] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the steps of the multi-dimensional track data association method disclosed above are implemented.
[0064] It can be seen that the present application discloses a multi-dimensional track data association method, including: performing trend analysis on the extreme value change direction of each track data in a multi-source data subset within a preset time range to obtain a first analysis result; performing trend analysis on the data difference change between each running gear data in the multi-source data subset to obtain a second analysis result; integrating and marking the first analysis result, the second analysis result, the track data and the running gear data in the multi-source data subset according to the common dimension to obtain a marked data set; performing layer-by-layer search and pruning processing on the marked data set to obtain a frequent item set composed of marked data items that meet the preset support threshold condition; generating a number of association rules based on the non-empty subset information of the frequent item set, calculating the confidence information of each association rule, and taking the association rules whose confidence information meets the preset confidence threshold condition as effective track data association rules. Thus, by performing trend analysis on the extreme value change direction of track data and trend analysis on the difference change of running gear data, the first analysis result and the second analysis result are obtained, deeply analyzing the change trend of the running state of the track and the running gear, timely discovering abnormal increases in track vibration or abnormal fluctuations in running gear data, providing a basis for predicting potential faults, integrating the above data according to the common dimension, breaking data barriers, and forming a comprehensive and systematic marked data set. It solves the problem of multi-source data fusion, provides a rich data basis for mining multi-dimensional association relationships, and ensures the accuracy and comprehensiveness of the analysis results. Using layer-by-layer search and pruning techniques to process the marked data set can quickly and accurately obtain frequent item sets that meet the preset support threshold, greatly improving the mining efficiency, accurately positioning key marked data items, and reducing computational resource consumption. Finally, association rules are generated based on the non-empty subsets of the frequent item set, and rules with confidence levels meeting the threshold are strictly calculated and screened to ensure that the obtained effective track data association rules are true and reliable, and the internal relationship between track data and running gear data can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0066] Figure 1 It is a flowchart of a multi-dimensional track data association method disclosed in the present application;
[0067] Figure 2 It is a flowchart of a frequent item set generation method disclosed in the present application;
[0068] Figure 3 Structural schematic diagram of a multi-dimensional track data association device disclosed in the present application;
[0069] Figure 4 Structural diagram of an electronic device disclosed in the present application. Specific implementation manners
[0070] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0071] With the continuous advancement of the urbanization process, the rail transit industry has developed rapidly. As an efficient, safe, and environmentally friendly transportation mode, rail transit occupies an important position in the urban transportation system. Track monitoring data and vehicle running gear trend data are crucial for the safe operation of rail transit. The interaction between tracks directly affects the running smoothness, safety, and comfort of trains. By analyzing track trend data such as track vibration, wavelength, and wave depth, the dynamic characteristics of the track can be understood, and potential problems can be discovered in a timely manner. The trend data of the vehicle running gear, such as impact and vibration, reflects the dynamic response during the operation of the vehicle and is of great significance for evaluating the running state and safety of the vehicle.
[0072] The track vibration and impact trend data are closely related to the vehicle running gear trend data. The intensification of track vibration often leads to an increase in the impact and vibration of the vehicle running gear, and by optimizing the track design parameters, this positive correlation can be reduced to a certain extent. The track trend data and the track design parameters influence each other. A smaller curve radius will increase the contact force between tracks, resulting in increased track vibration, and the wavelength and wave depth may also increase. Reasonable selection of track design parameters can reduce the adverse changes in the track trend data. The vehicle running gear trend data is closely related to the track design parameters. Unreasonable track design parameters may lead to the deterioration of the vehicle running gear trend data, while optimizing the track design parameters can improve the running state of the vehicle running gear.
[0073] At present, the analysis of track association requires the integration of various data such as the trend data of the running gear, the trend data of track vibration and impact, and track design parameters. However, these data vary in format, semantics, spatio-temporal scale, etc., posing great technical challenges to the integration work. The trend data of track vibration and impact and the trend data of the vehicle running gear are characterized by being massive, multi-source heterogeneous, and high-noise. Extracting effective information from them and performing correlation analysis requires complex data processing and analysis methods. Traditional data processing technologies are inefficient in processing big data, difficult to meet the real-time requirements, and have limitations in mining the deep-level features and laws of data. In the data processing process, different types of data need to adopt different processing methods and tools, and it is quite difficult to integrate and analyze structured and unstructured data. For the association analysis and mining work, accurately sorting out valuable association relationships from them is like precisely "fishing for needles" in the vast ocean of data, which requires a large amount of computing resources and time costs. The operating conditions of rail transit are complex and diverse, including different line conditions, train speeds, passenger loads, etc. These factors will affect the track interaction and the performance of the vehicle running gear. The current association analysis methods have insufficient adaptability under different operating conditions, resulting in the accuracy and reliability of the correlation analysis results being affected.
[0074] In the face of massive track data, data sampling methods such as random sampling are currently adopted. In the mining of track association analysis, the current practice in data cleaning and sorting is often to first conduct a preliminary review and screening of the track data by professional data processing personnel based on experience through manual screening, and then use statistical methods to mine track association relationships, such as correlation analysis (Pearson correlation coefficient, Spearman correlation coefficient, etc.). However, there are many uncertain factors in the operating environment and conditions of rail transit, such as line conditions, train speeds, passenger loads, climate changes, random distribution of track irregularities, uncertainties in passenger behavior, etc. These factors will affect the track interaction and the performance of the vehicle running gear, increasing the difficulty of association analysis and resulting in the accuracy and reliability of the association analysis results being affected.
[0075] Therefore, the present invention proposes a multi-dimensional track data association scheme, which can process and analyze complex non-linear problems and multi-factor coupling problems, accurately reveal the internal relationships between data such as track vibration and impact trend data, running gear data, and track design parameters, and mine the association relationships of rail transit data to obtain accurate association analysis results.
[0076] Refer to Figure 1 As shown, an embodiment of the present invention discloses a multi-dimensional track data association method, including:
[0077] Step S11: Perform a trend analysis on the extreme value change direction of each track data in the multi-source data subset within a preset time range to obtain a first analysis result.
[0078] In this embodiment, before performing the trend analysis on the extreme value change direction of each track data in the multi-source data subset within a preset time range, it further includes: screening a multi-source data subset related to the track association target from the multi-source data set including track vibration impact data, running gear data, track design parameters, and track operation and maintenance data. It can be understood that multi-source data from multiple parties is obtained. Specifically, a multi-source data set such as track vibration impact data, running gear data, track design parameters, and track operation and maintenance data (track status data, track maintenance data) is obtained. Among them, the track vibration impact data includes: time, city, line, train number, vehicle number, station, up / down direction, departure station, destination station, left / right rail, axle number, position number, starting kilometer mark, ending kilometer mark, sampling start time, sampling end time, starting rotational speed, ending rotational speed, vibration peak-to-peak value, vibration average value, wavelength, wave depth, vibration maximum value, vibration effective value, corrugation mark, impact SV effective value, etc.; the running gear data includes: time, city, line, train number, vehicle number, axle number, position number, impact dB, impact SV, temperature, wheel flatness, etc.; the track design parameters include: city, line, up / down direction, station, left / right rail, kilometer mark, curve, fastener, weld, turnout, slope, ballast, sleeper, curve radius, etc.; the track status data includes: time, city, line, up / down direction, station name, left / right rail, kilometer mark, wave depth, wavelength measurement value, damage, roughness, etc.; the track maintenance data includes: city, line, up / down direction, station name, left / right rail, track grinding time, grinding method, grinding kilometer mark information, etc.
[0079] Since there are various types of data in the multi-source data set, it is necessary to screen out the target data based on the track association target of the track association relationship to construct a multi-source data subset. Specifically, the track association target generally includes three different mining targets, namely 1. The mining target of the association rule between the track inspection result and track design parameters, speed, etc.; 2. The mining target of the correlation analysis between the trend change of the track vibration value, the change of the running gear data, and the track inspection result; 3. The mining target of the association rule between the large track vibration value and track design parameters, speed, etc.
[0080] When the mining target is the association rules between track inspection results and track design parameters, speed, etc., the target data to be screened should be track vibration and impact trend data (time, city, line, train number, vehicle number, station, up / down direction, left / right rail, axle number, position number, departure station, destination station, kilometer post, rotational speed, etc.), track design parameters (city, line, up / down direction, station, left / right rail, kilometer post, curve, fastener, weld, turnout, gradient, ballast bed, sleeper, curve radius, etc.), track status data (time, city, line, up / down direction, station name, left / right rail, kilometer post, wave depth, wavelength measurement value, damage, irregularity, etc.) and track maintenance data (city, line, up / down direction, station, left / right rail, track grinding time, grinding method, grinding kilometer post information, etc.).
[0081] When the mining target is the correlation analysis between the trend change of track vibration value, the change of running gear data and track inspection results, the target data to be screened are track trend data (time, city, line, train number, vehicle number, station, up / down direction, departure station, destination station, left / right rail, axle number, position number, kilometer post, vibration effective value, vibration peak-to-peak value, etc.), running gear data (time, city, line, train number, vehicle number, axle number, position number, time, tread dB, tread vibration effective value) and track maintenance data (city, line, up / down direction, station, left / right rail, track grinding time, grinding method, grinding kilometer post information, etc.).
[0082] When the mining target is the association rules between large track vibration value and track design parameters, speed, etc., the target data to be screened should be track trend data (time, city, line, train number, vehicle number, station, up / down direction, departure station, destination station, left / right rail, axle number, position number, kilometer post, vibration effective value, speed, etc.), track design parameters (city, line, station, up / down direction, left / right rail, kilometer post, curve, fastener, weld, turnout, gradient, ballast bed, sleeper, curve radius, etc.).
[0083] In this way, according to the actual track association target, the target data corresponding to the current track association target can be screened from numerous multi-source data sets. It can accurately locate key data based on different mining targets, eliminate the interference of irrelevant information, make the analysis direction clearer, and focus on the core mining target. In addition, if all the data in the multi-source data set are comprehensively analyzed, the computing resources and time costs are high. Screening the target data can significantly reduce the data volume and simplify the processing process; finally, screening the target data can ensure that the data participating in the analysis is highly relevant, reduce the influence of noise data and irrelevant factors, make the association analysis more targeted, the mined association relationship more accurate, and the analysis result more reliable.
[0084] In this embodiment, after screening a multi-source data subset related to the track association target from a multi-source data set including track vibration impact data, running gear data, track design parameters, and track operation and maintenance data, the following steps are further included: deleting abnormal data in the multi-source data subset to obtain a cleaned multi-source data subset; performing data filling processing on the cleaned multi-source data subset according to a preset data filling rule to obtain a preprocessed multi-source data subset; wherein, the preset data filling rule includes filling numerical missing data in the cleaned multi-source data subset by using the median method, and / or filling discrete type missing data in the cleaned multi-source data subset by using the mode method. It can be understood that after determining the multi-source data subset related to the track association target, the outliers in the current multi-source data subset are deleted to obtain a cleaned multi-source data subset; for the numerical missing values in the cleaned multi-source data subset, the median method is used for missing value filling, and for the missing values of discrete variables, the mode filling method is used to fill the missing values to ensure data integrity, and then a preprocessed multi-source data subset is obtained.
[0085] In a specific implementation manner, the steps of filling numerical missing data by using the median method are as follows:
[0086] The median refers to the data located in the middle position in a set of data, which divides the cleaned multi-source data subset into two equal parts, with half of the data greater than the median and half of the data less than the median. The basic principle of using the median to fill missing values is to replace the missing values with the median of this attribute to reduce the bias caused by missing values. First, sort the data by size, and assume the sorting result is:
[0087] ;
[0088] wherein, represents the value of the th data in the cleaned multi-source data subset, , is a positive integer, and the median calculation formula is as follows:
[0089] ;
[0090] In another specific implementation manner, the mode method is used to fill missing values for discrete data. Using the mode to fill missing values can retain the distribution characteristics of the data, especially suitable for nominal variables. The mode filling method is simple and does not require a complex statistical model.
[0091] In this embodiment, trend analysis of the extreme value change direction is performed on each track data in the preprocessed multi-source data subset within a preset time range to obtain a first analysis result. It can be understood that performing trend analysis of the extreme value change direction on each track data (filled discrete track data, filled numerical track data) in the preprocessed multi-source data subset within the preset time range to obtain a first analysis result can ensure the accuracy, continuity, and effectiveness of the track data and avoid errors in the trend analysis of the extreme value change direction of the track data.
[0092] In this embodiment, the track vibration values of each track data are decomposed to obtain a target vibration trend, a target seasonal component, and target residual information that change with time after decomposition as the first analysis result. Specifically, the period length of the seasonal component is set according to the periodic characteristics of the track vibration values of each track data; the vibration trend and the seasonal component are initialized to be empty to set the track vibration value as the residual information; the residual information is subjected to LOESS smoothing processing according to a preset smoothing window size, and the smoothed result is used as the current vibration trend; the current vibration trend is removed from the track vibration value to obtain detrended data; the detrended data is segmented according to the period length to obtain each detrended data segment, and each detrended data segment is subjected to LOESS smoothing processing to obtain a preliminary seasonal component; the preliminary seasonal component is subjected to in-period averaging processing to obtain a corrected seasonal component; the corrected seasonal component is removed from the track vibration value to obtain deseasonalized data; the deseasonalized data is subjected to LOESS smoothing processing according to a preset smoothing window size to update the current vibration trend, obtain a new current vibration trend, and jump to execute the step of removing the current vibration trend from the track vibration value until the change information of the vibration trend and the seasonal component is less than a preset change threshold, and the current vibration trend and the corrected seasonal component are output as the target vibration trend and the target seasonal component; the target residual information is determined by using the track vibration value, the target vibration trend, and the target seasonal component to obtain a first analysis result.
[0093] It can be understood that using STL (Seasonal-Trend Decomposition using LOESS) decomposition to decompose the track vibration value (vibration effective value, vibration peak-to-peak value) of the track data into three parts: a vibration trend (Trend), a seasonal component (Seasonal), and residual information (Residual), and its decomposition formula is expressed as follows:
[0094] ;
[0095] Among them, represents The target vibration trend at a moment, representing the long-term change trend of the data; Represent The target seasonal component at a moment, representing the periodic fluctuation with a fixed cycle length. Represent The target residual information at a moment, representing the random fluctuation or noise part.
[0096] First, average each trend component after STL decomposition by day to form a new daily average vibration effective value trend data, denoted as X. Secondly, calculate the extreme value change of X as the standard value, and the formula is:
[0097] ;
[0098] Where max and min respectively represent the maximum and minimum values in the whole segment of daily average vibration effective trend data, and argmax and argmin represent the serial numbers of the maximum and minimum values in the whole segment of data. Define the change value of each time period (for example, data of consecutive 2 days) as follows:
[0099] ;
[0100] Where is the starting point of each time period, is the ending point of each time period. A total of 3 discrete changes are defined as: decline, stable, and rise. If the value is less than or equal to the threshold (for example, 0.05), then the change is defined as stable. If the value is greater than (for example, 0.05), then it is defined as decline or rise. Among them, decline and rise are determined according to the magnitudes of the starting point and ending point of this time period, that is, if at this time, it is rise. If at this time, it is decline.
[0101] In this way, take the trend component obtained by averaging the STL decomposition by day as the trend data of each day, and then conduct trend analysis based on the extreme value change situation of the trend data to obtain three analysis results of trend rise / decline / stable.
[0102] Step S12: Conduct trend analysis on the data difference change between the data of each running part in the multi-source data subset to obtain the second analysis result.
[0103] In this embodiment, trend analysis is performed on the data difference changes among the data of each running gear in the preprocessed multi-source data subset to obtain a second analysis result. It can be understood that trend analysis of the data difference change directions among the data of each running gear (the filled discrete running gear data and the filled numerical running gear data) in the preprocessed multi-source data subset within the preset time range to obtain a second analysis result can ensure the accuracy, continuity, and effectiveness of the running gear data and avoid errors in the trend analysis of the data difference change directions of the running gear data.
[0104] In this embodiment, the data of each running gear are paired to obtain a number of running gear data pairs; a difference operation is performed on each of the running gear data pairs to obtain the difference results of each of the running gear data pairs; the difference results include positive difference results and negative difference results; the second analysis result of the running gear data is determined based on the quantity information of the positive difference results and the quantity information of the negative difference results. Specifically, the sum of the quantity of the positive difference results and the quantity of the negative difference results is used as the total number of data pairs; binomial distribution processing is performed using the total number of data pairs, the success probability of the positive / negative difference results of the running gear data pairs, and the number of successful times to obtain a success probability value, and the significance level information is calculated based on the success probability value; the trend of the data difference change of the running gear data is analyzed according to the size result of the quantity information of the positive difference results and the quantity information of the negative difference results and the significance level information to obtain a second analysis result. Among them, when the quantity of the positive difference results is greater than the quantity of the negative difference results and the significance level information is less than the preset significance level threshold, it is determined that the data difference change of the running gear data is an upward trend, and a second analysis result with an upward trend analysis result is obtained; when the quantity of the positive difference results is less than the quantity of the negative difference results and the significance level information is less than the preset significance level threshold, it is determined that the data difference change of the running gear data is a downward trend, and a second analysis result with a downward trend analysis result is obtained.
[0105] It can be understood that first, based on the time range of the track vibration data, the running gear data corresponding to the time range are selected, and then the Cox-stuart trend test is used to perform trend analysis on the running gear data (tread db, effective value of tread vibration) to obtain three analysis results: upward trend / downward trend / no trend; the specific steps of the running gear data trend analysis are as follows:
[0106] Select the running gear data with a data length of to form a data set .
[0107] Take and Form a pair of data ( , ∈ ), a total of pairs of running gear data pairs ( , ), ( , ), ( , ), etc., where is:
[0108] ;
[0109] Subtract the two running gear data of each pair, denoted as , set to record as the number of positive ones, that is, the number of positive difference results; Record as the number of negative ones, that is, the number of negative difference results.
[0110] Use the discrete distribution of the binomial to test whether the trend is significant. The specific calculation formula for the significance level value is as follows:
[0111] ;
[0112] where is the binomial distribution, expressed as follows:
[0113] ;
[0114] Among them, is the number of successes, that is, the test statistic, is the number of trials, that is, the total number of data pairs (excluding the number of pairs with a difference of 0), is the probability of success in each trial (for example, 0.5):
[0115] ;
[0116] When > , and is less than the significance level threshold (0.05), then there is an upward trend; when < , and is less than the significance level threshold (0.05), then there is a downward trend, and in other cases, there is no trend.
[0117] Step S13: Integrate and label the first analysis result, the second analysis result, the track data and the running gear data in the multi-source data subset according to the common dimension to obtain a labeled data set.
[0118] In this embodiment, the track data and the running gear data are integrated according to the common dimension to obtain a process integration data set; the first analysis result and the second analysis result are inserted into the process integration data set in the form of fields to obtain an integrated data set; the integrated data of the numerical type in the integrated data set is subjected to interval division processing, and the integrated data after division is discretely labeled according to the discrete labeling method to obtain a labeled data set. It can be understood that data integration needs to be carried out according to the common dimension of multi-source data and the mining target of the association relationship. For example, when there are the following three different mining targets, the integrated data is as follows:
[0119] If the mining target is the association rules between track inspection results and track design parameters, speed, etc., the integrated data set should be time, city, line, train number, vehicle number, up / down direction, departure station, destination station, left / right track, axle number, position number, kilometer mark, rotation speed, wavelength, wave depth, damage, unevenness, curve, fastener, weld, turnout, gradient, sleeper, ballast, etc.
[0120] If the mining target is the correlation analysis between the trend change of track vibration value, the trend data change of the running gear, and the track inspection results, the integrated data should be time, city, line, train number, vehicle number, up / down direction, departure station, destination station, left / right track, axle number, position number, kilometer mark, peak-to-peak value trend analysis result of track vibration, effective value trend analysis result of track vibration, effective value trend analysis result of running gear tread vibration, dB trend analysis result of running gear tread impact, whether to grind, etc.
[0121] If the mining target is the association rules between the large track vibration value and track design parameters, speed, etc., the integrated data set should be time, city, line, train number, vehicle number, up / down direction, departure station, destination station, left / right track, axle number, position number, kilometer mark, effective value trend rising analysis result of track vibration, rotation speed, curve, fastener, weld, turnout, gradient, sleeper, ballast, etc.
[0122] Furthermore, discrete data such as fastener types, sleeper types, and ballast types in track design parameters are marked by category (for example, if there are 3 types of fasteners, they are marked as Fastener 1, Fastener 2, and Fastener 3 respectively according to the type); for numerical data such as curve radius data, wavelength, wave depth, and rotational speed, first, equal-width interval division is used to convert continuous numerical values into discrete intervals, and then they are marked according to the discrete intervals (for example, if the curve radius is divided into two categories: 0 - 1000m and 1000 - 2000m, then the curve radius can be marked as Curve Radius 1 and Curve Radius 2).
[0123] In this way, the processed multi-source data is integrated and marked into a unified data set according to common dimensions (such as time, city, line, etc.), forming a marked data set.
[0124] Step S14: Perform layer-by-layer search and pruning processing on the marked data set to obtain a frequent item set composed of marked data items that meet the preset support threshold condition.
[0125] In this embodiment, scan the marked data set and set a minimum support threshold (such as 0.2) and a minimum confidence threshold (such as 0.2), and repeatedly mine the relationships between multi-dimensional data items to obtain a frequent item set composed of data items that meet the above threshold conditions.
[0126] Step S15: Generate a number of association rules based on the information of each non-empty subset of the frequent item set, calculate the confidence information of each association rule, and use the association rules whose confidence information meets the preset confidence threshold condition as valid track data association rules.
[0127] In this embodiment, use the combination method to generate each non-empty subset in the frequent item set, then use the data item information in each non-empty subset to generate the corresponding association rule, and calculate the association rule according to the confidence. If the calculated confidence meets the preset confidence threshold condition, then use the corresponding association rule as a valid track data association rule.
[0128] It can be seen that the present application discloses a multi-dimensional track data association method, including: performing trend analysis on the extreme value change directions of each track data in a multi-source data subset within a preset time range to obtain a first analysis result; performing trend analysis on the change of data difference between each running gear data in the multi-source data subset to obtain a second analysis result; integrating and marking the first analysis result, the second analysis result, the track data and the running gear data in the multi-source data subset according to the common dimension to obtain a marked data set; performing layer-by-layer search and pruning processing on the marked data set to obtain a frequent item set composed of marked data items that meet the preset support threshold condition; generating a number of association rules based on the non-empty subset information of the frequent item set, calculating the confidence information of each association rule, and taking the association rules whose confidence information meets the preset confidence threshold condition as effective track data association rules. Thus, by performing trend analysis on the extreme value change direction of track data and trend analysis on the difference change of running gear data, the first analysis result and the second analysis result are obtained, deeply analyzing the change trend of the running states of the track and the running gear, timely discovering the abnormal increase of track vibration or the abnormal fluctuation of running gear data, providing a basis for predicting potential faults, integrating the above data according to the common dimension, breaking the data barrier, and forming a comprehensive and systematic marked data set. It solves the problem of multi-source data fusion, provides a rich data basis for mining multi-dimensional association relationships, and ensures that the analysis results are more accurate and comprehensive. By using the layer-by-layer search and pruning technology to process the marked data set, the frequent item set that meets the preset support threshold can be quickly and accurately obtained, greatly improving the mining efficiency, accurately positioning the key marked data items, and reducing the consumption of computing resources. Finally, association rules are generated based on the non-empty subsets of the frequent item set, and the rules with confidence meeting the threshold are strictly calculated and screened to ensure that the obtained effective track data association rules are true and reliable, and the internal relationship between track data and running gear data can be obtained.
[0129] Referring to Figure 2 As shown, Embodiment of the present invention specifically discloses Step S14: A method for performing layer-by-layer search and pruning processing on the marked data set to obtain a frequent item set composed of marked data items that meet the preset support threshold condition. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically:
[0130] Step S21: Scanning the marked data set to count the number of times each marked data item appears in the marked data set.
[0131] In this embodiment, an association relationship mining model is established, such as an Apriori algorithm model, and the minimum support threshold (0.2) and minimum confidence threshold (0.2) of the model are set. First, the Apriori algorithm performs the first scan on the labeled dataset, and counts the number of times each single item (single labeled data item) appears in the labeled dataset. For example, when the labeled database has Tid 1.ACD, Tid 2.BCE, Tid 3.ABCE, and Tid 4.BE, the first scan operation is performed to count the number of times of each single labeled data item of ABCDE respectively. The statistical results are: the number of times A appears is 2, the number of times B appears is 3, the number of times C appears is 3, the number of times D appears is 1, the number of times E appears is 3, and the number of data records is 5.
[0132] Step S22: Calculate the support of the number information of each labeled data item in the data records of the labeled dataset.
[0133] In this embodiment, the Apriori algorithm model is used to calculate the support of all single items first. According to the above statistical results and the number of data records, the supports of ABCDE are calculated respectively as: 0.4, 0.6, 0.6, 0.2, 0.6.
[0134] Step S23: Prune the labeled data items that meet the condition that the support is less than the preset support threshold to obtain the updated labeled data items.
[0135] In this embodiment, then according to the preset minimum support threshold of 0.2, the single items that meet the conditions are selected as ABCE as the updated labeled data items, and the frequent item set 1 composed of these single items is {A, B, C, E}.
[0136] Step S24: Connect the updated labeled data items in each data record according to the preset connection rule, and use the connected labeled data items as the new labeled data items, and jump to execute the step of counting the number of times each labeled data item appears in the labeled dataset until pruning is completed, and output the frequent item set.
[0137] In this embodiment, after obtaining the frequent item set 1, candidate item sets are generated by self-joining the frequent item set 1. The joining rule is that if the first k - 2 items of two frequent item sets 1 are the same and the (k - 1)-th item is different, then they can be joined to generate a candidate item set. Therefore, the self-joining results of the above frequent item set 1 are: {A, B}, {A, C}, {A, E}, {B, C}, {B, E}, {C, E}. The occurrence times of the marked data items after joining in the candidate item sets are counted. The occurrence times of the joined and marked data items {A, B}, {A, C}, {A, E}, {B, C}, {B, E}, {C, E} are respectively: 1, 2, 1, 2, 3, 2. Then, the support degrees of the newly generated candidate item sets are calculated, and screening is performed again according to the screening requirements based on the minimum support degree threshold. The screening results are: {A, C}, {B, C}, {B, E}, {C, E}. For the generated candidate item sets, the data set is scanned again, the occurrence times of each candidate item set in the data set are counted, and its support degree is calculated. Then, according to the minimum support degree threshold, the candidate item sets that meet the conditions are screened out, and these candidate item sets constitute the frequent item set 2. Repeat the above steps until no new frequent item sets can be generated. Continuously repeat the above process to generate candidate k-item sets, calculate their support degrees and screen them until no new frequent item sets can be generated. Finally, the frequent item set is {B, C, E}, and the number of times is 2.
[0138] In this way, the frequent item set generation process adopts a layer-by-layer search and pruning technique, starting from the frequent item set 1 to gradually generate higher-order frequent item sets. After each new candidate item set is generated, support degree calculation and screening are performed to timely exclude the item sets that do not meet the conditions. Unnecessary calculations are reduced, the efficiency of the algorithm is improved, and good performance can be maintained when dealing with large-scale data sets.
[0139] In this embodiment, after the frequent item set is determined, for a frequent item set, all its non-empty subsets are generated by combination. For example, for the frequent item set {B, C, E}, its non-empty subsets are {B}, {C}, {E}, {B, C}, {B, E}, {C, E}.
[0140] For each generated rule , its confidence degree is calculated. For each generated rule, its confidence degree is calculated according to the confidence degree formula .
[0141] The calculated confidence degree is compared with the preset minimum confidence degree threshold. If the calculated confidence degree is greater than the minimum confidence degree threshold of 0.2, then this rule is a valid association rule. Correspondingly, the judgment methods for the rules corresponding to other non-empty subsets are the same as the above , and will not be elaborated here.
[0142] It can be seen that by calculating the support degrees of single items and candidate item sets and screening them according to the minimum support degree threshold, it is possible to accurately find the frequently occurring combinations of data items, i.e., frequent item sets, from a large amount of complex data sets. These frequent item sets are the core parts with representativeness and relevance in the data, which helps to focus on key information and avoid wasting computing resources and time on irrelevant data. Moreover, setting the minimum support degree and minimum confidence degree thresholds to screen the frequent item sets and association rules ensures that the mined rules have a certain degree of universality and reliability. Only those rules that frequently occur and have a high degree of association in the data set will be retained as valid association rules, avoiding mining some false rules that occur accidentally or have weak association, and improving the credibility and practicality of the analysis results.
[0143] Referring to Figure 3 as shown, the present invention also correspondingly discloses a multi-dimensional track data association device, including:
[0144] A first trend analysis module 11, configured to perform trend analysis on the extreme value change directions of each track data in a multi-source data subset within a preset time range to obtain a first analysis result;
[0145] A second trend analysis module 12, configured to perform trend analysis on the data difference change between each running gear data in the multi-source data subset to obtain a second analysis result;
[0146] An integration and marking module 13, configured to integrate and mark the first analysis result, the second analysis result, the track data and the running gear data in the multi-source data subset according to the common dimension to obtain a marked data set;
[0147] A search and processing module 14, configured to perform layer-by-layer search and pruning processing on the marked data set to obtain a frequent item set composed of marked data items that meet the preset support degree threshold condition;
[0148] An association generation module 15, configured to generate a number of association rules based on the non-empty subset information of the frequent item set, calculate the confidence information of each association rule, and use the association rules whose confidence information meets the preset confidence degree threshold condition as valid track data association rules.
[0149] It can be seen that the present application discloses performing trend analysis on the extreme value change directions of each track data in a multi-source data subset within a preset time range to obtain a first analysis result; performing trend analysis on the data difference change between each running gear data in the multi-source data subset to obtain a second analysis result; integrating and labeling the first analysis result, the second analysis result, the track data and the running gear data in the multi-source data subset according to a common dimension to obtain a labeled data set; performing layer-by-layer search and pruning processing on the labeled data set to obtain a frequent item set composed of labeled data items that meet the preset support threshold condition; generating a number of association rules based on the non-empty subset information of the frequent item set, calculating the confidence information of each association rule, and taking the association rules whose confidence information meets the preset confidence threshold condition as valid track data association rules. Thus, by performing trend analysis on the extreme value change directions of track data and trend analysis on the difference change of running gear data, the first analysis result and the second analysis result are obtained, deeply analyzing the change trends of the running states of the track and the running gear, timely discovering abnormal increases in track vibration or abnormal fluctuations in running gear data, providing a basis for predicting potential faults, integrating the above data according to the common dimension, breaking data barriers, and forming a comprehensive and systematic labeled data set. The problem of multi-source data fusion is solved, providing a rich data basis for mining multi-dimensional association relationships, and ensuring that the analysis results are more accurate and comprehensive. Using the layer-by-layer search and pruning technology to process the labeled data set can quickly and accurately obtain the frequent item set that meets the preset support threshold, greatly improving the mining efficiency, being able to accurately locate the key labeled data items, and reducing the consumption of computing resources. Finally, association rules are generated based on the non-empty subsets of the frequent item set, and the rules with confidence levels meeting the threshold are strictly calculated and screened to ensure that the obtained valid track data association rules are true and reliable, and the internal connections between track data and running gear data can be obtained.
[0150] Further, the embodiment of the present application also discloses an electronic device, Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation on the usage scope of the present application.
[0151] Figure 4 It is a structural schematic diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the multi-dimensional track data association method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0152] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed here; the input / output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0153] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one of the hardware forms of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0154] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0155] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, so as to implement the operation and processing of the massive data 223 in the memory 22 by the processor 21. It can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the multi-dimensional orbit data association method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks. The data 223 may include not only the data transmitted by the external device received by the electronic device, but also the data collected by its own input / output interface 25, etc.
[0156] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the multi-dimensional orbit data association method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0157] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts.
[0158] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application. The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, removable disk, CD-ROM (Compact Disc-Read Only Memory), or any other form of storage medium known in the technical field.
[0159] Finally, it should also be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0160] The above has introduced the solution provided by the present invention in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A multi-dimensional orbit data association method, characterized in that Including: Screening a multi-source data subset related to an orbit association target from a multi-source data set including orbit vibration impact data, running gear data, orbit design parameters, and orbit operation and maintenance data; Performing a trend analysis on the extreme value change direction of each orbit data in the multi-source data subset within a preset time range to obtain a first analysis result; Performing a trend analysis on the data difference change between each running gear data in the multi-source data subset to obtain a second analysis result; Integrating and labeling the first analysis result, the second analysis result, the orbit data and the running gear data in the multi-source data subset according to a common dimension to obtain a labeled data set; Performing a layer-by-layer search and pruning process on the labeled data set to obtain a frequent item set composed of labeled data items that meet the preset support threshold condition; Generating a number of association rules based on the information of each non-empty subset of the frequent item set, calculating the confidence information of each association rule, and taking the association rule whose confidence information meets the preset confidence threshold condition as an effective orbit data association rule.
2. The multi-dimensional track data association method according to claim 1, wherein After screening the multi-source data subset related to the orbit association target from the multi-source data set including orbit vibration impact data, running gear data, orbit design parameters, and orbit operation and maintenance data, it further includes: Deleting abnormal data in the multi-source data subset to obtain a cleaned multi-source data subset; Performing data filling processing on the cleaned multi-source data subset according to a preset data filling rule to obtain a preprocessed multi-source data subset; wherein, the preset data filling rule includes filling the numerical missing data in the cleaned multi-source data subset by using the median method and / or filling the discrete type missing data in the cleaned multi-source data subset by using the mode method; Correspondingly, performing a trend analysis on the extreme value change direction of each orbit data in the multi-source data subset within a preset time range to obtain a first analysis result; performing a trend analysis on the data difference change between each running gear data in the multi-source data subset to obtain a second analysis result, including: Performing a trend analysis on the extreme value change direction of each orbit data in the preprocessed multi-source data subset within a preset time range to obtain a first analysis result; Performing a trend analysis on the data difference change between each running gear data in the preprocessed multi-source data subset to obtain a second analysis result.
3. The multi-dimensional track data association method according to claim 1, wherein Performing a trend analysis on the extreme value change direction of each orbit data in the multi-source data subset within a preset time range to obtain a first analysis result, including: Decomposing the orbit vibration value of each orbit data to obtain a target vibration trend, a target seasonal component, and target residual information that change with time after decomposition, as the first analysis result.
4. The multi-dimensional track data association method according to claim 3, characterized in that The decomposing the orbit vibration value of each orbit data to obtain a target vibration trend, a target seasonal component, and target residual information that change with time after decomposition, as the first analysis result, includes: Setting the period length of the seasonal component according to the periodic characteristics of the orbit vibration value of each orbit data; Initialize the vibration trend and seasonal component as empty to set the track vibration value as residual information; Perform LOESS smoothing on the residual information according to a preset smoothing window size, and use the smoothed result as the current vibration trend; Remove the current vibration trend from the track vibration value to obtain detrended data; Segment the detrended data according to the cycle length to obtain detrended data segments, and perform LOESS smoothing on each detrended data segment to obtain a preliminary seasonal component; Perform in-cycle averaging on the preliminary seasonal component to obtain a corrected seasonal component; Remove the corrected seasonal component from the track vibration value to obtain deseasonalized data; Perform LOESS smoothing on the deseasonalized data according to a preset smoothing window size to update the current vibration trend, obtain a new current vibration trend, and jump to execute the step of removing the current vibration trend from the track vibration value until the change information of the vibration trend and the seasonal component is less than a preset change threshold, and output the current vibration trend and the corrected seasonal component as the target vibration trend and the target seasonal component; Use the track vibration value, the target vibration trend, and the target seasonal component to determine target residual information to obtain a first analysis result.
5. The multi-dimensional orbit data association method according to claim 4, characterized in that The determination of the target residual information using the track vibration value, the target vibration trend, and the target seasonal component includes: By determine the target residual information; Among them, represents the target vibration trend at the moment; represents the target seasonal component at the moment, represents the target residual information at the moment, represents the orbital vibration value at the moment.
6. The multi-dimensional track data association method according to claim 1, wherein The trend analysis of the data difference change between the running part data in the multi-source data subset to obtain a second analysis result includes: Pair the running part data to obtain several running part data pairs; Perform a difference operation on each running part data pair to obtain the difference results of each running part data pair; the difference results include positive difference results and negative difference results; Determine the second analysis result of the running part data based on the quantity information of the positive difference results and the quantity information of the negative difference results.
7. The multi-dimensional track data association method according to claim 6, characterized in that, The determination of the second analysis result of the running part data based on the quantity information of the positive difference results and the quantity information of the negative difference results includes: Use the sum of the quantity of the positive difference results and the quantity of the negative difference results as the total number of data pairs; Perform binomial distribution processing on the total number of data pairs, the success probability of the positive difference result / negative difference result of the running part data pair, and the number of successful times to obtain a success probability value, and calculate the significance level information based on the success probability value; Analyze the trend of the data difference change of the running part data according to the size result of the quantity information of the positive difference results and the quantity information of the negative difference results and the significance level information to obtain a second analysis result.
8. The multi-dimensional track data association method according to claim 7, wherein The analysis of the trend of the data difference change of the running part data according to the size result of the quantity information of the positive difference results and the quantity information of the negative difference results and the significance level information to obtain a second analysis result includes: When the number of positive difference results is greater than the number of negative difference results, and the significance level information is less than a preset significance level threshold, it is determined that the data difference change of the running gear data is an upward trend, so as to obtain a second analysis result with an upward trend analysis result; When the number of positive difference results is less than the number of negative difference results, and the significance level information is less than a preset significance level threshold, it is determined that the data difference change of the running gear data is a downward trend, so as to obtain a second analysis result with a downward trend analysis result.
9. The multi-dimensional track data association method according to claim 1, characterized in that The integration and marking process of the first analysis result, the second analysis result, the track data and the running gear data in the multi-source data subset according to the common dimension to obtain a marked data set includes: Integrating the track data and the running gear data according to the common dimension to obtain a process integration data set; Inserting the first analysis result and the second analysis result into the process integration data set in the form of fields to obtain an integrated data set; Performing interval division processing on the integrated data of the numerical type in the integrated data set, and performing discrete marking on the integrated data after the division interval according to the discrete marking method to obtain a marked data set.
10. The multi-dimensional track data association method according to claim 1, wherein The layer-by-layer search and pruning process of the marked data set to obtain a frequent item set composed of marked data items that meet the preset support threshold condition includes: Scanning the marked data set to count the occurrence times information of each marked data item in the marked data set; Calculating the support of the occurrence times information of each marked data item to the number of data records in the marked data set; Pruning the marked data items with support less than the preset support threshold condition to obtain updated marked data items; Connecting the updated marked data items in each data record according to the preset connection rule, and using the connected marked data items as new marked data items, and then jumping to execute the step of counting the occurrence times information of each marked data item in the marked data set until the pruning ends, and outputting the frequent item set.
11. A multi-dimensional orbit data association device, characterized in that including: The multi-dimensional track data association device is specifically used to screen a multi-source data subset related to the track association target from a multi-source data set including track vibration and impact data, running gear data, track design parameters, and track operation and maintenance data; The first trend analysis module is used to perform trend analysis on the extreme value change direction of each track data in the multi-source data subset within a preset time range to obtain a first analysis result; The second trend analysis module is used to perform trend analysis on the data difference change between each running gear data in the multi-source data subset to obtain a second analysis result; The integration and marking module is used to integrate and mark the first analysis result, the second analysis result, the track data and the running gear data in the multi-source data subset according to the common dimension to obtain a marked data set; A search processing module, configured to perform layer-by-layer search and pruning processing on the marked dataset to obtain a frequent item set composed of marked data items that meet the preset support threshold condition; An association generation module, configured to generate a number of association rules based on the information of each non-empty subset of the frequent item set, calculate the confidence information of each association rule, and use the association rules whose confidence information meets the preset confidence threshold condition as valid track data association rules.
12. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the steps of the multi-dimensional track data association method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, the steps of the multi-dimensional track data association method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Method and system for association rule analysis of moving object
CN104598566A
Collaborative optimization method for bus timetable based on big data
US20200027347A1