A railway infrastructure multi-source data fusion system and method

By performing distributed fusion and progressive processing of multi-source data in railway infrastructure, combined with time feature analysis and fusion rules determination, the problems of lag and inaccuracy of multi-source data fusion are solved, and the real-time and accuracy of data are improved.

CN119475223BActive Publication Date: 2025-05-09EAST CHINA JIAOTONG UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411512009.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-05-09
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

The direct fusion of multi-source data in the railway system has lag and inaccuracy, which affects the maintenance and management decisions of railway infrastructure.

Method used

A multi-source data fusion method for railway infrastructure is adopted. By acquiring and preprocessing multiple data sources, analyzing time characteristics to determine fusion rules, conducting progressive data fusion, and performing effectiveness evaluation and bias adjustment of fusion nodes.

Benefits of technology

It improves the real-time and accuracy of data fusion, ensures the scientificity and rationality of data, promptly discovers and deals with abnormal fusion nodes, and ensures the reliability and continuity of data fusion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119475223B_ABST
    Figure CN119475223B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of railway multi-source data fusion technology, specifically relating to a multi-source data fusion system and method for railway infrastructure. By analyzing the temporal characteristics of various benchmark data sources, this invention can accurately identify data sources with both centralized and decentralized characteristics, and formulate corresponding fusion rules accordingly. This not only ensures the scientific and rational nature of the data fusion process but also enables the timely detection and handling of abnormal fusion nodes, thereby guaranteeing the reliability of the data fusion results. Furthermore, this invention can dynamically adjust abnormal fusion nodes based on fusion deviation, ensuring the continuity and stability of the data fusion process. Through real-time monitoring and backtracking analysis, deviations in the data fusion process can be detected and corrected in a timely manner, thus avoiding decision-making errors caused by data fusion mistakes and providing accurate and timely data support for the maintenance and management of railway infrastructure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of railway multi-source data fusion, and in particular relates to a railway infrastructure multi-source data fusion system and method. Background Art

[0002] With the rapid development of railway infrastructure, a large amount of data resources have accumulated in the railway system. These data resources come from different data sources, such as sensors, monitoring systems, maintenance records, etc. In order to improve the efficiency and safety of railway operations, it is necessary to effectively integrate these multi-source data to obtain more accurate and comprehensive information, thereby supporting the maintenance and management decisions of railway infrastructure.

[0003] However, due to the diversity and complexity of data sources, directly fusing these data will face many challenges. For example, the collection frequency of different data sources may be different, and data fusion is mostly carried out periodically, which may cause lags and inaccuracies in the fused data, thereby affecting the accuracy of maintenance and management decisions of railway infrastructure. Based on this, this solution provides a railway infrastructure multi-source data fusion method to solve the above problems. Summary of the invention

[0004] The purpose of the present invention is to provide a railway infrastructure multi-source data fusion system and method, which can perform distributed fusion of railway infrastructure data, make the data fusion process progressive, thereby improving the real-time and accuracy of data fusion.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A railway infrastructure multi-source data fusion method, comprising:

[0007] Obtain various data sources of railway infrastructure and pre-process them to obtain benchmark data sources;

[0008] Collecting the time characteristics of each of the reference data sources, and determining the fusion rules of each of the reference data sources according to the time characteristics;

[0009] Performing fusion processing on each reference data source according to the fusion rule to obtain multiple groups of fusion data and fusion nodes of each group of the fusion data;

[0010] Performing effectiveness evaluation on each of the fusion nodes to obtain a fusion state of each fusion node, wherein the fusion state includes a normal fusion state and an abnormal fusion state;

[0011] In the normal fusion state, the reference data sources under each fusion node are fused progressively;

[0012] In the abnormal fusion state, the abnormal fusion node is located, and the fusion deviation of the reference data source under the abnormal fusion node is counted;

[0013] The abnormal fusion node is adjusted according to the fusion deviation until all the reference data sources are progressively fused and then stopped.

[0014] In a preferred embodiment, the step of obtaining multiple data sources of railway infrastructure and preprocessing them to obtain a reference data source includes:

[0015] Collecting operation data, maintenance data and environmental monitoring data of railway facilities, and summarizing them into independent data sets, and adding a unique data source identifier to each independent data set;

[0016] Cleaning the data in each of the independent data sets to remove invalid and erroneous data items;

[0017] The data in the cleaned independent data set is standardized, the data format in the independent data set is unified, and the independent data set with unified format is output as a reference data source.

[0018] In a preferred solution, the step of collecting the time characteristics of each of the reference data sources includes:

[0019] Obtaining the data collection nodes under each of the reference data sources, and arranging them in order of occurrence;

[0020] Collecting the time intervals between adjacent data collection nodes and marking them as evaluation condition parameters;

[0021] Acquire an allowable fluctuation range, and perform offset processing on each evaluation condition parameter according to the allowable fluctuation range to obtain multiple evaluation ranges;

[0022] Counting the number of evaluation condition parameters in each of the evaluation intervals and recording them as classification condition parameters, and determining the time characteristics of each benchmark data source respectively according to the classification condition parameters;

[0023] The time characteristics of the reference data source include a dispersion characteristic and a concentration characteristic, and the stability of the dispersion characteristic is lower than the stability of the concentration characteristic.

[0024] In a preferred solution, the step of determining the time characteristics of each reference data source according to the classification condition parameters includes:

[0025] Obtaining classification condition parameters under each of the reference data sources;

[0026] Arrange the classification condition parameters under the same reference data source in descending order, and calculate the proportion of the classification condition parameter with the highest ranking;

[0027] Obtaining a classification threshold, and comparing the classification threshold with the proportion of the classification condition parameter with the highest ranking;

[0028] If the proportion of the highest ranking classification condition parameter is higher than the classification threshold, it indicates that the reference data source corresponding to the classification condition parameter has a centralized feature;

[0029] If the proportion of the highest ranking classification condition parameter is lower than or equal to the classification threshold, it indicates that the reference data source corresponding to the classification condition parameter has a dispersion characteristic.

[0030] In a preferred solution, the step of determining the fusion rule of each reference data source according to the time feature includes:

[0031] Acquire the time characteristics of each of the reference data sources;

[0032] The maximum value of the evaluation interval corresponding to the central feature is used as the data fusion interval;

[0033] Obtaining data fusion intervals of the reference data source under all the centralized features, and arranging them in order from low to high to obtain a data fusion order;

[0034] The reference data sources are fused according to the fusion order, and after the fusion of the reference data sources under the centralized feature is completed, the reference data sources under the decentralized feature are supplemented and fused.

[0035] In a preferred solution, the step of evaluating the effectiveness of each fusion node includes:

[0036] Collecting the fusion results of the reference data sources under each of the fusion nodes, and performing vectorization conversion to obtain a plurality of vectors to be evaluated;

[0037] Obtaining an evaluation function, inputting the vector to be evaluated into the evaluation function, and recording the output result of the evaluation function as a parameter to be evaluated;

[0038] Obtaining an evaluation threshold, and comparing the evaluation threshold with the parameter to be evaluated;

[0039] If the parameter to be evaluated is less than the evaluation threshold, it indicates that the fusion node is in a normal fusion state;

[0040] If the parameter to be evaluated is greater than or equal to the evaluation threshold, it indicates that the fusion node is in an abnormal fusion state.

[0041] In a preferred solution, the step of counting the fusion deviation of the reference data source under the abnormal fusion node includes:

[0042] Obtaining an actual update node of a reference data source for which data fusion is not performed under the abnormal fusion node;

[0043] Calculating the time difference between the actual update node and the abnormal fusion node, and recording it as a real-time fusion deviation;

[0044] Backtracking and offsetting the abnormal fusion node to obtain the historical fusion deviation under the corresponding historical node;

[0045] Acquire the allowable deviation under the abnormal fusion node, and stop backtracking the abnormal fusion node when the historical fusion deviation is less than or equal to the allowable deviation;

[0046] A calculation function is obtained, and the real-time fusion deviation and the historical fusion deviation are input into the calculation function, and the output result of the calculation function is recorded as the fusion deviation.

[0047] In a preferred solution, the step of adjusting the abnormal fusion node according to the fusion deviation includes:

[0048] Collect the backtracking time of the abnormal fusion node and record it as a parameter to be evaluated;

[0049] Obtaining an evaluation threshold, and comparing the parameter to be evaluated with the evaluation threshold;

[0050] When the parameter to be evaluated is greater than or equal to the evaluation threshold, the abnormal fusion node is offset according to the fusion deviation to obtain an updated fusion node, and data fusion is performed under the updated fusion node;

[0051] When the parameter to be evaluated is less than the evaluation threshold, the fusion deviation is continuously collected until the parameter to be evaluated reaches or exceeds the evaluation threshold, and then the abnormal fusion node is adjusted.

[0052] The present invention also provides a railway infrastructure multi-source data fusion system, using the above-mentioned railway infrastructure multi-source data fusion method, comprising:

[0053] A data acquisition module, which is used to acquire multiple data sources of railway infrastructure and perform preprocessing to obtain a reference data source;

[0054] A rule output module, the rule output module is used to collect the time characteristics of each of the reference data sources, and determine the fusion rules of each of the reference data sources according to the time characteristics;

[0055] A data fusion module, the data fusion module is used to perform fusion processing on each reference data source according to the fusion rule to obtain multiple groups of fusion data and fusion nodes of each group of the fusion data;

[0056] A state evaluation module, the state evaluation module is used to evaluate the effectiveness of each fusion node to obtain the fusion state of each fusion node, wherein the fusion state includes a normal fusion state and an abnormal fusion state;

[0057] In the normal fusion state, the reference data sources under each fusion node are fused progressively;

[0058] In the abnormal fusion state, the abnormal fusion node is located, and the fusion deviation of the reference data source under the abnormal fusion node is counted;

[0059] A fusion optimization module is used to adjust the abnormal fusion node according to the fusion deviation until all the reference data sources are progressively fused and then stop.

[0060] And, an electronic device, the electronic device comprising:

[0061] at least one processor;

[0062] and a memory communicatively coupled to the at least one processor;

[0063] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned railway infrastructure multi-source data fusion method.

[0064] The technical effects achieved by the present invention are:

[0065] The present invention improves the accuracy and efficiency of data processing by effectively fusing multi-source data of railway infrastructure. By analyzing the time characteristics of each benchmark data source, it can accurately identify data sources with centralized and decentralized characteristics, and formulate corresponding fusion rules accordingly, which not only ensures the scientificity and rationality of the data fusion process, but also can timely discover and process abnormal fusion nodes, thereby ensuring the reliability of data fusion results. In addition, the present invention can also dynamically adjust abnormal fusion nodes according to the fusion deviation amount, ensuring the continuity and stability of the data fusion process. Through real-time monitoring and backtracking analysis, it can timely discover and correct deviations in the data fusion process, thereby avoiding decision-making errors caused by data fusion errors, providing more accurate and timely data support for the maintenance and management of railway infrastructure, and improving the safety and efficiency of railway operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a schematic flow chart of the method of the present invention;

[0067] Figure 2 It is a schematic diagram of the system module of the present invention;

[0068] Figure 3 It is a schematic diagram of the structure of an electronic device of the present invention. DETAILED DESCRIPTION

[0069] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0070] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0071] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure or characteristic that may be included in at least one implementation of the present invention. The phrase "in a preferred embodiment" that appears in different places in this specification does not refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments.

[0072] See also Figure 1 As shown, the present invention provides a railway infrastructure multi-source data fusion method, comprising:

[0073] S1. Obtain multiple data sources of railway infrastructure and perform preprocessing to obtain a benchmark data source;

[0074] S2, collecting the time characteristics of each benchmark data source, and determining the fusion rules of each benchmark data source according to the time characteristics;

[0075] S3, fusing each reference data source according to the fusion rule to obtain multiple groups of fused data and fusion nodes of each group of fused data;

[0076] S4, evaluating the effectiveness of each fusion node to obtain the fusion state of each fusion node, wherein the fusion state includes a normal fusion state and an abnormal fusion state;

[0077] In the normal fusion state, the reference data sources under each fusion node are progressively fused;

[0078] In the abnormal fusion state, locate the abnormal fusion node and count the fusion deviation of the benchmark data source under the abnormal fusion node;

[0079] S5. According to the fusion deviation, the abnormal fusion node is adjusted until all the reference data sources are progressively fused and then stopped.

[0080] As described in the above steps S1-S5, with the development of information technology, the maintenance and management of railway infrastructure has become increasingly dependent on the fusion analysis of multi-source data. In order to ensure the efficient operation of the railway system, real-time monitoring and data fusion of railway infrastructure are particularly important. In this embodiment, it is first necessary to obtain multiple data sources of railway infrastructure, including but not limited to sensor data, maintenance records, historical accident reports, etc. After obtaining these data, they need to be pre-processed synchronously to ensure the quality and consistency of the data. The pre-processing steps may include data cleaning, data standardization and data normalization, etc. Through these steps, a benchmark data source can be obtained as the basis for subsequent fusion processing, and then the time characteristics of each benchmark data source need to be collected. By analyzing these time characteristics, the inherent laws of the data can be better understood, and the fusion rules of each benchmark data source can be determined accordingly. The fusion rule refers to how to process and integrate information from different data sources in the data fusion process to achieve the best fusion effect. According to the determined fusion rule, each benchmark data source will be fused and processed, and then multiple groups of fused data can be obtained. Each group of fused data corresponds to a fusion node. The fusion node refers to the different data sources in the data fusion process. The key point where information needs to be intersected and integrated is that in order to ensure the quality and reliability of the fused data, it is necessary to evaluate the effectiveness of each fusion node. The purpose of the effectiveness evaluation is to determine whether the fusion node is in a normal state, that is, whether the data is fused in the expected way. The fusion state can be divided into a normal fusion state and an abnormal fusion state. In the normal fusion state, the benchmark data sources under each fusion node should present a progressive fusion relationship, that is, the data is gradually integrated during the fusion process, and finally high-quality fusion data is formed. However, in actual applications, abnormal fusion states may occur. In this case, it is necessary to locate the abnormal fusion node and count the fusion deviation of the benchmark data source under the abnormal fusion node. The fusion deviation refers to the difference between the actual fusion result and the expected fusion result. Finally, according to the fusion deviation, the abnormal fusion node needs to be adjusted. The adjustment measures may include modifying the fusion rules, optimizing the data preprocessing steps, adjusting the data fusion algorithm, etc. The ultimate goal is to make all benchmark data sources present a progressive fusion relationship during the fusion process through these adjustments. When all benchmark data sources have achieved progressive fusion, the adjustment process can be stopped to ensure the accuracy of railway infrastructure data fusion.

[0081] In a preferred embodiment, the steps of obtaining multiple data sources of railway infrastructure and preprocessing them to obtain a reference data source include:

[0082] S101. Collect operation data, maintenance data and environmental monitoring data of railway facilities, and summarize them into independent data sets, and add a unique data source identifier to each independent data set;

[0083] S102, cleaning the data in each independent data set to remove invalid and erroneous data items;

[0084] S103: Standardize the data in the cleaned independent data set, unify the data format in the independent data set, and output the independent data set with unified format as a reference data source.

[0085] As described in the above steps S101-S103, when determining the benchmark data source, it is first necessary to collect relevant data of railway infrastructure from multiple channels, including but not limited to operation data, maintenance data and environmental monitoring data of railway facilities, and summarize these data into independent data sets. In order to facilitate management and tracking, a unique data source identifier needs to be added to each independent data set, so that it can be clearly known which data source each data comes from, which is convenient for subsequent data processing and analysis. Then, the data in each independent data set is cleaned. In this process, invalid and erroneous data items will be eliminated. Invalid data may include data with incorrect format, missing key information or beyond a reasonable range. Erroneous data may be data distortion caused by sensor failure, human input error or other reasons. Through the cleaning process, the quality of the data set can be ensured, providing a reliable basis for subsequent data analysis and processing. After the cleaning is completed, the data in the independent data set will be standardized. The purpose of standardization is to unify the data format in each independent data set to ensure the consistency and comparability of the data. This may include unifying the timestamp format, unit, encoding method, etc. of the data. Through standardization, the independent data set with unified format can be output as a benchmark data source for further analysis and application.

[0086] In a preferred embodiment, the step of collecting the time characteristics of each reference data source includes:

[0087] S201, obtaining data collection nodes under each benchmark data source, and arranging them according to the occurrence time sequence;

[0088] S202, collecting the time intervals between adjacent data collection nodes and calibrating them as evaluation condition parameters;

[0089] S203, obtaining an allowable fluctuation range, and performing offset processing on each evaluation condition parameter according to the allowable fluctuation range to obtain multiple evaluation ranges;

[0090] S204, counting the number of evaluation condition parameters in each evaluation interval and recording them as classification condition parameters, and determining the time characteristics of each benchmark data source respectively according to the classification condition parameters;

[0091] Among them, the temporal characteristics of the benchmark data source include dispersion characteristics and concentration characteristics, and the stability of the dispersion characteristics is lower than the stability of the concentration characteristics.

[0092] As described in the above steps S201-S204, in order to determine the time characteristics of each benchmark data source, it is first necessary to obtain the data acquisition nodes under each benchmark data source and arrange them in the order of their occurrence time to ensure that the temporal relationship between the data acquisition nodes can be clearly defined. Then, the time intervals between adjacent data acquisition nodes are collected and recorded as evaluation condition parameters for subsequent analysis and comparison. Then, an allowable fluctuation interval is introduced. The allowable fluctuation interval defines the normal fluctuation range of the time interval between data acquisition nodes. According to the allowable fluctuation interval, each evaluation condition parameter is offset to obtain multiple evaluation intervals. The evaluation interval can help understand the fluctuation of the time interval between data acquisition nodes. Then, the number of evaluation condition parameters in each evaluation interval is counted, and the number of evaluation condition parameters is recorded as a classification condition parameter, so that the time characteristics of each benchmark data source can be accurately determined. It is necessary to clarify that the time characteristics of the benchmark data source include two main aspects: dispersion characteristics and concentration characteristics. The stability of dispersion characteristics is lower than that of concentration characteristics. In practical applications, the processing priority of concentration characteristics is higher than that of dispersion characteristics.

[0093] In a preferred embodiment, the step of determining the time characteristics of each reference data source respectively according to the classification condition parameter includes:

[0094] Obtain classification condition parameters under each benchmark data source;

[0095] Arrange the classification condition parameters under the same benchmark data source in descending order, and calculate the proportion of the classification condition parameter with the highest ranking;

[0096] Obtaining a classification threshold, and comparing the classification threshold with the proportion of the classification condition parameter with the highest ranking;

[0097] If the proportion of the highest ranking classification condition parameter is higher than the classification threshold, it indicates that the benchmark data source corresponding to the classification condition parameter has a centralized feature;

[0098] If the proportion of the highest ranked classification condition parameter is lower than or equal to the classification threshold, it indicates that the benchmark data source corresponding to the classification condition parameter has a dispersed characteristic.

[0099] In this implementation, when determining the category of the time feature of each reference data source, it is first necessary to obtain the classification condition parameters under each reference data source, and arrange the classification condition parameters under the same reference data source in order from large to small. Through this sorting, the importance distribution of each parameter in the data source can be intuitively seen. On this basis, the proportion of the classification condition parameter with the highest ranking is also calculated and recorded. The proportion reflects the relative importance of the parameter in the entire data source. Then, a classification threshold is introduced. The classification threshold is a pre-set standard for judging whether the time feature of the reference data source is centralized or decentralized. The classification threshold is compared with the proportion of the classification condition parameter with the highest ranking. If the proportion of the classification condition parameter with the highest ranking is higher than the classification threshold, then it is determined that the reference data source corresponding to the classification condition parameter has a centralized feature, which means that in the data source, some parameters occupy a dominant position and have a greater impact on the overall data. On the contrary, if the proportion of the classification condition parameter with the highest ranking is lower than or equal to the classification threshold, then it is considered that the reference data source corresponding to the classification condition parameter has a decentralized feature. In this case, the importance of each parameter in the data source is relatively balanced, and there is no obvious dominant factor.

[0100] In a preferred embodiment, the step of determining the fusion rules of each reference data source according to the time feature includes:

[0101] S205, obtaining time characteristics of each reference data source;

[0102] S206, taking the maximum value of the evaluation interval corresponding to the centrality feature as the data fusion interval;

[0103] S207, obtaining data fusion intervals of the reference data source under all centralized features, and arranging them in order from low to high to obtain a data fusion order;

[0104] S208, fusing each reference data source according to the fusing order, and after the fusing of the reference data sources under the centralized feature is completed, supplementing and fusing the reference data sources under the decentralized feature.

[0105] As described in the above steps S205-S208, in order to ensure that each benchmark data source can be effectively fused according to its time characteristics, it is first necessary to obtain the time characteristics of each benchmark data source, and first use the maximum value of the evaluation interval corresponding to each centralized feature as the interval for data fusion. By selecting the maximum value of the evaluation interval as the fusion interval, it can be ensured that the data fusion can be more precise and accurate, thereby increasing the corresponding fault tolerance. Then, it is necessary to obtain the data fusion intervals of the benchmark data sources under all centralized features, and arrange them in order from low to high to obtain a data fusion order. According to this order, each benchmark data source is fused in turn to ensure the sequentiality and logic of data fusion. Finally, according to the fusion order, each benchmark data source is fused. After the fusion of the benchmark data sources under the centralized feature is completed, it is necessary to supplement the fusion of the benchmark data sources under the decentralized feature to ensure the comprehensiveness and completeness of the data fusion.

[0106] In a preferred embodiment, the step of evaluating the effectiveness of each fusion node includes:

[0107] S401, collecting the fusion results of the benchmark data sources under each fusion node, and performing vectorization conversion to obtain multiple vectors to be evaluated;

[0108] S402, obtaining an evaluation function, inputting the vector to be evaluated into the evaluation function, and recording the output result of the evaluation function as a parameter to be evaluated;

[0109] S403, obtaining an evaluation threshold, and comparing the evaluation threshold with the parameter to be evaluated;

[0110] If the parameter to be evaluated is less than the evaluation threshold, it indicates that the fusion node is in a normal fusion state;

[0111] If the parameter to be evaluated is greater than or equal to the evaluation threshold, it indicates that the fusion node is in an abnormal fusion state.

[0112] As described in the above steps S401-S403, in order to evaluate the effectiveness of each fusion node, first, it is necessary to collect the fusion results of the benchmark data source under each fusion node, and record and organize the fusion results in detail to ensure the integrity and accuracy of the data. Next, the fusion results will be vectorized to convert the original data into multiple numerical vectors to be evaluated. Secondly, it is necessary to introduce a preset evaluation function, and the expression of the evaluation function is: In the formula, R represents the parameter to be evaluated, n represents the number of dimensional features of the vector to be evaluated, and a i represents the vector to be evaluated, b iRepresents the standard vector. The vector to be evaluated is input into the evaluation function to obtain the parameter to be evaluated. Then the evaluation threshold is introduced. The evaluation threshold is the key reference value used to judge whether the fusion node is in a normal state during the evaluation process. The evaluation threshold is compared with the parameter to be evaluated in detail, and a judgment is made based on the comparison result. If the parameter to be evaluated is greater than the evaluation threshold, then the fusion node is in a normal fusion state. On the contrary, if the parameter to be evaluated is less than or equal to the evaluation threshold, then the fusion node is judged to be in an abnormal fusion state.

[0113] In a preferred embodiment, the step of calculating the fusion deviation of the reference data source under the abnormal fusion node includes:

[0114] S404, obtaining the actual update node of the reference data source for which data fusion is not performed under the abnormal fusion node;

[0115] S405, calculating the time difference between the actual update node and the abnormal fusion node, and recording it as the real-time fusion deviation;

[0116] S406, backtracking and offsetting the abnormal fusion node to obtain the historical fusion deviation under the corresponding historical node;

[0117] S407, obtaining the allowable deviation amount under the abnormal fusion node, and stopping the backtracking of the abnormal fusion node when the historical fusion deviation amount is less than or equal to the allowable deviation amount;

[0118] S408. Obtain a calculation function, input the real-time fusion deviation and the historical fusion deviation into the calculation function, and record the output result of the calculation function as the fusion deviation.

[0119] As described in the above steps S404-S408, when counting the fusion deviation of the benchmark data source under the abnormal fusion node, first, it is necessary to obtain the actual update node of the benchmark data source under the abnormal fusion node that has not yet performed the data fusion operation, and synchronously calculate the time difference between the actual update node and the abnormal fusion node. This time difference reflects the time interval between data update and data fusion. This embodiment records the time difference as a real-time fusion deviation, and then backtracks the abnormal fusion node to find the historical fusion deviation under its corresponding historical node. In this way, it is possible to understand whether there are similar abnormal situations in the historical data, and The historical fusion deviation can be compared with the real-time fusion deviation to evaluate the persistence and impact range of the anomaly. After obtaining the historical fusion deviation, it is necessary to obtain the allowable deviation under the abnormal fusion node. The allowable deviation is a preset threshold used to determine whether the current deviation is within an acceptable range. If the historical fusion deviation is less than or equal to the allowable deviation, then stop further backtracking of the abnormal fusion node, because this indicates that the abnormal situation has been controlled within an acceptable range. Finally, it is necessary to introduce a measurement function, take the real-time fusion deviation and the historical fusion deviation as input parameters, and output the comprehensive fusion deviation. The expression of the measurement function is: In the formula, f p represents the fusion deviation, m represents the number of fusion nodes in the backtracking process, S j It represents the historical fusion deviation and the real-time fusion deviation. After the fusion deviation is output, the abnormal fusion node can be adjusted accordingly.

[0120] In a preferred embodiment, the step of adjusting the abnormal fusion node according to the fusion deviation includes:

[0121] S501. Collect the backtracking time of the abnormal fusion node and record it as a parameter to be evaluated;

[0122] S502, obtaining an evaluation threshold, and comparing the parameter to be evaluated with the evaluation threshold;

[0123] When the parameter to be evaluated is greater than or equal to the evaluation threshold, the abnormal fusion node is offset according to the fusion deviation to obtain an updated fusion node, and data fusion is performed under the updated fusion node;

[0124] When the parameter to be evaluated is less than the evaluation threshold, the fusion deviation is continuously collected until the parameter to be evaluated reaches or exceeds the evaluation threshold, and then the abnormal fusion node is adjusted.

[0125] In this implementation, in order to ensure the stability and accuracy of data fusion, it is necessary to make detailed adjustments to the abnormal fusion nodes. First, the backtracking time of the abnormal fusion nodes is collected. The backtracking time will be used as a pre-parameter for our subsequent evaluation and adjustment. This implementation records it as a parameter to be evaluated. Then, a pre-set evaluation threshold is introduced, and the parameter to be evaluated is compared in detail with the evaluation threshold. If it is found after comparison that the parameter to be evaluated is greater than or equal to the evaluation threshold, it means that the deviation of the abnormal fusion node has exceeded the set tolerance range. In this case, it is necessary to make corresponding offset adjustments to the abnormal fusion node based on the fusion deviation. Through this adjustment, an updated fusion node can be obtained. Under the updated fusion node, the data fusion operation is performed to ensure the accuracy and consistency of the data. However, if it is found after comparison that the collected backtracking time is less than the evaluation threshold, it means that the deviation of the abnormal fusion node is still within the tolerance range, which may be an occasional abnormality and does not need to be adjusted temporarily. In this case, the fusion deviation will continue to be collected and the status of the abnormal fusion node will be continuously monitored. When the parameter to be evaluated reaches or exceeds the evaluation threshold, the abnormal fusion node will be adjusted accordingly.

[0126] See also Figure 2 A railway infrastructure multi-source data fusion system, using the railway infrastructure multi-source data fusion method, comprises:

[0127] A data acquisition module is used to acquire various data sources of railway infrastructure and perform preprocessing to obtain a benchmark data source;

[0128] A rule output module, which is used to collect the time characteristics of each reference data source and determine the fusion rules of each reference data source according to the time characteristics;

[0129] The data fusion module is used to fuse various reference data sources according to the fusion rules to obtain multiple groups of fused data and fusion nodes of each group of fused data;

[0130] A state evaluation module, which is used to evaluate the effectiveness of each fusion node and obtain the fusion state of each fusion node, wherein the fusion state includes a normal fusion state and an abnormal fusion state;

[0131] In the normal fusion state, the reference data sources under each fusion node are progressively fused;

[0132] In the abnormal fusion state, locate the abnormal fusion node and count the fusion deviation of the benchmark data source under the abnormal fusion node;

[0133] The fusion optimization module is used to adjust the abnormal fusion nodes according to the fusion deviation until all the benchmark data sources are progressively fused and then stop.

[0134] In the above, the system includes a data acquisition module, a rule output module, a data fusion module, a state assessment module and a fusion optimization module. The main responsibility of the data acquisition module is to obtain various types of data sources from railway infrastructure. These data sources may include but are not limited to sensor data, monitoring videos, maintenance records, meteorological information, etc. The acquired data sources undergo preliminary preprocessing to ensure the quality and consistency of the data, thereby generating a benchmark data source to provide a reliable basis for subsequent data processing. The core function of the rule output module is to collect the time characteristics of each benchmark data source. Through the analysis of the time characteristics, the system can determine the fusion rules between each benchmark data source and guide the data fusion module on how to effectively integrate the information of different data sources at the corresponding time points to ensure the accuracy and efficiency of data fusion. The data fusion module fuses each benchmark data source according to the fusion rules to generate multiple sets of fused data, and can Identify the fusion nodes of each group of fused data. These fusion nodes are the key points in the data fusion process. They represent the intersection and integration of information from different data sources. The main task of the status assessment module is to evaluate the effectiveness of each fusion node. Through the evaluation, the system can determine the fusion state of each fusion node. These states can be divided into normal fusion state and abnormal fusion state. In the normal fusion state, the benchmark data sources under each fusion node show a progressive fusion trend, indicating that the data fusion process is proceeding smoothly. In the abnormal fusion state, the system can quickly locate the abnormal fusion node and count the fusion deviation of the benchmark data source under the node for further analysis and processing. The role of the fusion optimization module is to adjust and optimize the abnormal fusion node according to the fusion deviation, and gradually reduce the fusion deviation until all benchmark data sources can show a progressive fusion trend, providing strong data support for railway operation and maintenance.

[0135] See also Figure 3 , an electronic device, the electronic device comprising:

[0136] at least one processor;

[0137] and a memory communicatively coupled to the at least one processor;

[0138] The memory stores a computer program executable by at least one processor, and the computer program is executed by at least one processor so that the at least one processor can execute the above-mentioned railway infrastructure multi-source data fusion method.

[0139] The processors of the above electronic devices can be of various types, such as a central processing unit (CPU), a graphics processing unit (GPU) or an application-specific integrated circuit (ASIC). These processors can efficiently perform complex computing tasks, and the memory may include a random access memory (RAM), a read-only memory (ROM), a solid-state drive (SSD) or other types of storage devices, and also include input and output devices and an arithmetic unit. The input and output devices include a keyboard, a mouse, a touch screen, a display, etc., which are used to interact with the user, and the arithmetic unit is responsible for performing various mathematical and logical operations.

[0140] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0141] The above is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications should also be considered as the protection scope of the present invention. The structures, devices and operating methods not specifically described and explained in the present invention shall be implemented according to the conventional means in the art unless otherwise specified and limited.

Claims

1. A railway infrastructure multi-source data fusion method, characterized by: include: Obtain various data sources of railway infrastructure and pre-process them to obtain benchmark data sources; Collecting the time characteristics of each of the reference data sources, and determining the fusion rules of each of the reference data sources according to the time characteristics; Performing fusion processing on each reference data source according to the fusion rule to obtain multiple groups of fusion data and fusion nodes of each group of the fusion data; Performing effectiveness evaluation on each of the fusion nodes to obtain a fusion state of each fusion node, wherein the fusion state includes a normal fusion state and an abnormal fusion state; In the normal fusion state, the reference data sources under each fusion node are fused progressively; In the abnormal fusion state, the abnormal fusion node is located, and the fusion deviation of the reference data source under the abnormal fusion node is counted; According to the fusion deviation, the abnormal fusion node is adjusted until all the reference data sources are progressively fused and then stopped; The step of obtaining multiple data sources of railway infrastructure and preprocessing them to obtain a reference data source includes: Collecting operation data, maintenance data and environmental monitoring data of railway facilities, and summarizing them into independent data sets, and adding a unique data source identifier to each independent data set; Cleaning the data in each of the independent data sets to remove invalid and erroneous data items; Standardizing the data in the cleaned independent data set, unifying the data format in the independent data set, and outputting the independent data set with the unified format as a reference data source; The step of determining the fusion rule of each reference data source according to the time feature includes: Acquire the time characteristics of each of the reference data sources; The maximum value of the evaluation interval corresponding to the central feature is used as the data fusion interval; Obtaining data fusion intervals of the reference data source under all the centralized features, and arranging them in order from low to high to obtain a data fusion order; The reference data sources are fused according to the fusion order, and after the fusion of the reference data sources under the centralized feature is completed, the reference data sources under the decentralized feature are supplemented and fused; The step of evaluating the effectiveness of each fusion node includes: Collecting the fusion results of the reference data sources under each of the fusion nodes, and performing vectorization conversion to obtain a plurality of vectors to be evaluated; Obtaining an evaluation function, inputting the vector to be evaluated into the evaluation function, and recording the output result of the evaluation function as a parameter to be evaluated; Obtaining an evaluation threshold, and comparing the evaluation threshold with the parameter to be evaluated; If the parameter to be evaluated is less than the evaluation threshold, it indicates that the fusion node is in a normal fusion state; If the parameter to be evaluated is greater than or equal to the evaluation threshold, it indicates that the fusion node is in an abnormal fusion state.

2. The railway infrastructure multi-source data fusion method according to claim 1 is characterized by: The step of collecting the time characteristics of each of the reference data sources includes: Obtaining the data collection nodes under each of the reference data sources, and arranging them in order of occurrence; Collecting the time intervals between adjacent data collection nodes and marking them as evaluation condition parameters; Acquire an allowable fluctuation range, and perform offset processing on each evaluation condition parameter according to the allowable fluctuation range to obtain multiple evaluation ranges; Counting the number of evaluation condition parameters in each of the evaluation intervals and recording them as classification condition parameters, and determining the time characteristics of each benchmark data source respectively according to the classification condition parameters; The time characteristics of the reference data source include a dispersion characteristic and a concentration characteristic, and the stability of the dispersion characteristic is lower than the stability of the concentration characteristic.

3. The railway infrastructure multi-source data fusion method according to claim 2 is characterized by: The step of respectively determining the time characteristics of each reference data source according to the classification condition parameters comprises: Obtaining classification condition parameters under each of the reference data sources; Arrange the classification condition parameters under the same reference data source in descending order, and calculate the proportion of the classification condition parameter with the highest ranking; Obtaining a classification threshold, and comparing the classification threshold with the proportion of the classification condition parameter with the highest ranking; If the proportion of the highest ranking classification condition parameter is higher than the classification threshold, it indicates that the reference data source corresponding to the classification condition parameter has a centralized feature; If the proportion of the highest ranking classification condition parameter is lower than or equal to the classification threshold, it indicates that the reference data source corresponding to the classification condition parameter has a dispersion characteristic.

4. The railway infrastructure multi-source data fusion method according to claim 1 is characterized by: The step of counting the fusion deviation of the reference data source under the abnormal fusion node includes: Obtaining an actual update node of a reference data source for which data fusion is not performed under the abnormal fusion node; Calculating the time difference between the actual update node and the abnormal fusion node, and recording it as a real-time fusion deviation; Backtracking and offsetting the abnormal fusion node to obtain the historical fusion deviation under the corresponding historical node; Acquire the allowable deviation under the abnormal fusion node, and stop backtracking the abnormal fusion node when the historical fusion deviation is less than or equal to the allowable deviation; A calculation function is obtained, and the real-time fusion deviation and the historical fusion deviation are input into the calculation function, and the output result of the calculation function is recorded as the fusion deviation.

5. The railway infrastructure multi-source data fusion method according to claim 4 is characterized by: The step of adjusting the abnormal fusion node according to the fusion deviation comprises: Collect the backtracking time of the abnormal fusion node and record it as a parameter to be evaluated; Obtaining an evaluation threshold, and comparing the parameter to be evaluated with the evaluation threshold; When the parameter to be evaluated is greater than or equal to the evaluation threshold, the abnormal fusion node is offset according to the fusion deviation to obtain an updated fusion node, and data fusion is performed under the updated fusion node; When the parameter to be evaluated is less than the evaluation threshold, the fusion deviation is continuously collected until the parameter to be evaluated reaches or exceeds the evaluation threshold, and then the abnormal fusion node is adjusted.

6. A railway infrastructure multi-source data fusion system, characterized by: include: A data acquisition module, which is used to acquire multiple data sources of railway infrastructure and perform preprocessing to obtain a reference data source; A rule output module, the rule output module is used to collect the time characteristics of each of the reference data sources, and determine the fusion rules of each of the reference data sources according to the time characteristics; A data fusion module, the data fusion module is used to perform fusion processing on each reference data source according to the fusion rule to obtain multiple groups of fusion data and fusion nodes of each group of the fusion data; A state evaluation module, the state evaluation module is used to evaluate the effectiveness of each fusion node to obtain the fusion state of each fusion node, wherein the fusion state includes a normal fusion state and an abnormal fusion state; In the normal fusion state, the reference data sources under each fusion node are fused progressively; In the abnormal fusion state, the abnormal fusion node is located, and the fusion deviation of the reference data source under the abnormal fusion node is counted; A fusion optimization module, the fusion optimization module is used to adjust the abnormal fusion node according to the fusion deviation until all the reference data sources are progressively fused and then stop; The step of obtaining multiple data sources of railway infrastructure and preprocessing them to obtain a reference data source includes: Collecting operation data, maintenance data and environmental monitoring data of railway facilities, and summarizing them into independent data sets, and adding a unique data source identifier to each independent data set; Cleaning the data in each of the independent data sets to remove invalid and erroneous data items; Standardizing the data in the cleaned independent data set, unifying the data format in the independent data set, and outputting the independent data set with the unified format as a reference data source; The step of determining the fusion rule of each reference data source according to the time feature includes: Acquire the time characteristics of each of the reference data sources; The maximum value of the evaluation interval corresponding to the central feature is used as the data fusion interval; Obtaining data fusion intervals of the reference data source under all the centralized features, and arranging them in order from low to high to obtain a data fusion order; The reference data sources are fused according to the fusion order, and after the fusion of the reference data sources under the centralized feature is completed, the reference data sources under the decentralized feature are supplemented and fused; The step of evaluating the effectiveness of each fusion node includes: Collecting the fusion results of the reference data sources under each of the fusion nodes, and performing vectorization conversion to obtain a plurality of vectors to be evaluated; Obtaining an evaluation function, inputting the vector to be evaluated into the evaluation function, and recording the output result of the evaluation function as a parameter to be evaluated; Obtaining an evaluation threshold, and comparing the evaluation threshold with the parameter to be evaluated; If the parameter to be evaluated is less than the evaluation threshold, it indicates that the fusion node is in a normal fusion state; If the parameter to be evaluated is greater than or equal to the evaluation threshold, it indicates that the fusion node is in an abnormal fusion state.

7. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the railway infrastructure multi-source data fusion method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Geographic information analysis method for multi-source data fusion

    CN118568190A

  • Local and global decision fusion for cyber-physical system abnormality detection

    US20200089874A1