A multi-source heterogeneous data governance method for quad-network integration

Through the multi-source heterogeneous data governance method for four network fusion, through standardization, irreplaceability evaluation, interpolation method integration and replacement of real weight determination, data error and data distortion problems in data layer fusion technology are solved, and the authenticity and accuracy of data processing are improved.

CN119669658BActive Publication Date: 2025-05-13LANZHOU JIAOTONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510198718.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

When the prior art cleans and integrates data from different sources through data layer fusion technology, it is easy to lead to large errors between the data, resulting in a decrease in the authenticity of the final integrated data. At the same time, since different rail transit networks may adopt different data acquisition standards and methods, data distortion occurs during the integration and cleaning process.

Method used

A multi-source heterogeneous data governance method for four network fusion is proposed. By obtaining the historical data timing sequence of each transportation network source, standardizing and reducing multiple determination, evaluating the irreplaceability of each transportation network source, integrating the data sequence through interpolation method, determining the replacement real weight, and determining the final fusion data sequence through data cleaning combination and authenticity evaluation.

Benefits of technology

It effectively avoids the problem of large errors between data from different network sources, improves the authenticity and accuracy of multi-source heterogeneous data processing, avoids data distortion caused by different data acquisition standards and methods, and improves the authenticity and representativeness of the final integrated data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669658B_ABST
    Figure CN119669658B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data fusion technology, and specifically to a multi-source heterogeneous data governance method for quadruple network integration, which first preliminarily determines the reduction multiple of data from each traffic network source under the data type to be integrated when it is standardized, and then measures the irreplaceability of each traffic network source from the perspectives of uniqueness, sampling frequency accuracy and sampling frequency independence; further preliminarily integrates the data from all traffic network sources to determine the integrated data sequence; on the basis of the integrated data sequence, analyzes the replacement real weight that characterizes the rationality and authenticity of the replacement when selecting each reference network source data at each multi-source data time point; and then calculates the data cleaning authenticity based on the overall distribution of the replacement real weight and irreplaceability of the reference network source under different data cleaning combinations, thereby determining the final required fusion data sequence with higher authenticity and more accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data fusion, and in particular to a multi-source heterogeneous data governance method for quad-network integration. Background Art

[0002] With the rapid development of my country's transportation industry, the scale of the four networks, namely the trunk railway network, intercity railway network, urban (suburban) railway network and urban rail transit network, has continued to expand, forming a huge multi-source heterogeneous data system. These data come from different networks, devices and systems, with different data formats, data structures and data quality, which brings huge challenges to data governance;

[0003] Each transportation network may independently collect data of any data type through multiple transportation networks, thus generating multiple data sequences of the same data type. Existing technologies usually use data layer fusion technology to clean and integrate data sequences of each data type in all transportation network sources, and then perform multi-source heterogeneous data governance for the integration of four networks based on the obtained fused data sequences.

[0004] However, since the accuracy of monitoring required for any type of data varies from one transportation network source to another, that is, the actual information density of the same type of data in different networks is different, the quality of the same type of data from different data sources is uneven. At this time, when using the existing data layer fusion technology to clean and integrate data from different sources, it is easy to cause large errors between some data, resulting in a decrease in the authenticity of the final integrated data. At the same time, since different rail transit networks may adopt different data collection standards and methods, the data obtained by different networks are different and there are differences, resulting in data distortion during the integration and cleaning process. Summary of the invention

[0005] In order to solve the problem that in the process of cleaning and integrating data from different sources through data layer fusion technology in the existing technology, there is a large error between some data, resulting in a decrease in the authenticity of the final integrated data. At the same time, different rail transit systems may adopt different data collection standards and methods, resulting in data distortion during the integration and cleaning process, the purpose of this application is to provide a multi-source heterogeneous data governance method for quad-network integration, and the technical solution adopted is as follows.

[0006] The first aspect of the present application provides a multi-source heterogeneous data governance method for quad-network integration, including:

[0007] Obtain a historical data time series sequence of the data type to be integrated in each transportation network source; standardize the data according to the dimension size of the data in the historical data time series sequence, and determine the reduction multiple of the data type to be integrated in each transportation network source and the standardized data sequence;

[0008] Determine the irreplaceability of each transportation network source in the data type to be integrated based on the reduction factor, the distribution of the data type to be integrated in different transportation network sources, and the relative deviation of sampling frequencies and the overall size of sampling frequencies between the corresponding standardized data sequences;

[0009] Integrate the standardized data sequences of all traffic network sources by interpolation method to determine the integrated data sequence; obtain all multi-source data time points with overlapping time in all standardized data sequences and all reference network sources at each multi-source data time point; at each multi-source data time point, determine the replacement true weight of each reference network source according to the local fluctuation deviation between the standardized data sequence of each reference network source and the integrated data sequence;

[0010] All data cleaning combinations are determined by sequentially combining the reference network source types corresponding to each multi-source data time point; under each data cleaning combination, the data cleaning authenticity is determined according to the replacement true weight of the reference network source corresponding to each multi-source data time point and the irreplaceability; based on the data cleaning combination corresponding to the maximum value of the data cleaning authenticity, the fused data sequence of the data type to be integrated is determined.

[0011] Furthermore, the process of obtaining the reduction factor includes:

[0012] The largest historical data value in the historical data time series is rounded up to obtain a reduction multiple.

[0013] Furthermore, the process of acquiring the standardized data sequence includes:

[0014] All data in the historical data time series are maximized to determine a standardized data series.

[0015] Furthermore, the process of obtaining the irreplaceability includes:

[0016] Each traffic network source with a time series of historical data of the data type to be integrated is used as the target network source in turn; other traffic network sources with a time series of historical data of the data type to be integrated other than the target network source are used as the corresponding comparison network sources; the number of comparison network sources is used as the number of source features of the target network source;

[0017] When the number of source features is equal to 0, a preset replacement value is used as the non-replaceability of the target network source in the data type to be integrated;

[0018] When the number of source features is greater than 0:

[0019] The number of comparison network sources with historical data time series on the data type to be integrated is used as the number of source features of the target network source;

[0020] The mean of the sampling time intervals between all two adjacent data in the standardized data sequence corresponding to each traffic network source is used as the reference time interval corresponding to each traffic network source;

[0021] Determine the relative deviation of the sampling interval of the target network source according to the accumulated value of the difference between the reference time interval of the target network source and the reference time intervals of all the comparison network sources;

[0022] The irreplaceability of the target network source in the data type to be integrated is determined based on the reference time interval, the number of source features, the reduction factor, and the relative deviation of the sampling interval of the target network source; wherein the reference time interval and the number of source features are negatively correlated with the irreplaceability, and the reduction factor and the relative deviation of the sampling interval are positively correlated with the irreplaceability.

[0023] Furthermore, the process of determining the irreplaceability of the target network source in the data types to be integrated according to the reference time interval, the number of source features, the reduction factor, and the relative deviation of the sampling interval of the target network source includes:

[0024] The product of the negative correlation mapping value of the reference time interval of the target network source, the negative correlation mapping value of the source feature quantity, the reduction factor and the relative deviation of the sampling interval is normalized to determine the irreplaceability of the target network source in the data type to be integrated.

[0025] Furthermore, the process of obtaining all multi-source data time points that overlap in time in all standardized data sequences and all reference network sources for each multi-source data time point includes:

[0026] A time point having data in at least two standardized data sequences is used as a multi-source data time point; and a traffic network source corresponding to a standardized data sequence having data at the multi-source data time point is used as a reference network source for the multi-source data time point.

[0027] Furthermore, the process of obtaining the replacement real weight includes:

[0028] Each multi-source data time point is used as the target time point in turn; each reference network source corresponding to the target time point is used as the replacement network source in turn;

[0029] Using the data value of the target time point in the standardized data sequence of the replacement network source as the replacement data value; performing curve fitting on the standardized data sequence of the replacement network source to determine the standardized data curve of the replacement network source; performing curve fitting on the integrated data sequence to determine the integrated data curve;

[0030] In the standardized data sequence from the replacement network source, the time points of the nearest preset number of data adjacent to the target time point are used as adjacent time points of the target time point;

[0031] The difference between the data value of each adjacent time point on the standardized data curve and the replacement data value is used as the standard adjacent deviation of each adjacent time point; the cumulative value of the standard adjacent deviations of all adjacent time points is used as the standard volatility of the target time point;

[0032] The difference between the data value of each adjacent time point on the integrated data curve and the replacement data value is used as the integrated adjacent deviation of each adjacent time point; the cumulative value of the integrated adjacent deviations of all adjacent time points is used as the integrated volatility of the target time point;

[0033] The negative correlation mapping value of the difference between the standard volatility and the integrated volatility is used as the replacement real weight of the replacement network source at the target time point.

[0034] Furthermore, the process of obtaining the data cleaning combination includes:

[0035] At each multi-source data time point, a reference network source is selected to determine a data cleaning combination corresponding to all multi-source data time points under each selection method; all selection methods are traversed to obtain all data cleaning combinations.

[0036] Furthermore, the process of obtaining the authenticity of the data cleaning includes:

[0037] In each data cleaning combination, the product of the replacement authenticity weight of the reference network source corresponding to each multi-source data time point and the irreplaceability is used as the integrated replacement authenticity of each multi-source data time point; the data cleaning authenticity of each data cleaning combination is determined according to the mean of the integrated replacement authenticity of all multi-source data time points.

[0038] Furthermore, the process of acquiring the fused data sequence includes:

[0039] The data cleaning combination corresponding to the maximum value of data cleaning authenticity is used as the final cleaning combination; in the integrated data sequence, the data values ​​corresponding to each multi-source data time point are replaced with the data values ​​under the final cleaning combination to obtain a fused data sequence of the data type to be integrated.

[0040] In a second aspect, the present application provides a multi-source heterogeneous data governance system for quad-network integration, the system comprising:

[0041] A data acquisition preprocessing module is used to obtain the historical data time series sequence of the data type to be integrated in each transportation network source; standardize the data according to the dimension size of the data in the historical data time series sequence, determine the reduction multiple of the data type to be integrated in each transportation network source and the standardized data sequence;

[0042] A first determination module is used to determine the irreplaceability of each transportation network source in the data type to be integrated according to the reduction factor, the distribution of the data type to be integrated in different transportation network sources, and the relative deviation of the sampling frequency and the overall size of the sampling frequency between the corresponding standardized data sequences;

[0043] The second determination module is used to integrate the standardized data sequences of all traffic network sources by interpolation method to determine the integrated data sequence; obtain all multi-source data time points that overlap in time in all standardized data sequences and all reference network sources at each multi-source data time point; at each multi-source data time point, determine the replacement real weight of each reference network source according to the local fluctuation deviation between the standardized data sequence of each reference network source and the integrated data sequence;

[0044] The data fusion module is used to determine all data cleaning combinations by combining in sequence the reference network source types corresponding to each multi-source data time point; under each data cleaning combination, the data cleaning authenticity is determined according to the replacement real weight of the reference network source corresponding to each multi-source data time point and the said irreplaceability; according to the data cleaning combination corresponding to the maximum value of the data cleaning authenticity, the fused data sequence of the data type to be integrated is determined.

[0045] In a third aspect, the present application provides a computer device, comprising a memory and a processor. The memory is used to store computer program code, and the processor is used to call and run the computer program code from the memory to execute the method of the first aspect or any embodiment of the first aspect of the present application.

[0046] In a fourth aspect, the present application provides a computer program product, comprising a computer program code, which, when executed, performs the method of the first aspect of the present application or any embodiment of the first aspect.

[0047] In a fifth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program code. When the computer program code is executed, the method of the first aspect or any embodiment of the first aspect of the present application is performed.

[0048] This application has the following beneficial effects:

[0049] By obtaining the sampling time interval and the minimum precise unit of actual monitoring data in the standardized data sequence from any network source, and comparing them with the same type of data from different network sources with the same semantics, that is, according to the reduction factor, the relative deviation of the sampling frequency between the standardized data sequences from different traffic network sources and the overall size of the sampling frequency, the irreplaceability of data from any network source is obtained. This operation avoids the problem of reduced authenticity of the final integrated data due to differences in the actual information density of different networks for the same type of data, thereby improving the authenticity and accuracy of multi-source heterogeneous data processing.

[0050] By combining data from multiple sources at multiple time points, an integrated data sequence containing multiple different components is formed. By analyzing the change characteristics of the data sequence of actual monitoring data from different traffic network sources before and after integration, the real weight of the replacement of the data from the traffic network source at a certain sampling moment after integration is obtained. By traversing all combinations and determining the best integration and cleaning method, the cleaning and integration of multi-source heterogeneous data is completed. This operation avoids the problem of data distortion in the integration and cleaning process due to the fact that different networks may adopt different data collection standards and methods, thereby improving the authenticity and representativeness of the final integrated data. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0052] Figure 1 A flow chart of a multi-source heterogeneous data management method for quad-network integration provided by an embodiment of the present invention;

[0053] Figure 2 A structural diagram of a multi-source heterogeneous data management system for quad-network integration provided by an embodiment of the present invention;

[0054] Figure 3A schematic diagram of the structure of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of a multi-source heterogeneous data governance method for quad-network integration proposed by the present invention, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment, and the specific features, structures or characteristics in one or more embodiments may be combined in any suitable form. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.

[0056] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0057] The following is a detailed description of a specific solution of a multi-source heterogeneous data management method for quad-network integration provided by the present invention in conjunction with the accompanying drawings.

[0058] This application embodiment provides a multi-source heterogeneous data management method for quad-network integration. Figure 1 , which shows a flow chart of a multi-source heterogeneous data governance method for quad-network integration provided by an embodiment of the present invention, the method comprising:

[0059] Step S101: Obtain the historical data time series sequence of the data type to be integrated in each transportation network source; standardize the data according to the dimension size in the historical data time series sequence, and determine the reduction multiple of the data type to be integrated in each transportation network source and the standardized data sequence.

[0060] The four networks refer to the trunk railway network, the intercity railway network, the urban (suburban) railway network, and the urban rail transit network. These networks have the following functions: The trunk railway network is the main skeleton of the national railway network, which is mainly responsible for connecting major cities and important transportation hubs and providing long-distance rapid transportation services. The intercity railway network mainly connects adjacent cities or urban agglomerations, providing medium-distance rapid transportation services. It is a supplement to the trunk railway network and helps to ease the traffic pressure between large cities. The urban (suburban) railway network mainly serves cities and their surrounding areas, providing medium- and long-distance rapid transportation services. Urban rail transit mainly serves the city and provides short-distance rapid transportation services. Among them, the types of data to be integrated can include train operation data, passenger flow data, equipment status data, etc. The analysis methods for all data types to be integrated in this application are the same, so this application will only analyze the data types to be integrated in the future, so as to perform multi-source heterogeneous data governance on all data types; these data will be collected at every moment of every day through various corresponding types of sensors and monitoring equipment in the four networks, and transmitted to the data platform through the network where they are located for centralized processing, among which the sampling frequencies of different transportation network sources are usually different.

[0061] In a specific implementation of an embodiment of the present invention, the sampling duration of the historical data time series is set to one month before the current moment. Since different transportation networks have different regulatory departments in real life, for example, the trunk railway network usually regulates EMU trains and ordinary trains, while the urban rail transit network usually regulates the operation of light rail and subway, etc., different regulatory departments have different requirements for data quality. Therefore, different networks have different dimensions for different types of data in the process of obtaining information. In order to avoid large errors in the results caused by the inconsistency of units and dimensions between data from different sources during subsequent operations, different types of data are standardized.

[0062] Preferably, in some possible implementation methods of the embodiments of the present invention, the process of obtaining the standardized data sequence includes: maximizing all the data in the historical data time series sequence to determine the standardized data sequence; maximization is to divide each data in each historical data time series sequence by the maximum value therein, so that the value range of each data obtained is within 0 to 1.

[0063] It is further necessary to consider that, for each traffic network source, the larger the dimension in its historical data time series sequence, that is, the larger the reduction factor in the standardization process, the higher the accuracy of the traffic network source data; therefore, in order to facilitate the subsequent irreplaceable calculations, the reduction factor in the standardization process of each traffic network source data is further calculated, wherein the process of obtaining the reduction factor includes: rounding up the largest historical data value in the historical data time series sequence to obtain the reduction factor; since the embodiment of the present application implements standardization through maximization, that is, each data in each historical data time series sequence is divided by the corresponding maximum value, the corresponding reduction factor is determined according to the largest historical data value in each historical data time series sequence; in addition, the purpose of rounding up here is to reduce the amount of calculation, which can be adjusted according to the specific implementation environment.

[0064] Step S102: Determine the irreplaceability of each transportation network source in the data type to be integrated based on the reduction factor, the relative deviation of the sampling frequencies between the standardized data sequences from different transportation network sources, and the overall size of the sampling frequencies.

[0065] According to prior knowledge, for any data type to be integrated, multiple transportation networks usually collect data independently, thus generating multiple data sequences of the same type. Since the data types to be integrated have different monitoring accuracy under different network sources, that is, different networks have different actual information densities for the same type of data, resulting in uneven quality of the same type of data from different data sources. At this time, in the process of cleaning and integrating data from different sources, it is easy to cause large errors between some data, resulting in a decrease in the authenticity of the final integrated data. There is an obvious inclusion relationship between the various transportation networks. Therefore, for any type of data to be integrated, in the process of progressing from the innermost layer, that is, from the urban rail transit network to the outer layer, the fewer transportation networks that have data of the same category in other transportation networks, the data of the type to be integrated from the current transportation network source is obviously irreplaceable; that is, if the type of data to be integrated can only be collected in one of the transportation network sources, then the transportation network source is irreplaceable for the type of data to be integrated; the fewer transportation network sources corresponding to the corresponding data type to be integrated, that is, the fewer transportation network sources with historical data time series sequences of the data type to be integrated, the higher the irreplaceability of the transportation network sources with historical data time series sequences of the data type to be integrated.

[0066] In addition, for each transportation network source, if the sampling time interval of the historical data time series of the data type to be integrated is smaller as a whole and the reduction factor is larger, it means that the data collected by the transportation network source is more accurate and the data covers more extensive information, and its irreplaceability should be stronger; in addition, the greater the difference between the sampling time interval of the transportation network source and the sampling time interval of other transportation network sources, it means that the sampling frequency of the transportation network source is more unique or independent, which may also reflect the corresponding irreplaceability to a certain extent.

[0067] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining irreplaceability includes:

[0068] Each traffic network source with a time series of historical data of the data type to be integrated is used as the target network source in turn; other traffic network sources with a time series of historical data of the data type to be integrated other than the target network source are used as the corresponding comparison network sources; the number of comparison network sources is used as the number of source features of the target network source;

[0069] When the number of source features is equal to 0, the preset replacement value is used as the irreplaceability of the target network source in the data type to be integrated. When the value of the number of source features is 0, it means that the data type to be integrated only exists in the target network source, so no subsequent integration process is required, and it itself is irreplaceable. Therefore, the preset replacement number is used as the corresponding irreplaceability, and during the subsequent integration and data cleaning, since there is no data from other transportation network sources, the standardized data sequence of the target network source is directly used as the fused data sequence ultimately required for this application, and no further details are given here. In a specific implementation method of an embodiment of the present invention, the preset replacement value is set to 1.

[0070] When the number of source features is greater than 0: the number of comparison network sources with historical data time series on the data type to be integrated is used as the number of source features of the target network source; the mean of the sampling time intervals between all two adjacent data in the standardized data sequence corresponding to each traffic network source is used as the reference time interval corresponding to each traffic network source; the relative deviation of the sampling interval of the target network source is determined based on the cumulative value of the difference between the reference time interval of the target network source and the reference time interval of all comparison network sources; the larger the number of source features, the fewer traffic network sources with historical data time series of the data type to be integrated, and the greater the irreplaceability of the target network source; the reference time interval is the same as the reference time interval of the target network source. The smaller the interval and the larger the reduction factor, the higher the accuracy of the data collected by the target network source and the wider the information covered by the data, and the greater its irreplaceability should be; and the larger the relative deviation of the sampling interval, the more unique or independent the sampling frequency of the target network source is, which may also reflect the corresponding irreplaceability to a certain extent; therefore, the irreplaceability of the target network source in the data type to be integrated is further determined based on the reference time interval, the number of source features, the reduction factor and the relative deviation of the sampling interval of the target network source; among them, the reference time interval and the number of source features are negatively correlated with the irreplaceability, and the reduction factor and the relative deviation of the sampling interval are positively correlated with the irreplaceability.

[0071] Preferably, in some possible implementations of the embodiments of the present invention, the process of determining the irreplaceability of the target network source in the data type to be integrated according to the reference time interval, the number of source features, the reduction factor, and the relative deviation of the sampling interval of the target network source includes:

[0072] Normalize the product of the negative correlation mapping value of the reference time interval of the target network source, the negative correlation mapping value of the source feature quantity, the reduction factor, and the relative deviation of the sampling interval to determine the irreplaceability of the target network source in the data type to be integrated. In a specific implementation of the embodiment of the present invention, the process of obtaining irreplaceability is expressed by the formula: ;in, Target network source Irreplaceability in the type of data to be integrated; Target network source The reference time interval; Target network source The number of source features, that is, the number of corresponding comparison network sources; Target network source The corresponding A reference time interval for comparison network sources; Target network source The relative deviation of the sampling interval; The reduction factor of the data type to be integrated in the target network source; is an exponential function with a natural constant as base; is a linear normalization function; wherein the specific process of obtaining the reference time interval includes: in the standardized data sequence from each traffic network source, the time interval between each data and the previous data is used as the local interval of each data; the mean of the local intervals of all data except the first data is used as the corresponding reference time interval.

[0073] Step S103: Integrate the standardized data sequences of all traffic network sources by interpolation method to determine the integrated data sequence; obtain all multi-source data time points that overlap in time in all standardized data sequences and all reference network sources at each multi-source data time point; at each multi-source data time point, determine the replacement true weight of each reference network source based on the local fluctuation deviation between the standardized data sequence of each reference network source and the integrated data sequence.

[0074] Due to the different sampling time intervals, the data at some time points are missing, so it is necessary to integrate the data through other networks. At the same time, there are also data of the same type that obtain multiple data with different values ​​through different networks at the same time, so the data needs to be cleaned to eliminate the impact of duplicate data. However, since different railway networks may adopt different data collection standards and methods, the data obtained by different networks when obtaining corresponding information are different, resulting in data distortion in the process of integration and cleaning. Therefore, the same type of data from multiple different transportation network sources is first integrated using the existing interpolation method, specifically by arranging the standardized data corresponding to the same sampling time point, that is, the multi-source data time point, in chronological order as a new data set, that is, an integrated data sequence; specifically, if a sampling time point, that is, a multi-source data time point, has data from multiple different transportation network sources, then the mean of the data from all transportation network sources at the sampling time point is used as the data at the sampling time point in the integrated data sequence. However, different transportation networks adopt different data collection standards and methods, and the integrated data sequence that integrates the multi-source data time points with the mean value usually has data distortion, so it cannot be directly analyzed as a fused data sequence, and further adjustments need to be made on this basis.

[0075] In the data cleaning process, there are multiple data at multi-source data time points, that is, time points where data exists from different traffic network sources. In order to make the obtained fused data sequence more accurate, it is necessary to select the most authentic one from the multiple data corresponding to the multi-source data time points as the fusion result for replacement; therefore, it is first necessary to determine the multi-source data time points and the reference network sources corresponding to each multi-source data time point. Therefore, in a specific implementation of an embodiment of the present invention, the process of obtaining all multi-source data time points with overlapping time in all standardized data sequences and all reference network sources for each multi-source data time point includes:

[0076] A time point having data in at least two standardized data sequences is used as a multi-source data time point; and a traffic network source corresponding to a standardized data sequence having data at the multi-source data time point is used as a reference network source for the multi-source data time point.

[0077] Since the same type of data from different network sources has different sampling time intervals in the time series, and due to different data collection methods, the change trend of the acquired data itself is also different from the data series corresponding to other sources. Therefore, by judging the changes between the new data series obtained by integrating data from various sources and the data series corresponding to the original network sources of each data, the real weight of replacement of any current network source at a certain sampling time point can be judged, so that the larger the real weight of replacement, the more real the data of the corresponding reference network source is at the multi-source data time point, and the greater the possibility of replacement.

[0078] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the replacement real weight includes:

[0079] Each multi-source data time point is used as the target time point in turn; each reference network source corresponding to the target time point is used as the replacement network source in turn;

[0080] The data value of the target time point in the standardized data sequence of the replacement network source is used as the replacement data value; the standardized data sequence of the replacement network source is subjected to curve fitting to determine the standardized data curve of the replacement network source; the integrated data sequence is subjected to curve fitting to determine the integrated data curve;

[0081] In the standardized data sequence of the replacement network source, the time points of the nearest preset number of data adjacent to the target time point are used as the adjacent time points of the target time point; the difference between the data value of each adjacent time point on the standardized data curve and the replacement data value is used as the standard adjacent deviation of each adjacent time point; the cumulative value of the standard adjacent deviations of all adjacent time points is used as the standard volatility of the target time point; the difference between the data value of each adjacent time point on the integrated data curve and the replacement data value is used as the integrated adjacent deviation of each adjacent time point; the cumulative value of the integrated adjacent deviations of all adjacent time points is used as the integrated volatility of the target time point; the negative correlation mapping value of the difference between the standard volatility and the integrated volatility is used as the replacement real weight of the replacement network source at the target time point. In a specific implementation of the embodiment of the present invention, the preset number is set to 4, which can be adjusted according to the specific implementation environment.

[0082] For the target time point, if the local fluctuation of the replacement data value in the standardized data curve of the replacement network source is more similar to the local fluctuation of the integrated data curve, it means that the replacement data value is more consistent with the actual changes after the overall integration, then the impact after the replacement will be smaller, the rationality and authenticity of the replacement data value will be higher, and the real weight of the replacement should be larger; and the smaller the difference between the standard volatility and the integrated volatility, the more similar the local fluctuation of the standardized data curve is to the local fluctuation of the integrated data curve, and the corresponding real weight of the replacement should be larger.

[0083] In a specific implementation of the embodiment of the present invention, the process of obtaining the replacement real weight is expressed by the formula: ;

[0084] in, Target time point Replace the network source Replace the real weight of is the number of adjacent time points of the target time point; To replace the network source The target time point in the standardized data series The data value of , that is, the replacement data value; To replace the network source The target time point on the standardized data curve No. The data values ​​of adjacent time points; is the absolute value symbol; Target time point No. The standard deviation of neighboring time points; Target time point Standard volatility of To integrate the target time point on the data curve No. The data values ​​of adjacent time points; Target time point No. The integrated adjacency deviation of adjacent time points; Target time point The integrated volatility of is an exponential function with a natural constant as its base.

[0085] Step S104: Determine all data cleaning combinations by combining the reference network source types corresponding to each multi-source data time point in sequence; Under each data cleaning combination, determine the data cleaning authenticity according to the replacement real weight and irreplaceability of the reference network source corresponding to each multi-source data time point; Determine the fused data sequence of the data type to be integrated according to the data cleaning combination corresponding to the maximum value of the data cleaning authenticity.

[0086] Since there are usually multiple multi-source data time points, in order to determine the most reasonable data cleaning result, the selection of reference network sources for all multi-source data time points is further combined, and then the irreplaceability of the reference network sources selected for each multi-source data time point under each combination and the replacement true weight are analyzed, so as to comprehensively evaluate the data cleaning authenticity of each combination, and then select the most reasonable and authentic combination according to the data cleaning authenticity, and determine the final required fusion data sequence.

[0087] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the data cleaning combination includes:

[0088] At each multi-source data time point, a reference network source is selected, and a data cleaning combination corresponding to all multi-source data time points under each selection method is determined; all selection methods are traversed to obtain all data cleaning combinations. All data cleaning combinations can include the combination of all reference network sources at all multi-source data time points, and each data cleaning combination is further calculated to select the most preferred data cleaning combination for cleaning and obtain the final fused data sequence.

[0089] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the authenticity of data cleaning includes:

[0090] In each data cleaning combination, the product of the replacement authenticity weight and irreplaceability of the reference network source corresponding to each multi-source data time point is used as the integrated replacement authenticity of each multi-source data time point; the data cleaning authenticity of each data cleaning combination is determined according to the mean of the integrated replacement authenticity of all multi-source data time points. In each data cleaning combination, a corresponding reference network source is selected for each multi-source data time point; each reference network source usually corresponds to different irreplaceability, and the greater the irreplaceability, the greater the possibility that the data from the reference network source is used as the final fused data; and when the replacement authenticity weight is larger, the corresponding reference network source data is more authentic at the multi-source data time point, and the possibility of substitution is also greater; therefore, when the overall replacement authenticity of each multi-source data time point obtained by the product of the replacement authenticity weight and irreplaceability is larger, that is, when the data cleaning authenticity is larger, the corresponding data cleaning combination is more in line with the authenticity requirements, and the data sequence obtained under the corresponding data cleaning combination is more likely to correspond to the final integrated cleaning result.

[0091] In a specific implementation of the embodiment of the present invention, the process of obtaining the authenticity of data cleaning is expressed by the formula: ;

[0092] in, For the The authenticity of data cleaning for each data cleaning combination; is the number of time points of multi-source data; For the Data cleaning combination The replacement true weights of the reference network sources at each multi-source data point in time; For the Data cleaning combination The irreplaceability of the reference network source for multiple source data time points; For the Data cleaning combination The integration of multiple source data time points replaces authenticity.

[0093] Since the greater the authenticity of data cleaning, the more the corresponding data cleaning combination meets the authenticity requirement; therefore, the final required fusion data sequence can be determined according to the maximum data cleaning authenticity. Preferably, in some possible implementations of the embodiments of the present invention, the acquisition process of the fusion data sequence includes:

[0094] The data cleaning combination corresponding to the maximum value of data cleaning authenticity is used as the final cleaning combination; in the integrated data sequence, the data values ​​corresponding to each multi-source data time point are replaced with the data values ​​under the final cleaning combination to obtain a fused data sequence of the data type to be integrated; in the integrated data sequence, the data at other time points other than the multi-source data time point have unique data values ​​after integration, and the multi-source data time points actually correspond to different data values ​​in different traffic network sources. The integrated data sequence is not accurate enough to obtain new values ​​only by the method of averaging, so each data value in the selected final cleaning combination is replaced in turn to obtain the fused data sequence ultimately required by this application. Finally, according to the method for obtaining the fused data sequence of the data type to be integrated, the fused data sequence is obtained for other data types, thereby conducting governance of multi-source heterogeneous data for the integration of four networks.

[0095] In summary, the multi-source heterogeneous data governance method for the integration of four networks proposed in this application first preliminarily determines the reduction multiple of the data from each transportation network source under the data type to be integrated when it is standardized, and then measures the irreplaceability of each transportation network source from the perspective of uniqueness, sampling frequency accuracy and sampling frequency independence; further integrates the data from all transportation network sources to determine the integrated data sequence; on the basis of the integrated data sequence, analyzes the replacement real weight that characterizes the rationality and authenticity of the replacement when selecting each reference network source data at each multi-source data time point; and then calculates the authenticity of data cleaning based on the overall distribution of the replacement real weight and irreplaceability of the reference network source under different data cleaning combinations, so as to determine the final required fusion data sequence with higher authenticity and more accuracy.

[0096] This application also provides a multi-source heterogeneous data governance system for quad-network integration. Figure 2 , which shows a structural diagram of a multi-source heterogeneous data governance system for quad-network integration provided by an embodiment of the present invention, the system comprising: a data acquisition preprocessing module 201, a first determination module 202, a second determination module 203 and a data fusion module 204.

[0097] The data acquisition preprocessing module 201 is used to obtain the historical data time series sequence of the data type to be integrated in each transportation network source; standardize the data according to the dimension size of the historical data time series sequence, determine the reduction multiple of the data type to be integrated in each transportation network source and the standardized data sequence;

[0098] The first determination module 202 is used to determine the irreplaceability of each transportation network source in the data type to be integrated according to the reduction factor, the distribution of the data type to be integrated in different transportation network sources, and the relative deviation of the sampling frequency and the overall size of the sampling frequency between the corresponding standardized data sequences;

[0099] The second determination module 203 is used to integrate the standardized data sequences of all traffic network sources by interpolation method to determine the integrated data sequence; obtain all multi-source data time points that overlap in time in all standardized data sequences and all reference network sources at each multi-source data time point; at each multi-source data time point, determine the replacement real weight of each reference network source according to the local fluctuation deviation between the standardized data sequence of each reference network source and the integrated data sequence;

[0100] The data fusion module 204 is used to determine all data cleaning combinations by combining in sequence the reference network source types corresponding to each multi-source data time point; under each data cleaning combination, determine the data cleaning authenticity according to the replacement real weight and irreplaceability of the reference network source corresponding to each multi-source data time point; and determine the fused data sequence of the data type to be integrated according to the data cleaning combination corresponding to the maximum value of the data cleaning authenticity.

[0101] It should be noted that the system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the multi-source heterogeneous data governance system for quadruple network convergence and the multi-source heterogeneous data governance method embodiment for quadruple network convergence provided in the above embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0102] The present application also provides a computer device. Figure 3 , which shows a schematic diagram of the structure of a computer device provided by an embodiment of the present invention, the computer device includes a memory 301, a processor 302, and a computer program 303 stored in the memory 301 and running on the processor 302, wherein when the processor 302 executes the computer program 303, the computer device can execute any of the multi-source heterogeneous data governance methods for quad-network integration introduced above.

[0103] An embodiment of the present application also provides a computer program product. When the computer program product runs on a computer device, the computer device can execute any of the multi-source heterogeneous data governance methods for quad-network integration described above.

[0104] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program code is stored. When the computer program code runs on a computer device, the computer device can execute any of the multi-source heterogeneous data governance methods for quadruple-network integration described above.

[0105] In the embodiments provided in the present application, it should be understood that the provided computer device, computer program product and computer-readable storage medium are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the methods provided above and will not be repeated here.

[0106] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0107] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

Claims

1. A multi-source heterogeneous data governance method for quad-network integration, characterized in that: The method comprises: Obtain a historical data time series sequence of the data type to be integrated in each transportation network source; standardize the data according to the dimension size of the data in the historical data time series sequence, and determine the reduction multiple of the data type to be integrated in each transportation network source and the standardized data sequence; Determine the irreplaceability of each transportation network source in the data type to be integrated based on the reduction factor, the distribution of the data type to be integrated in different transportation network sources, and the relative deviation of sampling frequencies and the overall size of sampling frequencies between the corresponding standardized data sequences; Integrate the standardized data sequences of all traffic network sources by interpolation method to determine the integrated data sequence; obtain all multi-source data time points with overlapping time in all standardized data sequences and all reference network sources at each multi-source data time point; at each multi-source data time point, determine the replacement true weight of each reference network source according to the local fluctuation deviation between the standardized data sequence of each reference network source and the integrated data sequence; All data cleaning combinations are determined by sequentially combining the reference network source types corresponding to each multi-source data time point; under each data cleaning combination, the data cleaning authenticity is determined according to the replacement true weight of the reference network source corresponding to each multi-source data time point and the irreplaceability; based on the data cleaning combination corresponding to the maximum value of the data cleaning authenticity, the fused data sequence of the data type to be integrated is determined.

2. According to the method for managing multi-source heterogeneous data for quadruple network integration according to claim 1, it is characterized in that: The process of obtaining the reduction factor includes: The largest historical data value in the historical data time series is rounded up to obtain a reduction multiple.

3. According to the method for managing multi-source heterogeneous data for quadruple network integration according to claim 1, it is characterized in that: The process of obtaining the standardized data sequence includes: All data in the historical data time series are maximized to determine a standardized data series.

4. The method for managing multi-source heterogeneous data for quadruple network integration according to claim 1 is characterized in that: The process of obtaining the irreplaceability includes: Each traffic network source with a time series of historical data of the data type to be integrated is used as the target network source in turn; other traffic network sources with a time series of historical data of the data type to be integrated other than the target network source are used as the corresponding comparison network sources; the number of comparison network sources is used as the number of source features of the target network source; When the number of source features is equal to 0, a preset replacement value is used as the non-replaceability of the target network source in the data type to be integrated; When the number of source features is greater than 0: The number of comparison network sources with historical data time series on the data type to be integrated is used as the number of source features of the target network source; The mean of the sampling time intervals between all two adjacent data in the standardized data sequence corresponding to each traffic network source is used as the reference time interval corresponding to each traffic network source; Determine the relative deviation of the sampling interval of the target network source according to the accumulated value of the difference between the reference time interval of the target network source and the reference time intervals of all the comparison network sources; The irreplaceability of the target network source in the data type to be integrated is determined based on the reference time interval, the number of source features, the reduction factor, and the relative deviation of the sampling interval of the target network source; wherein the reference time interval and the number of source features are negatively correlated with the irreplaceability, and the reduction factor and the relative deviation of the sampling interval are positively correlated with the irreplaceability.

5. The method for managing multi-source heterogeneous data for quadruple network integration according to claim 4 is characterized in that: The process of determining the irreplaceability of the target network source in the data types to be integrated according to the reference time interval, the number of source features, the reduction factor, and the relative deviation of the sampling interval of the target network source includes: The product of the negative correlation mapping value of the reference time interval of the target network source, the negative correlation mapping value of the source feature quantity, the reduction factor and the relative deviation of the sampling interval is normalized to determine the irreplaceability of the target network source in the data type to be integrated.

6. The method for managing multi-source heterogeneous data for quadruple network integration according to claim 1 is characterized in that: The process of obtaining all multi-source data time points that coincide in time in all standardized data sequences and all reference network sources for each multi-source data time point comprises: A time point having data in at least two standardized data sequences is used as a multi-source data time point; and a traffic network source corresponding to a standardized data sequence having data at the multi-source data time point is used as a reference network source for the multi-source data time point.

7. The method for managing multi-source heterogeneous data for quadruple network integration according to claim 1 is characterized in that: The process of obtaining the replacement real weight includes: Each multi-source data time point is used as the target time point in turn; each reference network source corresponding to the target time point is used as the replacement network source in turn; Using the data value of the target time point in the standardized data sequence of the replacement network source as the replacement data value; performing curve fitting on the standardized data sequence of the replacement network source to determine the standardized data curve of the replacement network source; performing curve fitting on the integrated data sequence to determine the integrated data curve; In the standardized data sequence from the replacement network source, the time points of the nearest preset number of data adjacent to the target time point are used as adjacent time points of the target time point; The difference between the data value of each adjacent time point on the standardized data curve and the replacement data value is used as the standard adjacent deviation of each adjacent time point; the cumulative value of the standard adjacent deviations of all adjacent time points is used as the standard volatility of the target time point; The difference between the data value of each adjacent time point on the integrated data curve and the replacement data value is used as the integrated adjacent deviation of each adjacent time point; the cumulative value of the integrated adjacent deviations of all adjacent time points is used as the integrated volatility of the target time point; The negative correlation mapping value of the difference between the standard volatility and the integrated volatility is used as the replacement real weight of the replacement network source at the target time point.

8. The method for managing multi-source heterogeneous data for quadruple network integration according to claim 1 is characterized in that: The process of obtaining the data cleaning combination includes: At each multi-source data time point, a reference network source is selected to determine a data cleaning combination corresponding to all multi-source data time points under each selection method; all selection methods are traversed to obtain all data cleaning combinations.

9. The method for managing multi-source heterogeneous data for quadruple network integration according to claim 1, characterized in that: The process of obtaining the authenticity of data cleaning includes: In each data cleaning combination, the product of the replacement authenticity weight of the reference network source corresponding to each multi-source data time point and the irreplaceability is used as the integrated replacement authenticity of each multi-source data time point; the data cleaning authenticity of each data cleaning combination is determined according to the mean of the integrated replacement authenticity of all multi-source data time points.

10. The method for managing multi-source heterogeneous data for quadruple network integration according to claim 1, characterized in that: The acquisition process of the fused data sequence includes: The data cleaning combination corresponding to the maximum value of data cleaning authenticity is used as the final cleaning combination; in the integrated data sequence, the data values ​​corresponding to each multi-source data time point are replaced with the data values ​​under the final cleaning combination to obtain a fused data sequence of the data type to be integrated.

Citation Information

Patent Citations

  • Intelligent processing system for relay protection multi-source heterogeneous information

    CN109617001A

  • Enterprise multi-source heterogeneous data fusion visualization system

    CN117743450A