Data fusion method and system for urban physical examination spatiotemporal big data

By acquiring and constructing the time domain and spatiotemporal weights of traffic data, the accuracy problem when fusing different data sampling levels is solved, and the efficient and accurate fusion of urban physical examination data is achieved.

CN120561866BActive Publication Date: 2025-10-03HEBEI GEOGRAPHIC INFORMATION GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511028831.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-03
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In the existing technology, traffic data from different data sampling levels are fused with the same weight, resulting in low data accuracy, which affects the accuracy of urban physical examinations.

Method used

By obtaining the temporal and spatial weights of traffic data at each data sampling level, combining the dynamic time warping algorithm and Kriging interpolation, constructing constraint conditions, and performing adaptive data fusion, the temporal and spatial weights of traffic data in each node area are determined.

Benefits of technology

The accuracy of data fusion is improved, ensuring that the traffic fusion data of each node area is more accurate, and supporting efficient evaluation of urban physical examinations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561866B_ABST
    Figure CN120561866B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data fusion technology, and in particular to a data fusion method and system for urban physical examination spatiotemporal big data. The method performs preliminary characterization of data authenticity based on the matching of data change value trends between upper and lower nodes in a node area at each data sampling level, thereby determining the temporal weight of traffic data in each node area. Then, constraints are constructed based on the differences in data at different levels within the same area in the city at the fusion moment to perform spatial autocorrelation characterization, thereby determining the spatiotemporal weight of traffic data corresponding to urban data in each node area at different data sampling levels. Finally, the temporal weight of traffic data and the spatiotemporal weight of traffic data are integrated to perform more robust adaptive data fusion, thereby making the traffic fusion data of each node area obtained after data fusion more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data fusion technology, and in particular to a data fusion method and system for urban physical examination spatiotemporal big data. Background Art

[0002] Urban health checks are a comprehensive approach to assessing a city's functions, environmental quality, infrastructure, and socioeconomic activities. These areas can all be reflected to some extent in traffic data, so urban health checks are conducted by combining big data with traffic data. Because urban health checks require analysis based on multiple data sampling levels (such as provincial, municipal, and district levels), existing technologies typically fuse traffic data from different sampling levels with equal weights to generate fused traffic data, which is then used to conduct urban health checks.

[0003] However, since different data sampling levels focus on different data scales, the collected traffic sampling data have different resolutions. When the same weight is used for data fusion, not only will some data sampling levels be unable to be fused at some moments due to data synchronization issues, but the different sources and credibility of data at different data sampling levels are not taken into account. As a result, the existing technology has a poor effect in fusing traffic data from different data sampling levels with the same weight, resulting in low accuracy of the obtained traffic fusion data, which in turn affects the accuracy of urban physical examinations based on traffic fusion data. Summary of the Invention

[0004] In order to solve the technical problem that the existing technology of fusing traffic data from different data sampling levels with the same weight has poor effect, resulting in low accuracy of the obtained traffic fusion data, the purpose of this application is to provide a data fusion method and system for urban physical examination spatiotemporal big data. The technical solutions adopted are as follows:

[0005] The first aspect of the present application provides a data fusion method for urban physical examination spatiotemporal big data, comprising:

[0006] Acquire all traffic sampling data of each monitoring node in each node area at each data sampling level in the urban traffic historical data; wherein all monitoring nodes in each node area include one upper node and at least one lower node;

[0007] Determine the corresponding time domain weight of traffic data according to the traffic sampling data association between the upper node and each lower node in each node area under each data sampling level and the sampling frequency;

[0008] At the fusion moment, the constraint conditions are constructed based on the temporal weight of the traffic data combined with the standard deviation distribution of the traffic sampling data of each node area in each data sampling layer, and the spatiotemporal weight of the traffic data of each node area in each data sampling layer is determined;

[0009] Data fusion is performed based on the time domain weight and spatiotemporal weight of traffic data of each node area at each data sampling level to obtain the traffic fusion data of each node area at the fusion time.

[0010] Furthermore, the process of obtaining the time-domain weight of traffic data includes:

[0011] Obtain all traffic multi-source data for each monitoring node; in each node area, determine the trend similarity of each node area at each data sampling level based on the similarity of the time series change trends between the traffic multi-source data of the upper node and the traffic sampling data of each lower node at each data sampling level;

[0012] The product of the total number of all traffic sampling data collected by the upper node of each node area at each data sampling level and the trend similarity is normalized to determine the time domain weight of the traffic data of each node area at each data sampling level.

[0013] Furthermore, the process of obtaining the trend similarity includes:

[0014] In each node area, the mean of the traffic sampling data of all lower-level nodes at each sampling time of each data sampling layer is used as the corresponding lower-level traffic data;

[0015] Arrange the lower-level traffic data of each node area at all sampling moments at each data sampling level in chronological order and perform curve fitting to obtain the lower-level fitting curve;

[0016] Arrange all traffic multi-source data of each upper node in chronological order and perform curve fitting to determine the upper fitting curve;

[0017] Based on the dynamic time warping algorithm, the DTW distance between the lower-level fitting curve and the upper-level fitting curve is negatively correlated to determine the corresponding trend similarity.

[0018] Furthermore, the process of obtaining the spatiotemporal weight of traffic data includes:

[0019] Based on all traffic sampling data of each node area in each data sampling layer, the traffic fitting data of each monitoring node at the fusion time is determined by Kriging interpolation; at the fusion time, based on the actual deviation of the traffic fitting data, the residual of each monitoring node in each data sampling layer at the fusion time is determined;

[0020] Obtain the initial spatiotemporal weight of each node area in each data sampling layer at the fusion time; wherein, the traffic prediction data obtained by weighted averaging the corresponding traffic fitting data of the initial spatiotemporal weights in each data sampling layer corresponding to each node area at the fusion time satisfies the conditions of unbiasedness and minimum variance;

[0021] Determine the first constraint condition based on the discrete residuals in each node region at each data sampling level and the deviation of the initial spatiotemporal weight of each node region compared to the initial spatiotemporal weights of all node regions;

[0022] Get node area and node area , the node area and node area are any nodes in all node regions, and the node regions and node area Can be the same node area; according to the node area and node area The second constraint condition is determined based on the similarity of the residual space distribution between each data sampling level and the initial spatiotemporal weight deviation;

[0023] The first and second constraints are embedded in the joint kriging system to construct an extended kriging matrix equation. The alternating direction multiplier method is used to decompose the kriging matrix equation into a two-step iterative solution of intra-level weight update and cross-level constraint coordination to determine the spatiotemporal weight of traffic data for each node area in each data sampling level.

[0024] Furthermore, the residual acquisition process includes:

[0025] In each node area, the difference between the traffic multi-source data of each monitoring node at the fusion moment and the corresponding traffic fitting data is used as the residual of each monitoring node in each data sampling layer at the fusion moment.

[0026] Furthermore, the process of obtaining the first constraint condition includes:

[0027] At the fusion moment, the standard deviation of the residuals of all monitoring nodes in all node areas in each data sampling level is used as the corresponding allocation fluctuation tolerance threshold; the mean of the initial spatiotemporal weights of all node areas in each data sampling level is used as the corresponding reference average weight; and the difference between the initial spatiotemporal weight of each node area in each data sampling level and the reference average weight is less than or equal to the fluctuation tolerance threshold as the first constraint condition.

[0028] Furthermore, the process of obtaining the second constraint condition includes:

[0029] Using any data sampling level as the reference level, and using data sampling levels other than the reference level as the comparison level of the reference level;

[0030] Calculate the node area in the reference layer at the time of fusion The residual space and the corresponding node area in each comparison level Moran index between the residual spaces; according to the node area in the reference level The initial spatiotemporal weights and the corresponding node areas in each comparison level The deviation between the initial spatiotemporal weights is used to determine the corresponding reference weight deviation value; the product between the Moran index and the reference weight deviation value is used as the node area in the reference layer at the fusion moment The node area in each comparison level The local constraint value between them; according to the node area in the reference layer at the fusion time And the corresponding node areas in all comparison levels The cumulative value of the local constraint values ​​between the two determines the reference level in the node area at the fusion moment With node area Reference constraint values ​​between; based on all data sampling levels in the node area With node area The cumulative value of the reference constraint values ​​between the nodes determines the node area With node area The overall constraint value between the two levels is less than or equal to the preset cross-level constraint threshold as the second constraint condition.

[0031] Furthermore, the process of obtaining the reference weight deviation value includes:

[0032] Based on the node area in the reference hierarchy The initial spatiotemporal weights and the corresponding node areas in each comparison level The square of the difference between the initial spatiotemporal weights of the reference layer is used to determine the overall weight deviation value; according to the node area in the reference layer The initial spatiotemporal weights and the corresponding node areas in each comparison level The product of the initial spatiotemporal weights of the reference correction parameters is used to determine the reference correction parameters; the node area in the reference hierarchy is determined according to the ratio between the overall weight deviation value and the reference correction parameters. The node area in each comparison level The reference weight deviation value between .

[0033] Furthermore, the process of acquiring the traffic fusion data includes:

[0034] Each node area is taken as the target area in turn, and the product of the target area's traffic data temporal weight and traffic data spatiotemporal weight at each data sampling level is normalized to serve as the overall weight of the target area at each data sampling level.

[0035] The local fusion value of the target area at each data sampling level is determined according to the product of the overall weight and the traffic fitting data of the target area at each data sampling level; the traffic fusion data of the target area at the fusion moment is determined according to the accumulated value of the local fusion values ​​of the target area at all data sampling levels.

[0036] In a second aspect, the present application provides a data fusion system for urban physical examination spatiotemporal big data, the system comprising:

[0037] A data acquisition and preprocessing module is used to obtain all traffic sampling data of each monitoring node in each node area at each data sampling level in the urban traffic historical data; wherein all monitoring nodes in each node area include one upper node and at least one lower node;

[0038] The first determination module is used to determine the corresponding time domain weight of the traffic data according to the traffic sampling data association between the upper node and each lower node in each node area under each data sampling level and the sampling frequency;

[0039] The second determination module is used to construct constraint conditions based on the temporal weight of the traffic data and the standard deviation distribution of the traffic sampling data of each node area in each data sampling layer at the fusion time, and determine the spatiotemporal weight of the traffic data of each node area in each data sampling layer;

[0040] The data fusion module is used to fuse the traffic data of each node area at each data sampling level according to the time domain weight and the spatiotemporal weight of the traffic data, and obtain the traffic fusion data of each node area at the fusion time.

[0041] In a third aspect, the present application provides a computer device comprising a memory and a processor. The memory is configured to store computer program code, and the processor is configured to call and execute the computer program code from the memory to perform the method of the first aspect or any embodiment of the first aspect of the present application.

[0042] In a fourth aspect, the present application provides a computer program product, comprising a computer program code. When the computer program code is executed, the method of the first aspect or any embodiment of the first aspect of the present application is performed.

[0043] In a fifth aspect, the present application provides a computer-readable storage medium, which stores computer program code. When the computer program code is executed, it performs the method of the first aspect of the present application or any embodiment of the first aspect.

[0044] This application has the following beneficial effects:

[0045] The present invention first performs a preliminary characterization of data authenticity based on the matching of data change value trends between upper and lower nodes in the node area at each data sampling level, thereby determining the time-domain weight of traffic data in each node area; then, based on the differences in data at different levels in the same area of ​​the city at the fusion time, constraints are constructed to perform spatial autocorrelation characterization, thereby determining the spatiotemporal weight of traffic data corresponding to the urban data of each node area at different data sampling levels; finally, the time-domain weight of traffic data and the spatiotemporal weight of traffic data are integrated to perform more robust adaptive data fusion, so that the traffic fusion data of each node area obtained after data fusion is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] Figure 1 A flow chart of a data fusion method for urban physical examination spatiotemporal big data provided by one embodiment of the present invention;

[0048] Figure 2 This is a structural diagram of a data fusion system for urban physical examination spatiotemporal big data provided by one embodiment of the present invention;

[0049] Figure 3 The present invention provides a schematic diagram of a computer device structure according to an embodiment of the present invention. DETAILED DESCRIPTION

[0050] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following, in combination with the accompanying drawings and preferred embodiments, a data fusion method and system for urban physical examination spatiotemporal big data proposed by the present invention, its specific implementation method, structure, characteristics and effects are described in detail as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment, and the specific features, structures or characteristics in one or more embodiments may be combined in any suitable form. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.

[0051] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0052] The following describes in detail a data fusion method and system for urban physical examination spatiotemporal big data provided by the present invention with reference to the accompanying drawings.

[0053] This application embodiment provides a data fusion method for urban physical examination spatiotemporal big data. Figure 1 , which shows a flow chart of a data fusion method for urban physical examination spatiotemporal big data provided by one embodiment of the present invention, the method comprising:

[0054] Step S101: Acquire all traffic sampling data of each monitoring node in each node area at each data sampling level in urban traffic historical data; wherein all monitoring nodes in each node area include one upper node and at least one lower node.

[0055] In a specific implementation of an embodiment of the present invention, the intersection where a first-class highway in a city intersects with other lower-class highways is taken as an upper-level node, and each upper-level node is directly connected to an intersection through other lower-class highways and the road distance is less than the preset intersection distance as the corresponding lower-level node; the area composed of each upper-level node and the corresponding lower-level nodes is taken as a node area; that is, the node area corresponds to the area corresponding to the intersection, and the traffic data of the intersection is monitored to more accurately reflect the traffic conditions in the city; it should be noted that the implementer can adjust the highway level according to the specific implementation environment, for example, the intersection where a second-class highway in a city intersects with other lower-class highways is taken as the upper-level node; and the traffic data analyzed by the embodiment of the present invention is vehicle flow data, and the traffic data type can be adjusted according to the specific implementation environment, for example, to vehicle speed data, pedestrian flow data, etc., which will not be further elaborated here. In a specific implementation of an embodiment of the present invention, the preset intersection distance is set to 1km, which can be adjusted according to the specific implementation environment so that each node area includes at least one lower-level node.

[0056] In a specific implementation of an embodiment of the present invention, the data sampling levels include the provincial transportation department sampling level, the municipal transportation department sampling level, and the district transportation department sampling level, which can be adjusted according to the specific implementation environment; different data sampling levels have different data collection purposes and storage pressures, resulting in different resolutions or frequencies of data collection for the same node area; that is, the embodiment of the present invention obtains all traffic sampling data of each monitoring node in each node area at each data sampling level through a database for storing urban traffic historical data at each data sampling level.

[0057] Step S102: determining the corresponding time domain weight of the traffic data according to the traffic sampling data association between the upper node and each lower node in each node area at each data sampling level and the sampling frequency.

[0058] Considering that the collected data between a parent node and its corresponding child nodes may differ, for example, the intersection of a parent node corresponding to a first-level highway may connect to child nodes corresponding to second-level, third-level, or even numerous fourth-level highways. The traffic flow at these child nodes can more accurately describe the traffic conditions in the area where the parent node is located. However, since a parent node approximates the traffic conditions in its area with a single node, it is necessary to smooth the traffic flow data of the numerous child nodes. Furthermore, based on the data association between the parent and child nodes and the sampling frequency that reflects the precision of the collected information, a preliminary determination of the temporal weight of the traffic data for each node area at each data sampling level is made.

[0059] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the time-domain weight of traffic data includes:

[0060] Obtain all traffic multi-source data for each monitoring node. In a specific implementation of an embodiment of the present invention, considering that different departments may simultaneously conduct traffic data statistics at each monitoring node, all traffic flow data collected by the traffic data statistics department of the street where each monitoring node is located and the traffic data departments of other streets are further used as traffic multi-source data; traffic multi-source data is the data with the highest resolution for each monitoring node, that is, the most detailed data. It should be noted that the corresponding time period of all traffic multi-source data and all traffic sampling data in the embodiment of the present invention is within one month before the current moment.

[0061] In each node region, based on the similarity of the time series change trends between the multi-source traffic data of the upper node and the traffic sampling data of each lower node at each data sampling level, the trend similarity of each node region at each data sampling level is determined. In a specific implementation of the embodiment of the present invention, the process of obtaining the trend similarity includes:

[0062] In each node region, the mean of the traffic sampling data of all lower-level nodes at each sampling moment of each data sampling layer is used as the corresponding lower-level traffic data; the lower-level traffic data of all sampling moments of each data sampling layer in each node region are arranged in chronological order and then curve-fitted to obtain the lower-level fitting curve; all the multi-source traffic data of each upper-level node are arranged in chronological order and then curve-fitted to determine the upper-level fitting curve; based on the dynamic time warping algorithm, the DTW distance between the lower-level fitting curve and the upper-level fitting curve is negatively correlated to determine the corresponding trend similarity. It should be noted that after the multi-source traffic data are arranged in chronological order, if all the multi-source traffic data at the same moment exist, the mean of all the multi-source traffic data at that moment is used as the final multi-source traffic data at the corresponding moment for curve fitting.

[0063] The lower-level traffic data refers to the data features after smoothing of the traffic sampling data of all lower-level nodes in each node area at each data sampling level; the lower-level traffic data is fitted into a curve and compared with the upper-level fitting curve. The higher the similarity between the two corresponding curves, that is, the smaller the DTW distance, the more consistent the trend of the traffic data of the lower-level node at the corresponding data sampling level is with that of the upper-level node, indicating that the data provided by the node area in the corresponding data sampling level is more real, that is, the greater the trend similarity, the greater the time domain weight of the traffic data in the corresponding node area at the corresponding data sampling level should be. In a specific implementation of an embodiment of the present invention, dynamic time warping and curve fitting are technical means well known to those skilled in the art, and are not further defined or elaborated here.

[0064] For each data sampling level, since the data collection time period in the embodiment of the present invention is fixed, the greater the amount of traffic sampling data, the higher the information density of the node area at that data sampling level, the more precise the collected data, and therefore the closer it is to the real data. Therefore, the corresponding traffic data temporal weight should be larger. Therefore, the product of the total amount of traffic sampling data collected by the parent node of each node area at each data sampling level and the trend similarity is normalized to determine the traffic data temporal weight of each node area at each data sampling level.

[0065] In a specific implementation of the embodiment of the present invention, the process of obtaining the time-domain weight of traffic data is expressed by the formula: ;in, For the The node area is in The time domain weight of traffic data at each data sampling level; For the The node area is in The DTW distance between the lower fitting curve at each data sampling level and the upper fitting curve of the corresponding upper node; is an exponential function with a natural constant as its base; For the The node area is in Trend similarity at each data sampling level; For the The parent node of the node area is in The total number of all traffic sampling data collected at each data sampling level; is a linear normalization function.

[0066] Step S103: At the fusion moment, constraint conditions are constructed based on the temporal weight of the traffic data combined with the standard deviation distribution of the traffic sampling data of each node area in each data sampling layer to determine the spatiotemporal weight of the traffic data of each node area in each data sampling layer.

[0067] At a higher data sampling level, the focus is on the overall data curve within a node area, while at a lower level, the data describes more detailed changes. Therefore, we further integrate these data through hierarchical modeling and coordination of multi-level spatial data to determine the spatiotemporal weight of traffic data in each node area at each data sampling level.

[0068] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the spatiotemporal weight of traffic data includes:

[0069] Based on all traffic sampling data of each node area in each data sampling layer, the traffic fitting data of each monitoring node at the fusion time is determined by Kriging interpolation; at the fusion time, based on the actual deviation of the traffic fitting data, the residual of each monitoring node in each data sampling layer at the fusion time is determined.

[0070] In one specific implementation of an embodiment of the present invention, the residual acquisition process includes: within each node region, taking the difference between the multi-source traffic data and the corresponding traffic fitting data at each monitoring node at the time of fusion as the residual for each monitoring node at each data sampling level at the time of fusion. It should be noted that Kriging interpolation is a well-known technical method for those skilled in the art and will not be further described here. The traffic fitting data represents the traffic data at the monitoring node under the global distribution, while the multi-source traffic data represents the accurate traffic data. Therefore, the residual at each monitoring node at each data sampling level is calculated based on the deviation between the two. For each data sampling level, the more similar the overall residuals of each node region, the more similar the spatiotemporal weights of the traffic data should be distributed at the corresponding data sampling level. Furthermore, for each node region at different data sampling levels, if the spatial distribution of the node regions is similar, that is, the spatial autocorrelation is high, then the weight during fusion should be higher, that is, the spatiotemporal weight of the traffic data should be greater. Therefore, after combining the above analysis to construct constraints, the kriging equation is combined to determine the spatiotemporal weight of the traffic data for each node region at each data sampling level.

[0071] Obtain the initial spatiotemporal weights for each node region at each data sampling level at the fusion time. The traffic prediction data obtained by weighted averaging the corresponding traffic fitting data at each data sampling level corresponding to each node region at the fusion time satisfies the conditions of unbiasedness and minimum variance. Construct mathematical constraints between the traffic fitting data and the traffic prediction data. Based on the conditions of unbiasedness and minimum variance, the determination of the initial spatiotemporal weights can avoid systematic overestimation or underestimation of traffic data, minimize the fluctuation range of prediction errors, and achieve statistical optimality. Then, based on the initial spatiotemporal weights, analysis is performed to determine the final required spatiotemporal weights for traffic data. Furthermore, the first constraint condition is determined based on the residual dispersion of each node region at each data sampling level and the deviation of the initial spatiotemporal weight of each node region compared to the initial spatiotemporal weights of all node regions.

[0072] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the first constraint condition includes:

[0073] At the fusion moment, the standard deviation of the residuals of all monitoring nodes in all node areas in each data sampling layer is used as the corresponding distribution fluctuation tolerance threshold; the mean of the initial spatiotemporal weights of all node areas in each data sampling layer is used as the corresponding reference average weight; for all node areas, the smaller the standard deviation of the residuals of all monitoring nodes, the more similar the overall residual sizes of each node area are, and the corresponding node areas should have similar traffic data spatiotemporal weights distributed under the corresponding data sampling layers; and under each data sampling layer, the smaller the deviation between the initial spatiotemporal weights of each node area and the overall initial spatiotemporal weights is, the closer the initial spatiotemporal weights between the node areas are. Therefore, the first constraint condition is that the difference between the initial spatiotemporal weight of each node area in each data sampling layer and the reference average weight is less than or equal to the fluctuation tolerance threshold.

[0074] Further obtain node area and node area , node area and node area are any nodes in all node regions, and the node regions and node area Can be the same node area; according to the node area and node area The second constraint condition is determined based on the similarity of the residual space distribution between each data sampling level and the initial spatiotemporal weight deviation.

[0075] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the second constraint condition includes:

[0076] Take any data sampling level as the reference level and the data sampling level outside the reference level as the comparison level of the reference level; calculate the node area in the reference level at the fusion time The residual space and the corresponding node area in each comparison level It should be noted that the calculation of the Moran index is a technical means well known to those skilled in the art and will not be further described here.

[0077] The larger the Moran index is, the smaller the node area in the reference level is. The residual space and the corresponding node area in each comparison level The more similar the residual space distribution is, the greater the weight of the corresponding node area should be, that is, the node area in the reference level Compare with the corresponding node area in the hierarchy The higher the spatial autocorrelation between them, the higher the spatial autocorrelation between them. Under the constraint of the cross-level threshold, the node area needs to be With node area The closer the spatiotemporal weights are to the same level at different data sampling levels, the larger the spatiotemporal weights will be assigned. The initial spatiotemporal weights and the corresponding node areas in each comparison level The deviation between the initial spatiotemporal weights is used to determine the corresponding reference weight deviation value.

[0078] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the reference weight deviation value includes:

[0079] Based on the node area in the reference hierarchy The initial spatiotemporal weights and the corresponding node areas in each comparison level The square of the difference between the initial spatiotemporal weights of the reference layer is used to determine the overall weight deviation value; according to the node area in the reference layer The initial spatiotemporal weights and the corresponding node areas in each comparison level The reference correction parameter is determined by multiplying the initial spatiotemporal weights of the reference layer by the product of the overall weight deviation value and the reference correction parameter; the node area in the reference layer is determined according to the ratio between the overall weight deviation value and the reference correction parameter. The node area in each comparison level The reference weight deviation value between .

[0080] Among them, the smaller the overall weight deviation value, the better the node area in the reference level. The initial spatiotemporal weights and the corresponding node areas in each comparison level The closer the initial spatiotemporal weights are to the reference level, the larger the reference correction parameter is, indicating that the node area in the reference level is closer to the reference level. The initial spatiotemporal weights and the corresponding node areas in each comparison level The larger the values ​​of the two parameters of the initial spatiotemporal weight are, the larger the reference weight deviation value obtained by the ratio between the overall weight deviation value and the reference correction parameter is. The initial spatiotemporal weights and the corresponding node areas in each comparison level The larger the value of the initial spatiotemporal weight is, the closer the value is.

[0081] Then, according to the larger the Moran index, the node area should be With node area The more consistent the spatiotemporal weights are between different data sampling levels, the larger the spatiotemporal weights are assigned. The larger the Moran index is, the smaller the reference weight deviation should be. Therefore, the product of the Moran index and the reference weight deviation is further used as the node area in the reference layer at the fusion moment. The node area in each comparison level The local constraint value between them; according to the node area in the reference layer at the fusion time And the corresponding node areas in all comparison levels The cumulative value of the local constraint values ​​between the two determines the reference level in the node area at the fusion moment With node area Reference constraint values ​​between; based on all data sampling levels in the node area With node area The cumulative value of the reference constraint values ​​between the nodes determines the node area With node area The overall constraint value between them; the overall constraint value is less than or equal to the preset cross-level constraint threshold as the second constraint condition.

[0082] In a specific implementation of an embodiment of the present invention, the preset cross-level constraint threshold is set to 0.1, which can be adjusted according to the specific implementation environment. With the help of the preset cross-level constraint threshold, the second constraint condition is satisfied. When the Moran index is larger, the spatiotemporal weights between the two node areas are more consistent between different data sampling levels, and the larger the spatiotemporal weights are given. The first constraint condition and the second constraint condition are further embedded in the joint kriging system to construct an extended kriging matrix equation; the alternating direction multiplier method is used to decompose the kriging matrix equation into two-step iterative solutions of intra-level weight update and cross-level constraint coordination to determine the final traffic data spatiotemporal weight corresponding to each node area in each data sampling level. It should be noted that the construction of the kriging matrix equation and the alternating direction multiplier method are technical means well known to those skilled in the art and will not be further elaborated here.

[0083] Step S104: performing data fusion based on the time domain weight and the spatiotemporal weight of the traffic data of each node area at each data sampling level to obtain the traffic fusion data of each node area at the fusion time.

[0084] After determining the time domain weight of traffic data and the spatiotemporal weight of traffic data that represent the fusion weight, the data fusion process is further carried out; that is, data fusion is performed according to the time domain weight of traffic data and the spatiotemporal weight of traffic data of each node area at each data sampling level to obtain the traffic fusion data of each node area at the fusion time.

[0085] Preferably, in some possible implementation methods of the embodiments of the present invention, the process of acquiring traffic fusion data includes: taking each node area as the target area in turn, normalizing the product between the time domain weight of the traffic data and the spatiotemporal weight of the traffic data of the target area at each data sampling level, and using it as the overall weight of the target area at each data sampling level; determining the local fusion value of the target area at each data sampling level based on the product between the overall weight and the traffic fitting data of the target area at each data sampling level; and determining the traffic fusion data of the target area at the fusion moment based on the accumulated value of the local fusion values ​​of the target area at all data sampling levels.

[0086] The temporal weights and spatiotemporal weights of the traffic data of the target area at each data sampling level are combined and normalized by multiplication, thereby comprehensively representing the fusion weights in the data fusion process through the overall weight to perform data fusion; and traffic fitting data is used during data fusion to avoid the situation where some data sampling levels do not have traffic sampling data at the fusion time due to different sampling frequencies, resulting in the inability to perform data fusion. In a specific implementation of an embodiment of the present invention, the process of obtaining traffic fusion data is expressed by the formula: ;in, For the fusion moment Lower target area Traffic fusion data; is the number of data sampling levels; Target area In the The time domain weight of traffic data at each data sampling level; For the fusion moment Lower target area In the The spatiotemporal weight of traffic data at each data sampling level; For the fusion moment Lower target area In the The overall weight of each data sampling level; for function, through The function is normalized so that all data sampling levels are in the target area The sum of all corresponding overall weights is 1, which makes the traffic fusion data obtained after weighting in combination with the cumulative method more accurate; For the fusion moment Lower target area In the Traffic fitting data of data sampling levels; For the fusion moment Lower target area In the The local fusion value of the data sampling level.

[0087] Further determine the traffic fitting data of each node area at each fusion moment; the frequency of the fusion moment can be set according to the needs of the specific implementation environment. This application is set to collect data every 5 minutes; further store the traffic fitting data of each node area at all fusion moments in the database as a data sample for urban physical examination.

[0088] In summary, a data fusion method for urban physical examination spatiotemporal big data performs a preliminary characterization of data authenticity based on the matching of data change value trends between the upper and lower nodes in the node area at each data sampling level, thereby determining the temporal weight of traffic data in each node area; then, based on the differences in data at different levels in the same area of ​​the city at the fusion moment, constraints are constructed to perform spatial autocorrelation characterization, thereby determining the spatiotemporal weight of traffic data corresponding to the urban data of each node area at different data sampling levels; finally, the temporal weight of traffic data and the spatiotemporal weight of traffic data are integrated to perform more robust adaptive data fusion, so that the traffic fusion data of each node area obtained after data fusion is more accurate.

[0089] This application also provides a data fusion system for urban physical examination spatiotemporal big data, please refer to Figure 2 , which shows a structural diagram of a data fusion system for urban physical examination spatiotemporal big data provided by an embodiment of the present invention. The system includes: a data acquisition and preprocessing module 201, a first determination module 202, a second determination module 203 and a data fusion module 204.

[0090] The data collection and preprocessing module 201 is used to obtain all traffic sampling data of each monitoring node in each node area at each data sampling level in the urban traffic historical data; wherein all monitoring nodes in each node area include one upper node and at least one lower node;

[0091] The first determination module 202 is configured to determine the corresponding time domain weight of the traffic data according to the traffic sampling data association between the upper node and each lower node in each node area at each data sampling level and the sampling frequency;

[0092] The second determination module 203 is used to construct constraint conditions based on the temporal weight of the traffic data and the standard deviation distribution of the traffic sampling data of each node area in each data sampling layer at the fusion time, and determine the spatiotemporal weight of the traffic data of each node area in each data sampling layer;

[0093] The data fusion module 204 is used to perform data fusion based on the time domain weight and space-time weight of traffic data of each node area at each data sampling level to obtain traffic fusion data of each node area at the fusion time.

[0094] It should be noted that the system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the data fusion system for urban physical examination spatiotemporal big data and the data fusion method embodiment for urban physical examination spatiotemporal big data provided in the above embodiment are of the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0095] The present application also provides a computer device. Figure 3, which shows a schematic diagram of the structure of a computer device provided by an embodiment of the present invention, the computer device includes a memory 301, a processor 302, and a computer program 303 stored in the memory 301 and running on the processor 302, wherein when the processor 302 executes the computer program 303, the computer device can execute any of the data fusion methods for urban physical examination spatiotemporal big data introduced above.

[0096] An embodiment of the present application also provides a computer program product. When the computer program product is run on a computer device, the computer device can execute any of the data fusion methods for urban physical examination spatiotemporal big data introduced above.

[0097] An embodiment of the present application also provides a computer-readable storage medium, which stores computer program code. When the computer program code runs on a computer device, the computer device can execute any of the data fusion methods for urban physical examination spatiotemporal big data introduced above.

[0098] In the embodiments provided in the present application, it should be understood that the provided computer devices, computer program products and computer-readable storage media are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the methods provided above and will not be repeated here.

[0099] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0100] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A data fusion method for urban physical examination spatiotemporal big data, characterized by: The method comprises: Obtain all traffic sampling data for each monitoring node in each node area at each data sampling level in urban traffic historical data; wherein all monitoring nodes in each node area include one upper node and at least one lower node; intersections where a first-level highway in the city intersects with other lower-level highways are regarded as upper nodes, and intersections to which each upper node is directly connected by other lower-level highways and the road distance is less than a preset intersection distance are regarded as corresponding lower nodes; and the area formed by each upper node and its corresponding lower nodes is regarded as a node area; Determine the corresponding time domain weight of traffic data according to the traffic sampling data association between the upper node and each lower node in each node area under each data sampling level and the sampling frequency; At the fusion moment, the constraint conditions are constructed based on the temporal weight of the traffic data combined with the standard deviation distribution of the traffic sampling data of each node area in each data sampling layer, and the spatiotemporal weight of the traffic data of each node area in each data sampling layer is determined; Data fusion is performed based on the time domain weight and spatiotemporal weight of traffic data of each node area at each data sampling level to obtain the traffic fusion data of each node area at the fusion time; The process of obtaining the time domain weight of the traffic data includes: Obtain all traffic multi-source data for each monitoring node; in each node area, determine the trend similarity of each node area at each data sampling level based on the similarity of the time series change trends between the traffic multi-source data of the upper node and the traffic sampling data of each lower node at each data sampling level; The product of the total number of all traffic sampling data collected by the upper node of each node area at each data sampling level and the trend similarity is normalized to determine the time domain weight of the traffic data of each node area at each data sampling level.

2. The data fusion method for urban physical examination spatiotemporal big data according to claim 1 is characterized in that: The process of obtaining the trend similarity includes: In each node area, the mean of the traffic sampling data of all lower-level nodes at each sampling time of each data sampling layer is used as the corresponding lower-level traffic data; Arrange the lower-level traffic data of each node area at all sampling moments at each data sampling level in chronological order and perform curve fitting to obtain the lower-level fitting curve; Arrange all traffic multi-source data of each upper node in chronological order and perform curve fitting to determine the upper fitting curve; Based on the dynamic time warping algorithm, the DTW distance between the lower-level fitting curve and the upper-level fitting curve is negatively correlated to determine the corresponding trend similarity.

3. The data fusion method for urban physical examination spatiotemporal big data according to claim 1 is characterized in that: The process of obtaining the spatiotemporal weight of traffic data includes: Based on all traffic sampling data of each node area in each data sampling layer, the traffic fitting data of each monitoring node at the fusion time is determined by Kriging interpolation; at the fusion time, based on the actual deviation of the traffic fitting data, the residual of each monitoring node in each data sampling layer at the fusion time is determined; Obtain the initial spatiotemporal weight of each node area in each data sampling layer at the fusion time; wherein, the traffic prediction data obtained by weighted averaging the corresponding traffic fitting data of the initial spatiotemporal weights in each data sampling layer corresponding to each node area at the fusion time satisfies the conditions of unbiasedness and minimum variance; Determine the first constraint condition based on the discrete residuals in each node region at each data sampling level and the deviation of the initial spatiotemporal weight of each node region compared to the initial spatiotemporal weights of all node regions; Get node area and node area , the node area and node area are any nodes in all node regions, and the node regions and node area Can be the same node area; according to the node area and node area The second constraint condition is determined based on the similarity of the residual space distribution between each data sampling level and the initial spatiotemporal weight deviation; The first and second constraints are embedded in the joint kriging system to construct an extended kriging matrix equation. The alternating direction multiplier method is used to decompose the kriging matrix equation into a two-step iterative solution of intra-level weight update and cross-level constraint coordination to determine the spatiotemporal weight of traffic data for each node area in each data sampling level.

4. The data fusion method for urban physical examination spatiotemporal big data according to claim 3 is characterized in that: The process of obtaining the residual includes: In each node area, the difference between the traffic multi-source data of each monitoring node at the fusion moment and the corresponding traffic fitting data is used as the residual of each monitoring node in each data sampling layer at the fusion moment.

5. The data fusion method for urban physical examination spatiotemporal big data according to claim 3 is characterized in that: The process of obtaining the first constraint condition includes: At the fusion moment, the standard deviation of the residuals of all monitoring nodes in all node areas in each data sampling level is used as the corresponding allocation fluctuation tolerance threshold; the mean of the initial spatiotemporal weights of all node areas in each data sampling level is used as the corresponding reference average weight; and the difference between the initial spatiotemporal weight of each node area in each data sampling level and the reference average weight is less than or equal to the fluctuation tolerance threshold as the first constraint condition.

6. The data fusion method for urban physical examination spatiotemporal big data according to claim 3 is characterized in that: The process of obtaining the second constraint condition includes: Using any data sampling level as the reference level, and using data sampling levels other than the reference level as the comparison level of the reference level; Calculate the node area in the reference layer at the time of fusion The residual space and the corresponding node area in each comparison level Moran index between the residual spaces; according to the node area in the reference level The initial spatiotemporal weights and the corresponding node areas in each comparison level The deviation between the initial spatiotemporal weights is used to determine the corresponding reference weight deviation value; the product between the Moran index and the reference weight deviation value is used as the node area in the reference layer at the fusion moment The node area in each comparison level The local constraint value between them; according to the node area in the reference layer at the fusion time And the corresponding node areas in all comparison levels The cumulative value of the local constraint values ​​between the two determines the reference level in the node area at the fusion moment With node area Reference constraint values ​​between; based on all data sampling levels in the node area With node area The cumulative value of the reference constraint values ​​between the nodes determines the node area With node area The overall constraint value between the two levels is less than or equal to the preset cross-level constraint threshold as the second constraint condition.

7. The data fusion method for urban physical examination spatiotemporal big data according to claim 6 is characterized in that: The process of obtaining the reference weight deviation value includes: Based on the node area in the reference hierarchy The initial spatiotemporal weights and the corresponding node areas in each comparison level The square of the difference between the initial spatiotemporal weights of the reference layer is used to determine the overall weight deviation value; according to the node area in the reference layer The initial spatiotemporal weights and the corresponding node areas in each comparison level The product of the initial spatiotemporal weights of the reference correction parameters is used to determine the reference correction parameters; the node area in the reference hierarchy is determined according to the ratio between the overall weight deviation value and the reference correction parameters. The node area in each comparison level The reference weight deviation value between .

8. The data fusion method for urban physical examination spatiotemporal big data according to claim 3 is characterized in that: The process of acquiring the traffic fusion data includes: Each node area is taken as the target area in turn, and the product of the target area's traffic data temporal weight and traffic data spatiotemporal weight at each data sampling level is normalized to serve as the overall weight of the target area at each data sampling level. The local fusion value of the target area at each data sampling level is determined according to the product of the overall weight and the traffic fitting data of the target area at each data sampling level; the traffic fusion data of the target area at the fusion moment is determined according to the accumulated value of the local fusion values ​​of the target area at all data sampling levels.

9. A data fusion system for urban physical examination spatiotemporal big data, characterized by: The system comprises: The data acquisition and preprocessing module is used to obtain all traffic sampling data from each monitoring node in each node area at each data sampling level in the urban traffic historical data; wherein all monitoring nodes in each node area include one upper node and at least one lower node; the intersections where the first-level highway in the city intersects with other lower-level highways are regarded as upper nodes, and the intersections where each upper node is directly connected to other lower-level highways and the road distance is less than a preset intersection distance are regarded as corresponding lower nodes; the area formed by each upper node and its corresponding lower nodes is regarded as a node area; The first determination module is used to determine the corresponding time domain weight of the traffic data according to the traffic sampling data association between the upper node and each lower node in each node area under each data sampling level and the sampling frequency; The process of obtaining the time domain weight of the traffic data includes: Obtain all traffic multi-source data for each monitoring node; in each node area, determine the trend similarity of each node area at each data sampling level based on the similarity of the time series change trends between the traffic multi-source data of the upper node and the traffic sampling data of each lower node at each data sampling level; Normalizing the product of the total number of all traffic sampling data collected by the upper node of each node area at each data sampling level and the trend similarity to determine the time domain weight of the traffic data of each node area at each data sampling level; The second determination module is used to construct constraint conditions based on the temporal weight of the traffic data and the standard deviation distribution of the traffic sampling data of each node area in each data sampling layer at the fusion time, and determine the spatiotemporal weight of the traffic data of each node area in each data sampling layer; The data fusion module is used to fuse the traffic data of each node area at each data sampling level according to the time domain weight and the spatiotemporal weight of the traffic data, and obtain the traffic fusion data of each node area at the fusion time.

Citation Information

Patent Citations

  • A RNN based Spatio Temporal Data Mining model for urban road Planning

    AU2020104112A4

  • A traffic data fusion system and the related method for providing a traffic state for a network of roads

    WO2016096226A1