Spatiotemporal co-occurrence analysis method and system based on multi-source data fusion

By constructing a spatiotemporal reference coordinate system and a co-occurrence pattern mining model, the spatiotemporal difference problem of multi-source heterogeneous data is solved, efficient spatiotemporal co-occurrence analysis and anomaly detection are achieved, data collection parameters are optimized, and data utilization efficiency and accuracy are improved.

CN120354374BActive Publication Date: 2025-09-09杨群鹏
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510846427.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-09
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively handling the spatiotemporal differences of multi-source heterogeneous data, and are unable to fully explore the potential relationships and patterns between data. In addition, anomaly detection and data acquisition parameter optimization suffer from high false alarm rates and waste of resources.

Method used

By constructing a spatiotemporal reference coordinate system for data alignment, generating a spatiotemporal weight matrix and a set of associated feature vectors, using a spatiotemporal sliding window for feature fusion, building a co-occurrence pattern mining model and performing anomaly detection, and optimizing data collection parameters.

Benefits of technology

It achieves effective integration of multi-source heterogeneous data and spatiotemporal co-occurrence analysis, improves data availability and analytical value, reduces the false alarm rate of anomaly detection, optimizes the data collection process, and improves resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354374B_ABST
    Figure CN120354374B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing and analysis technology, and discloses a spatiotemporal co-occurrence analysis method and system for multi-source data fusion. The method first acquires multi-source heterogeneous data in a target area and performs spatiotemporal calibration, constructs a spatiotemporal reference coordinate system for spatiotemporal alignment, and generates a spatiotemporal weight matrix and a set of associated feature vectors. Spatiotemporal sliding window segmentation and feature fusion are used to obtain a spatiotemporal correlation feature map, and a co-occurrence pattern mining model is trained to extract a spatiotemporal co-occurrence pattern feature set. An anomaly detection model is constructed and verified to obtain an abnormal co-occurrence pattern set, and a parameter optimization model is constructed to dynamically adjust data acquisition parameters. Finally, the coordinate system is updated based on the optimal parameters and real-time data, and an analysis report is generated. The method and system can effectively process multi-source heterogeneous data, mine spatiotemporal co-occurrence patterns, accurately detect anomalies, optimize data acquisition parameters, and provide strong support for decision-making in related fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing and analysis, and in particular to a spatiotemporal co-occurrence analysis method and system for multi-source data fusion. Background Art

[0002] In today's digital age, data, as a key production factor, is experiencing explosive growth in both volume and variety. In numerous fields, such as urban management, environmental monitoring, and intelligent transportation, the need for in-depth understanding and analysis of diverse phenomena and activities within a specific area necessitates the acquisition and processing of multi-source data. However, this multi-source data is often heterogeneous, presenting numerous challenges to its effective utilization.

[0003] Geospatial data describes the location, shape, and distribution of geographic entities, but it comes in a variety of formats, such as vector and raster data, with significant differences in data structure and storage methods. Device perception data is collected by various sensor devices, with varying frequency, accuracy, and data types. For example, the data collected by temperature and humidity sensors vary in format and meaning. Event log data typically exists in text or structured tables, with inconsistent levels of detail and time stamping.

[0004] The problem of spatiotemporal discrepancies in multi-source heterogeneous data is prominent. Due to the varying deployment locations and timings of data collection equipment, data cannot be directly aligned and correlated across time and space. For example, in urban traffic monitoring, cameras at different intersections may collect data at different times and cover varying spatial ranges, making it difficult to comprehensively analyze spatiotemporal variations in traffic flow. Traditional data processing methods struggle to effectively handle this complex, multi-source heterogeneous data and are unable to fully explore the underlying relationships and patterns between the data.

[0005] In terms of data fusion, existing technologies primarily focus on simple data splicing or fusion based on a single feature, failing to fully consider the spatiotemporal characteristics and multi-source heterogeneity of data. In spatiotemporal co-occurrence analysis, there is a lack of efficient methods to identify temporal and spatial co-occurrence patterns between different data sources, making it difficult to uncover valuable information hidden within the data. For example, in urban environmental monitoring, it is impossible to quickly and accurately identify the spatiotemporal co-occurrence patterns between air quality anomalies and specific industrial activities or meteorological conditions.

[0006] When it comes to anomaly detection, traditional methods are mostly based on single data sources or simple threshold judgments, resulting in high false alarm rates and inability to adapt to complex and changing data environments. Existing technologies are particularly inadequate for detecting anomalies in complex spatiotemporal co-occurrence patterns. For example, in smart grid monitoring, it is difficult to accurately detect abnormal operating conditions in the power system through comprehensive analysis of data from multiple power equipment and power event data.

[0007] When it comes to optimizing data collection parameters, existing technologies often rely on manual experience or simple rule-based implementations, failing to fully leverage the information in the data for dynamic adjustments. This leads to wasted resources and incomplete data collection during the data collection process. For example, in IoT device data collection, improperly set collection frequencies can lead to the generation of large amounts of redundant data while preventing timely acquisition of critical information. In summary, there is an urgent need for a method and system that can effectively process multi-source heterogeneous data, enabling efficient spatiotemporal co-occurrence analysis, accurate anomaly detection, and intelligent data collection parameter optimization. Summary of the Invention

[0008] The purpose of the present invention is to provide a spatiotemporal co-occurrence analysis method and system for multi-source data fusion to solve the problems raised in the above background technology.

[0009] To achieve the above objectives, the present invention provides the following technical solution: a spatiotemporal co-occurrence analysis method for multi-source data fusion, the method comprising:

[0010] Acquire multi-source heterogeneous data of the target area and perform spatiotemporal calibration, where the multi-source heterogeneous data includes at least geospatial data, device perception data, and event record data;

[0011] Construct a spatiotemporal reference coordinate system and perform spatiotemporal alignment on multi-source heterogeneous data to obtain a spatiotemporal aligned data sequence. Based on the spatiotemporal aligned data sequence, a spatiotemporal weight matrix and a set of associated feature vectors are generated.

[0012] The spatiotemporal weight matrix is ​​divided into multiple spatiotemporal data sub-blocks through a spatiotemporal sliding window. The spatiotemporal data sub-blocks are fused based on the associated feature vector set to obtain a spatiotemporal correlation feature map.

[0013] Construct a co-occurrence pattern mining model and train it through the spatiotemporal correlation feature map to obtain a co-occurrence pattern recognition model, and then extract the spatiotemporal co-occurrence pattern feature set;

[0014] An anomaly detection model is constructed based on the spatiotemporal co-occurrence pattern feature set, and the anomaly detection model is verified through a sliding verification mechanism to obtain an anomaly co-occurrence pattern set;

[0015] Based on the abnormal co-occurrence pattern set and data collection parameters, a parameter optimization model is constructed and the data collection parameters are dynamically adjusted to obtain the optimal data collection parameters for the target area;

[0016] Based on the optimal data collection parameters and real-time monitoring data, the spatiotemporal reference coordinate system is updated and a spatiotemporal co-occurrence analysis report is generated.

[0017] Preferably, the spatiotemporal weight matrix includes at least time dimension weight, space dimension weight and data source weight; the associated feature vector set includes at least spatial proximity vector, time synchronization vector and data association vector; the spatiotemporal co-occurrence pattern feature set includes at least spatial aggregation pattern, time continuity pattern and cross-source association pattern.

[0018] Preferably, the step of constructing a spatiotemporal reference coordinate system and performing spatiotemporal alignment on multi-source heterogeneous data to obtain a spatiotemporal aligned data sequence comprises the following steps:

[0019] Define the grid division rules and timestamp alignment rules of the spatiotemporal reference coordinate system based on the geographic boundaries and time span of the target area;

[0020] Perform spatial projection transformation on multi-source heterogeneous data to make its coordinate parameters consistent with the spatiotemporal reference coordinate system;

[0021] Normalize the timestamps of multi-source heterogeneous data to generate time series data with a unified time base;

[0022] Based on the spatial proximity threshold and temporal synchronization threshold, data interpolation and conflict resolution are performed on the normalized multi-source heterogeneous data to obtain a spatiotemporal aligned data sequence.

[0023] Preferably, generating a spatiotemporal weight matrix and a set of associated feature vectors based on the spatiotemporal alignment data sequence comprises the following steps:

[0024] Calculate the time dimension weight and space dimension weight based on the data source collection frequency, accuracy and coverage;

[0025] Calculate the data source weight based on the attribute similarity and historical correlation of the data source;

[0026] Combine the time dimension weight, space dimension weight and data source weight into a spatiotemporal weight matrix;

[0027] Perform spatial proximity calculation on the spatiotemporal aligned data sequence to generate a spatial proximity vector;

[0028] Calculate the time synchronization degree of the spatiotemporal alignment data sequence and generate a time synchronization degree vector;

[0029] The data association degree vector is generated based on the attribute association of the data source, and then combined into a set of association feature vectors.

[0030] Preferably, generating the spatial proximity vector comprises the following steps:

[0031] Calculate the spatial neighborhood coverage of each data point in the spatiotemporally aligned data series;

[0032] Count the number of data points within the preset neighborhood radius of each data point to obtain the spatial density value;

[0033] Generate a spatial proximity vector based on the spatial density value and neighborhood overlap.

[0034] Preferably, the construction of the co-occurrence pattern mining model and training through the spatiotemporal correlation feature map to obtain the co-occurrence pattern recognition model includes the following steps:

[0035] Construct a co-occurrence pattern mining model based on convolutional neural networks;

[0036] Input the spatiotemporal correlation feature map into the co-occurrence pattern mining model for feature extraction to obtain potential co-occurrence pattern features;

[0037] The potential co-occurrence pattern features are classified by unsupervised clustering algorithm to obtain the spatiotemporal co-occurrence pattern feature set;

[0038] The classification results are fed back to the co-occurrence pattern mining model for parameter optimization to obtain the co-occurrence pattern recognition model.

[0039] Preferably, the pattern classification of potential co-occurrence pattern features by an unsupervised clustering algorithm comprises the following steps:

[0040] Identify high-density regions in potential co-occurrence pattern features based on density clustering algorithm;

[0041] Merge patterns in high-density areas to generate a set of candidate co-occurrence patterns;

[0042] The clustering effect of the candidate co-occurrence pattern set is evaluated by the silhouette coefficient, and the optimal clustering result is selected as the spatiotemporal co-occurrence pattern feature set.

[0043] Preferably, the construction of an anomaly detection model based on a spatiotemporal co-occurrence pattern feature set comprises the following steps:

[0044] Divide the spatiotemporal co-occurrence pattern feature set into a training set and a validation set;

[0045] Build an initial anomaly detection model based on the isolation forest algorithm and train the model using the training set;

[0046] Use the validation set to perform sliding window validation on the initial anomaly detection model and calculate the anomaly confidence score;

[0047] Adjust the model threshold parameters according to the anomaly confidence score to obtain the optimized anomaly detection model;

[0048] The real-time monitoring data is input into the optimized anomaly detection model, and a set of anomaly co-occurrence patterns is output.

[0049] Preferably, the construction of the parameter optimization model and the dynamic adjustment of the data acquisition parameters include the following steps:

[0050] Defining decision variables of a parameter optimization model, wherein the decision variables include at least acquisition frequency, storage capacity, and transmission interval;

[0051] Constructing an optimization objective function based on the historical distribution of the abnormal co-occurrence pattern set, wherein the objective function at least includes data coverage integrity and resource consumption cost;

[0052] Setting constraints, wherein the constraints include at least a device workload upper limit and a data transmission delay threshold;

[0053] The parameter optimization model is iteratively solved through a heuristic search algorithm to obtain the optimal data acquisition parameters.

[0054] Preferably, the present invention further includes a spatiotemporal co-occurrence analysis system for multi-source data fusion, wherein the shampooing system includes:

[0055] A data acquisition and calibration module is used to acquire multi-source heterogeneous data of the target area, including at least geospatial data, device perception data, and event record data, and perform spatiotemporal calibration;

[0056] The spatiotemporal alignment processing module is used to construct a spatiotemporal reference coordinate system, perform spatiotemporal alignment on multi-source heterogeneous data, obtain a spatiotemporal aligned data sequence, and generate a spatiotemporal weight matrix and a set of associated feature vectors based on the spatiotemporal aligned data sequence;

[0057] The feature fusion graph construction module divides the spatiotemporal weight matrix into multiple spatiotemporal data sub-blocks through a spatiotemporal sliding window, and performs feature fusion on the spatiotemporal data sub-blocks based on the associated feature vector set to obtain a spatiotemporal correlation feature graph;

[0058] The co-occurrence pattern mining module builds a co-occurrence pattern mining model and trains it through the spatiotemporal correlation feature map to obtain a co-occurrence pattern recognition model, and then extracts the spatiotemporal co-occurrence pattern feature set;

[0059] The anomaly detection module builds an anomaly detection model based on the spatiotemporal co-occurrence pattern feature set, verifies the anomaly detection model through a sliding verification mechanism, and obtains an anomaly co-occurrence pattern set;

[0060] The parameter optimization module builds a parameter optimization model based on the set of abnormal co-occurrence patterns and data collection parameters, and dynamically adjusts the data collection parameters to obtain the optimal data collection parameters for the target area;

[0061] The report generation module updates the spatiotemporal reference coordinate system and generates a spatiotemporal co-occurrence analysis report based on the optimal data collection parameters and real-time monitoring data.

[0062] Compared with the prior art, the present invention has the following beneficial effects:

[0063] During data processing, the problem of data heterogeneity and spatiotemporal differences is effectively addressed by acquiring multi-source heterogeneous data from the target area, including geospatial data, device perception data, and event log data, and performing spatiotemporal calibration and alignment. This processing approach can integrate previously dispersed and heterogeneous data into a unified spatiotemporally aligned data sequence, laying a solid foundation for subsequent analysis. Compared to traditional data splicing or single feature fusion methods, this invention fully considers the spatiotemporal characteristics and multi-source heterogeneity of the data, greatly improving the data's usability and analytical value.

[0064] To mine the potential value of data, a spatiotemporal weight matrix and a set of associated feature vectors are generated based on spatiotemporal aligned data sequences. A spatiotemporal sliding window and feature fusion are then used to generate a spatiotemporal correlation feature map. A co-occurrence pattern mining model is then constructed for training and pattern classification, extracting spatiotemporal co-occurrence pattern feature sets such as spatial aggregation patterns, temporal persistence patterns, and cross-source correlation patterns. This enables in-depth exploration of the temporal and spatial co-occurrence patterns of different data sources, uncovering valuable information hidden within the data. For example, in urban environmental monitoring, the co-occurrence patterns between air quality anomalies and specific industrial activities and meteorological conditions can be quickly and accurately identified, providing strong data support for environmental governance.

[0065] Anomaly detection is a highlight of the present invention. An anomaly detection model is constructed based on a set of spatiotemporal co-occurrence pattern features, and is verified and optimized through a sliding verification mechanism. Compared with traditional anomaly detection methods based on a single data source or simple threshold judgment, this method has a lower false alarm rate and can better adapt to complex and changing data environments. In smart grid monitoring, with the help of the method of the present invention, it is possible to comprehensively analyze a variety of power equipment data and power consumption event data, accurately detect abnormal operating conditions of the power system, promptly discover potential fault risks, and ensure the stable operation of the power system.

[0066] Regarding data collection parameter optimization, a parameter optimization model is constructed based on the set of abnormal co-occurrence patterns and data collection parameters, dynamically adjusting parameters such as collection frequency, storage capacity, and transmission interval. This optimization process fully utilizes the information contained in the data, avoiding the resource waste and incomplete data collection caused by manual experience or simple rule-based settings. In IoT device data collection, the optimization method of this invention can rationally set the collection frequency, reduce the generation of redundant data, ensure timely acquisition of key information, improve the efficiency and quality of data collection, and achieve efficient resource utilization.

[0067] Finally, based on the optimal data collection parameters and real-time monitoring data, the spatiotemporal reference coordinate system is updated and a spatiotemporal co-occurrence analysis report is generated, providing users with comprehensive and accurate analysis results, helping them to gain an in-depth understanding of the situation in the target area so that they can make scientific and reasonable decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a working principle diagram of the spatiotemporal co-occurrence analysis method for multi-source data fusion according to the present invention;

[0069] Figure 2 Construct a flow chart for aligning data with a spatiotemporal reference coordinate system;

[0070] Figure 3 Flowchart for building and training a co-occurrence pattern mining model;

[0071] Figure 4 A flowchart for anomaly detection model based on co-occurrence patterns;

[0072] Figure 5 Flowchart for model building for data acquisition parameter optimization. DETAILED DESCRIPTION

[0073] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0074] See also Figure 1-5 The present invention provides a technical solution: a spatiotemporal co-occurrence analysis method for multi-source data fusion, the method comprising:

[0075] Acquire multi-source heterogeneous data from the target area and perform spatiotemporal alignment. This data includes at least geospatial data, device perception data, and event log data. In practice, specialized data acquisition equipment and interfaces are used to acquire data from various sources. For example, geospatial data can be acquired from a geographic information system (GIS), device perception data can be acquired from various sensor devices, and event log data can be acquired from an event management system. After acquisition, specific calibration algorithms and techniques are used to calibrate the temporal and spatial attributes of the data to ensure consistency across both spatial and temporal dimensions.

[0076] A spatiotemporal reference coordinate system is constructed and spatiotemporally aligned across multi-source heterogeneous data to produce a spatiotemporally aligned data sequence. A spatiotemporal weight matrix and associated feature vector set are generated based on the spatiotemporally aligned data sequence. The gridding rules and timestamp alignment rules for the spatiotemporal reference coordinate system are determined based on the actual geographic boundaries and time span of the target area. Spatial projection transformation techniques are used to convert the coordinate parameters of the multi-source heterogeneous data into a format consistent with the spatiotemporal reference coordinate system. The timestamps of the data are normalized to form time series data with a unified time reference. Data interpolation and conflict resolution are then performed on the normalized data based on predefined spatial proximity and temporal synchronization thresholds to produce a spatiotemporally aligned data sequence. Subsequently, temporal and spatial dimension weights are calculated based on factors such as data source acquisition frequency, accuracy, and coverage. Data source weights are calculated based on attribute similarity and historical correlation between data sources, and these weights are combined into a spatiotemporal weight matrix. At the same time, spatial proximity calculation is performed on the spatiotemporal aligned data sequence to generate a spatial proximity vector, time synchronization calculation is performed to generate a time synchronization vector, and data association vector is generated based on the attribute association of the data source, which are then combined into a set of association feature vectors.

[0077] The spatiotemporal weight matrix is ​​segmented using a spatiotemporal sliding window to form multiple spatiotemporal data sub-blocks. These sub-blocks are then fused based on a set of associated feature vectors to generate a spatiotemporal correlation feature map. A preset spatiotemporal sliding window is used to slide across the spatiotemporal weight matrix at a specific step size, segmenting it into multiple spatiotemporal data sub-blocks. For each spatiotemporal data sub-block, a specific feature fusion algorithm is applied, combining the individual vectors in the set of associated feature vectors, to fuse features from different dimensions, ultimately generating a spatiotemporal correlation feature map.

[0078] A co-occurrence pattern mining model is constructed and trained using the spatiotemporal correlation feature map to obtain a co-occurrence pattern recognition model, which then extracts a spatiotemporal co-occurrence pattern feature set. The co-occurrence pattern mining model is built based on a convolutional neural network. The generated spatiotemporal correlation feature map is input into the model for feature extraction, resulting in potential co-occurrence pattern features. An unsupervised clustering algorithm is used to classify the potential co-occurrence pattern features, resulting in a spatiotemporal co-occurrence pattern feature set. The classification results are fed back into the co-occurrence pattern mining model to optimize the model parameters, ultimately resulting in a co-occurrence pattern recognition model.

[0079] An anomaly detection model is constructed based on the spatiotemporal co-occurrence pattern feature set. This model is validated using a sliding validation mechanism to obtain a set of anomaly co-occurrence patterns. The spatiotemporal co-occurrence pattern feature set is divided into a training set and a validation set. An initial anomaly detection model is constructed using the isolation forest algorithm and trained using the training set. The initial anomaly detection model is validated using a sliding window on the validation set, and an anomaly confidence score is calculated. The model's threshold parameters are adjusted based on the anomaly confidence score to obtain an optimized anomaly detection model. Real-time monitoring data is input into the optimized anomaly detection model, which outputs a set of anomaly co-occurrence patterns.

[0080] Based on the set of abnormal co-occurrence patterns and data collection parameters, a parameter optimization model is constructed and the data collection parameters are dynamically adjusted to obtain the optimal data collection parameters for the target area. The decision variables of the parameter optimization model are defined, including at least the collection frequency, storage capacity, and transmission interval. An optimization objective function is constructed based on the historical distribution of the set of abnormal co-occurrence patterns. The objective function includes at least data coverage completeness and resource consumption cost. Constraints are set, including at least the upper limit of device workload and the data transmission delay threshold. The parameter optimization model is iteratively solved using a heuristic search algorithm to obtain the optimal data collection parameters.

[0081] Based on optimized data collection parameters and real-time monitoring data, the spatiotemporal reference coordinate system is updated and a spatiotemporal co-occurrence analysis report is generated. Using the acquired optimized data collection parameters and real-time monitoring data, the spatiotemporal reference coordinate system is updated to better reflect the current data situation. Based on the updated coordinate system and relevant data, a specialized report generation algorithm and template is used to generate a spatiotemporal co-occurrence analysis report.

[0082] The present invention will be further described below in conjunction with Examples 1 to 5: Example

[0083] After performing spatiotemporal alignment on multi-source heterogeneous data, the spatiotemporal reference coordinate system is constructed and then aligned. First, the gridding rules and timestamp alignment rules for the spatiotemporal reference coordinate system are determined based on the geographic boundaries and time span of the target area. Assume the target area is a city, with geographic boundaries encompassing all administrative divisions, and a time span of one month. For the gridding rules, the city can be divided into several appropriately sized grids based on its functional areas and geographic characteristics, with each grid representing a spatial unit. The timestamp alignment rule can be set to align data in hourly units, aligning all data timestamps to this precision.

[0084] Perform spatial projection conversion on multi-source heterogeneous data. If geospatial data uses a specific projection coordinate system, while the spatiotemporal reference coordinate system uses a universal geographic coordinate system, specialized geographic information processing software or algorithms are required to convert the coordinate parameters of the geospatial data into coordinates in the spatiotemporal reference coordinate system. Device perception data and event log data, if they contain spatial information, also require conversion using the same method.

[0085] Normalize the timestamps of multi-source heterogeneous data. If the timestamp format of geospatial data is "year-month-day hour:minute:second," the timestamp format of device perception data is "seconds (calculated from a certain starting point)," and the timestamp format of event log data is "YYYYMMDDHHMMSS," a unified time conversion method is needed to convert them all into time series data in hours. For example, extract the hours from data in the "year-month-day hour:minute:second" format, convert the seconds to the corresponding hours, and parse the hours from data in the "YYYYMMDDHHMMSS" format.

[0086] Based on spatial proximity thresholds and temporal synchronization thresholds, data interpolation and conflict resolution are performed on the normalized multi-source heterogeneous data. Assume that the spatial proximity threshold is set to 500 meters and the temporal synchronization threshold is set to 1 hour. If, at a certain point in time, a device senses that there are other data points within 500 meters of a data point, but the time difference is more than 1 hour, data interpolation is required based on the surrounding data points, such as using linear interpolation or other appropriate algorithms to supplement the missing data. If there is a data conflict, such as two data points at the same spatiotemporal location but different data values, conflict resolution is required based on data credibility or other rules to select the more reliable data value. Through these steps, a spatiotemporally aligned data sequence is ultimately obtained.

[0087] Taking a city's traffic management scenario as an example, this city's traffic management involves a variety of multi-source heterogeneous data, such as geospatial data (city road map data), device perception data (data collected by various sensors installed on the road, including traffic flow sensors, vehicle speed sensors, etc.) and event record data (traffic accident records, road construction records, etc.).

[0088] After acquiring this multi-source, heterogeneous data, we begin the process of temporal and spatial calibration. Geospatial data, obtained from urban geographic information databases, may have different map projections. Device perception data, collected in real time from sensors distributed across roads, may have inconsistent timestamps due to differences in the sensor's own clocks. Event record data, obtained from the traffic management department's event registration system, may also have different time formats and recording standards. Using professional geographic information processing tools and time synchronization algorithms, we calibrate the temporal and spatial attributes of this data to ensure initial consistency across the spatial and temporal dimensions.

[0089] Construct a spatiotemporal reference coordinate system and perform spatiotemporal alignment. Based on the city's geographic boundaries, including its various urban and suburban areas, and the time span, and assuming a week of data for analysis, determine the grid division rules and timestamp alignment rules for the spatiotemporal reference coordinate system. Spatially, the city is divided into 1-kilometer square grids based on its road layout and regional functions. Temporally, alignment is performed at 15-minute intervals.

[0090] Perform spatial projection conversion on multi-source heterogeneous data. If the city road map data originally uses the Gauss-Krüger projection, but the set spatiotemporal reference coordinate system uses the Web Mercator projection, it is necessary to use the projection conversion tool in the Geographic Information System (GIS) to convert the coordinate parameters of the geospatial data from the Gauss-Krüger projection to the Web Mercator projection to ensure consistency with the spatiotemporal reference coordinate system. The spatial information contained in device perception data and event log data is also processed using the same projection conversion method.

[0091] Normalize timestamps for heterogeneous data from multiple sources. For example, timestamps recorded by traffic flow sensors may be relative to device startup, while timestamps from speed sensors may be accurate to the second. Traffic accident records may use the "year / month / day hour:minute:second" format, while road construction records may use the "YYYY-MM-DDHH:MM" format. Develop a dedicated time conversion program to convert these different timestamp formats into unified time series data with 15-minute intervals. For example, convert relative times to absolute times corresponding to the system time, and then round them to the nearest 15-minute interval.

[0092] Based on spatial proximity thresholds and temporal synchronization thresholds, data interpolation and conflict resolution are performed on the normalized multi-source heterogeneous data. Assume that the spatial proximity threshold is set to 500 meters and the temporal synchronization threshold is set to 30 minutes. If, within a 15-minute interval, data from a speed sensor within 500 meters of a traffic flow sensor is missing, but data from other surrounding sensors is normal, linear interpolation or other appropriate algorithms can be used to supplement the missing speed data based on the changing trends of the surrounding sensor data. If traffic flow data recorded by two sensors at the same spatiotemporal location (the same grid and the same 15-minute interval) differ significantly, the more reliable data value is selected by comparing the historical data accuracy and device status of the two sensors to resolve the conflict. These steps ultimately yield a spatiotemporally aligned data sequence suitable for urban traffic management analysis, providing an accurate data foundation for subsequent traffic condition analysis, congestion prediction, and other tasks. Example

[0093] When generating the spatiotemporal weight matrix and associated feature vector set based on the spatiotemporally aligned data sequence, the time and space dimension weights are calculated. The time dimension weight is considered based on the data source's acquisition frequency. If one sensor acquires data every 10 minutes and another acquires data every hour, the sensor with the higher acquisition frequency will have a higher time dimension weight. This is because data with a higher acquisition frequency can reflect data changes more promptly and therefore has greater importance in the time dimension. The accuracy of the data source is also considered; data with higher accuracy will have a higher weight in the time dimension weight calculation. For example, a temperature sensor with an accuracy of 0.1°C will have a higher time dimension weight than one with an accuracy of 1°C. The spatial dimension weight is determined based on the data source's coverage. If one piece of geospatial data covers an entire city, while another piece of device-sensed data covers only a small area within the city, the geospatial data with the larger coverage will have a higher spatial dimension weight.

[0094] Data source weights are calculated based on the attribute similarity and historical correlation between the data sources. Assume that the data sources include weather station data, traffic flow data, and environmental monitoring data. If the temperature attributes in the weather station data and the air temperature in the environmental monitoring data are similar, and if, in the historical data, changes in temperature are accompanied by corresponding changes in some related indicators in the environmental monitoring data, this indicates a high historical correlation between the two data sources, and their data source weights will be relatively high.

[0095] The time dimension weights, space dimension weights, and data source weights are combined into a spatiotemporal weight matrix. Following certain rules, such as assigning the time dimension weights, space dimension weights, and data source weights to different columns or rows of a matrix, a multidimensional matrix is ​​formed. This matrix is ​​the spatiotemporal weight matrix.

[0096] When generating a spatial proximity vector, the spatial neighborhood coverage of each data point in the spatiotemporal alignment data sequence is calculated. A specific radius can be set with each data point as the center, and the area covered by this radius is the spatial neighborhood coverage of the data point. Count the number of data points within the preset neighborhood radius for each data point to obtain the spatial density value. For example, the preset neighborhood radius is 300 meters, and the number of statistical points within this range is counted. The greater the number, the greater the spatial density value. Generate a spatial proximity vector based on the spatial density value and neighborhood overlap. If the spatial density values ​​of two data points are both large and their neighborhood overlap is also high, it means that the two data points are highly spatially adjacent, and the corresponding element value in the spatial proximity vector will be larger.

[0097] Calculate the temporal synchronization of spatiotemporally aligned data sequences to generate a temporal synchronization vector. If the timestamp difference between two data points is small and within the set temporal synchronization threshold, their temporal synchronization is high, and the corresponding element value in the temporal synchronization vector is large. Generate a data association vector based on the attribute correlation of the data source. For example, if the temperature attribute in weather station data is related to the air temperature attribute in environmental monitoring data, the corresponding element value in their data association vector is large. Combine the spatial proximity vector, temporal synchronization vector, and data association vector into a set of association feature vectors.

[0098] The implementation of Example 2 is described in detail using a monitoring scenario in a smart agricultural park as an example. In this smart agricultural park, there are multiple data sources, including meteorological data provided by a weather station (a type of geospatial data), soil data collected by various sensors distributed throughout the fields (device perception data), and agricultural activity records (event record data).

[0099] When generating the spatiotemporal weight matrix and associated eigenvector set based on the spatiotemporally aligned data sequence, the temporal and spatial weights are first calculated. Regarding the temporal weight, the weather station collects data every 15 minutes, while the soil moisture sensor collects data every hour. Because meteorological data is collected more frequently and can more promptly reflect the impact of weather changes on crops, the weather station data has a relatively higher weight in the temporal dimension. Furthermore, considering accuracy, if the temperature measurement accuracy of the weather station is 0.1°C, while the soil temperature sensor's accuracy is 0.5°C, the higher-accuracy weather station temperature data will receive a higher weight in the temporal weight calculation. Regarding the spatial weight, the weather station monitors the entire agricultural park, while the soil sensor monitors only a small area of ​​farmland. Clearly, the weather station data, with its wider coverage, has a higher spatial weight.

[0100] Data source weights are calculated based on attribute similarity and historical correlation between data sources. Precipitation in meteorological data and soil moisture data are correlated. When precipitation increases, soil moisture generally also increases, indicating a high degree of historical correlation between these two data sources. Furthermore, temperature in meteorological data is linked to the crop growth stage in agricultural activity records. Certain temperature requirements exist for crops at specific growth stages, and these two attributes have a high degree of similarity. Based on these factors, meteorological data, soil moisture data, and agricultural activity record data are given relatively high weights in the data source weight calculation.

[0101] Combine the time dimension weights, space dimension weights, and data source weights into a spatiotemporal weight matrix. For example, set the time dimension weights to the first row of the matrix, the space dimension weights to the second row, and the data source weights to the third row, forming a matrix with three rows and multiple columns (the number of columns depends on the number of data points). This matrix is ​​the spatiotemporal weight matrix.

[0102] To generate a spatial proximity vector, the spatial neighborhood coverage of each data point in the spatiotemporally aligned data sequence is calculated. For example, a circular area with a radius of 100 meters, centered at the soil sensor's location, is defined as the spatial neighborhood coverage. For each data point, the number of data points within the preset neighborhood radius is counted to obtain a spatial density value. For example, if a soil sensor has three other soil sensors within its 100-meter neighborhood, its spatial density value is 3. A spatial proximity vector is generated based on the spatial density values ​​and neighborhood overlap. If both soil sensors have high spatial density values ​​and high neighborhood overlap, this indicates a high spatial proximity between the two sensors, resulting in a larger value for the corresponding element in the spatial proximity vector.

[0103] Temporal synchronization is calculated for spatiotemporally aligned data sequences to generate a temporal synchronization vector. If data collected by a weather station and a soil sensor at the same time have identical timestamps and are within the set temporal synchronization threshold (assuming the threshold is 5 minutes), their temporal synchronization is high, and the corresponding element value in the temporal synchronization vector is large. A data correlation vector is generated based on the attribute associations of the data sources. For example, if the light intensity in meteorological data is related to the light requirements of crops in agricultural activity records, the corresponding element value in their data correlation vector is large. Finally, the spatial proximity vector, temporal synchronization vector, and data correlation vector are combined into a set of correlation feature vectors, providing rich feature information for subsequent data analysis. Example

[0104] When building a co-occurrence pattern mining model and training it on spatiotemporal correlation feature maps, the co-occurrence pattern mining model is constructed based on a convolutional neural network. First, select an appropriate convolutional neural network architecture. For example, classic architectures such as LeNet and AlexNet can be used, and these architectures can also be modified according to actual needs. Based on the characteristics and size of the spatiotemporal correlation feature map, adjust the convolutional neural network parameters, such as the size and number of convolution kernels and the parameters of the pooling layer.

[0105] The spatiotemporal correlation feature map is input into the co-occurrence pattern mining model for feature extraction. The convolutional layer in the convolutional neural network performs a convolution operation on the spatiotemporal correlation feature map to extract local features. Through the combination of multiple convolutional and pooling layers, higher-level and more abstract potential co-occurrence pattern features are gradually extracted. For example, the convolutional layer can extract local spatial clustering features and temporal variation features of the data.

[0106] An unsupervised clustering algorithm is used to classify latent co-occurrence pattern features. A density-based clustering algorithm is used to identify high-density regions within the latent co-occurrence pattern features. This density-based clustering algorithm groups high-density regions into clusters based on the density of data points. Within the latent co-occurrence pattern feature dataset, regions with a high density of surrounding data points are identified; these regions may represent distinct co-occurrence patterns. Patterns in these high-density regions are merged to generate a candidate co-occurrence pattern set. If two high-density regions have similar features or are close to each other, they can be merged into a single candidate co-occurrence pattern. The silhouette coefficient is used to evaluate the clustering effectiveness of the candidate co-occurrence pattern set, and the optimal clustering result is selected as the spatiotemporal co-occurrence pattern feature set. The silhouette coefficient is a metric used to evaluate clustering effectiveness that comprehensively considers both intra-cluster compactness and inter-cluster separation. The silhouette coefficient is calculated for each candidate co-occurrence pattern set, and the candidate co-occurrence pattern set with the largest silhouette coefficient is selected as the final spatiotemporal co-occurrence pattern feature set.

[0107] The classification results are fed back into the co-occurrence pattern mining model for parameter optimization, resulting in a co-occurrence pattern recognition model. Based on the resulting spatiotemporal co-occurrence pattern feature set, the model's feature extraction process is analyzed for deficiencies, such as insufficient extraction of certain features or noise in the extracted features. The model is optimized by adjusting convolutional neural network parameters, such as the learning rate and weights, enabling more accurate co-occurrence pattern recognition. Ultimately, the co-occurrence pattern recognition model is obtained.

[0108] Take, for example, a city public security monitoring scenario. Data sources include surveillance cameras (providing video image data, considered device perception data), public security incident records (event log data), and urban geographic information data (geospatial data). After preliminary processing such as spatiotemporal calibration and alignment, this data is transformed into a spatiotemporal correlation feature map. The next step is to build and train a co-occurrence pattern mining model.

[0109] A co-occurrence pattern mining model was constructed based on a convolutional neural network. A modified AlexNet network architecture was selected for its excellent performance in image feature extraction and its ability to better process the spatial and temporal features in surveillance video image data. The convolutional neural network parameters were adjusted based on the size of the spatiotemporal correlation feature map (assuming the map is 224×224 pixels in the spatial dimension and contains data for 10 time steps in the temporal dimension). For example, the convolution kernel size of the first convolutional layer was set to 11×11, the stride to 4, and the number of convolution kernels to 96. This configuration effectively extracted local features from the map.

[0110] The spatiotemporal correlation feature map is input into the co-occurrence pattern mining model for feature extraction. Assume that the spatiotemporal correlation feature map contains surveillance video image data from different times and locations, as well as corresponding public security incident records. The convolutional layer of the convolutional neural network performs a convolution operation on the map. For example, the first convolutional layer slides the convolution kernel across the map and performs a weighted summation calculation on the local area of ​​the image. The formula is:

[0111] in, Indicates the Convolution kernels in the atlas The convolution output value of the position; It is The convolution kernel is The weight of the position; It is in the atlas The pixel value of the position; and are the height and width of the convolution kernel respectively; It is Through such convolution operations, potential co-occurrence pattern features such as areas where people gather and abnormal behavior movements are gradually extracted.

[0112] The latent co-occurrence pattern features are classified by unsupervised clustering algorithm. The density-based clustering algorithm (DBSCAN) identifies high-density areas in the latent co-occurrence pattern features. Assume that in the latent co-occurrence pattern feature dataset extracted by convolutional neural network, each data point represents a monitoring scene feature vector at a specific time and place. The DBSCAN algorithm is based on the set neighborhood radius. (Assume ) and the minimum number of points (Assume ), divide the density-connected data points into different clusters. For example, in a certain area, when the radius around a data point is The number of data points in the range is greater than or equal to , this area may be identified as a high-density area.

[0113] High-density regions are merged to generate a set of candidate co-occurrence patterns. If two high-density regions have similar features, such as clusters of people and unusual behavior, and are close to each other (within a certain distance threshold), they are merged into one candidate co-occurrence pattern.

[0114] The silhouette coefficient is used to evaluate the clustering effect of the candidate co-occurrence pattern set, and the optimal clustering result is selected as the spatiotemporal co-occurrence pattern feature set. The calculation formula of the silhouette coefficient is:

[0115] in, Indicates the The silhouette coefficient of the data points; is a data point The average distance to other data points of the same category, measuring the compactness within the category; is a data point The minimum average distance to data points of other different categories measures the degree of separation between classes. The silhouette coefficient is calculated for each candidate co-occurrence pattern set, and the candidate co-occurrence pattern set with the largest silhouette coefficient is selected as the final spatiotemporal co-occurrence pattern feature set. The closer the silhouette coefficient is to 1, the better the clustering effect, that is, the data points within a class are tightly clustered and the data points between classes are highly separated.

[0116] The classification results are fed back into the co-occurrence pattern mining model for parameter optimization, resulting in a co-occurrence pattern recognition model. Based on the resulting spatiotemporal co-occurrence pattern feature set, the model's feature extraction process is analyzed for deficiencies. For example, it was discovered that certain abnormal behavior features were not fully extracted, or that the extracted features contained noise. The model was optimized by adjusting convolutional neural network parameters, such as the learning rate (assuming an initial learning rate of 0.001, adjusted to 0.0005 based on feedback) and weights, enabling more accurate co-occurrence pattern recognition. Ultimately, a co-occurrence pattern recognition model was obtained, which was used for subsequent analysis and prediction of urban public safety conditions.

[0117] Example 4:

[0118] When building an anomaly detection model based on a spatiotemporal co-occurrence pattern feature set, divide the feature set into a training set and a validation set. A specific ratio can be used, such as 70% of the data for the training set and 30% for the validation set. This division should ensure that the data distribution of the training and validation sets is representative and reflects the overall characteristics of the spatiotemporal co-occurrence pattern feature set.

[0119] An initial anomaly detection model is constructed based on the Isolation Forest algorithm and trained using the training set. The Isolation Forest algorithm is a tree-based anomaly detection algorithm that partitions data by constructing multiple isolation trees. During training, the spatiotemporal co-occurrence pattern features from the training set are fed into the Isolation Forest model, which learns the distribution patterns of normal data. Each isolation tree randomly partitions the data and evaluates the degree of anomaly of a data point by calculating metrics such as the path length from the data point to the root node.

[0120] Use the validation set to perform sliding window validation on the initial anomaly detection model and calculate anomaly confidence scores. Set a sliding window size and step size and slide the window over the validation set. The data within each window is fed into the initial anomaly detection model, which outputs an anomaly score for each data point. Based on these anomaly scores, an anomaly confidence score is calculated. For example, statistical methods can be used to calculate the proportion of data points within the window whose anomaly score exceeds a certain threshold. This proportion serves as the anomaly confidence score.

[0121] Adjust the model threshold parameters based on the anomaly confidence score to obtain an optimized anomaly detection model. If the anomaly confidence score is too high, it means that the model may be misclassifying too many normal data as anomalies, and the model threshold needs to be appropriately increased. If the anomaly confidence score is too low, it means that the model may not detect enough anomalies, and the model threshold needs to be appropriately lowered. By continuously adjusting the threshold parameters, the model achieves the best detection effect on the validation set, resulting in an optimized anomaly detection model.

[0122] Real-time monitoring data is fed into the optimized anomaly detection model, which outputs a set of anomalous co-occurrence patterns. The real-time monitoring data is processed using the same preprocessing methods as the training data and then fed into the optimized anomaly detection model. The model analyzes this data and determines whether each data point belongs to an anomalous co-occurrence pattern. All data points identified as anomalous co-occurrence patterns are collected to form a set of anomalous co-occurrence patterns.

[0123] Taking the cargo flow monitoring scenario of a large e-commerce warehouse as an example, a variety of data collection devices are deployed in this e-commerce warehouse to obtain data related to cargo flow. At the same time, various warehouse operation event data are also recorded. After processing, this data is obtained to obtain a set of spatiotemporal co-occurrence pattern features, which is used as the basis for building an anomaly detection model.

[0124] When partitioning the spatiotemporal co-occurrence pattern feature set into training and validation sets, we assume that data such as the hourly inbound and outbound shipments over the past month, as well as changes in the storage locations of goods in different areas of the warehouse, is collected. This data is then processed to generate the spatiotemporal co-occurrence pattern feature set. The data is partitioned in a 7:3 ratio, with the first 21 days of data serving as the training set and the last 9 days of data serving as the validation set. During the partitioning process, the training and validation sets are ensured to be chronologically continuous and cover a variety of scenarios from daily warehouse operations, such as promotional events and regular sales periods, to ensure a representative data distribution.

[0125] An initial anomaly detection model is constructed based on the isolation forest algorithm and trained using the training set. In this warehouse scenario, the isolation forest algorithm processes the spatiotemporal co-occurrence pattern feature data in the training set. Each isolation tree randomly partitions the data. For example, assume that the incoming goods quantity in the training set ranges from 0 to 1000 pieces / hour. The isolation tree randomly selects a split point, such as 500 pieces / hour, to divide the data into two parts: those with incoming goods less than 500 pieces / hour and those with greater than or equal to 500 pieces / hour. Similar random partitioning is then performed within the subsets until only one data point remains in each subset or a stopping condition is met. The degree of anomaly is assessed by calculating the path length from each data point to the root node. If a data point is quickly isolated during tree construction (i.e., its path length is short), it is considered more likely to be an anomaly.

[0126] The initial anomaly detection model is validated using a sliding window on the validation set to calculate an anomaly confidence score. The sliding window size is set to 3 hours, with a step size of 1 hour. Starting from the first hour of the first day, data from hours 1-3 is fed into the initial anomaly detection model. The model outputs an anomaly score for each data point within these three hours (e.g., the number of goods shipped out per hour, the number of times a specific shelf is moved, etc.). Assuming anomaly scores range from 0 to 1, scores closer to 1 indicate a higher likelihood of an anomaly. The proportion of data points within the window with an anomaly score exceeding 0.5 (this can be adjusted based on actual conditions) is calculated as the anomaly confidence score for that window. For example, if there are 10 data points within a 3-hour window, and 3 of them have an anomaly score exceeding 0.5, the anomaly confidence score for that window is 30%.

[0127] The model threshold parameters are adjusted based on the anomaly confidence score to obtain an optimized anomaly detection model. If the anomaly confidence score is too high, it means that the model may be misclassifying too much normal data as anomalies. For example, if the anomaly confidence scores of multiple windows are all above 50%, but such a high proportion of anomalies is not expected in actual warehouse operations, the model threshold should be appropriately increased, such as raising the anomaly score judgment threshold from 0.5 to 0.7. Conversely, if the anomaly confidence score is too low, it means that the model may not detect enough anomaly data, and the threshold should be lowered. After multiple adjustments and verifications, the model achieves optimal detection results on the validation set, resulting in the optimized anomaly detection model.

[0128] Real-time monitoring data is input into the optimized anomaly detection model, which outputs a set of abnormal co-occurrence patterns. In daily warehouse operations, hourly cargo flow data is collected in real time, such as the current hour's cargo inflow and the number of times items are picked from a specific shelf. This real-time data is processed using the same preprocessing methods as the training data and then input into the optimized anomaly detection model. The model analyzes and judges this data. For example, if the model detects an abnormally high number of pickups from a shelf within a certain hour, coupled with a sudden drop in inventory in the area of ​​that shelf within a short period of time, and the spatiotemporal co-occurrence pattern of these two conditions differs significantly from the normal pattern in the training set, the model identifies this as an abnormal co-occurrence pattern. All data points identified as abnormal co-occurrence patterns are collected to form a set of abnormal co-occurrence patterns. Warehouse managers can use this set of abnormal co-occurrence patterns to promptly identify potential issues in warehouse operations, such as lost goods and operational errors, and take appropriate measures to address them.

[0129] Example 5:

[0130] When building a parameter optimization model and dynamically adjusting data collection parameters, define the model's decision variables. These decision variables include at least the collection frequency, storage capacity, and transmission interval. For the collection frequency, you can set a range of values, such as once per hour at the minimum and once per minute at the maximum. The storage capacity can be determined based on the device's actual storage capacity and data volume. For example, the minimum storage capacity can be 1GB, while the maximum is 10GB. The transmission interval can also be set within a range, such as 5 minutes at the minimum and 1 hour at the maximum.

[0131] An optimization objective function is constructed based on the historical distribution of anomalous co-occurrence patterns. This objective function includes at least data coverage completeness and resource consumption costs. Data coverage completeness can be measured by calculating the coverage ratio of anomalous co-occurrence patterns in different time periods and spatial regions. If anomalous co-occurrence patterns can be accurately detected within a given time period or spatial region, data coverage completeness is high. Resource consumption costs include device energy consumption, storage costs, and transmission costs. For example, a higher acquisition frequency increases device energy consumption; a larger storage capacity increases storage costs; and a shorter transmission interval increases transmission costs. Taking these factors into consideration, an objective function is constructed that balances data coverage completeness and resource consumption costs.

[0132] Set constraints, which should include at least the device workload limit and the data transmission delay threshold. The device workload limit refers to the maximum operating pressure the device can withstand within a certain period of time. If the acquisition frequency is too high, the storage capacity is too large, or the transmission interval is too short, the device workload may exceed the limit, affecting normal operation of the device. The data transmission delay threshold is the maximum allowable delay from data acquisition to data transmission to the target location. If the transmission interval is too long, the data transmission delay may exceed the threshold, affecting the real-time performance of the data. These constraints should be set appropriately based on the actual performance of the device and application requirements.

[0133] The parameter optimization model is iteratively solved using a heuristic search algorithm to obtain the optimal data acquisition parameters. Heuristic search algorithms can include genetic algorithms and particle swarm optimization algorithms. For example, a genetic algorithm first randomly generates an initial set of decision variable values ​​representing different combinations of acquisition frequency, storage capacity, and transmission interval, called individuals. The objective function value for each individual is calculated and evaluated. New individuals are then generated through genetic operations such as selection, crossover, and mutation. This process is continuously iterated, allowing the individuals in the population to gradually approach the optimal solution. During the iterative process, it is important to ensure that each individual satisfies the set constraints. When certain stopping conditions are met, such as when the objective function value stops changing or changes very little, the optimal data acquisition parameters are obtained.

[0134] Take the freight transportation management of a large logistics park as an example. In this logistics park, there are a large number of freight transportation vehicles, warehouses and various data collection equipment. In order to optimize data collection and resource utilization, it is necessary to build a parameter optimization model and dynamically adjust the data collection parameters.

[0135] Define the decision variables of the parameter optimization model. In the logistics park scenario, the collection frequency refers to how often information such as vehicle location and cargo status is obtained. For example, the current vehicle location collection frequency range is set to every 5 to 30 minutes. Storage capacity is related to storing data such as vehicle driving trajectories and cargo transportation records. The minimum storage capacity of data storage devices in the logistics park is 500GB, which can be expanded to a maximum of 5TB. The transmission interval determines the time interval for data transmission from the collection device to the management system, and its value range is set to 10 to 60 minutes. These collection frequency, storage capacity, and transmission interval are the decision variables of the parameter optimization model.

[0136] The optimization objective function is constructed based on the historical distribution of the set of abnormal co-occurrence patterns. In the logistics park, abnormal co-occurrence patterns may include vehicles that are stuck in non-designated areas for a long time and the cargo status has not been updated for a long time, and a large number of vehicles gathered at the door of a warehouse at the same time to cause congestion. In terms of data coverage integrity, it is measured by calculating the proportion of abnormal co-occurrence patterns that are accurately monitored in different time periods and regions. For example, during the busy transportation period of the logistics park (such as 10 am to 2 pm), if more than 90% of the abnormal co-occurrence patterns can be monitored, it means that the data coverage integrity of this period is relatively high. Resource consumption costs include equipment energy consumption, storage costs and transmission costs. The higher the acquisition frequency, the greater the energy consumption of the positioning equipment and cargo status monitoring equipment on the vehicle; the larger the storage capacity, the higher the purchase and maintenance costs of the storage equipment; the shorter the transmission interval, the higher the network traffic cost consumed by data transmission. Incorporating data coverage integrity and resource consumption costs into the objective function, the goal is to minimize resource consumption costs while ensuring a certain level of data coverage integrity. Assume that data coverage integrity is expressed in Indicates (value range , indicates full coverage), resource consumption cost is The objective function can be expressed as ,in and] It is a weight coefficient set according to actual needs, which is used to balance the importance of data coverage integrity and resource consumption cost.

[0137] Set constraints. The upper limit of equipment workload is a key constraint. Data collection equipment and transmission equipment in the logistics park have certain workload limitations. For example, the positioning equipment on the vehicle may overheat, degrade performance, and other problems when the collection frequency is too high. Assuming that its maximum tolerable collection frequency is once every 5 minutes, this is the upper limit constraint of the collection frequency. The data transmission delay threshold is also very important. Cargo transportation data needs to be transmitted to the management system in a timely manner for scheduling decisions. If the transmission delay is too long, it may lead to wrong decisions. Set the data transmission delay threshold to 30 minutes, that is, the time from data collection to transmission to the management system cannot exceed 30 minutes, otherwise it is considered to have exceeded the threshold.

[0138] The parameter optimization model is iteratively solved using a heuristic search algorithm to obtain the optimal data acquisition parameters. The particle swarm optimization algorithm is employed here. First, a set of initial decision variable values ​​is randomly generated, namely, different combinations of acquisition frequency, storage capacity, and transmission interval. Each combination is considered a particle. For example, the parameter combination for particle 1 is an acquisition frequency of 15 minutes, a storage capacity of 1 TB, and a transmission interval of 30 minutes; the parameter combination for particle 2 is an acquisition frequency of 20 minutes, a storage capacity of 2 TB, and a transmission interval of 40 minutes. The objective function value of each particle is calculated to evaluate its performance. The particle swarm optimization algorithm uses information sharing and collaboration among particles to guide particles toward a more optimal path. In each iteration, particles adjust their speed and position—that is, adjust the acquisition frequency, storage capacity, and transmission interval—based on their own historical optimal position and the swarm's global optimal position. After multiple iterations, the optimal data acquisition parameters are obtained when the objective function value stabilizes or meets specific convergence criteria. For example, the optimal parameter combination finally obtained is a collection frequency of once every 20 minutes, a storage capacity of 1.5TB, and a transmission interval of 40 minutes. The logistics park can adjust the settings of the data collection equipment according to these parameters to achieve more efficient data collection and resource utilization, while better monitoring abnormal situations during cargo transportation and improving the overall operational management level.

[0139] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0140] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A spatiotemporal co-occurrence analysis method for multi-source data fusion, characterized in that: The following steps are involved: Acquire multi-source heterogeneous data of the target area and perform spatiotemporal calibration, where the multi-source heterogeneous data includes at least geospatial data, device perception data, and event record data; Construct a spatiotemporal reference coordinate system and perform spatiotemporal alignment on multi-source heterogeneous data to obtain a spatiotemporal aligned data sequence. Based on the spatiotemporal aligned data sequence, a spatiotemporal weight matrix and a set of associated feature vectors are generated. The spatiotemporal weight matrix is ​​divided into multiple spatiotemporal data sub-blocks through a spatiotemporal sliding window. The spatiotemporal data sub-blocks are fused based on the associated feature vector set to obtain a spatiotemporal correlation feature map. Construct a co-occurrence pattern mining model and train it through the spatiotemporal correlation feature map to obtain a co-occurrence pattern recognition model, and then extract the spatiotemporal co-occurrence pattern feature set; An anomaly detection model is constructed based on the spatiotemporal co-occurrence pattern feature set, and the anomaly detection model is verified through a sliding verification mechanism to obtain an anomaly co-occurrence pattern set; Based on the abnormal co-occurrence pattern set and data collection parameters, a parameter optimization model is constructed and the data collection parameters are dynamically adjusted to obtain the optimal data collection parameters for the target area; Based on the optimal data collection parameters and real-time monitoring data, the spatiotemporal reference coordinate system is updated and a spatiotemporal co-occurrence analysis report is generated.

2. The spatiotemporal co-occurrence analysis method for multi-source data fusion according to claim 1, characterized in that: The spatiotemporal weight matrix includes at least the time dimension weight, the space dimension weight and the data source weight; the association feature vector set includes at least the spatial proximity vector, the time synchronization vector and the data association vector; the spatiotemporal co-occurrence pattern feature set includes at least the spatial aggregation pattern, the time continuity pattern and the cross-source association pattern.

3. The spatiotemporal co-occurrence analysis method for multi-source data fusion according to claim 1, characterized in that: The process of constructing a spatiotemporal reference coordinate system and performing spatiotemporal alignment on multi-source heterogeneous data to obtain a spatiotemporal aligned data sequence includes the following steps: Define the grid division rules and timestamp alignment rules of the spatiotemporal reference coordinate system based on the geographic boundaries and time span of the target area; Perform spatial projection transformation on multi-source heterogeneous data to make its coordinate parameters consistent with the spatiotemporal reference coordinate system; Normalize the timestamps of multi-source heterogeneous data to generate time series data with a unified time base; Based on the spatial proximity threshold and temporal synchronization threshold, data interpolation and conflict resolution are performed on the normalized multi-source heterogeneous data to obtain a spatiotemporal aligned data sequence.

4. The spatiotemporal co-occurrence analysis method for multi-source data fusion according to claim 1, characterized in that: The method of generating a spatiotemporal weight matrix and a set of associated feature vectors based on a spatiotemporal alignment data sequence comprises the following steps: Calculate the time dimension weight and space dimension weight based on the data source collection frequency, accuracy and coverage; Calculate the data source weight based on the attribute similarity and historical correlation of the data source; Combine the time dimension weight, space dimension weight and data source weight into a spatiotemporal weight matrix; Perform spatial proximity calculation on the spatiotemporal aligned data sequence to generate a spatial proximity vector; Calculate the time synchronization degree of the spatiotemporal alignment data sequence and generate a time synchronization degree vector; The data association degree vector is generated based on the attribute association of the data source, and then combined into a set of association feature vectors.

5. The spatiotemporal co-occurrence analysis method for multi-source data fusion according to claim 4 is characterized in that: Generating the spatial proximity vector comprises the following steps: Calculate the spatial neighborhood coverage of each data point in the spatiotemporally aligned data series; Count the number of data points within the preset neighborhood radius of each data point to obtain the spatial density value; Generate a spatial proximity vector based on the spatial density value and neighborhood overlap.

6. The spatiotemporal co-occurrence analysis method for multi-source data fusion according to claim 1, characterized in that: The co-occurrence pattern mining model is constructed and trained through the spatiotemporal correlation feature map to obtain a co-occurrence pattern recognition model, including the following steps: Construct a co-occurrence pattern mining model based on convolutional neural networks; Input the spatiotemporal correlation feature map into the co-occurrence pattern mining model for feature extraction to obtain potential co-occurrence pattern features; The potential co-occurrence pattern features are classified by unsupervised clustering algorithm to obtain the spatiotemporal co-occurrence pattern feature set; The classification results are fed back to the co-occurrence pattern mining model for parameter optimization to obtain the co-occurrence pattern recognition model.

7. The spatiotemporal co-occurrence analysis method for multi-source data fusion according to claim 6, characterized in that: The method of performing pattern classification on potential co-occurrence pattern features by using an unsupervised clustering algorithm comprises the following steps: Identify high-density regions in potential co-occurrence pattern features based on density clustering algorithm; Merge patterns in high-density areas to generate a set of candidate co-occurrence patterns; The clustering effect of the candidate co-occurrence pattern set is evaluated by the silhouette coefficient, and the optimal clustering result is selected as the spatiotemporal co-occurrence pattern feature set.

8. The spatiotemporal co-occurrence analysis method for multi-source data fusion according to claim 1, characterized in that: The method of constructing an anomaly detection model based on a spatiotemporal co-occurrence pattern feature set includes the following steps: Divide the spatiotemporal co-occurrence pattern feature set into a training set and a validation set; Build an initial anomaly detection model based on the isolation forest algorithm and train the model using the training set; Use the validation set to perform sliding window validation on the initial anomaly detection model and calculate the anomaly confidence score; Adjust the model threshold parameters according to the anomaly confidence score to obtain the optimized anomaly detection model; The real-time monitoring data is input into the optimized anomaly detection model, and a set of anomaly co-occurrence patterns is output.

9. The spatiotemporal co-occurrence analysis method for multi-source data fusion according to claim 1, characterized in that: The method of constructing a parameter optimization model and dynamically adjusting data acquisition parameters includes the following steps: Defining decision variables of a parameter optimization model, wherein the decision variables include at least acquisition frequency, storage capacity, and transmission interval; Constructing an optimization objective function based on the historical distribution of the abnormal co-occurrence pattern set, wherein the objective function at least includes data coverage integrity and resource consumption cost; Setting constraints, wherein the constraints include at least a device workload upper limit and a data transmission delay threshold; The parameter optimization model is iteratively solved through a heuristic search algorithm to obtain the optimal data acquisition parameters.

10. A spatiotemporal co-occurrence analysis system for multi-source data fusion, characterized by: include: A data acquisition and calibration module is used to acquire multi-source heterogeneous data of the target area, including at least geospatial data, device perception data, and event record data, and perform spatiotemporal calibration; The spatiotemporal alignment processing module is used to construct a spatiotemporal reference coordinate system, perform spatiotemporal alignment on multi-source heterogeneous data, obtain a spatiotemporal aligned data sequence, and generate a spatiotemporal weight matrix and a set of associated feature vectors based on the spatiotemporal aligned data sequence; The feature fusion graph construction module divides the spatiotemporal weight matrix into multiple spatiotemporal data sub-blocks through a spatiotemporal sliding window, and performs feature fusion on the spatiotemporal data sub-blocks based on the associated feature vector set to obtain a spatiotemporal correlation feature graph; The co-occurrence pattern mining module builds a co-occurrence pattern mining model and trains it through the spatiotemporal correlation feature map to obtain a co-occurrence pattern recognition model, and then extracts the spatiotemporal co-occurrence pattern feature set; The anomaly detection module builds an anomaly detection model based on the spatiotemporal co-occurrence pattern feature set, verifies the anomaly detection model through a sliding verification mechanism, and obtains an anomaly co-occurrence pattern set; The parameter optimization module builds a parameter optimization model based on the set of abnormal co-occurrence patterns and data collection parameters, and dynamically adjusts the data collection parameters to obtain the optimal data collection parameters for the target area; The report generation module updates the spatiotemporal reference coordinate system and generates a spatiotemporal co-occurrence analysis report based on the optimal data collection parameters and real-time monitoring data.

Citation Information

Patent Citations

  • Technical list generation method and system based on multi-source data and topic model

    CN114780617A

  • Intelligent clothing industry production regulation and control method and system based on data analysis

    CN118798494A