Method, device and storage medium for detecting abnormal events in spatiotemporal big data
By standardizing the processing of urban spatiotemporal big data and building distributed storage strategies, combining cosine similarity and multi-dimensional evaluation index optimization methods, time series and spatial distribution characteristics are extracted, and the support vector machine classifier of twin network structure is used to identify abnormal events, which solves the problems of low computing efficiency and insufficient feature extraction in traditional methods, and achieves efficient and accurate abnormal event detection.
Patent Information
- Application Number
- CN202510209715.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-25
AI Technical Summary
When dealing with massive space-time big data, traditional anomaly event detection methods face problems such as low computational efficiency, insufficient feature extraction and difficulty in dealing with complex nonlinear relationships in high-dimensional feature space.
By unifying the time stamps and standardizing spatial coordinates of urban spatiotemporal big data, a distributed storage strategy and computing node coordination mechanism are built, a neighborhood selection method with cosine similarity and multi-objective optimization operation of multi-dimensional evaluation indicators are used to extract time series and spatial distribution characteristics, and an abnormal event recognition is used to use the support vector machine classifier of the twin network structure.
It improves the accuracy and efficiency of abnormal event detection, enhances the model's ability to express complex spatial and temporal relationships, improves the ability to classify high-dimensional nonlinear features, and realizes the all-round feature portrayal of abnormal events.
Smart Images

Figure CN119691580B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a method, device and storage medium for detecting abnormal events in spatiotemporal big data. Background Art
[0002] With the rapid development of smart city construction, urban spatiotemporal big data presents the characteristics of high dimension, multi-source heterogeneity, and real-time dynamics. These data contain rich information about urban operation and are of great value for discovering urban abnormal events and preventing urban safety hazards. However, traditional abnormal event detection methods face problems such as low computational efficiency and insufficient feature extraction when processing massive spatiotemporal big data.
[0003] The current abnormal event detection methods have the following main problems: in the neighborhood selection process, the neighborhood division is not accurate enough due to the lack of in-depth exploration of the spatiotemporal correlation relationship of data points; secondly, the existing feature extraction methods often only focus on the features of a single dimension and cannot effectively capture the multi-dimensional correlation characteristics of the data; in the classification and recognition stage, traditional classifiers find it difficult to handle complex nonlinear relationships in high-dimensional feature spaces. Summary of the invention
[0004] The present invention provides a method, device and storage medium for detecting abnormal events in spatiotemporal big data. The present invention improves the classification capability of high-dimensional nonlinear features and realizes all-round feature characterization and detection of abnormal events.
[0005] In a first aspect, the present invention provides a method for detecting abnormal events in spatiotemporal big data, the method comprising:
[0006] Perform time stamp unification and spatial coordinate standardization processing on urban spatiotemporal big data to obtain a standardized spatiotemporal data set, and distribute the standardized spatiotemporal data set to multiple computing nodes through a distributed storage strategy;
[0007] In the computing node, the cosine similarity of the feature vectors of each data point in the standardized spatiotemporal data set with other data points is calculated to obtain a target similarity value, and data points whose target similarity values are greater than the similarity threshold are selected according to a preset similarity threshold to obtain an initial neighborhood data set;
[0008] Calculating the spatiotemporal density distribution of the initial neighborhood data set, and obtaining a core neighborhood data set through a multi-objective optimization operation of multi-dimensional neighborhood evaluation indicators;
[0009] Extracting time series features and spatial distribution features from the core neighborhood data set to obtain a fusion feature weight matrix;
[0010] The fused feature weight matrix is input into a support vector machine classifier to perform abnormal event recognition to obtain an abnormal event detection result.
[0011] In a second aspect, the present invention provides a device for detecting abnormal events of spatiotemporal big data, the device comprising:
[0012] A standardization processing module is used to unify the timestamps and standardize the spatial coordinates of the urban spatiotemporal big data to obtain a standardized spatiotemporal data set, and distribute the standardized spatiotemporal data set to multiple computing nodes through a distributed storage strategy;
[0013] A similarity calculation module is used to calculate the cosine similarity of the feature vectors of each data point in the standardized spatiotemporal data set with other data points in the calculation node to obtain a target similarity value, and select data points whose target similarity values are greater than the similarity threshold according to a preset similarity threshold to obtain an initial neighborhood data set;
[0014] A distribution calculation module is used to calculate the spatiotemporal density distribution of the initial neighborhood data set, and obtain a core neighborhood data set through a multi-objective optimization operation of multi-dimensional neighborhood evaluation indicators;
[0015] A feature extraction module is used to extract time series features and spatial distribution features from the core neighborhood data set to obtain a fusion feature weight matrix;
[0016] The event recognition module is used to input the fusion feature weight matrix into the support vector machine classifier to perform abnormal event recognition and obtain abnormal event detection results.
[0017] The third aspect of the present invention provides a spatiotemporal big data abnormal event detection device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the spatiotemporal big data abnormal event detection device executes the above-mentioned spatiotemporal big data abnormal event detection method.
[0018] A fourth aspect of the present invention provides a computer-readable storage medium, which stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned method for detecting abnormal events in spatiotemporal big data.
[0019] In the technical solution provided by the present invention, by constructing a distributed storage strategy and a computing node coordination mechanism, the efficiency problem of large-scale spatiotemporal data processing is solved, so that the system can efficiently process massive data; a neighborhood selection method based on cosine similarity is combined with an optimization strategy of multi-dimensional evaluation indicators to improve the accuracy of abnormal event detection; heterogeneous graph learning and multi-view hypergraph structure are introduced to achieve effective extraction of deep-level features of data and enhance the model's ability to express complex spatiotemporal relationships; a support vector machine classifier with a twin network structure is used to improve the model's classification ability for high-dimensional nonlinear features; multi-layer feature fusion and dynamic threshold strategy are used to improve the reliability of abnormal event detection results; a method for extracting temporal features and spatial features is integrated to achieve a comprehensive characterization of abnormal events; a multi-objective optimization algorithm is used for neighborhood evaluation to ensure the accuracy and interpretability of detection results. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0021] Figure 1 A schematic diagram of a process for detecting abnormal events in spatiotemporal big data provided in an embodiment of the present application;
[0022] Figure 2 A schematic block diagram of the structure of a device for detecting abnormal events in spatiotemporal big data provided in an embodiment of the present application;
[0023] Figure 3 A schematic block diagram of the structure of a device for detecting abnormal events in spatiotemporal big data provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change based on actual conditions.
[0026] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0027] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0028] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0029] See also Figure 1 , Figure 1 A flow chart of a method for detecting abnormal events in spatiotemporal big data provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the spatiotemporal big data abnormal event detection method provided in the embodiment of the present application includes steps S100 to S500.
[0030] Step S100: performing time stamp unification and spatial coordinate standardization processing on the urban spatiotemporal big data to obtain a standardized spatiotemporal data set, and distributing the standardized spatiotemporal data set to multiple computing nodes through a distributed storage strategy;
[0031] It is understandable that the execution subject of the present invention may be a spatiotemporal big data abnormal event detection device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0032] Specifically, the time stamps of the urban spatiotemporal big data of each data collection point are converted, and the inconsistent or different-format time stamps in the original data are converted into a unified standard time representation. By processing the time field, the format conflicts and inconsistent time intervals in the time records are solved, and a time-normalized data set is generated to ensure seamless integration of data from different sources in the time dimension. The spatial coordinate information in the time-normalized data set is converted into a unified geographic coordinate system to obtain a coordinate-normalized data set. The outliers and noise data in the coordinate-normalized data set are processed, including detecting and removing coordinate points that deviate from the normal range, using statistical methods or clustering analysis algorithms to remove outliers, and correcting the offset values caused by equipment errors or data collection errors to obtain a preprocessed data set. A hierarchical index structure of the time dimension and the space dimension is constructed based on the preprocessed data set to optimize the storage and retrieval efficiency of subsequent data. The hierarchical index structure organizes the time and space dimensions in layers, such as constructing an index with the granularity of year, month, and day in the time dimension, and hierarchically dividing the spatial coordinates based on the grid partitioning method or quadtree structure in the space dimension. Through the index, data within a specific time and space range can be quickly located. The preprocessed data set is divided into blocks according to the hierarchical index structure. The preprocessed data set is divided into multiple data slices according to the spatiotemporal proximity relationship, so that the data in each slice has a strong correlation in time and space. For example, based on the grid division method, the geographical area is divided into multiple small grids, and combined with the time interval division, the data in each grid is subdivided into multiple fragments according to the time dimension to generate multiple standardized spatiotemporal data sets. A distributed file system is established based on the standardized spatiotemporal data set. In the distributed file system, an efficient data storage and access mechanism is designed to support large-scale parallel computing. For example, HDFS or other distributed storage frameworks are used to distribute the standardized data set to multiple computing nodes, and the location information of each data block is recorded through the metadata management mechanism. At the same time, in order to utilize the computing resources in the distributed environment, the load calculation of the standardized spatiotemporal data set in the distributed storage structure is performed. Considering the computing power, storage capacity and current load status of each computing node, a reasonable data processing task allocation scheme is designed. For example, a load balancing algorithm is used to dynamically allocate tasks to ensure that the computing resources of each node are fully utilized while avoiding the overload problem of a single node. After the task allocation plan is generated, it is sent to each computing node, and a data synchronization channel is established between the nodes to support cross-node data exchange and collaborative processing, and complete the distributed storage of standardized spatiotemporal data sets among computing nodes.
[0033] Step S200: In the computing node, for each data point in the standardized spatiotemporal data set, the cosine similarity of the feature vector with other data points is calculated to obtain a target similarity value, and data points whose target similarity value is greater than the similarity threshold are selected according to a preset similarity threshold to obtain an initial neighborhood data set;
[0034] Specifically, in each computing node, a feature matrix is constructed for each data point based on the standardized spatiotemporal data set. The feature matrix reflects the spatiotemporal characteristics of the data point, including the time feature dimension and the spatial feature dimension. The time feature dimension contains information such as timestamp, time period category, and time period, while the spatial feature dimension contains features such as the geographical location, regional distribution, and spatial topological relationship of the data point. The feature matrix is segmented into time windows, and the data is divided into multiple time segments according to predefined time intervals, so that each segment can more accurately reflect the behavioral characteristics of the data point in a specific time period. For example, the data is divided into hours, days, or shorter time periods to capture the dynamic changes of the data in the time dimension. Based on the segmented feature vectors, the spatial distance and time distance between each data point are calculated to measure the similarity of the data points in the spatial position and time dimension respectively. The Euclidean distance and the time distance are weighted and combined to generate a comprehensive distance matrix to reflect the degree of spatiotemporal association between the data points. The comprehensive distance matrix is input into the similarity calculation module. By analyzing the relationship between the features of data points, the similarity between each pair of data points is calculated to form a matrix in which each value represents the degree of similarity between two data points. The similarity matrix is normalized so that its value range is within a normalized interval to obtain the target similarity value. The target similarity value is compared with the preset similarity threshold, and the data point pairs that meet the threshold condition are screened out. These point pairs constitute the preliminary neighborhood candidate point set. However, the neighborhood candidate point set is subjected to density clustering operation. The local density value and distance factor of each data point are calculated to measure the degree of aggregation of the data point in its local area and its relative position to other data points. Based on these indicators, the data points are sorted to form a density sorting sequence. Based on the density sorting sequence, the data point groups are hierarchically processed to identify data point groups with similar spatiotemporal distribution characteristics. The hierarchical grouping structure allows the clustering characteristics of the data to be observed and analyzed at different levels, thereby revealing the multi-level association relationship of the data. The data points in these groups are subjected to association analysis based on the principle of spatiotemporal continuity. For example, analyzing the temporal continuity of data points to identify the start and end time of an event, or analyzing the spatial distribution characteristics to determine the area where an event occurs. Through these operations, a subset of data points with significant spatiotemporal correlations is extracted from the hierarchical neighborhood structure to generate an initial neighborhood dataset.
[0035] Step S300, calculating the spatiotemporal density distribution of the initial neighborhood data set, and obtaining a core neighborhood data set through multi-objective optimization calculation of multi-dimensional neighborhood evaluation indicators;
[0036] Specifically, the temporal and spatial distribution relationship of the data points in the initial neighborhood data set is calculated, and the Gaussian kernel function is selected as the basic function for density calculation. The Gaussian kernel function can reflect the distribution characteristics of the data points in its neighborhood due to its smoothness and locality. By substituting the temporal and spatial coordinates of each data point into the Gaussian kernel function, a density value matrix is generated, and each element of the matrix represents the density value of the corresponding data point in its local neighborhood. Based on the density value matrix, a neighborhood evaluation vector containing spatial uniformity, temporal continuity and boundary compactness is constructed. Among them, spatial uniformity is used to measure the uniformity of the spatial distribution of data points, temporal continuity is used to evaluate the temporal correlation between data points, and boundary compactness reflects the compactness of the data point group on the boundary and the sparsity of the boundary points. Based on the neighborhood evaluation vector, the optimization objectives of maximizing spatial uniformity, temporal continuity and boundary compactness are set to obtain a multi-objective optimization function. This function improves the accuracy and stability of neighborhood division through optimization operations. Based on the multi-objective optimization function, a multi-objective optimization method based on genetic algorithm or particle swarm algorithm is adopted to find a non-dominated solution set through population iteration and solution space search. The non-dominated solution set is the result set of multi-objective optimization, which contains solutions that achieve the best trade-off between different optimization objectives. In order to improve the applicability of the optimization results, the non-dominated solution set is weighted. The solutions are weighted and combined according to the spatiotemporal correlation to obtain the optimal neighborhood solution. By combining the needs of actual applications, such as focusing on sudden events in the time dimension or the distribution characteristics in the spatial dimension, the weights of each optimization objective are dynamically adjusted so that the final optimal neighborhood solution can adapt to the spatiotemporal data characteristics in a specific scenario. Based on the optimal neighborhood solution, the initial neighborhood data set is repartitioned, and the core data points that meet the spatiotemporal density threshold are screened to form a core member set. The core member set consists of data points that are significantly higher than the average level of the neighborhood in spatiotemporal density. These points are usually high-probability areas of abnormal events or key nodes of data distribution. The core member set is optimized by heterogeneous graph learning to enhance the representation ability of the core neighborhood data set. The spatiotemporal relationship in the core member set is modeled by constructing a heterogeneous graph using the heterogeneous graph learning method. Heterogeneous graph learning can effectively integrate different types of relationships (such as temporal correlation and spatial adjacency), extract deep connections between core data points through the characteristics of graph structure, and improve the quality and applicability of core neighborhood datasets.
[0037] A heterogeneous spatiotemporal information network is constructed based on the core member set. Each data point in the core member set is modeled as a node in the network. At the same time, multi-type edges are created according to the spatiotemporal association between the data points. These edges represent relationships of different dimensions such as time association, spatial proximity and attribute similarity, forming a heterogeneous network structure with multi-type nodes and edges. A multi-path set containing time association, spatial proximity and attribute similarity is designed based on the heterogeneous network structure. These path sets are used to capture the complex semantic relationships between different types of nodes and edges in the network. By semantically encoding each meta-path, the information transmission pattern and contextual relevance in the path are extracted to generate a path feature set. The path feature set is input into the heterogeneous graph neural network to perform message transmission and feature aggregation on different types of nodes and edges in the network. Through this process, the graph neural network learns the representation vector of each node in the heterogeneous network. These representation vectors integrate the node's own attributes and its association information with surrounding nodes, and can reflect the multi-dimensional characteristics of the node. A multi-view hypergraph structure is constructed based on the node representation vector. The association relationship of the time dimension, space dimension and attribute dimension is encoded as a hyperedge to form a hierarchical hypergraph. Hierarchical hypergraphs can capture correlation information from different perspectives and reveal potential high-order relationships by connecting multiple nodes through hyperedges. In order to analyze the information in the hierarchical hypergraph, it is decomposed into multiple levels. Subgraph structural features are extracted from levels of different granularity to generate multi-scale feature representations. Based on the multi-scale features, a graph attention network is constructed. By calculating the importance weights of features at different levels, features of different granularities are weightedly fused to obtain feature fusion weights. The feature fusion weights are weightedly combined with the features at each level to generate an enhanced feature set. The enhanced feature set is subjected to dimensionality reduction. By reducing the feature dimension to remove redundant information while retaining the most important feature patterns, the data is made more compact and easy to use, and a core neighborhood dataset is obtained.
[0038] Step S400, extracting time series features and spatial distribution features from the core neighborhood data set to obtain a fusion feature weight matrix;
[0039] Specifically, the time series data in the core neighborhood dataset are transformed in the frequency domain and statistical features are extracted. The frequency domain transformation maps the time series from the time domain to the frequency domain, revealing its hidden periodicity, trend and amplitude characteristics; at the same time, the statistical feature extraction can capture the basic characteristics of the time series, including mean, variance, skewness, kurtosis, etc. These characteristics constitute the core content of the time feature sequence. A time series association graph is constructed based on the time feature sequence. The time series association graph converts the time features into a topological structured representation by representing the key nodes in the time series as vertices in the graph and the time association relationship between nodes as the edges of the graph. The time series association graph is topologically analyzed, including the calculation of indicators such as degree distribution, clustering coefficient and connectivity, to reveal the local and global relationships between nodes in the time series. At the same time, the path feature extraction technology is used to analyze the structural information such as the key path and the shortest path in the graph to form a time series association matrix to express the topological relationship and dynamic change pattern of the time series features. At the same time, cluster analysis is performed on the spatial coordinate data contained in the core neighborhood dataset to reveal the spatial distribution characteristics of the data points. Cluster analysis models the aggregation behavior of data points, identifies high-density and sparse areas in space, and calculates the spatial distribution density and distribution form of data points. This information forms a spatial feature sequence. The spatial feature sequence is input into the spatial autocorrelation function to calculate the strength and directional characteristics of the spatial association of data points at different scales. The spatial autocorrelation function quantifies the association of data points in different spatial ranges and generates a spatial association matrix, which reflects the spatial distribution law and association pattern of data points to a certain extent. The temporal association matrix and the spatial association matrix are combined with multiple layers of features to form a hierarchical feature set. The initial weight coefficient is assigned to each layer of feature nodes in the hierarchical feature set to generate an initial weight matrix. The allocation of the initial weight coefficient is based on the importance and influence range of the feature. By giving higher weights to high-impact features, it ensures that the feature combination can show higher expressiveness in optimization. The initial weight matrix is back-propagated for training. Back-propagation training continuously adjusts the weight coefficient so that the final target weight matrix can more accurately reflect the intrinsic correlation of data features. During the training process, the weights are iteratively optimized using the error gradient descent method, gradually approaching the global optimal solution, and obtaining the target weight matrix. Perform matrix multiplication on the target weight matrix and the hierarchical feature set to generate a fused feature weight matrix.
[0040] Step S500: input the fused feature weight matrix into a support vector machine classifier to perform abnormal event recognition to obtain abnormal event detection results.
[0041] Specifically, the fusion feature weight matrix is divided into the main feature submatrix and the auxiliary feature submatrix. The division is based on the order of feature importance. Features with stronger distinguishing ability are classified into the main feature submatrix, while other auxiliary information is stored in the auxiliary feature submatrix. The main feature submatrix and the auxiliary feature submatrix are respectively input into the left and right branches of the twin network for processing. Each branch contains three layers of RBF kernel function mapping layers, which are used to perform nonlinear transformation on the input features. The RBF kernel function captures more complex feature patterns and enhances the distinguishability between features at different scales by mapping the original features to a high-dimensional space. In the twin network, the left and right branches perform independent nonlinear mapping on the main features and auxiliary features, respectively, to generate a pair of high-dimensional feature maps, and these mapping results constitute a high-dimensional feature map pair. After obtaining the high-dimensional feature map pair, the Euclidean distance between the mapping pairs is calculated to construct the twin metric space. The twin metric space can capture the potential correlation of abnormal events at the feature level by quantifying the similarity between sample pairs. The similarity matrix is generated by calculating the similarity score of the sample pairs. Based on the similarity matrix, a dual support vector machine classifier is constructed. The dual support vector machine uses a symmetric positive definite kernel matrix to map features and classifies samples by finding the best hyperplane in high-dimensional space. The dual classification hyperplane is constructed with the goal of maximizing the interval between categories, so that samples of different categories can be clearly divided in the feature space. By inputting the classification hyperplane into the nonlinear decision function, the sample points in the feature space are divided by boundaries to generate initial classification labels. The initial classification labels are evaluated for confidence. The confidence evaluation combines the similarity score of the twin network and the distance between the support vector and the classification hyperplane to quantify the credibility of the classification result of each sample. The label confidence is sorted in descending order according to the numerical value, and high-confidence sample points are filtered based on the dynamic threshold strategy. The dynamic threshold strategy dynamically adjusts the filtering conditions according to the distribution characteristics of the specific data set to ensure that the filtered abnormal event candidate set has both high confidence and can cover the potential abnormal patterns in the data. Through this step, a candidate set containing potential abnormal events is obtained. The abnormal event candidate set is clustered by spatiotemporal features to extract abnormal event groups with significant spatiotemporal correlation. By analyzing the spatiotemporal characteristics and abnormality degree of events in each group, the final abnormal event detection results are generated, including the time range, spatial location and quantitative description of the abnormality degree of the event.
[0042] In the embodiment of the present invention, by constructing a distributed storage strategy and a computing node collaboration mechanism, the efficiency problem of large-scale spatiotemporal data processing is solved, so that the system can efficiently process massive data; a neighborhood selection method based on cosine similarity is combined with an optimization strategy of multi-dimensional evaluation indicators to improve the accuracy of abnormal event detection; heterogeneous graph learning and multi-view hypergraph structure are introduced to achieve effective extraction of deep-level features of data and enhance the model's ability to express complex spatiotemporal relationships; a support vector machine classifier with a twin network structure is used to improve the model's classification ability for high-dimensional nonlinear features; multi-layer feature fusion and dynamic threshold strategy are used to improve the reliability of abnormal event detection results; a method for extracting temporal features and spatial features is integrated to achieve a comprehensive characterization of abnormal events; a multi-objective optimization algorithm is used for neighborhood evaluation to ensure the accuracy and interpretability of detection results.
[0043] In a specific embodiment, the process of executing step S100 may specifically include the following steps:
[0044] Perform timestamp conversion on the urban spatiotemporal big data of each data collection point to obtain a time-normalized data set;
[0045] The spatial coordinate information in the time-normalized data set is converted into a unified geographic coordinate system to obtain a coordinate-normalized data set, and the coordinate-normalized data set is subjected to noise elimination and outlier screening to obtain a preprocessed data set;
[0046] Based on the preprocessed data set, a hierarchical index structure of time dimension and space dimension is constructed, and the preprocessed data set is divided into blocks according to the hierarchical index structure, and the preprocessed data set is divided into multiple data slices according to the temporal and spatial proximity relationship to obtain multiple standardized spatiotemporal data sets;
[0047] Establish a distributed file system based on multiple standardized spatiotemporal data sets, build data storage and access mechanisms on multiple computing nodes, and obtain a distributed storage structure;
[0048] Perform load calculation on the standardized spatiotemporal dataset in the distributed storage structure, allocate data processing tasks according to the computing power and storage capacity of each computing node, obtain the task allocation plan, and send the task allocation plan to each computing node, establish a data synchronization channel between nodes, and complete the distributed storage of standardized spatiotemporal datasets between computing nodes.
[0049] Specifically, the timestamps of each data collection point are converted to unify the time data from different sources and in various formats into a consistent standardized time representation. Assuming that the original time data is input in different forms, such as YYYY-MM-DD HH:mm:ss or non-standard time formats, it is uniformly converted to the standard time format . Assume the original representation of the timestamp is , the time offset is (For example, conversion between different time zones), the formula for unified time standards is:
[0050] ;
[0051] in, Represents the normalized timestamp. It is used to correct the deviation of different time zones and different time bases. After the time normalization is completed, the spatial coordinates of the time normalized data set are processed. The data in different coordinate reference systems are uniformly mapped to the standard geographic coordinate system. Let the original coordinates be , the standardized coordinates are , the transformation of the reference coordinate system is performed by the projection transformation formula:
[0052] ;
[0053] in, is the projection transformation matrix, whose parameters depend on the mapping relationship between the original coordinate system and the target coordinate system. After completing the spatial coordinate standardization, the noise and outliers in the data are processed. Noise elimination is achieved by setting upper and lower thresholds. To filter data points:
[0054] ;
[0055] Combined with clustering algorithms (such as DBSCAN), isolated points and abnormal density areas are removed. After preprocessing, a hierarchical index structure is constructed based on time and space dimensions. In the time dimension, hierarchical time index is used to divide the time range into By time step Divide into multiple intervals:
[0056] ;
[0057] In terms of spatial dimension, a grid-based spatial partitioning method is used to divide the coordinate range Divide into several grid units, each grid size is . The data is divided into blocks according to the spatiotemporal proximity relationship, thereby dividing the preprocessed data set into multiple data shards. The sharded standardized spatiotemporal data set is stored in a distributed file system. The core of a distributed file system is to distribute data blocks to multiple computing nodes and generate copies for each data block to ensure fault tolerance. For example, assuming that the storage capacity of each computing node is , the task is to Storage on nodes Data shards, the shard distribution strategy is defined as:
[0058] ;
[0059] Among them, Allocation Represents fragmentation The allocation target node, Size is the shard size. While storing data, load calculation is performed based on the computing power and storage capacity of each node to balance the task allocation of the nodes. Assume that the computing power of each node is , the storage capacity is , the allocation goal is to balance the load so that the computing pressure and storage utilization of all nodes are optimal. The objective function is expressed as:
[0060] ;
[0061] The task allocation plan is sent to each computing node, and a data synchronization channel is established between nodes to support cross-node data query and collaborative processing.
[0062] In a specific embodiment, the process of executing step S200 may specifically include the following steps:
[0063] In the computing node, a feature matrix containing time feature dimensions and space feature dimensions is constructed based on each data point in the standardized spatiotemporal data set, and the feature matrix is segmented into time windows to obtain segmented feature vectors;
[0064] Based on the segmented feature vector, the Euclidean distance and time distance between each data point are calculated, and the Euclidean distance and time distance are weightedly combined to obtain a comprehensive distance matrix;
[0065] The comprehensive distance matrix is input into the cosine similarity calculation function, the similarity coefficient matrix is calculated according to the angle relationship between the data points, and the similarity coefficient matrix is normalized to obtain the target similarity value;
[0066] Compare the target similarity value with the similarity threshold, filter out the data point pairs that meet the threshold condition, obtain the neighborhood candidate point set, and perform density clustering operation on the neighborhood candidate point set, calculate the local density value and distance factor of each data point, and obtain the density sorting sequence;
[0067] The data points are hierarchically processed based on the density sorting sequence to identify groups of data points with similar spatiotemporal distribution characteristics to obtain a hierarchical neighborhood structure. The data points in the hierarchical neighborhood structure are then subjected to association analysis according to the principle of spatiotemporal continuity to extract a subset of data points with significant correlation to obtain an initial neighborhood data set.
[0068] Specifically, for each data point in the standardized spatiotemporal dataset, a feature matrix is constructed, the rows of which represent the unique identifier of the data point, and the columns represent the characteristics of the temporal feature dimension and the spatial feature dimension, including timestamp, time interval, spatial coordinates (such as longitude and latitude), and context attributes. A collection of data points , each data point The characteristic expression is:
[0069] ;
[0070] in, is the timestamp, is the time interval, are spatial coordinates, It is The data point By combining the features of all data points, a complete feature matrix is formed :
[0071] ;
[0072] The feature matrix is segmented into time windows, the time axis is divided into windows of fixed size, and the features of the data points in each window are combined into a segment feature vector. Assume that the size of the time window is , data points The time window belongs to:
[0073] ;
[0074] in, is the starting point of time, is the time window size, express The window number to which it belongs. After aggregating the features of the data points in the same time window, the segmented feature vector is obtained. ,in Indicates the window number. Based on the segmented feature vector, the Euclidean distance and time distance between each data point are calculated. For data points and , whose Euclidean distance and time distance The calculation formulas are:
[0075] ;
[0076] ;
[0077] The two distances are weighted and combined to obtain the comprehensive distance matrix :
[0078] ;
[0079] in, and is the weight coefficient used to adjust the importance of spatial distance and temporal distance. The comprehensive distance matrix is input into the cosine similarity calculation function, and the similarity coefficient matrix is calculated according to the angular relationship of the feature vectors of the data points. :
[0080] ;
[0081] in, represents the vector dot product, and is the modulus of the vector. After the calculation is completed, the similarity coefficient matrix is normalized so that The value of is in the range of [0, 1] to make it easier to filter data points. The target similarity value is compared with the preset similarity threshold. Compare and select data point pairs with similarity higher than the threshold to form a neighborhood candidate point set Perform density clustering on the neighborhood candidate point set and calculate the local density value of each data point and distance factor The local density value is defined as The number of data points with high similarity:
[0082] ;
[0083] in, is the indicator function. Distance factor Indicate point Minimum distance to a data point with higher density. Sort by density , hierarchical processing of data points, identification of data point groups with similar spatiotemporal distribution characteristics, and construction of a hierarchical neighborhood structure. According to the principle of spatiotemporal continuity, correlation analysis is performed on the data points in the hierarchical neighborhood structure to extract a subset of data points with significant correlation. This subset contains key data points that meet the density and similarity conditions, and ultimately forms the initial neighborhood dataset.
[0084] In a specific embodiment, the process of executing step S300 may specifically include the following steps:
[0085] Calculate the spatial-temporal distribution relationship of the data points in the initial neighborhood data set, select the Gaussian kernel function as the basic function for density calculation, and substitute the spatial-temporal coordinates of each data point into the Gaussian kernel function to obtain the density value matrix;
[0086] Based on the density value matrix, a neighborhood evaluation vector including spatial uniformity, temporal continuity and boundary compactness is constructed, and based on the neighborhood evaluation vector, maximizing spatial uniformity, temporal continuity and boundary compactness is set as the optimization goal to obtain a multi-objective optimization function;
[0087] Based on the multi-objective optimization function, population iteration and solution space search are performed to obtain a non-dominated solution set, and weights are assigned to the non-dominated solution set. Solutions of different optimization objectives are weightedly combined according to spatiotemporal correlation to obtain the optimal neighborhood solution.
[0088] Based on the optimal neighborhood solution, the initial neighborhood data set is re-divided, and the core data points that meet the spatiotemporal density threshold are screened out to obtain the core member set;
[0089] Heterogeneous graph learning optimization is performed on the core member set to obtain the core neighborhood dataset.
[0090] Specifically, the temporal and spatial distribution relationship of the data points in the initial neighborhood data set is analyzed. The Gaussian kernel function is selected as the basic function for density calculation, and the smoothness and locality of the Gaussian kernel function are used to map the temporal and spatial coordinates of each data point to a density value. Assume that the initial neighborhood data set contains data points, and the space-time coordinates of each data point are expressed as , then The density value calculation formula centered on a point is:
[0091] ;
[0092] in, Indicates The density value of the data points; It is The space-time coordinates of the data points; and are smoothing factors for spatial and temporal scales, respectively, which are used to control the decay rate of distance weight. This formula generates a density value matrix by considering both spatial distance and temporal distance at the same time: , where each element corresponds to the local density of a data point. Based on the density value matrix, a neighborhood evaluation vector containing spatial uniformity, temporal continuity, and boundary tightness is constructed. Spatial uniformity measures the smoothness of the distribution of data points in space, and its formula is:
[0093] ;
[0094] in, represents the average density value of all data points, Smaller values indicate more uniform spatial distribution. Temporal continuity measures the consistency of data point density in the time dimension and is defined as:
[0095] ;
[0096] in, is the indicator function, is the time window threshold, The larger the value, the stronger the temporal continuity. Boundary density measures the density of data points in the boundary area, and the formula is:
[0097] ;
[0098] Among them, Boundary is a set of boundary points, is the number of boundary points, Higher values indicate tighter boundaries. To maximize spatial uniformity (minimize ), time continuity (maximizing ) and boundary tightness (maximizing ) is the optimization objective, and a multi-objective optimization function is constructed:
[0099] ;
[0100] in, is the optimization variable, which represents the partitioning scheme of the data points. The multi-objective optimization method is used to solve this problem through population iteration and solution space search. The non-dominated sorting genetic algorithm or other multi-objective optimization algorithms are used to generate a non-dominated solution set. , where each solution is a balance between different objectives. After obtaining the non-dominated solution set, each solution in the solution set is weighted and combined according to the temporal and spatial correlation to determine the optimal neighborhood solution. Let the weight vector be , which correspond to the weights of spatial uniformity, temporal continuity and boundary compactness respectively. The objective function of the optimal neighborhood solution is:
[0101] ;
[0102] Among them, the weight The setting is dynamically adjusted according to the specific application scenario. Through this weighted combination, the optimal neighborhood solution is obtained. Based on the optimal neighborhood solution, the initial neighborhood data set is re-divided, and the core data points that meet the spatiotemporal density threshold are screened out to form a core member set. The screening conditions for core members are:
[0103] ;
[0104] in, is a preset density threshold, and CoreSet contains all data points with density values above the threshold. Perform heterogeneous graph learning optimization on the core member set, and construct a heterogeneous graph to model the spatiotemporal relationship between core members. Each core data point is a node in the graph, and the edges between nodes represent the proximity in time or space. Based on the heterogeneous graph, a path set containing time and space semantics is designed, and the graph neural network is used to learn the features of nodes and edges to generate the representation vector of the core data points. By learning the heterogeneous graph, the structure of the core member set is optimized, and finally a core neighborhood data set is formed.
[0105] In a specific embodiment, the execution step performs heterogeneous graph learning optimization on the core member set to obtain a core neighborhood data set, which may specifically include the following steps:
[0106] Based on the core member set, a heterogeneous spatiotemporal information network is constructed, and the spatiotemporal correlation between data points is modeled as multi-type edges to obtain a heterogeneous network structure.
[0107] Based on the heterogeneous network structure, a multi-path set with temporal association, spatial proximity and attribute similarity is designed, and each meta-path in the multi-path set is semantically encoded to obtain a path feature set.
[0108] The path feature set is input into the heterogeneous graph neural network, and message passing and feature aggregation are performed on different types of nodes and edges to obtain node representation vectors;
[0109] A multi-view hypergraph structure is constructed based on the node representation vector, and the association between the time dimension, space dimension and attribute dimension is encoded as a hyperedge to obtain a hierarchical hypergraph.
[0110] Perform multi-level decomposition on the hierarchical hypergraph, extract subgraph structural features at different granularity levels, obtain multi-scale feature representation, and build a graph attention network based on the multi-scale feature representation, calculate the importance weights of features at different levels, and obtain feature fusion weights;
[0111] The feature fusion weights are weightedly combined with the features at each level to obtain an enhanced feature set, and the enhanced feature set is subjected to dimensionality reduction to obtain a core neighborhood data set.
[0112] Specifically, based on the data points in the core member set, the data points are regarded as nodes in a heterogeneous network, and multi-type edges are established through spatiotemporal associations. Assume that the core member set contains data points, each with Indicates that are spatial coordinates, is the timestamp, is the attribute feature vector. The edge types between nodes include time-related edges , spatial adjacent edge Similar edges with attributes The calculation formulas for edge weights are:
[0113] ;
[0114] ;
[0115] ;
[0116] in, , and Represent the edge weights of temporal association, spatial proximity and attribute similarity, respectively. and are the attenuation parameters in time and space, is the modulus of the vector. Through these edge weights, a heterogeneous network structure containing multiple types of nodes and edges is constructed. Based on this heterogeneous network, a multi-path set is designed to capture the complex correlations between time, space and attributes. Each path in the path set is composed of different types of edges. For example, a path is represented as , indicating a node and There is a time relationship, and and There is a spatial proximity relationship. For each path , which is converted into feature representation through semantic encoding:
[0117] ;
[0118] in, is the feature mapping function of the edge type, is the weight of the edge. These path features constitute the path feature set The path feature set is input into the heterogeneous graph neural network to perform message passing and feature aggregation on different types of nodes and edges. In each layer of the graph neural network, the nodes The feature update formula is:
[0119] ;
[0120] in, Is a node In the The characteristics of the layer, is a set of edge types, Is of type The set of neighbor nodes of It is The type weight matrix of the layer, is an activation function. Through multi-layer message passing, a high-dimensional feature representation vector of each node is generated. After completing the node representation learning, a multi-view hypergraph structure is constructed based on the node representation vector, and the association relationship between the time dimension, space dimension and attribute dimension is encoded as a hyperedge. Based on the hypergraph, it is decomposed into multiple levels, and the subgraph structure features are extracted from different granularity levels to generate multi-scale feature representations. The graph attention network is constructed through multi-scale features, and the importance weights of features at different levels are calculated. The weight calculation formula is:
[0121] ;
[0122] in, is the attention vector, It is The feature representation of the layer, It is The weight of the layer features. The feature fusion weights are weighted combined with the features of each layer to generate an enhanced feature set:
[0123] ;
[0124] The enhanced feature set is subjected to dimensionality reduction processing, and the principal component analysis or linear discriminant analysis method is used to reduce the high-dimensional features to a low-dimensional space to generate the final core neighborhood data set.
[0125] In a specific embodiment, the process of executing step S400 may specifically include the following steps:
[0126] Perform frequency domain transformation and statistical feature extraction on the time series data in the core neighborhood data set to obtain a time feature sequence, build a time series association graph based on the time feature sequence, and perform topological structure analysis and path feature extraction on the time series association graph to obtain a time series association matrix;
[0127] Cluster analysis is performed on the spatial coordinate data in the core neighborhood data set to calculate the spatial distribution density and distribution form of the data points to obtain the spatial feature sequence, which is then input into the spatial autocorrelation function to calculate the spatial correlation intensity and directional characteristics at different scales to obtain the spatial correlation matrix.
[0128] Perform multi-layer feature combination on the temporal correlation matrix and the spatial correlation matrix to obtain a hierarchical feature set, and based on the hierarchical feature set, assign initial weight coefficients to the feature nodes of each layer to obtain an initial weight matrix;
[0129] The initial weight matrix is trained by back propagation, the weight coefficients are iteratively optimized to obtain the target weight matrix, and matrix multiplication is performed on the target weight matrix and the hierarchical feature set to obtain the fused feature weight matrix.
[0130] Specifically, the time series data is transformed into the frequency domain and statistical features are extracted to obtain the time feature sequence. Assume that the core neighborhood data set contains data points, and the time series of each data point is expressed as ,in For the The data point in The value at a time point, is the number of time points. In order to reveal the frequency characteristics of the time series, Perform fast Fourier transform, the frequency domain is expressed as:
[0131] ;
[0132] in, is the frequency The complex coefficients under describe the amplitude and phase information of the time series at this frequency. By calculating the amplitude of the frequency component , extract the frequency domain features of the time series, such as the main frequency, frequency bandwidth and spectrum energy. Combined with statistical feature extraction (such as mean, variance, skewness and kurtosis), generate a time feature series:
[0133] ;
[0134] in, is the mean, is the kurtosis. A time series association graph is constructed based on the time feature sequence. Each data point is used as a node of the graph, and the edge weights between nodes are calculated based on the similarity of the time series. Suppose two data points and The time characteristics are and , the edge weight calculation formula is:
[0135] ;
[0136] in, represents the Euclidean distance, is the adjustment parameter of the time scale. By calculating the edge weights between all nodes, the adjacency matrix of the time series association graph is generated. The topological structure of the time series association graph is analyzed, and the global properties of the graph (such as clustering coefficient, average shortest path length) and local path characteristics (such as the shortest path distribution of nodes) are calculated to form a time series association matrix. For the spatial coordinate data in the core neighborhood dataset, cluster analysis is performed to identify the spatial distribution pattern. Assume that the spatial coordinates are , the spatial distribution density and distribution form of data points are calculated by density clustering algorithm (such as DBSCAN). The spatial distribution density is defined as:
[0137] ;
[0138] in, Yes The local density of is the distance threshold, is an indicator function, which indicates the number of points in the neighborhood. Spatial morphological features (such as cluster compactness, boundary point distribution, etc.) are extracted based on the clustering results. The generated spatial feature sequence is expressed as:
[0139] ;
[0140] in, is the cluster compactness, is the boundary feature. The spatial feature sequence is input into the spatial autocorrelation function to calculate the spatial correlation strength and directional characteristics at different scales. The calculation formula for spatial autocorrelation strength (Moran's I) is:
[0141] ;
[0142] in, is the number of data points, is the weight value in the spatial weight matrix, and is the normalized value of the data point, is the average value. By adjusting the parameters of different scales, the multi-scale spatial correlation matrix is calculated. The temporal correlation matrix and the spatial correlation matrix are combined with multiple layers of features to form a hierarchical feature set. Assume that the feature node set is , the hierarchical feature set is expressed as:
[0143]
[0144] in, Indicates Layer feature subset. Assign initial weight coefficients to each layer feature node to generate the initial weight matrix:
[0145] ;
[0146] Perform back propagation training on the initial weight matrix, adjust the weight coefficient by optimizing the objective function (such as minimizing the classification error), and iteratively update the weight matrix:
[0147] ;
[0148] in, is the learning rate, is the gradient of the weight. Finally, the target weight matrix is obtained Perform matrix multiplication on the target weight matrix and the hierarchical feature set to generate the fused feature weight matrix:
[0149] H.
[0150] In a specific embodiment, the process of executing step S500 may specifically include the following steps:
[0151] Perform feature space segmentation on the fused feature weight matrix to obtain a main feature sub-matrix and an auxiliary feature sub-matrix;
[0152] The main feature submatrix and the auxiliary feature submatrix are input into the left and right branches of the twin network respectively. Each branch contains three layers of RBF kernel function mapping layers, and nonlinear transformation of features is performed to obtain high-dimensional feature mapping pairs.
[0153] Calculate the Euclidean distance metric for high-dimensional feature mapping pairs, construct a twin metric space, and calculate the similarity scores between sample pairs in the twin metric space to obtain a similarity matrix;
[0154] Based on the similarity matrix, a dual support vector machine classifier is constructed, and a symmetric positive definite kernel matrix is used for feature mapping to obtain a dual classification hyperplane. The dual classification hyperplane is input into a nonlinear decision function to perform boundary division and category labeling on sample points in the feature space to obtain an initial classification label.
[0155] The confidence of the initial classification label is evaluated, and the classification reliability of the sample is calculated by combining the similarity score of the twin network and the support vector distance to obtain the label confidence. The label confidence is then sorted in descending order according to the numerical value. High-confidence sample points are screened based on the dynamic threshold strategy to obtain the abnormal event candidate set.
[0156] The spatiotemporal feature clustering analysis is performed on the abnormal event candidate set to extract abnormal event groups with significant spatiotemporal correlations, and generate abnormal event detection results including event spatiotemporal features and abnormality levels.
[0157] Specifically, the fusion feature weight matrix is segmented in feature space to extract two sub-matrices: main feature and auxiliary feature. Assume that the fusion feature weight matrix is , whose dimensions are ,in is the sample size, As feature dimensions, the first The important features are divided into main feature sub-matrices , the remaining Auxiliary feature sub-matrix The segmentation formula is as follows:
[0158] ;
[0159] in, , The main feature submatrix and the auxiliary feature submatrix are input into the left and right branches of the twin network respectively. Each branch contains three layers of RBF kernel function mapping layers for nonlinear transformation of features. Assume that the main feature and the auxiliary feature are in the first The representation of the samples is and , the mapping formula of RBF kernel function is:
[0160] ;
[0161] in, For the The mapping value of each core, It is the nuclear center. is the kernel width, In the three-layer mapping, nonlinear features are generated in sequence , and The output of each layer is used as the input of the next layer. After three layers of transformation, the main features and auxiliary features are mapped into high-dimensional feature vectors and . For high-dimensional feature mapping Calculate the Euclidean distance and construct the twin metric space. The calculation formula of the Euclidean distance is:
[0162] ;
[0163] Generate a similarity score matrix in the twin metric space by calculating the distance of all sample pairs , where each element By distance Transformed from:
[0164] ;
[0165] Based on the similarity matrix Construct a dual support vector machine classifier. Assume that the sample label is ,in , the optimization objective function of the dual support vector machine is:
[0166] ;
[0167] in, is the Lagrange multiplier, Is a symmetric positive definite kernel matrix, representing the feature mapping between samples. The dual classification hyperplane obtained by optimization is:
[0168] ;
[0169] Used for classification decisions. is the weight vector of the classification hyperplane. is the bias of the classification hyperplane. The classification result generates the initial classification label through a nonlinear decision function To evaluate the reliability of the initial classification label, a confidence evaluation is performed on each sample. The confidence is calculated by combining the similarity score of the twin network and the support vector distance, and the formula is:
[0170] ;
[0171] Among them, Support is the support vector set, Representation sample The confidence levels are sorted in descending order, and high confidence sample points are screened based on the dynamic threshold strategy:
[0172] ;
[0173] in, is a dynamic threshold, which is adjusted according to the sample confidence distribution. For the selected abnormal event candidate set HighConfSet, a spatiotemporal feature clustering analysis is performed to extract abnormal event groups with significant spatiotemporal correlation. Assume that the spatiotemporal coordinates of the candidate set samples are , calculate clustering by time-space distance and generate abnormal event groups The degree of abnormality of each group is calculated by the density within the group and the time and space span:
[0174] ;
[0175] in, Yes The density of is the size of the group, is the time scale. Generates detection results that include event spatiotemporal characteristics and abnormality levels.
[0176] See also Figure 2 , Figure 2 A schematic block diagram of the structure of the spatiotemporal big data abnormal event detection device 200 provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the spatiotemporal big data abnormal event detection device 200 includes:
[0177] The standardization processing module 210 is used to perform time stamp unification and spatial coordinate standardization processing on the urban spatiotemporal big data to obtain a standardized spatiotemporal data set, and distribute the standardized spatiotemporal data set to multiple computing nodes through a distributed storage strategy;
[0178] The similarity calculation module 220 is used to calculate the cosine similarity of the feature vectors of each data point in the standardized spatiotemporal data set with other data points in the calculation node to obtain a target similarity value, and select the data points whose target similarity value is greater than the similarity threshold according to a preset similarity threshold to obtain an initial neighborhood data set;
[0179] The distribution calculation module 230 is used to calculate the spatiotemporal density distribution of the initial neighborhood data set, and obtain the core neighborhood data set through multi-objective optimization calculation of multi-dimensional neighborhood evaluation indicators;
[0180] The feature extraction module 240 is used to extract time series features and spatial distribution features from the core neighborhood data set to obtain a fusion feature weight matrix;
[0181] The event recognition module 250 is used to input the fused feature weight matrix into the support vector machine classifier to perform abnormal event recognition and obtain abnormal event detection results.
[0182] Through the collaborative cooperation of the above-mentioned components, by building a distributed storage strategy and a computing node collaboration mechanism, the efficiency problem of large-scale spatiotemporal data processing is solved, enabling the system to efficiently process massive data; the neighborhood selection method based on cosine similarity, combined with the optimization strategy of multi-dimensional evaluation indicators, improves the accuracy of abnormal event detection; the introduction of heterogeneous graph learning and multi-view hypergraph structure realizes the effective extraction of deep-level features of data and enhances the model's ability to express complex spatiotemporal relationships; the support vector machine classifier with a twin network structure improves the model's classification ability for high-dimensional nonlinear features; the reliability of abnormal event detection results is improved through multi-layer feature fusion and dynamic threshold strategy; the extraction methods of temporal features and spatial features are integrated to achieve a comprehensive characterization of abnormal events; the multi-objective optimization algorithm is used for neighborhood evaluation to ensure the accuracy and interpretability of detection results.
[0183] See also Figure 3 , Figure 3 A schematic block diagram of the structure of a spatiotemporal big data abnormal event detection device 300 provided in an embodiment of the present application, wherein the spatiotemporal big data abnormal event detection device 300 includes a processor 301 and a memory 302, wherein the processor 301 and the memory 302 are connected via a device bus 303, wherein the memory 302 may include a non-volatile storage medium and an internal memory.
[0184] The non-volatile storage medium can store a computer program. The computer program includes program instructions, and when the program instructions are executed by the processor 301, the processor 301 can execute any of the above-mentioned spatiotemporal big data abnormal event detection methods.
[0185] The processor 301 is used to provide computing and control capabilities to support the operation of the entire spatiotemporal big data abnormal event detection device 300.
[0186] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor 301, the processor 301 can execute any of the above-mentioned spatiotemporal big data abnormal event detection methods.
[0187] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a partial structure related to the present application scheme, and does not constitute a limitation on the spatiotemporal big data abnormal event detection device 300 involved in the present application scheme. The specific spatiotemporal big data abnormal event detection device 300 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0188] It should be understood that the processor 301 may be a central processing unit (CPU), and the processor 301 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0189] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the spatiotemporal big data abnormal event detection device 300 described above can refer to the corresponding process of the aforementioned spatiotemporal big data abnormal event detection method, and will not be repeated here.
[0190] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by one or more processors, the one or more processors implement the method for detecting abnormal events in spatiotemporal big data as provided in the embodiment of the present application.
[0191] The computer-readable storage medium may be an internal storage unit of the spatiotemporal big data abnormal event detection device 300 of the aforementioned embodiment, such as a hard disk or memory of the spatiotemporal big data abnormal event detection device 300. The computer-readable storage medium may also be an external storage device of the spatiotemporal big data abnormal event detection device 300, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped with the spatiotemporal big data abnormal event detection device 300.
[0192] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0194] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting abnormal events in spatiotemporal big data, characterized in that: include: Perform time stamp unification and spatial coordinate standardization processing on urban spatiotemporal big data to obtain a standardized spatiotemporal data set, and distribute the standardized spatiotemporal data set to multiple computing nodes through a distributed storage strategy; In the computing node, the cosine similarity of the feature vectors of each data point in the standardized spatiotemporal data set with other data points is calculated to obtain a target similarity value, and data points whose target similarity values are greater than the similarity threshold are selected according to a preset similarity threshold to obtain an initial neighborhood data set; The spatiotemporal density distribution of the initial neighborhood data set is calculated, and a core neighborhood data set is obtained through a multi-objective optimization operation of a multi-dimensional neighborhood evaluation index; specifically, the method includes: calculating the spatiotemporal distribution relationship of data points on the initial neighborhood data set, selecting a Gaussian kernel function as a basic function for density calculation, and substituting the spatiotemporal coordinates of each data point into the Gaussian kernel function to obtain a density value matrix; constructing a neighborhood evaluation vector including spatial uniformity, temporal continuity and boundary compactness based on the density value matrix, and setting the maximization of spatial uniformity, temporal continuity and boundary compactness as the optimization goal based on the neighborhood evaluation vector to obtain a multi-objective optimization function; performing population iteration and solution space search based on the multi-objective optimization function to obtain a non-dominated solution set, and weighting the non-dominated solution set, and weighting and combining solutions of different optimization goals according to spatiotemporal correlation to obtain an optimal neighborhood solution; re-dividing the initial neighborhood data set based on the optimal neighborhood solution, screening out core data points that meet the spatiotemporal density threshold, and obtaining a core member set; performing heterogeneous graph learning optimization on the core member set to obtain a core neighborhood data set; Extracting time series features and spatial distribution features from the core neighborhood data set to obtain a fusion feature weight matrix; The fused feature weight matrix is input into a support vector machine classifier to perform abnormal event recognition to obtain an abnormal event detection result.
2. The method for detecting abnormal events in spatiotemporal big data according to claim 1, characterized in that: The time stamp unification and spatial coordinate standardization processing of the urban spatiotemporal big data are performed to obtain a standardized spatiotemporal data set, and the standardized spatiotemporal data set is distributed to multiple computing nodes through a distributed storage strategy, including: Perform timestamp conversion on the urban spatiotemporal big data of each data collection point to obtain a time-normalized data set; The spatial coordinate information in the time-normalized data set is converted into a unified geographic coordinate system to obtain a coordinate-normalized data set, and the coordinate-normalized data set is subjected to noise elimination and outlier screening to obtain a preprocessed data set; Constructing a hierarchical index structure of time dimension and space dimension based on the preprocessed data set, and performing block processing on the preprocessed data set according to the hierarchical index structure, dividing the preprocessed data set into multiple data slices according to the spatiotemporal proximity relationship, and obtaining multiple standardized spatiotemporal data sets; Establishing a distributed file system based on the multiple standardized spatiotemporal data sets, building a data storage and access mechanism on multiple computing nodes, and obtaining a distributed storage structure; Perform load calculation on the standardized spatiotemporal data set in the distributed storage structure, allocate data processing tasks according to the computing power and storage capacity of each computing node, obtain a task allocation plan, and send the task allocation plan to each computing node, establish a data synchronization channel between the nodes, and complete the distributed storage of the standardized spatiotemporal data set between the computing nodes.
3. The method for detecting abnormal events in spatiotemporal big data according to claim 1, characterized in that: In the computing node, the cosine similarity of the feature vectors of each data point in the standardized spatiotemporal data set with other data points is calculated to obtain a target similarity value, and data points whose target similarity values are greater than the similarity threshold are selected according to a preset similarity threshold to obtain an initial neighborhood data set, including: In the computing node, a feature matrix including a time feature dimension and a space feature dimension is constructed based on each data point in the standardized spatiotemporal data set, and the feature matrix is segmented into time windows to obtain segmented feature vectors; Calculate the Euclidean distance and the temporal distance between each data point based on the segmented feature vector, and perform weighted combination of the Euclidean distance and the temporal distance to obtain a comprehensive distance matrix; The comprehensive distance matrix is input into the cosine similarity calculation function, a similarity coefficient matrix is calculated according to the angle relationship between the data points, and the similarity coefficient matrix is normalized to obtain a target similarity value; Compare the target similarity value with the similarity threshold, filter out data point pairs that meet the threshold condition, obtain a neighborhood candidate point set, perform density clustering operation on the neighborhood candidate point set, calculate the local density value and distance factor of each data point, and obtain a density sorting sequence; The data points are hierarchically processed based on the density sorting sequence, and groups of data points with similar spatiotemporal distribution characteristics are identified to obtain a hierarchical neighborhood structure. The data points in the hierarchical neighborhood structure are subjected to association analysis according to the principle of spatiotemporal continuity, and a subset of data points with significant correlation is extracted to obtain an initial neighborhood data set.
4. The method for detecting abnormal events in spatiotemporal big data according to claim 1, characterized in that: The performing heterogeneous graph learning optimization on the core member set to obtain a core neighborhood data set includes: Building a heterogeneous spatiotemporal information network based on the core member set, modeling the spatiotemporal association relationship between data points as multi-type edges, and obtaining a heterogeneous network structure; Based on the heterogeneous network structure, a multi-path set including time association, spatial proximity, and attribute similarity is designed, and each meta-path in the multi-path set is semantically encoded to obtain a path feature set; The path feature set is input into a heterogeneous graph neural network, message passing and feature aggregation are performed on different types of nodes and edges to obtain a node representation vector; Building a multi-view hypergraph structure based on the node representation vector, encoding the association relationship among the time dimension, the space dimension and the attribute dimension into a hyperedge, and obtaining a hierarchical hypergraph; Decomposing the hierarchical hypergraph at multiple levels, extracting subgraph structural features at different granularity levels, obtaining multi-scale feature representations, constructing a graph attention network based on the multi-scale feature representations, calculating the importance weights of features at different levels, and obtaining feature fusion weights; The feature fusion weights are weighted and combined with the features at each level to obtain an enhanced feature set, and the enhanced feature set is subjected to dimensionality reduction processing to obtain a core neighborhood data set.
5. The method for detecting abnormal events in spatiotemporal big data according to claim 1, characterized in that: The extracting of time series features and spatial distribution features from the core neighborhood data set to obtain a fusion feature weight matrix includes: Performing frequency domain transformation and statistical feature extraction on the time series data in the core neighborhood data set to obtain a time feature sequence, constructing a time series association graph based on the time feature sequence, and performing topological structure analysis and path feature extraction on the time series association graph to obtain a time series association matrix; Performing cluster analysis on the spatial coordinate data in the core neighborhood data set, calculating the spatial distribution density and distribution form of the data points, obtaining a spatial feature sequence, and inputting the spatial feature sequence into a spatial autocorrelation function, calculating the spatial correlation intensity and directional characteristics at different scales, and obtaining a spatial correlation matrix; Performing multi-layer feature combination on the temporal correlation matrix and the spatial correlation matrix to obtain a hierarchical feature set, and assigning initial weight coefficients to feature nodes of each layer based on the hierarchical feature set to obtain an initial weight matrix; The initial weight matrix is trained by back propagation, the weight coefficients are iteratively optimized to obtain a target weight matrix, and a matrix multiplication operation is performed on the target weight matrix and the hierarchical feature set to obtain a fused feature weight matrix.
6. The method for detecting abnormal events in spatiotemporal big data according to claim 1, characterized in that: The step of inputting the fusion feature weight matrix into a support vector machine classifier to identify abnormal events and obtain abnormal event detection results includes: Performing feature space segmentation on the fused feature weight matrix to obtain a main feature sub-matrix and an auxiliary feature sub-matrix; The main feature submatrix and the auxiliary feature submatrix are respectively input into the left and right branches of the twin network, each branch includes three layers of RBF kernel function mapping layers, and nonlinear transformation of features is performed to obtain a high-dimensional feature mapping pair; Calculating the Euclidean distance metric for the high-dimensional feature mapping pair, constructing a twin metric space, and calculating the similarity scores between the sample pairs in the twin metric space to obtain a similarity matrix; Based on the similarity matrix, a dual support vector machine classifier is constructed, a symmetric positive definite kernel matrix is used for feature mapping to obtain a dual classification hyperplane, and the dual classification hyperplane is input into a nonlinear decision function to perform boundary division and category labeling on sample points in the feature space to obtain an initial classification label; Perform confidence evaluation on the initial classification labels, calculate the classification reliability of the samples by combining the similarity score of the twin network and the support vector distance, obtain label confidence, and arrange the label confidence in descending order according to the numerical value, filter high-confidence sample points based on the dynamic threshold strategy, and obtain the abnormal event candidate set; A spatiotemporal feature clustering analysis is performed on the abnormal event candidate set to extract abnormal event groups with significant spatiotemporal correlation, and an abnormal event detection result including event spatiotemporal features and abnormality degree is generated.
7. A device for detecting abnormal events in spatiotemporal big data, characterized in that: The method for detecting abnormal events of spatiotemporal big data according to any one of claims 1 to 6 comprises: A standardization processing module is used to unify the timestamps and standardize the spatial coordinates of the urban spatiotemporal big data to obtain a standardized spatiotemporal data set, and distribute the standardized spatiotemporal data set to multiple computing nodes through a distributed storage strategy; A similarity calculation module is used to calculate the cosine similarity of the feature vectors of each data point in the standardized spatiotemporal data set with other data points in the calculation node to obtain a target similarity value, and select data points whose target similarity values are greater than the similarity threshold according to a preset similarity threshold to obtain an initial neighborhood data set; A distribution calculation module is used to calculate the spatiotemporal density distribution of the initial neighborhood data set, and obtain a core neighborhood data set through a multi-objective optimization operation of multi-dimensional neighborhood evaluation indicators; A feature extraction module is used to extract time series features and spatial distribution features from the core neighborhood data set to obtain a fusion feature weight matrix; The event recognition module is used to input the fusion feature weight matrix into the support vector machine classifier to perform abnormal event recognition and obtain abnormal event detection results.
8. A device for detecting abnormal events in spatiotemporal big data, characterized in that: The spatiotemporal big data abnormal event detection device comprises: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory so that the spatiotemporal big data abnormal event detection device executes the spatiotemporal big data abnormal event detection method as described in any one of claims 1-6.
9. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the method for detecting abnormal events in spatiotemporal big data as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Abnormity detection method and device based on unsupervised and hierarchical clustering
CN117235554A
Customer portrait key data mining method and system based on space-time big data
CN118797542A