Urban anomaly detection method and device based on spatiotemporal data, medium and equipment
By combining spatiotemporal feature data in both temporal and spatial dimensions for anomaly detection, the problem of low anomaly detection accuracy in existing technologies is solved, achieving higher detection accuracy and spatiotemporal correlation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JINGDONG CITY BEIJING DIGITS TECH CO LTD
- Filing Date
- 2023-02-09
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies for anomaly detection based on single-source data have a high misclassification rate, while anomaly detection based on multi-source data fails to fully explore the potential relationships between data and has low spatiotemporal correlation, resulting in low anomaly detection accuracy.
By acquiring spatiotemporal feature data of various urban areas, anomaly detection is performed in both temporal and spatial dimensions. Anomaly judgment is made from multiple perspectives by combining anomaly information in both temporal and spatial dimensions, and spatiotemporal correlation feature data is used for processing.
It improves the accuracy and spatiotemporal correlation of abnormal crowd flow detection, thereby enhancing the accuracy of anomaly detection.
Smart Images

Figure CN116340871B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of smart city technology and urban governance technology, and in particular to urban anomaly detection methods, devices, media and equipment based on spatiotemporal data. Background Technology
[0002] Crowd movement is a common and frequent migration phenomenon in cities. Commuting and traveling are concrete carriers of crowd movement, while the stillness, gathering, dispersion, and movement of crowds are all macroscopic forms of crowd movement. Anomaly detection of crowd movement can also be used in multiple fields such as traffic control, risk assessment, epidemic prevention and control, and public safety, such as predicting urban road traffic congestion and predicting and alerting crowd stampede events.
[0003] Anomalies in population movement typically refer to changes in population flow within a specific area of a city over a period of time that deviate from the expected population movement patterns for that area. Currently, anomaly detection methods based on single-source data and multi-source data fusion can be used to detect anomalies in population movement.
[0004] In the process of realizing this invention, it was found that at least the following technical problems exist in the prior art: Anomaly detection based on single-source data can usually only reflect abnormal events from a limited perspective, and often has a high misclassification rate in anomaly detection tasks; Anomaly detection methods based on multi-source data fail to fully explore and utilize the potential relationships between multi-source data, the spatiotemporal correlation between data is low, and the deep semantic information of the data is ignored, resulting in low anomaly detection accuracy. Summary of the Invention
[0005] This invention provides a method, apparatus, medium, and equipment for urban anomaly detection based on spatiotemporal data, which improves the accuracy of anomaly detection by detecting anomalies in population flow in both temporal and spatial dimensions based on spatiotemporal data.
[0006] According to one aspect of the present invention, a method for urban anomaly detection based on spatiotemporal data is provided, comprising:
[0007] Acquire spatiotemporal feature data of the city to be detected, wherein the spatiotemporal feature data includes the temporal feature data of each region of the city to be detected;
[0008] For any spatiotemporal characteristic data of a region, determine the temporal flow anomaly information of the region, and determine the spatial flow anomaly information of the region;
[0009] Based on the temporal and spatial traffic anomaly information of each region, abnormal regions are identified in each region.
[0010] According to another aspect of the present invention, an urban anomaly detection device based on spatiotemporal data is provided, comprising:
[0011] The spatiotemporal data acquisition module is used to acquire spatiotemporal feature data of the city to be detected, wherein the spatiotemporal feature data includes the temporal feature data of each region of the city to be detected;
[0012] The first anomaly information determination module is used to determine the temporal traffic anomaly information of any region based on the spatiotemporal characteristic data of that region.
[0013] The second anomaly information determination module is used to determine the spatial flow anomaly information of any region based on the spatiotemporal characteristic data of that region.
[0014] The abnormal region determination module is used to determine abnormal regions in each region based on the temporal and spatial traffic anomaly information of each region.
[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the urban anomaly detection method based on spatiotemporal data according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the urban anomaly detection method based on spatiotemporal data as described in any embodiment of the present invention.
[0020] The technical solution of this invention acquires spatiotemporal feature data of the city to be detected, performs anomaly detection in both the temporal and spatial dimensions, and makes anomaly judgments on each region based on the anomaly information in the temporal and spatial dimensions to identify areas with abnormal population flow. In the processing, the feature data with temporal and spatial correlation are processed to take into account both temporal and spatial features, improve the spatiotemporal correlation in the processing, and improve the accuracy of anomaly detection through multi-view anomaly detection.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of an urban anomaly detection method based on spatiotemporal data provided in an embodiment of the present invention;
[0024] Figure 2 This is a schematic diagram of a region division provided in an embodiment of the present invention;
[0025] Figure 3 This is a schematic diagram of the three-dimensional dataset provided in an embodiment of the present invention;
[0026] Figure 4 This is a schematic diagram of the inflow prediction results provided in an embodiment of the present invention;
[0027] Figure 5 This is a schematic diagram of the outflow prediction results provided in an embodiment of the present invention;
[0028] Figure 6 This is a schematic diagram of the structure of a traffic prediction model provided in an embodiment of the present invention;
[0029] Figure 7 This is a schematic diagram of abnormal region points within the metric space provided in an embodiment of the present invention;
[0030] Figure 8 This is a flowchart of an urban anomaly detection method based on spatiotemporal data provided in an embodiment of the present invention;
[0031] Figure 9 This is a schematic diagram of a region clustering result provided in an embodiment of the present invention;
[0032] Figure 10 This is a schematic diagram illustrating how to optimize clustering results based on spatial relationship constraints, as provided in an embodiment of the present invention.
[0033] Figure 11 This is a schematic diagram of the system structure for implementing a spatiotemporal data-based urban anomaly detection method, as improved in this embodiment of the invention.
[0034] Figure 12 This is a flowchart of a spatiotemporal data urban anomaly detection method provided by an embodiment of the present invention;
[0035] Figure 13This is a schematic diagram of the structure of an urban anomaly detection device based on spatiotemporal data provided in an embodiment of the present invention;
[0036] Figure 14 This is a schematic diagram of the structure of an electronic device that implements an embodiment of the present invention. Detailed Implementation
[0037] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0039] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0040] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0041] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0042] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0043] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0044] Figure 1 This is a flowchart of an urban anomaly detection method based on spatiotemporal data provided in an embodiment of the present invention. This embodiment is applicable to detecting abnormal population movement within a city. The method can be executed by an urban anomaly detection device based on spatiotemporal data, which can be implemented in hardware and / or software and can be configured in electronic devices such as mobile phones, tablets, computers, or servers. Figure 1 As shown, the method includes:
[0045] S110. Obtain the spatiotemporal feature data of the city to be detected, wherein the spatiotemporal feature data includes the temporal feature data of each region of the city to be detected.
[0046] S120. For any spatiotemporal characteristic data of a region, determine the temporal flow anomaly information of the region, and determine the spatial flow anomaly information of the region.
[0047] S130. Based on the temporal and spatial traffic anomaly information of each region, identify the abnormal regions in each region.
[0048] Spatiotemporal characteristic data can include time information, spatial information, and characteristic data associated with time and spatial information. This characteristic data includes, but is not limited to, meteorological data and crowd trajectory data. Meteorological data can include, but is not limited to, time, longitude, latitude, temperature, humidity, visibility, precipitation, and dew point, and can be obtained from meteorological data databases. Crowd trajectory data can be determined by acquiring signaling data from the electronic devices carried by each user in the crowd. For example, for a certain area within a certain time period, by using the signaling data of newly added and departing electronic devices, the inflow and outflow information of the crowd in that area during that time period can be determined.
[0049] In cities, spatial location information is typically identified using latitude and longitude. Each anomalous event is objectively confined to a certain spatial range. In this embodiment, the city to be detected is divided into regions for anomaly detection, facilitating better representation and modeling of anomalous events. The regional division of the city can be predetermined. During anomaly detection, the regional distribution of the city to be detected is obtained, and anomaly detection is performed simultaneously on each region within that distribution. Correspondingly, obtaining the spatiotemporal feature data of the city to be detected includes: obtaining the regional distribution of the city to be detected and obtaining feature data of each region within a preset time range.
[0050] In this embodiment, the granularity of the city division is determined with detectable accuracy. Optionally, the city to be detected is divided into regions based on a preset granularity, resulting in multiple regions. Each region can be a regular shape, and the area of each region can be the same. For example, the latitude and longitude range of the city to be detected is iteratively bisected to obtain multiple rectangular regions. See [example description]. Figure 2 , Figure 2 This is a schematic diagram of a region division provided by an embodiment of the present invention. Figure 2 This is merely one example. Optionally, the city to be detected can be divided into regions based on regional functions. For example, the city to be detected may include residential areas, commercial areas, industrial areas, leisure areas, etc. The division can be based on each functional area within the city to be detected, resulting in different areas for each region. Each region can be a regular-shaped region or an irregular-shaped region; there is no limitation on this. Optionally, the city to be detected can be divided into regions based on preset levels of administrative regions, such as streets, communities, towns, etc. Optionally, the city to be detected can be divided into regions based on road lines. In this embodiment, the method of dividing the city to be detected into regions is not limited. Regions can be divided according to anomaly detection requirements. The regional distribution of the divided city to be detected is stored. When an anomaly detection command is detected for the city to be detected, the stored regional distribution can be called, and the temporal feature data of each region, i.e., spatiotemporal feature data, can be obtained based on this regional distribution.
[0051] Optionally, each region is assigned a unique coded index for identification and lookup. The coded index can be a region code or coded data based on the region's spatial coordinates, such as a hash value; there are no restrictions on this.
[0052] Based on spatial data, temporal data, and corresponding feature data from each region, a three-dimensional spatiotemporal feature dataset is formed. For example, see [link to example dataset]. Figure 3 , Figure 3This is a schematic diagram illustrating the process of matching temporal feature data to a raster after region segmentation, as provided in an embodiment of the present invention, to form a three-dimensional dataset. Figure 3 This includes spatiotemporal coordinates and corresponding feature data for each coordinate, such as, but not limited to, crowd flow and meteorological data. Accordingly, the spatiotemporal feature dataset can be S = { <r i ,t j ,v k >, 1≤i≤m, 1≤j≤t, 1≤k≤n}, where r i This could be spatial data, such as the encoded index of a region. j This refers to time-based data, such as time segments. k For each spatiotemporal coordinate<r,t> The corresponding unique feature data vector v = <e1,e2,…,e n >
[0053] In this embodiment, based on spatiotemporal feature data, anomalies are detected in each area of the city to be tested in both temporal and spatial dimensions. Based on the anomaly detection results in both dimensions, it is determined whether any anomalies exist in each area. It should be noted that anomaly detection for each area is performed by predicting and detecting the flow of people in the area to determine whether there is abnormal inflow and / or outflow of people. Each area has potential population flow patterns; for example, commercial areas experience large inflows of people during morning and evening peak hours, while residential areas experience large outflows of people during the same periods.
[0054] In this embodiment, by predicting the crowd flow in each area at each time period, and comparing the predicted crowd flow with the actual crowd flow, abnormal crowd flow information can be determined. In some embodiments, determining the abnormal flow information of the area in time includes: inputting the feature data of the area within a preset time range into a pre-trained flow prediction model to obtain the flow prediction data of the area; and determining the abnormal flow information of the area in time based on the actual flow data and the flow prediction data of the area.
[0055] The feature data of the region within the preset time series range can be feature data of the region in multiple consecutive time periods before the current time. For example, the duration of each time period can be preset, such as one hour or half an hour, etc., without limitation. Based on the preset traffic prediction model, the feature data of the region in multiple consecutive time periods before the current time are predicted to obtain traffic prediction data of the region in the current time period and / or multiple future time periods of the current time.
[0056] Traffic forecasting models can predict traffic data for one or more time periods, depending on the forecasting requirements. When the model is used to predict traffic for the current time period, it performs time-series traffic forecasting based on the time intervals corresponding to each period. Given the actual traffic data for the current time period, the difference between the actual and predicted traffic data is used to determine traffic anomalies in the current time period. When the model is used to predict traffic for multiple future time periods, it performs time-series traffic forecasting based on the time intervals corresponding to each period. For the current time period, the final traffic forecast for the current time period is determined based on traffic forecasts from multiple historical time periods. For example, this could be achieved by averaging the traffic forecasts from multiple historical time periods, or by weighting the traffic forecasts from multiple historical time periods. The difference between the actual and final traffic forecast data is then used to determine traffic anomalies in the current time period.
[0057] In some embodiments, the traffic prediction data includes inflow prediction data and outflow prediction data, and the actual traffic data includes actual inflow data and actual outflow data; correspondingly, the time-series traffic anomaly information of the region includes inflow anomaly information and outflow anomaly information, that is, inflow anomaly information is determined by the difference between inflow prediction data and actual inflow data, and outflow anomaly information is determined based on the difference between outflow prediction data and actual outflow data. For example, see... Figure 4 and Figure 5 ,in, Figure 4 This is a schematic diagram of the inflow prediction results provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of the outflow prediction results provided in an embodiment of the present invention. Figure 4 and Figure 5 The solid line represents the traffic forecast curve, the dashed line represents the actual traffic data, and the solid area corresponds to the difference between the forecast and actual traffic data, i.e., traffic anomaly information. Figure 4 and Figure 5 Anomalies in inflow and outflow can be displayed separately, facilitating the viewing of trends and detailed changes in crowd flow data. In this embodiment, predicted flow data, actual flow data, and flow anomaly information for each area can be displayed in the form of visual curves or charts, allowing inspection personnel to intuitively view the anomaly detection results for each area.
[0058] Based on the above embodiments, the traffic prediction model can be a neural network model, a logistic regression model, etc. The structure and type of the traffic prediction model are not limited, as long as it has traffic prediction functionality. This traffic prediction model can be trained based on traffic prediction data and actual traffic data from historical periods. For example, feature data of any region within a preset time series prior to a historical period is obtained, and this feature data is input into the traffic prediction model during training to obtain traffic prediction data. A loss function is determined based on the traffic prediction data and actual traffic data from that historical period, and the model parameters are adjusted in reverse. This training process is iteratively executed until the training termination condition is met, resulting in a completed traffic prediction model.
[0059] In some embodiments, the traffic prediction model includes a vector transformation layer, an encoding layer, a decoding layer, and a fully connected layer; for example, see [link to example]. Figure 6 , Figure 6 This is a schematic diagram of a traffic prediction model provided in an embodiment of the present invention. The vector transformation layer converts the input feature data into vector data, which is then input to the encoding layer. The encoding layer includes multiple loop units, each of which sequentially processes the vector data corresponding to each time-series feature data and inputs the processing result to the decoding layer. The decoding layer includes multiple loop units, each of which is connected to a fully connected layer and outputs the processing result to the fully connected layer. The fully connected layer outputs the corresponding traffic prediction data.
[0060] The vector transformation layer can be an enbadding layer, used to map the input data into vector data. Optionally, the traffic prediction model may include two parallel vector transformation layers. One vector transformation layer is used to perform vector transformation on the crowd traffic data in the spatiotemporal feature data, and the other vector transformation layer is used to perform vector transformation on the time static data and meteorological data in the spatiotemporal feature data. The vector data obtained from the transformations are concatenated and input into the encoding layer.
[0061] The coding layer includes multiple recurrent units. The feature data of the predicted region falls within a preset time range. The feature data of each time period can serve as input information for one recurrent unit. Simultaneously, the processing result of the previous recurrent unit on the feature data of the previous time period is input to the next recurrent unit. (See [link]). Figure 6 The recurrent units in the encoding layer can be, for example, LSTM (Long Short-Term Memory) network blocks, RNN (Recurrent Neural Network) blocks, etc., without limitation. The output information of the last recurrent unit in the encoding layer serves as the input information for the decoding layer.
[0062] The decoding layer comprises multiple recurrent units, which can be, for example, LSTM or RNN network blocks, without limitation. Each recurrent unit in the decoding layer is connected to a fully connected layer. The output information of each recurrent unit in the decoding layer is output to the connected fully connected layer and the next recurrent unit, respectively. The fully connected layer converts the implicit state information output by the recurrent units into two-dimensional traffic prediction data.
[0063] For example, see Table 1, Table 1 is Figure 6 Parameters of each network layer.
[0064]
[0065] Building upon the above embodiments, anomaly detection is performed on the spatial dimensions of each region of the city to be detected based on spatiotemporal feature data, determining anomaly detection information in the spatial dimension. Specifically, this can be achieved by using a spatial detection model to predict and process the spatiotemporal feature data corresponding to each region, obtaining anomaly information in the spatial dimension. This can involve classifying regions to determine their region types, then calling the corresponding spatial detection model based on the region type to process the spatiotemporal feature data of that region and obtain anomaly detection information. The region type can be based on regional function classification, such as commercial areas, residential areas, etc.; or it can be obtained by clustering traffic data of each region over historical periods. Historical spatiotemporal feature data corresponding to different region types are obtained, and spatiotemporal detection models corresponding to each region type are trained to achieve spatial anomaly detection for each region.
[0066] For any region, temporal and spatial anomaly information are combined to obtain the anomaly vectors for each region, and anomaly determination is performed on the region. In some embodiments, the temporal and spatial anomaly information can be calculated based on preset calculation rules, such as weighted processing, to obtain anomaly calculation results. Anomaly determination is then made based on the anomaly determination threshold. If the anomaly calculation result is greater than the anomaly determination threshold, it is determined that there is anomaly in the population flow in the region. In some embodiments, the temporal and spatial anomaly information can be determined separately based on preset determination conditions. If one or more of the temporal and spatial anomaly information meet the anomaly determination conditions, it is determined that there is anomaly in the population flow in the region. The preset determination conditions corresponding to the temporal and spatial anomaly information can be different.
[0067] In some embodiments, determining abnormal regions in each region based on temporal and spatial traffic anomaly information includes: mapping the temporal and spatial traffic anomaly information of each region to a metric space; and determining abnormal regions based on the location information of each region in the metric space.
[0068] For example, the metric space can be a three-dimensional space, where the three dimensions can be spatial dimension anomalies, temporal dimension inflow anomalies, and temporal dimension outflow anomalies. Based on the anomaly information vectors of each region, the position of the region within the aforementioned metric space is determined. Based on the positional information of each region point in the metric space, anomaly regions are identified; for example, edge region points among the region points in the metric space are identified as anomaly region points. Optionally, for the current region point, other region points are traversed, and the positional information of the current region point is compared sequentially with the positional information of each other region point. If at least one dimension value in the positional information of any other location point is less than the corresponding dimension value of the current region point's positional information, the comparison continues with the positional information of the next other location point. If there exists a location point whose positional information has all dimension values greater than the corresponding dimension values of the current region point's positional information, then the current region point is determined not to be an anomaly region point, and the comparison with other region points is stopped. If there are no other location points whose positional information has all dimension values greater than the corresponding dimension values of the current region point's positional information, then the current region point is determined to be an anomaly region point.
[0069] In some embodiments, the temporal and spatial traffic anomaly information of each region can be mapped to a skyline metric space based on a skyline algorithm. Within the skyline metric space, skyline points are identified among the region points, and the regions corresponding to these skyline points are designated as anomaly regions. A skyline point can be a point among all region points where no other point has a value greater than it in all dimensions. In this embodiment, a skyline point identification rule can be pre-set, and the skyline point is identified in the skyline metric space by invoking this rule. Furthermore, skyline points can be displayed distinctly within the skyline metric space for easy and intuitive viewing. For example, see [link to example]. Figure 7 , Figure 7 This is a schematic diagram of abnormal region points in the metric space provided in the embodiments of the present invention.
[0070] The technical solution of this embodiment acquires spatiotemporal feature data of the city to be detected, performs anomaly detection in both the temporal and spatial dimensions, and makes anomaly judgments on each region based on the anomaly information in the temporal and spatial dimensions to identify areas with abnormal population flow. In the processing, the feature data with temporal and spatial correlation is processed to take into account both temporal and spatial features, improve the spatiotemporal correlation in the processing, and improve the accuracy of anomaly detection through multi-view anomaly detection.
[0071] Figure 8 This is a flowchart of a method for detecting urban anomalies based on spatiotemporal data provided by an embodiment of the present invention. It is a refinement of the above embodiment. Optionally, determining the spatial traffic anomaly information of the region includes: calling the regional anomaly identification model corresponding to each regional type based on the regional category to which the region belongs at each time scale; processing the spatiotemporal feature data of the region based on the regional anomaly identification model corresponding to each regional type to determine the anomaly information corresponding to each regional type; and fusing the anomaly information corresponding to the regional types to obtain the spatial traffic anomaly information of the region. Figure 8 As shown, the method includes:
[0072] S210. Obtain the spatiotemporal feature data of the city to be detected, wherein the spatiotemporal feature data includes the temporal feature data of each region of the city to be detected.
[0073] S220. For the spatiotemporal characteristic data of any region, determine the traffic anomaly information of the region in time series.
[0074] S230. Based on the region category to which the region belongs at each time scale, call the region anomaly identification model corresponding to each region type, process the spatiotemporal feature data of the region based on the region anomaly identification model corresponding to each region type, and determine the anomaly information corresponding to each region type.
[0075] S240. The abnormal information corresponding to the region type is fused to obtain the spatial traffic abnormality information of the region.
[0076] S250. Based on the temporal and spatial traffic anomaly information of each region, identify the abnormal regions in each region.
[0077] Different regions exhibit different population flow patterns, which can be used to classify regions into regional types. It should be noted that the population flow pattern of each region is not solely related to regional functional division. In this embodiment, regional types can be classified based on spatiotemporal feature data corresponding to the region across multiple time scales, improving the accuracy of regional classification. These time scales may include, but are not limited to, temporal proximity, periodicity, and trends. Accordingly, each region can correspond to a regional type at each of the aforementioned time scales. Historical spatiotemporal feature data can be pre-classified across these multiple time scales to obtain and store the regional types for each region at these time scales.
[0078] Optionally, the method for determining the region category of each region at each time scale includes: performing clustering processing on each region at each time scale based on the historical spatiotemporal feature data of each region, and determining the region category of each region at each time scale, wherein each cluster in the clustering processing result corresponds to a region type.
[0079] The time scale may include, but is not limited to, temporal proximity, periodicity, and trend. For any time scale, based on the characteristics of the time scale, historical spatiotemporal feature data for clustering processing of each region at that time scale are determined. Based on the determined historical spatiotemporal feature data of each region, the regions are clustered to obtain multiple clusters at that time scale. Each cluster may include at least one region, and each cluster corresponds to a region type. Accordingly, clustering processing is performed on the temporal proximity of each region in the city to be tested to determine the region type of each region at the proximity scale; clustering processing is performed on the temporal periodicity of each region in the city to be tested to determine the region type of each region at the periodic scale; and clustering processing is performed on the temporal trend of each region in the city to be tested to determine the region type of each region at the trend scale.
[0080] Specifically, clustering is performed on each region at each time scale to determine the region category to which each region belongs at each time scale. This includes: clustering each region based on traffic characteristic data of each region within a preset adjacent time period to obtain a region clustering result based on proximity; clustering each region based on traffic characteristic data of each region within a preset periodic time period to obtain a region clustering result based on periodicity; and clustering each region based on traffic characteristic data of each region within a preset continuous time period to obtain a region clustering result based on trend.
[0081] On a temporal proximity scale, traffic characteristic data for each region within preset adjacent time periods are determined for clustering processing. For example, for each region, characteristic data from multiple adjacent time periods are determined to avoid the problem of randomness in characteristic data from a single adjacent time period. For example, the traffic characteristic data for each region within the preset adjacent time periods could be characteristic data from 7:00-8:00 AM and 8:00-9:00 AM, and characteristic data from 6:00-7:00 PM and 7:00-8:00 PM. The preset adjacent time periods are the same for all regions to improve the accuracy of clustering.
[0082] To address the periodicity of time, traffic characteristic data for each region within a preset periodic time frame are determined for clustering on a periodic scale. For each region, multiple sets of characteristic data for periodic time periods are identified to avoid the problem of randomness in characteristic data for a single periodic time period. For example, within a three-month time frame, characteristic data from every Monday can be used as a set of characteristic data for periodic time periods; within a one-month time frame, characteristic data from 7:00 to 8:00 AM each day can be used as a set of characteristic data for periodic time periods. Here, there are no restrictions on the duration of the period or the length of the time period.
[0083] To assess the temporal trend, traffic characteristic data for each region within a preset continuous time period is determined and used for clustering on a trend scale. For each region, multiple sets of traffic characteristic data within preset continuous time periods are identified to avoid the problem of randomness in characteristic data from a single continuous time period. For example, characteristic data from multiple consecutive time periods within a day can be used as multiple sets of characteristic data within a continuous time period, or characteristic data from multiple consecutive time periods within a week can be used as multiple sets of characteristic data within a continuous time period. The number and duration of the continuous time periods are not limited here.
[0084] In this embodiment, clustering processing can be implemented based on one or more of the following clustering algorithms: bi-kmeans clustering algorithm, kmene clustering algorithm, kmenes++ clustering algorithm, bi-kmeans clustering algorithm, dbscan clustering algorithm, and optics clustering algorithm, etc., without limitation. Clustering processing at different time scales can be implemented using the same clustering algorithm or different clustering algorithms, which can be set according to user needs.
[0085] Based on the aforementioned arbitrary clustering algorithm, the historical spatiotemporal feature data at each time scale are clustered separately to obtain the clustering results for each region at each time scale. For example, see [link to example]. Figure 9 , Figure 9 This is a schematic diagram of a region clustering result provided in an embodiment of the present invention. Figure 9Regions of the same color belong to the same cluster and correspond to the same region type. Each region type can be assigned a type identifier to facilitate the storage and retrieval of region types.
[0086] Building upon the above embodiments, the clustering of regions is based on a time scale, ignoring spatial information within intervals. In this embodiment, the clustering results at each time scale are optimized based on spatial relationships to update the region types of each region, thereby improving the accuracy of the region types. Figure 9 Taking the clustering results as an example, regions with the same color indicate that they belong to the same cluster. Region 'a' has two different region categories within its four neighboring regions, and region 'a' does not belong to either category. This region 'a' is spatially isolated. If determining this isolated region based on spatial relationships is unreasonable, it can be optimized based on the clusters of its neighboring regions.
[0087] Optionally, after obtaining the clustering results at each time scale, the process further includes: determining whether the isolated region is reasonable, and updating the region type of the isolated region if it is unreasonable. The determination of the reasonableness of the isolated region may include: determining whether the error distribution between the isolated region and the cluster centers of its neighboring regions is within the range of a standard normal distribution 3δ. If not, the isolated region is determined to be reasonable; if so, the isolated region is determined to be unreasonable, and its region type needs to be optimized. The mean of the standard normal distribution is the average Euclidean distance from each point in the cluster to the center point, and the standard deviation is the standard deviation of the difference between the distance from each point in the cluster to the center point and the average Euclidean distance.
[0088] Optionally, after clustering each region at each time scale, if it is determined that an isolated region is spatially unreasonable, the method further includes: for an isolated region in the clustering results at any time scale, updating the cluster to which the isolated region belongs based on the cluster centers of the isolated region and the clusters to which its neighboring regions belong. Specifically, this can involve determining the distance information from the isolated region to the centers of the clusters to which its neighboring regions belong, and determining the cluster to which the isolated region belongs as the cluster corresponding to the minimum distance information. For example, see [link to example]. Figure 9 An isolated region a has two neighboring clusters. The distances d1 and d2 from isolated region a to the centers of the two clusters are determined. If d1 < d2, then isolated region a is updated to the cluster corresponding to d2, and the region type of the isolated region is updated to the region type corresponding to the cluster corresponding to d2. See also... Figure 10 , Figure 10 This is a schematic diagram illustrating the optimization of clustering results based on spatial relationship constraints according to an embodiment of the present invention. The left image shows the clustering results of region clustering in a real dataset, and the right image shows the optimization results based on spatial constraints. Figure 10 It can be seen that there are a large number of isolated region categories in the initial clustering results. After spatial constraint optimization, the number of isolated region categories is reduced.
[0089] Based on the above embodiments, the clustering process at each time scale can include clustering based on inflow characteristic data and clustering based on outflow characteristic data. That is, the regional clustering result at any time scale includes inflow clustering results and outflow clustering results. Correspondingly, at a single time scale, each region belongs to two regional categories in the inflow cluster and the outflow cluster. In this embodiment, the inflow clustering results and the outflow clustering results are merged so that each region corresponds to a unique regional category at a single time scale.
[0090] Optionally, after obtaining the clustering results at each time scale, or after optimizing the clustering results at each time scale based on spatial relationship constraints, the method further includes: at any time scale, merging the inflow clustering results and the outflow clustering results to obtain the target region type of each region at the aforementioned time scale.
[0091] In this embodiment, the regions to be merged can be determined based on the similarity between the clusters in the inflow clustering results and the outflow clustering results. Optionally, the merging process of the inflow clustering results and the outflow clustering results includes: determining the similarity between each inflow cluster in the inflow clustering results and each outflow cluster in the outflow clustering results; and forming a target cluster based on the intersection of regions in the inflow clusters and outflow clusters that meet the similarity condition.
[0092] The inflow clustering results include *cin* inflow clusters, and the outflow clustering results include *con* outflow clusters. The similarity between each inflow and outflow cluster is calculated. This similarity can be achieved by using the region identifiers within each cluster as clustering information, or by using the flow characteristic data corresponding to each region within each cluster as clustering information. Based on this similarity calculation, a similarity matrix is obtained, where each similarity value corresponds to one inflow and one outflow cluster. If the similarity meets a similarity condition, such as exceeding a similarity threshold, the inflow and outflow clusters corresponding to that similarity value are merged.
[0093] The merging of inflow and outflow clusters can be achieved by taking the intersection of the regions included in the inflow and outflow clusters and defining a new target cluster, which corresponds to a new target region type. For example, if inflow cluster A and outflow cluster B meet a similarity condition, and inflow cluster A includes regions 1, 2, and 3, while outflow cluster B includes regions 2 and 3, then regions 2 and 3 are determined to belong to a new target cluster, corresponding to a new target region type. All inflow and outflow clusters that meet the similarity condition are traversed to form new target clusters, and regions not belonging to any target cluster are designated as undetermined regions. Accordingly, in the example above, region 1 is designated as an undetermined region.
[0094] Based on the above embodiments, for the undetermined region outside the target cluster, distance information between it and each target cluster is determined, and the target cluster to which the undetermined region belongs is determined. In some embodiments, this may involve determining the spatial distance information between the undetermined region and each region in each target cluster. For example, determining the spatial distance information between the undetermined region and the spatial center of each region in each target cluster, and determining the target cluster corresponding to the minimum distance information as the cluster to which the undetermined region belongs. In some embodiments, distance information between the traffic characteristic data of the undetermined region and the traffic characteristic data of each region in each cluster is determined. Based on the mapping of each region to the clustering space, the cluster center of each target cluster is determined, and the distance information between the undetermined region and the cluster center of each target cluster is determined. The target cluster corresponding to the minimum distance information is determined as the cluster to which the undetermined region belongs.
[0095] By merging the inflow clustering results and the outflow clustering results, the number of region types can be reduced, thus lowering the difficulty and computational load of anomaly detection.
[0096] Based on the above embodiments, each region type corresponds to a region anomaly detection model, which can be trained based on the traffic characteristic data of each region within the corresponding region type. The type of region anomaly detection model is not limited here; it can be a neural network model, etc. Optionally, the region anomaly detection model includes an isolated forest model, which comprises multiple isolated trees. By setting up an isolated forest model, unsupervised anomaly detection of spatiotemporal characteristic data in space can be achieved.
[0097] The creation process of the isolated forest model can be as follows: For any region type, obtain the traffic feature dataset {x1,x2,…,x} for each region of that region type. nc}, where x i = <inflow i outflow i >, inflowi For inflow data, outflow i This is the outflow data. ψ sample points are extracted from the dataset to form a subset X′ of X, which is then treated as an isolated tree. The root node is randomly selected from the inflow and outflow dimensions q, and a random split point p is generated in the current data, satisfying min(x) = q / q. i [j],j=q,x i ∈X) <q<max(x i [j],j=q,x i When p ∈ X), the split point p forms a hyperplane in the current node data, dividing the current data space into two subspaces. Data with dimension p less than q is assigned to the left child node, and the rest are assigned to the right child node. This process is repeated until all leaf nodes contain only one sample point or the current isolated tree has reached the specified height. This process is repeated to construct multiple isolated trees, forming an isolated forest model.
[0098] For any region, the corresponding region type at each time scale is determined. The corresponding isolated forest model for each region type is then invoked. Based on each isolated forest model, the spatiotemporal feature data for that region is processed to identify anomalous information for that region across each region type. Specifically, processing the spatiotemporal feature data of the region based on the region anomaly identification model corresponding to each region type to determine the anomalous information for each region type includes: inputting the spatiotemporal feature data of the region into multiple isolated forest models corresponding to that region, obtaining the output results of each isolated forest model; and determining the anomalous information for each region type based on the output results of the isolated forest models.
[0099] In this embodiment, the spatiotemporal feature data of a region can be the real-time inflow and outflow data of that region. The real-time inflow and outflow data of a region is input into each isolated tree in an isolated forest model corresponding to each region type, and each isolated tree in the isolated forest model is traversed. The output of the isolated forest model includes the height information of the spatiotemporal feature data of the region in each isolated tree. Specifically, the height information of the real-time inflow and outflow data in each isolated tree is determined. This height information can be determined as the height information between the real-time inflow and outflow data at the leaf node (from the root node to the leaf node, where it cannot be included in the next level of classification) and the corresponding leaf node.
[0100] Accordingly, determining the anomaly information corresponding to the region type based on the output of the isolated forest model includes: determining the anomaly information corresponding to the region type based on the mean height information of the region's spatiotemporal feature data in each isolated tree in the output of the isolated forest model. Specifically, an anomaly index can be determined based on the ratio of the mean height information of the region's spatiotemporal feature data in each isolated tree to the average path length when the preset sample size is ψ; the anomaly information is then determined based on the index base and the anomaly index. For example, the anomaly information of a region in any region type can be calculated based on the following formula:
[0101]
[0102] Where E(h(x)) is the mean of the spatiotemporal feature data of the region in each isolated tree, and c(ψ) is the average path length when the number of samples is ψ.
[0103] Each region can correspond to multiple region types, and the anomaly information corresponding to each region type is fused. The fusion processing method can be a weighted processing with preset weights.
[0104] Based on the above embodiments, anomalies can be determined for each region based on temporal and spatial anomaly information to identify abnormal regions.
[0105] In this embodiment, since regions can obtain different clustering results at different time scales (i.e., regions correspond to different region types at different time scales), after determining the anomaly information for each region type, the anomaly information corresponding to the different region types is fused. This process accommodates anomaly information from different region classifications and improves the accuracy of spatial anomaly information determination. Furthermore, by using temporal and spatial anomaly information to determine anomalies in each region, the accuracy of anomaly region detection is improved.
[0106] The method further includes: visually displaying traffic prediction data for each area and one or more abnormal areas on a display interface. Based on detection information (e.g., crowd traffic prediction data) and abnormal detection results (e.g., abnormal areas) during the anomaly detection process, an anomaly display page can be generated. This page can be displayed on the detection display device, facilitating intuitive access to detection information by detection personnel. Upon request from any client, the anomaly display page can be sent to the client for display.
[0107] Based on the above embodiments, this invention also provides a preferred example of a method for urban anomaly detection based on spatiotemporal data, see [link to example]. Figure 11 , Figure 11This is a schematic diagram of the system structure for implementing a spatiotemporal data-based urban anomaly detection method, as improved in this embodiment of the invention. The system structure may include a data layer, a processing layer, a platform layer, an application layer, and a user layer. The data layer acquires and stores data, which may include data from various dimensions, such as environmental data, traffic data, and population data. This data layer can be stored using a database. The processing layer processes the data, including but not limited to data cleaning, data alignment, data segmentation, and data statistics. Based on this data and the distribution across different regions, spatiotemporal feature data is formed. The platform layer performs anomaly detection on the spatiotemporal feature data, and the application layer monitors the anomaly detection results and displays them on an anomaly display page. The user layer handles requests for anomaly detection results from various clients and provides feedback to users.
[0108] See Figure 12 , Figure 12 This is a flowchart of a spatiotemporal data-based urban anomaly detection method provided by an embodiment of the present invention. The city to be detected is segmented into regions by iteratively binary-splitting the latitude and longitude range, dividing the city into multiple rectangular regions, each corresponding to a unique coded index. The address code of the region can be set using the Geohash method. Feature data within the city to be detected, such as signaling data from electronic devices and meteorological data, is acquired, and the acquired data is aligned with the geographical scope. After matching the data to the raster after region segmentation, a three-dimensional dataset S is obtained, S = { <r i ,t j ,v k >, 1≤i≤m, 1≤j≤t, 1≤k≤n}.
[0109] Based on a pre-defined traffic prediction model, time-series prediction detection is performed. The model's input mainly consists of two parts: dynamic crowd flow data and static meteorological and temporal data. Part of the latter data is remapped through two embedding layers before being concatenated with the time-series traffic data and used as the model's input. The output of the last LTM cell in the Encoder layer serves as the input to the first LSTM cell in the Decoder layer. Each LSTM cell in the Decoder layer is followed by a fully connected layer to convert its output hidden state into predicted two-dimensional traffic data. By predicting crowd flow in each region using this model and comparing the deviation between the model's predicted values and the actual crowd flow data for each region, the degree of temporal anomaly in each region can be captured, thus obtaining the temporal anomaly information for each region. The dynamic crowd flow data can be r for each region. i The top 24 groups of crowd flow data at current time t The output of the model is the predicted value of the current time population flow Compare with the true value The deviation between the predicted value can be used to obtain the time series outliers for each region Among them, the inflow anomaly information is Outflow anomaly information
[0110] For anomaly detection in space:
[0111] The population flow pattern is often reflected in three time scales: temporal proximity, periodicity, and trend. Using clustering methods on these three time scales respectively, the population flow patterns M of the city at different time scales can be captured c 、M p 、M t And their corresponding regions, where M c Can be the regional classification on the proximity scale, M p Can be the regional classification on the periodicity scale, M t Can be the regional classification on the trend scale: M = {M c ,M p ,M t}.
[0112] For example, if the urban area shows nc population flow patterns in temporal proximity, and each population flow pattern corresponds to several regions. For the population flow pattern [[ID=3...]]Because clustering is performed using inflow and outflow data separately, each region belongs to two categories in the inflow and outflow clusters at a single time scale. To align the clustering results so that each region corresponds to a unique category at a single time scale, the clustering results need to be merged. The Jaccard similarity matrix Simi between the inflow and outflow clusters is calculated. cin×con Where cin represents the number of inflow data clusters and con represents the number of outflow data clusters, the Jaccard similarity can be calculated as follows:
[0116] Here, A and B are different clusters. From Simi cin×con Extract all cluster pairs (cluster_pair) with similarity greater than the α threshold. Calculate the region intersection of each pair. For example, if the pair [ci1,co2] represents a match between inflow cluster 1 and outflow cluster 2, then the region intersection of this pair is the set of all regions that belong to both inflow cluster 1 and outflow cluster 2.
[0117] After merging the inflow and outflow clustering results, |cluster_pair| new clusters are formed. The size of |cluster_pair| is usually smaller than the original number of inflow and outflow clusters, so many regions will not belong to any of the new clusters. The cluster center is calculated based on the region set corresponding to each cluster, and then the remaining regions that do not belong to any cluster are added to the nearest cluster. In this way, all regions are matched with their appropriate clusters.
[0118] Each region type trains an isolated forest model. The actual inflow and outflow data for each region are input into the corresponding isolated forest model. Each isolated tree in the isolated forest for the current region is traversed, and the value of point x is calculated. i The average height h(x) in each forest i The outlier for each point can be obtained by normalizing the average height of all current real-time data points. The outlier is calculated as follows, where c(ψ) is the average path length when the number of samples is ψ:
[0119]
[0120] In summary, the spatial outliers in each region under the temporal proximity clustering results Similarly, spatial outliers in clustering results under time periodicity and time trend can be calculated separately to obtain AD. p and AD t Since this patent focuses more on anomalies in crowd flow in real-time scenarios, it is necessary to address AD (Advanced Data Intervention) issues.c AD p and AD t Assign different weights β c β p and β t This approach aims to enhance the sensitivity of spatial outliers to real-time data without neglecting large-scale population movement patterns within a region. In this way, the model obtains a temporal outlier vector for each region. in
[0121] Based on this, outliers from both temporal and spatial perspectives can be obtained. The normalized outliers for each region from both perspectives are then combined into a vector, resulting in an outlier vector set.
[0122] Outliers should have large values in all three outlier dimensions. In other words, outliers should be located at the outer edge of the distribution within the outlier vector space formed by the anomalous vector representation (AD). The AD is represented in space using the skyline algorithm, and the final outliers are identified by calculating their spatial relationships. For example, skyline points in the metric space can be identified as outliers. A skyline point is a point in the dataset where no other point has a value greater than its own in all dimensions.
[0123] In the data collection process described in the above embodiments, information display pages are formed, such as a crowd flow situation awareness page, a regional crowd flow prediction page, and a regional crowd anomaly detection page. The crowd flow situation awareness page is used to visualize and statistically analyze real-time crowd flow data in the city, including displaying real-time crowd flow using a map, displaying changes in urban area inflow and outflow using a line chart, displaying the city's traffic ranking in the past hour using a bar chart, displaying the population transfer volume of adjacent areas in the past hour using a bar chart, and displaying the current proportion of inflow and outflow in the region using a pie chart. The regional crowd flow prediction page is used to visualize and statistically analyze the system's predicted traffic values, including displaying the predicted trend of regional crowd flow using a map, displaying the inflow and outflow of people in each region in the next four hours using a line chart, displaying the deviation between the system's predicted and actual values in the past twelve hours using a line chart, displaying the city's traffic ranking in the next hour using a bar chart, and displaying the population transfer volume of adjacent areas in the next hour using a bar chart. The crowd anomaly detection page is used to visualize and perform corresponding statistical analysis on the anomaly information detected by the system, including using a map to display regional anomaly classification information, using a line chart to display historical anomaly data for regions, and using a bar chart to display the ranking of anomalies in urban areas.
[0124] Figure 13 This is a schematic diagram of the structure of an urban anomaly detection device based on spatiotemporal data provided in an embodiment of the present invention. Figure 13As shown, the device includes: a spatiotemporal data acquisition module 310, a first anomaly information determination module 320, a second anomaly information determination module 330, and an anomaly region determination module 340.
[0125] The spatiotemporal data acquisition module 310 is used to acquire spatiotemporal feature data of the city to be detected, wherein the spatiotemporal feature data includes the temporal feature data of each region of the city to be detected.
[0126] The first anomaly information determination module 320 is used to determine the temporal traffic anomaly information of any region based on the spatiotemporal characteristic data of any region.
[0127] The second anomaly information determination module 330 is applied to determine the spatial flow anomaly information of any region based on the spatiotemporal characteristic data of that region.
[0128] The abnormal region determination module 340 is used to determine abnormal regions in each region based on the temporal and spatial traffic anomaly information of each region.
[0129] Based on the above embodiments, optionally, the spatiotemporal data acquisition module 310 is used for:
[0130] Obtain the regional distribution of the city to be detected, and obtain the feature data of each region within a preset time range.
[0131] Based on the above embodiments, optionally, the first anomaly information determination module 320 is used for:
[0132] The feature data of the region within a preset time range are input into a pre-trained traffic prediction model to obtain traffic prediction data for the region.
[0133] Based on the actual traffic data and the traffic prediction data of the region, the time-series traffic anomaly information of the region is determined.
[0134] Based on the above embodiments, optionally, the flow prediction data includes inflow prediction data and outflow prediction data;
[0135] The time-series flow anomaly information of the region includes inflow anomaly information and outflow anomaly information.
[0136] Based on the above embodiments, optionally, the traffic prediction model includes a vector transformation layer, an encoding layer, a decoding layer, and a fully connected layer. The vector transformation layer converts the input feature data into vector data, which is then input to the encoding layer. The encoding layer includes multiple loop units, each processing the vector data corresponding to each time-series feature data sequentially and inputting the processing result to the decoding layer. The decoding layer includes multiple loop units, each connected to a fully connected layer, outputting the processing result to the fully connected layer. The fully connected layer outputs the corresponding traffic prediction data.
[0137] Based on the above embodiments, optionally, the second anomaly information determination module 330 includes:
[0138] The model invocation unit is used to invoke the region anomaly identification model corresponding to each region type based on the region category to which the region belongs at each time scale.
[0139] An anomaly information determination unit is used to process the spatiotemporal feature data of the region based on the region anomaly identification model corresponding to each region type, and determine the anomaly information corresponding to each region type.
[0140] An anomaly information fusion unit is used to fuse the anomaly information corresponding to the region type to obtain the spatial traffic anomaly information of the region.
[0141] Optionally, based on the above embodiments, the device further includes:
[0142] The region type determination module is used to perform clustering processing on each region at each time scale based on the historical spatiotemporal feature data of each region, and determine the region category to which each region belongs at each time scale. Each cluster in the clustering processing result corresponds to a region type.
[0143] Based on the above embodiments, optionally, the time scale includes temporal proximity, periodicity, and trend;
[0144] The region type determination module is used for:
[0145] Based on the traffic characteristic data of each region within a preset adjacent time period, the regions are clustered to obtain the region clustering results based on proximity.
[0146] Based on the traffic characteristic data of each region within a preset periodic time period, the regions are clustered to obtain the regional clustering results in terms of periodicity.
[0147] Based on the traffic characteristic data of each region within a preset continuous time period, the regions are clustered to obtain the regional clustering results in terms of trend.
[0148] Optionally, based on the above embodiments, the device further includes:
[0149] The first optimization module is used to update the cluster cluster to which the isolated region belongs based on the cluster centers of the isolated region and the clusters of the adjacent regions in the clustering results at any time scale.
[0150] Based on the above embodiments, optionally, the region clustering results at any time scale include inflow clustering results and outflow clustering results;
[0151] The device also includes:
[0152] The second optimization module is used to merge the inflow clustering results and the outflow clustering results at any time scale to obtain the target region type of each region at the above time scale.
[0153] Based on the above embodiments, optionally, the second optimization module is used for:
[0154] Determine the similarity between each inflow cluster in the inflow clustering results and each outflow cluster in the outflow clustering results;
[0155] The target cluster is formed based on the intersection of regions in the inflow and outflow clusters that meet the similarity condition.
[0156] Optionally, based on the above embodiments, the second optimization module is further configured to:
[0157] For the undetermined region outside the target cluster, determine the distance information to each target cluster, and determine the target cluster to which the undetermined region belongs.
[0158] Based on the above embodiments, optionally, the regional anomaly identification model includes an isolated forest model;
[0159] The anomaly information determination unit is used to: input the spatiotemporal feature data of the region into multiple isolated forest models corresponding to the region, and obtain the output results of each isolated forest model; and determine the anomaly information corresponding to the region type based on the output results of the isolated forest models.
[0160] Based on the above embodiments, optionally, the output of the isolated forest model includes the spatiotemporal feature data of the region and the height information of each isolated tree in the isolated forest model.
[0161] The anomaly information determination unit is used to determine the anomaly information corresponding to the region type based on the mean height information of the spatiotemporal feature data of the region in each isolated tree in the output of the isolated forest model.
[0162] Based on the above embodiments, optionally, the abnormal region determination module 340 is used for:
[0163] The temporal and spatial traffic anomaly information of each region is mapped into the metric space, respectively.
[0164] Based on the location information of each region in the metric space, abnormal regions are determined.
[0165] Optionally, based on the above embodiments, the device further includes:
[0166] The visualization module is used to visually display traffic forecast data for each region and one or more abnormal regions on the display interface.
[0167] The urban anomaly detection device based on spatiotemporal data provided in the embodiments of the present invention can execute the urban anomaly detection method based on spatiotemporal data provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0168] Figure 14 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0169] like Figure 14As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0170] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0171] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as urban anomaly detection methods based on spatiotemporal data.
[0172] In some embodiments, the spatiotemporal data-based urban anomaly detection method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the spatiotemporal data-based urban anomaly detection method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the spatiotemporal data-based urban anomaly detection method by any other suitable means (e.g., by means of firmware).
[0173] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0174] Computer programs for implementing the spatiotemporal data-based urban anomaly detection method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0175] This invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute a method for detecting urban anomalies based on spatiotemporal data. The method includes:
[0176] Acquire spatiotemporal feature data of the city to be detected, wherein the spatiotemporal feature data includes the temporal feature data of each region of the city to be detected; for the spatiotemporal feature data of any region, determine the temporal traffic anomaly information of the region and determine the spatial traffic anomaly information of the region; based on the temporal and spatial traffic anomaly information of each region, identify abnormal regions in each region.
[0177] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0178] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0179] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0180] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0181] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0182] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for detecting urban anomalies based on spatiotemporal data, characterized in that, include: Acquire spatiotemporal feature data of the city to be detected, wherein the spatiotemporal feature data includes the temporal feature data of each region of the city to be detected; For any spatiotemporal characteristic data of a region, determine the temporal flow anomaly information of the region, and determine the spatial flow anomaly information of the region; Based on the temporal and spatial traffic anomaly information of each region, abnormal regions are identified in each region; The step of determining the temporal traffic anomaly information of the region includes: The feature data of the region within a preset time series range are input into a pre-trained traffic prediction model to obtain the traffic prediction data of the region; based on the actual traffic data of the region and the traffic prediction data, the traffic anomaly information of the region in the time series is determined. The determination of spatial traffic anomaly information in the region includes: Based on the region category of the region at each time scale, the region anomaly identification model corresponding to each region type is invoked; the spatiotemporal feature data of the region is processed based on the region anomaly identification model corresponding to each region type to determine the anomaly information corresponding to each region type; the anomaly information corresponding to the region type is fused to obtain the spatial traffic anomaly information of the region. The process of identifying abnormal regions based on temporal and spatial traffic anomaly information in each region includes: The temporal and spatial traffic anomaly information of each region is mapped into a metric space; the abnormal region is determined based on the location information of each region in the metric space.
2. The method according to claim 1, characterized in that, The acquisition of spatiotemporal feature data of the city to be detected includes: Obtain the regional distribution of the city to be detected, and obtain the feature data of each region within a preset time range.
3. The method according to claim 1, characterized in that, The flow prediction data includes inflow prediction data and outflow prediction data; The time-series flow anomaly information of the region includes inflow anomaly information and outflow anomaly information.
4. The method according to claim 1, characterized in that, The traffic prediction model includes a vector transformation layer, an encoding layer, a decoding layer, and a fully connected layer. The vector transformation layer converts the input feature data into vector data, which is then input to the encoding layer. The encoding layer includes multiple loop units, each of which processes the vector data corresponding to each time-series feature data sequentially and inputs the processing result to the decoding layer. The decoding layer includes multiple loop units, each of which is connected to a fully connected layer and outputs the processing result to the fully connected layer. The fully connected layer outputs the corresponding traffic prediction data.
5. The method according to claim 1, characterized in that, The method for determining the region category at each time scale includes: Based on the historical spatiotemporal characteristic data of each region, clustering is performed on each region at each time scale to determine the region category to which each region belongs at each time scale. Each cluster in the clustering result corresponds to a region type.
6. The method according to claim 5, characterized in that, The time scale includes the proximity, periodicity, and trend of time; Clustering is performed on each of the aforementioned regions at each time scale to determine the region category to which each region belongs at each time scale, including: Based on the traffic characteristic data of each region within a preset adjacent time period, the regions are clustered to obtain the region clustering results based on proximity. Based on the traffic characteristic data of each region within a preset periodic time period, the regions are clustered to obtain the regional clustering results in terms of periodicity. Based on the traffic characteristic data of each region within a preset continuous time period, the regions are clustered to obtain the regional clustering results in terms of trend.
7. The method according to claim 5, characterized in that, After clustering the regions at each time scale, the process further includes: For any isolated region in the clustering results at any time scale, the cluster to which the isolated region belongs is updated based on the clustering of the cluster centers of the clusters to which the isolated region belongs and the clusters to which the adjacent regions belong.
8. The method according to any one of claims 5-7, characterized in that, The region clustering results at any time scale include inflow clustering results and outflow clustering results; The method further includes: At any given time scale, the inflow clustering results and the outflow clustering results are merged to obtain the target region type for each region at the aforementioned time scale.
9. The method according to claim 8, characterized in that, The merging process of the inflow clustering results and the outflow clustering results includes: Determine the similarity between each inflow cluster in the inflow clustering results and each outflow cluster in the outflow clustering results; The target cluster is formed based on the intersection of regions in the inflow and outflow clusters that meet the similarity condition.
10. The method according to claim 9, characterized in that, The method further includes: For the undetermined region outside the target cluster, determine the distance information to each target cluster, and determine the target cluster to which the undetermined region belongs.
11. The method according to claim 1, characterized in that, The regional anomaly identification model includes the isolated forest model; The process of processing the spatiotemporal feature data of the region based on the region anomaly identification model corresponding to each region type to determine the anomaly information corresponding to each region type includes: The spatiotemporal feature data of the region are input into multiple isolated forest models corresponding to the region to obtain the output results of each isolated forest model; Anomaly information corresponding to the region type is determined based on the output of the isolated forest model.
12. The method according to claim 11, characterized in that, The output of the isolated forest model includes the spatiotemporal feature data of the region and the height information of each isolated tree in the isolated forest model. The determination of anomaly information corresponding to the region type based on the output of the isolated forest model includes: Based on the output of the isolated forest model, the mean value of the height information of the spatiotemporal feature data of the region in each isolated tree determines the anomaly information corresponding to the region type.
13. The method according to claim 1, characterized in that, The method further includes: The display interface visually shows traffic forecast data for each region, as well as one or more abnormal regions.
14. A city anomaly detection device based on spatiotemporal data, characterized in that, include: The spatiotemporal data acquisition module is used to acquire spatiotemporal feature data of the city to be detected, wherein the spatiotemporal feature data includes the temporal feature data of each region of the city to be detected; The first anomaly information determination module is used to determine the temporal traffic anomaly information of any region based on the spatiotemporal characteristic data of that region. The second anomaly information determination module is used to determine the spatial flow anomaly information of any region based on the spatiotemporal characteristic data of that region. The abnormal region determination module is used to determine abnormal regions in each region based on the temporal and spatial traffic anomaly information of each region. The first anomaly information determination module is used to: input the feature data of the region within a preset time series range into a pre-trained traffic prediction model to obtain traffic prediction data for the region; and determine the traffic anomaly information of the region in time series based on the actual traffic data of the region and the traffic prediction data. The second anomaly information determination module includes: a model invocation unit, used to invoke the regional anomaly identification model corresponding to each regional type based on the regional category to which the region belongs at each time scale; an anomaly information determination unit, used to process the spatiotemporal feature data of the region based on the regional anomaly identification model corresponding to each regional type, and determine the anomaly information corresponding to each regional type; and an anomaly information fusion unit, used to fuse the anomaly information corresponding to the regional types to obtain the spatial flow anomaly information of the region. The abnormal region determination module is used to: map the time-series traffic anomaly information and the spatial traffic anomaly information of each region into the metric space respectively; and determine the abnormal region based on the location information of each region in the metric space.
15. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the urban anomaly detection method based on spatiotemporal data as described in any one of claims 1-13.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the urban anomaly detection method based on spatiotemporal data as described in any one of claims 1-13.