Space-time crime risk prediction method based on multi-dimensional geographical environment factors

By introducing a combination method of multi-dimensional geographical and environmental factors and Mamba structures in the space-time crime prediction technology, the problems of space-time crime prediction and space-time heterogeneity modeling under fine-grained grids are solved, and a higher precision and practical space-time crime prediction is achieved.

CN120197938APending Publication Date: 2025-06-24BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510273124.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing space-time crime prediction technology is difficult to effectively predict crime events under fine-grained grids, and it lacks the ability to model space-time heterogeneity and long sequence contexts, making it difficult to apply to actual tasks.

Method used

The space-time crime risk prediction method based on multi-dimensional geographical environment factors is adopted, and the time feature calculation module, the space-time feature calculation module and the space-time attention module are combined with the Mamba structure to improve the model's long-sequence context modeling ability and perception of space-time heterogeneity.

Benefits of technology

The prediction of space-time crime under fine-grained grids is realized, which improves the accuracy and practicality of predictions, and can better cope with the complexity of space-time heterogeneity and long sequence context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197938A_ABST
    Figure CN120197938A_ABST
Patent Text Reader

Abstract

The invention discloses a spatio-temporal crime risk prediction method based on multi-dimensional geographical environment factors, which is finer in prediction granularity for spatio-temporal crime events: for the problems of data sparseness and data volume dramatic increase in fine granularity, a Mama structure is adopted to deal with the challenges; the Mama has very good memory and processing capabilities for long sequence contexts and has very excellent training speed, and by adopting the Mama structure, the expression and performance of the model can be improved in a fine-grained space-time crime event prediction task; criminal events are not uniformly distributed in different time and spaces, multi-modal spatial features are adopted, the features are fused from different scales, and perception of different spatial factors during criminal event prediction is improved; attention distribution of Mama on long sequence context also improves perception of time features in different time scales. The spatio-temporal attention mechanism finally and reasonably fuses spatio-temporal features to obtain a prediction result of spatio-temporal criminal events, and spatio-temporal heterogeneity is effectively dealt with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and specifically relates to a spatio-temporal crime risk prediction method based on multi-dimensional geographical environment factors. Background Art

[0002] Crimes that endanger public safety and cause economic losses have received great attention worldwide. Spatio-temporal crime prediction can effectively evaluate the occurrence of crimes at a certain moment in a region, which is very useful for long-term decision-making, such as optimizing the allocation of facilities and police resources, assisting in the formulation of police patrol routes, and preventing the occurrence of crime events.

[0003] In the research of spatio-temporal crime hotspots, the prediction area is often divided into grid cells of the same size. Generally speaking, the larger the spatial unit, the higher the prediction accuracy. Some studies set the grid accuracy to 3000m×3000m, and these studies have achieved good results in prediction accuracy. However, the prediction accuracy is too macroscopic, resulting in difficulty in applying the above studies in practical tasks such as patrol route planning, and it is necessary to conduct crime prediction for more refined regional grids.

[0004] Crime is affected by various factors such as time, environment, weather, and network, and has strong spatio-temporal correlation. Some studies have incorporated environmental data such as points of interest (POIs), road network density, meteorological data, population flow data, public safety data, and public service complaint data into spatio-temporal crime prediction to improve the effect of spatio-temporal crime prediction. However, crime events are not evenly distributed in space and time, showing spatio-temporal heterogeneity. Crime data in different times and regions often show differences, and it is difficult for the same model to capture crime patterns in different times and regions simultaneously. Geographical environment factors can well express spatial characteristics, which is very effective for the model to capture the relationship between crime events and spatial feature distributions, and then to cope with spatio-temporal heterogeneity and improve the prediction effect. Previous spatio-temporal crime predictions have paid little attention to geographical environment factors. If the goal is to improve the performance of spatio-temporal crime prediction models under uneven spatio-temporal distributions, geographical information must be considered. Population flow data and public data are difficult to obtain accurately in most regions, which limits the application of spatio-temporal crime technologies relying on the above data. Geographical environment features such as digital maps and remote sensing maps are easily available in most regions, and it is more universal to use these factors for spatio-temporal crime prediction.

[0005] In order to capture effective features in spatio-temporal sequence information, previous spatio-temporal crime event prediction technologies have adopted methods such as improved deep spatio-temporal 3D convolutional neural networks, the neural network model ST-ACLCrime based on ConvLSTM and SE blocks, the CLSTM-NN combining LSTM and CNN, and the Transformer prediction network based on the multi-head spatio-temporal attention mechanism. However, the common problem with the above models is that it is difficult to handle the large amount of context and long sequence information of spatio-temporal crime events. The occurrence of crime events is relatively sparse and there are many types of crimes. When the model makes predictions, it needs to be based on longer sequences and more context to better capture the occurrence pattern of spatio-temporal crime events. In order to improve the long sequence modeling ability of the model, it is necessary to optimize the model structure specifically. We notice that Mamba, as an excellent variant of SSM, demonstrates feature analysis strength comparable to that of the Transformer model. Its uniqueness lies in its ability to efficiently process long sequence data and achieve linear expansion of the sequence length. This feature has made it shine in multiple fields, including computer vision, natural language processing, and time series analysis. By precisely focusing on key information, it effectively avoids the blind traversal of the entire sequence in traditional methods, especially excelling in summarizing and refining the content of long sequences. Reasonably integrating the Mamba structure is of great significance for improving the model's modeling ability in the context of long sequences and enhancing spatio-temporal crime prediction performance. Figure 1 Shows the general process of spatio-temporal crime prediction.

[0006] Most existing works and technologies have the following defects:

[0007] The spatio-temporal crime prediction granularity is coarse. If the grid division of the prediction area is too large, it may not be able to accurately capture the subtle changes in criminal activities, covering up the real situation in hot spots and making it difficult to accurately identify hot spots. This may lead to unreasonable police deployment and the inability to effectively contain criminal activities. When formulating patrol routes or deploying police forces, if based on the results of rough grid analysis, it may lead to unreasonable police distribution and the inability to effectively respond to criminal activities. The problem to be solved by the present invention is spatio-temporal crime prediction under a fine-grained grid (50m×50m), and solving the effective feature extraction of a large amount of sparse data and long sequence data under fine-grained prediction.

[0008] Difficulty in data acquisition. The occurrence of spatio-temporal crime events has spatio-temporal heterogeneity. Obviously, the number of crime events in areas such as forests and parks is much less than that in areas such as apartments and shopping malls. Multiple environmental factors such as points of interest, population, geographical features, and roads are effective features for coping with spatio-temporal heterogeneity. Existing work has explored the impact of factors such as weather, urban villages, public service complaints, and public safety conditions on spatio-temporal crime time, but the accessibility of these data is not strong, and it is difficult to obtain clear, complete, and reliable above data in most areas. The problem to be solved by the present invention is to better cope with the spatio-temporal heterogeneity of spatio-temporal crimes and better predict spatio-temporal crimes through easily accessible multi-modal geographical data such as digital maps, remote sensing maps, street view maps, and points of interest.

[0009] The existing models have insufficient long-sequence context modeling ability for spatio-temporal crime events. Since the occurrence probability of crime events is relatively low compared to other types of risk events, the number of crime events in the same time span is relatively small, which requires a longer time series when modeling spatio-temporal crime events. At the same time, the large number of crime types leads to a further increase in the prediction context, and traditional structures such as LSTM and Transformer are not satisfactory in coping with the above challenges. The problem to be solved by the present invention is to adopt the Mamba structure to improve the long-sequence context modeling ability of the model and improve the accuracy of predicting spatio-temporal crime events. Summary of the Invention

[0010] In view of this, the purpose of the present invention is to provide a spatio-temporal crime risk prediction method based on multi-dimensional geographical environment factors, optimize the model structure to better cope with long-sequence context modeling, and improve the accuracy and practicality of the spatio-temporal crime prediction method.

[0011] A spatio-temporal crime risk prediction method based on multi-dimensional geographical environment factors is used for prediction by a spatio-temporal crime speculation model, and the model includes a time feature calculation module, a spatial feature calculation module, and a spatio-temporal attention module;

[0012] The time feature calculation module includes a plurality of stacked Mamba modules; the input is the historical moment crime event sequence The output is the intermediate time feature of the crime event at time t n+1 The intermediate time feature of the crime event at time t Among them, R is the total number of grids evenly divided for the area to be predicted, R = I × J, I is the total number of grid rows, and J is the total number of grid columns; represents the number of historical moments; C represents the total number of crime types;

[0013] The inputs of the spatial feature calculation module are unstructured data and structured data respectively; among them, the unstructured data is satellite images Digital map images Street view images Structured data includes the distribution characteristics of points of interest after gridification

[0014] The spatial feature calculation module first performs convolution operations on satellite images, digital map images, and street view images at different scales, and obtains the interlayer features of the three respectively and Mark the distribution characteristics of points of interest on the I×J grid Set the grid containing point-of-interest data to 1, otherwise 0, and finally obtain the point-of-interest feature Then the data H G ,H R ,H S ,H P are concatenated into Input the data H V into the multi-level downsampling module to extract multi-dimensional features with different receptive fields; fuse the outputs of different-level downsampling modules, and first upsample the features of different sizes to adjust the feature sizes to be the same, and then concatenate the features from the channel dimension to obtain Then re-split H V to obtain the intermediate features corresponding to each spatial factor As the input of the spatio-temporal attention module;

[0015] The spatio-temporal attention module uses the intermediate time feature H M as the query feature, passes through a linear layer to obtain the Q feature, U is the intermediate feature dimension in the attention module; the multi-modal intermediate features are used as the "searched" feature and the content feature respectively, pass through two linear layers to obtain their respective K features and V features, After that, calculate the attention weight and use it to calculate the attention feature map; finally, reshape the participating feature map into the original shape and pass it through the GeGLU layer and the linear layer Linear to obtain the crime event at time t n+1 Specific formulas are as follows: Specific formulas are as follows:

[0016]

[0017] where W Q , are learnable parameters; W Q is the learnable parameter of the time query feature, are the learnable parameters of the "searched" features of the digital map, satellite image, street view image, and points of interest respectively, Learnable parameters for the content features of digital maps, satellite images, street view images, and points of interest respectively; Re() represents adjusting the feature dimension.

[0018] Furthermore, it also includes the training process of the spatio-temporal crime prediction model and the learnable parameters. Among them, the loss used is binary cross-entropy loss:

[0019]

[0020] Where is the true value corresponding to each sample, is the model prediction value corresponding to each sample, and N is the number of samples participating in the model training.

[0021] Preferably, the Adam optimizer is used for parameter update in the training process of the spatio-temporal crime prediction model.

[0022] The present invention has the following beneficial effects:

[0023] The present invention has a finer prediction granularity for spatio-temporal crime events: Facing the problems of data sparsity and huge increase in data volume in fine-grained scenarios, we adopt the Mamba structure to address the above challenges. Mamba has very good memory and processing capabilities for long-sequence contexts, and at the same time has an excellent training speed. Adopting the Mamba structure can improve the performance and capabilities of the model in the prediction task of fine-grained spatio-temporal crime events.

[0024] Better cope with the spatio-temporal heterogeneity of spatio-temporal crime events: Crime events are not evenly distributed in different times and spaces. This technology adopts multi-modal spatial features and fuses features at different scales to enhance the perception of different spatial factors during crime event prediction. The attention allocation of Mamba to long-sequence contexts also improves the perception of time features at different time scales. The spatio-temporal attention mechanism finally reasonably fuses spatio-temporal features to obtain the prediction results of spatio-temporal crime events, effectively coping with spatio-temporal heterogeneity. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is the existing spatio-temporal crime flow chart;

[0026] Figure 2 is the prediction model structure diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] The following are specific embodiments in conjunction with the accompanying drawings to describe the present invention in detail.

[0028] 1. Problem Definition

[0029] Definition 1 (Geographical Region): The research area is evenly divided into a grid area of \(R = I\times J\), where \(R\) is the total number of grids, \(I\) is the total number of rows, and \(J\) is the total number of columns. Each area is the target spatial unit for predicting crime occurrences.

[0030] Definition 2 (Crime Event): The crime event \(X\) that occurs at time \(T\) T can be represented as a two-dimensional vector where \(C\) represents the total number of crime types. Specifically, \(X\) T represents the number of crimes under each of the \(C\) different crime types in the \(R\) grid areas at time \(T\).

[0031] The research problem of the present invention is to predict the crime events at time \(t\) given the historical crime event sequence in a region satellite images digital map images street view images and the distribution characteristics of points of interest The goal is to predict the crime events at time \(t\) n+1 Satellite images, digital map images, and street view images usually have higher resolutions than risk observation maps and crime risk maps. Here, we use \(r\) to represent the magnification factor.

[0032] 2. Introduction to the Method

[0033] As Figure 2 shown, the multi-dimensional data fusion spatio-temporal crime speculation model of the present invention mainly consists of three modules: a time feature calculation module for analyzing time series; a spatial feature calculation module for analyzing multi-modal spatial features; and a spatio-temporal attention module for weighting the importance of spatio-temporal features. The three modules will be introduced separately below.

[0034] 1) Time Feature Calculation Module

[0035] The input of the time feature calculation module is the historical crime event sequence Since the granularity of the crime area studied in the present invention is relatively fine, the total number of grids \(R\) is large, which in turn makes the features large, which undoubtedly increases the computational amount and the model fitting difficulty during the model training process. Mamba has unique advantages in long sequence modeling and training speed, and is very suitable for the crime event sequence feature extraction task.

[0036] Several Mamba modules are stacked in the module. The activation layer selects the SiLU activation function, and the state space model selects the selective state space model with hardware-aware state expansion. After using 5 Mamba modules for feature processing in the present invention, we obtain the intermediate time features of the crime events at time \(t\) n+1 To distinguish the spatial hiding features in the following text, the subscript M represents the temporal feature, and the subscript V represents the spatial feature.

[0037] 2) Spatial Feature Calculation Module

[0038] The inputs of the spatial feature calculation module are respectively unstructured data: satellite images Digital map images Street view images and structured data: the distributed features of grid-based points of interest To ensure that different multi-modal features are convenient for processing and analysis in the subsequent process, we perform different convolution operations on digital maps, satellite maps, and street view images respectively to adjust their feature sizes. Among them, digital maps and satellite maps respectively pass through r / 2 layers of convolution operations with a convolution kernel size of 3 and a stride of 2 to obtain inter-layer features and Since each grid area can be represented by a street view image, the number of street view images input each time is R, that is We pass the street view images through a convolution operation with a convolution kernel size of rI and a stride of rI to obtain Since the points of interest feature is structured data, we mark it on the I×J grid, set the cells containing the points of interest data to 1, otherwise to 0, and finally obtain the points of interest feature To facilitate subsequent feature splicing calculation, expand the distributed feature P of the points of interest into

[0039] To perform parallel calculation on multi-modal spatial features, we perform G , H R , H S , H P Splice them into H V Input it into the multi-level downsampling module to extract multi-dimensional features with different receptive fields. Each downsampling module consists of a group normalization module, an activation layer, and a convolution layer. To better improve the representation ability of spatial features, capture environmental features at different macro and micro scales, and better cope with the spatial heterogeneity shown by crime events, we fuse the outputs of different-level downsampling modules. Since the output feature sizes of different-level downsampling modules are different, in the multi-scale spatial feature fusion module, features of different sizes are first upsampled to adjust the different feature sizes to be the same, and then the features are spliced from the channel dimension to obtain

[0040] During the occurrence of crime events, different spatial factors have different influences on the occurrence of crimes, which requires flexible allocation of weights to different spatial factors. Accordingly, we perform VRe - split to obtain the intermediate features corresponding to each spatial factor As the input of the spatio - temporal attention module.

[0041] 3) Spatio - temporal attention module

[0042] Based on the temporal characteristics of the crime, the spatio - temporal attention module assigns weights to different spatial features and selects the spatial features that have the most impact on the occurrence of the crime event at the current moment for modality fusion. The operation process of this module is shown in formula (1), and the Re operation represents dimension adjustment of the feature matrix. Specifically, the intermediate temporal feature H M is used as the query feature, and after passing through the linear layer, we get Figure 2 the Q feature in where U is the intermediate feature dimension in the attention module. The multi - modal intermediate features are used as the "searched" feature and the content feature respectively, and after passing through two linear layers, we get Figure 2 the respective K feature and V feature in After that, calculate the attention weight and use it to calculate the attention feature map. Finally, we reshape the participating feature map into the original shape and pass it through the GeGLU layer and the linear layer Linear, t n+1 of the crime event

[0043] Q = W Q H M

[0044]

[0045] where W Q , is a learnable parameter. W Q is the learnable parameter for the temporal query feature, are the learnable parameters for the "searched" features of the digital map, satellite image, street view image and point of interest respectively, are the learnable parameters for the content features of the digital map, satellite image, street view image and point of interest respectively.

[0046] During the training process, the predicted crime event at time t n+1 is compared with the real crime event (label) to calculate the loss value. We use Binary Cross - Entropy Loss to measure the difference between the predicted value and the real value. The smaller the loss value, the more accurate the prediction of the model. The formula for Binary Cross - Entropy Loss is:

[0047]

[0048] where is the true value corresponding to each sample, is the model prediction value corresponding to each sample, and N is the number of samples participating in model training.

[0049] During the optimization process, the Adam optimizer (Adaptive Moment Estimation) is used for parameter update. Adam combines the advantages of momentum optimization and adaptive learning rate, dynamically adjusts the learning rate through the estimation of first-order and second-order momentum, and adapts to the gradient changes of different features. During training, Adam calculates the gradient of each parameter and updates the model parameters according to the goal of minimizing the loss value, so that the model gradually improves the speculation accuracy. This process is achieved through backpropagation, and the weights of the model will be continuously updated until the loss converges to a smaller value or reaches the preset number of training epochs. Through multiple iterations and optimizations, the crime risk speculation ability of the model will be gradually improved.

[0050] The trained model can directly predict crime events at a specific moment through digital maps, satellite images, street view images, points of interest, and known crime event sequences, which has practical value for urban planning, police force deployment, and other work.

[0051] It should be noted that the spatio-temporal attention module can be replaced by a simple feature fusion mechanism, such as feature concatenation, feature addition, etc. The spatial feature calculation module can be replaced by a simple convolutional stacking mechanism, but the multi-scale spatial perception ability will decline. The temporal feature calculation module can be replaced by the Transformer mechanism, but the training time and training difficulty will increase, and it performs poorly in large-scale crime event sequences. The multi-modal auxiliary speculation data considered in the present invention are digital maps, satellite maps, historical data, points of interest, and street view maps. Other data such as trajectory maps and road maps can be used as auxiliary speculation data in the speculation model.

[0052] In summary, the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A spatiotemporal crime risk prediction method based on multi-dimensional geographical environmental factors, characterized in that: The spatiotemporal crime inference model is used for prediction, which includes a temporal feature calculation module, a spatial feature calculation module, and a spatiotemporal attention module; The time feature calculation module includes multiple stacked Mamba modules; the input is a sequence of crime events at historical moments The output is t n+1 The temporal characteristics of criminal events at different moments Where R is the total number of grids evenly divided into the prediction area, R = I × J, I is the total number of grid rows, and J is the total number of grid columns; represents the number of historical moments; C represents the total number of crime types; The input of the spatial feature calculation module is unstructured data and structured data; among them, the unstructured data satellite image Digital map image Street View Imagery Structured data includes the distribution characteristics of points of interest after gridding The spatial feature calculation module first performs convolution operations on satellite images, digital map images, and street view images at different scales to obtain the inter-layer features of the three images. and Distribution features of interest points on an I×J grid Marking, the grid containing the point of interest data is set to 1, otherwise it is 0, and finally the point of interest feature is obtained Then the data H G , H R , H S , H P Splice to The data H V Input into the multi-level downsampling module to extract multi-dimensional features with different receptive fields; fuse the outputs of downsampling modules at different levels, and first upsample the features of different sizes to adjust the feature sizes to the same, and then splice the features from the channel dimension to obtain Then H V Re-split to obtain the intermediate features corresponding to each spatial factor As the input of the spatiotemporal attention module; The spatiotemporal attention module takes the intermediate temporal features H m As the query feature, after the linear layer, we get the Q feature. U is the intermediate feature dimension in the attention module; the multimodal intermediate features As the "checked" feature and content feature respectively, after two linear layers, the respective K features and V features are obtained. Afterwards, the attention weights are calculated And used to calculate the attention feature map; finally, the attended feature map is reshaped into its original shape and passed through the GeGLU layer and the linear layer to obtain t n+1 Crime incidents at the moment The specific formula is as follows: Where W Q , is a learnable parameter; W Q is the learnable parameter of the temporal query feature, are the learnable parameters of the "checked" features of digital maps, satellite images, street view images and points of interest, respectively. are the learnable parameters of the content features of digital maps, satellite images, street view images and points of interest respectively; Re() represents the adjustment of the feature dimension.

2. A method for predicting spatiotemporal crime risk based on multi-dimensional geographical environmental factors as claimed in claim 1, characterized in that: It also includes the training process of the spatiotemporal crime inference model and learnable parameters, where the loss used is the binary cross entropy loss: in is the true value corresponding to each sample, is the model prediction value corresponding to each sample, and N is the number of samples involved in model training.

3. A method for predicting spatiotemporal crime risk based on multi-dimensional geographical environmental factors as claimed in claim 2, characterized in that: The Adam optimizer is used to update parameters in the training process of the spatiotemporal crime speculation model.