Urban criminal event real-time prediction method based on multi-target reinforcement learning

Through the multi-objective reinforcement learning method, the spatiotemporal interaction features are extracted using node data, combined with gated loop units and deep Q networks, the problems of insufficient spatiotemporal dynamic feature capture and neglected relationships in urban crime incident prediction are solved, and more efficient crime incident prediction is achieved.

CN120163283APending Publication Date: 2025-06-17XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224938.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively capture the dynamic characteristics of space-time characteristics of urban crime events, which makes it difficult to balance prediction accuracy and real-time performance, and the complex relationship between nodes is not fully considered, affecting the prediction effect.

Method used

The multi-objective reinforcement learning method is adopted, and the node data is extracted and processed, and the gated loop unit and the multi-objective spatiotemporal prediction decoder are used to combine the node embedding function and the deep Q network of the ε-greedy strategy to dynamically adjust the prediction timing, determine the action decision variables, and obtain the final prediction result.

Benefits of technology

It improves the real-time and accuracy of criminal incident prediction, can provide strong data support and decision-making basis in urban public safety management, dynamically adjust the best prediction timing, and improve the accuracy and timeliness of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163283A_ABST
    Figure CN120163283A_ABST
Patent Text Reader

Abstract

The invention relates to an urban criminal event real-time prediction method based on multi-target reinforcement learning, and the method comprises the steps: carrying out the feature extraction of a plurality of pieces of obtained node data, obtaining a plurality of spatial-temporal features, and enabling each piece of node data to comprise crime occurrence time information, crime occurrence place information and crime type information; inputting the plurality of spatial-temporal characteristics into a gating circulation unit for processing to obtain a plurality of node hidden state data; inputting the hidden state data of the plurality of nodes into a pre-trained multi-target space-time prediction decoder to obtain an initial prediction result; respectively carrying out operation processing on the plurality of pieces of node hidden state data by utilizing a node embedding function to obtain a plurality of pieces of node state data; determining an action decision variable based on the state data of the plurality of nodes and a deep Q network applying an epsilon-greedy strategy; and performing Hadamard product calculation on the initial prediction result and the action decision variable to obtain a final prediction result. According to the method, the real-time performance and the accuracy of criminal event prediction can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of urban public safety services, and particularly relates to a real-time prediction method for urban crime events based on multi-objective reinforcement learning. Background Art

[0002] Urban crime events have obvious spatio-temporal distribution characteristics, and their occurrence is often affected by various factors such as time and space. For example, the crime rates in different regions show highly dynamic changes at different time periods, and traditional crime prediction methods are difficult to capture such complex spatio-temporal relationships.

[0003] Existing methods have obvious deficiencies in the face of the following problems. Insufficient capture of spatio-temporal dynamic features: The occurrence of urban crime events is affected by multiple spatio-temporal factors, and existing models cannot effectively extract spatio-temporal dynamic features of multiple time steps, resulting in unsatisfactory prediction effects. Difficulty in balancing accuracy and real-time performance: Early prediction requires making decisions within a limited time. Premature prediction will lead to a decrease in accuracy, while late prediction may miss the best intervention opportunity. Ignoring complex relationships between nodes: Crime events within urban areas have strong spatial dependence and interactive characteristics between nodes, and existing models do not fully consider these relationships when constructing crime prediction networks. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a real-time prediction method for urban crime events based on multi-objective reinforcement learning. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] The present invention provides a real-time prediction method for urban crime events based on multi-objective reinforcement learning, including:

[0007] Performing feature extraction on the obtained multiple node data to obtain corresponding multiple spatio-temporal features, wherein each node data includes crime occurrence time information, crime occurrence location information, and crime type information; inputting the multiple spatio-temporal features into a gated recurrent unit for processing to obtain corresponding multiple node hidden state data; inputting the multiple node hidden state data into a pre-trained multi-objective spatio-temporal prediction decoder to obtain an initial prediction result; performing arithmetic processing on the multiple node hidden state data respectively by using a node embedding function to obtain corresponding multiple node state data; determining an action decision variable based on the multiple node state data and a deep Q network applying an ε-greedy strategy; performing a Hadamard product calculation on the initial prediction result and the action decision variable to obtain a final prediction result.

[0008] Compared with the prior art, the beneficial effects of the present invention:

[0009] The data mining ability of the existing crime prediction method is insufficient, resulting in low accuracy of crime early warning and inability to be well applied to practical life. The present invention provides a real-time prediction method for urban crime events based on multi-objective reinforcement learning. The method obtains node data, extracts spatio-temporal interaction features between nodes by using the node data, deeply mines hidden data information in the spatio-temporal interaction features through a gated recurrent unit, and makes a preliminary prediction by using the mined hidden data information. Subsequently, combined with a deep Q-network applying an ε-greedy strategy, an action decision variable for determining whether to detect is determined, and the final prediction result is obtained by combining the action decision variable and the preliminary prediction result. During the whole calculation process, the accuracy and timeliness of the prediction can be balanced, the best prediction time can be dynamically adjusted, and the real-time performance and accuracy of crime event prediction can be effectively improved, providing strong data support and decision-making basis for urban public security management. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 FIG. is a schematic flowchart of a real-time prediction method for urban crime events based on multi-objective reinforcement learning provided by an embodiment of the present invention;

[0011] Figure 2 FIG. is an application example diagram of a real-time prediction method for urban crime events based on multi-objective reinforcement learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The present invention will be further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0013] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0014] In the description of this specification, the description of reference terms such as "an embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.

[0015] Although the present invention has been described in connection with various embodiments, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims during the implementation of the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0016] A real-time prediction method for urban crime events based on multi-objective reinforcement learning provided by an embodiment of the present invention will be described in detail below with reference to the accompanying drawings.

[0017] Figure 1 It is a flowchart of a real-time prediction method for urban crime events based on multi-objective reinforcement learning provided by an embodiment of the present invention. As Figure 1 shown, the method includes:

[0018] Step 110: Extract features from the obtained multiple node data to obtain a corresponding plurality of spatio-temporal features, where each node data includes crime occurrence time information, crime occurrence location information, and crime type information.

[0019] First, it should be noted that the basic graph structure constructed using multiple node data can be expressed as where V is a set of n nodes of the basic graph structure, expressed as E is the set of edges between nodes, expressed as That is, the distance between node i and node j, and the values of node i and node j are both 1 to n. In the present invention, the adjacency matrix A is mainly used to capture the associations between nodes, such as distance relationships, and the observed data at time t is recorded as

[0020] It should be noted that the obtained node data has timeliness, that is, the obtained node data is the data in a certain period, and the quantity and corresponding information content of the node data corresponding to different periods will be different.

[0021] Here, step 110 specifically includes: (1) obtaining the i-th node data and the j-th node data among multiple node data, where the crime occurrence time of the j-th node data is earlier than or equal to that of the i-th node data, both i and j are positive integers, and i = j or i ≠ j; (2) using the i-th node data and the j-th node data to establish an inter-node similarity matrix for characterizing the distance similarity between the i-th node and the j-th node, and a sequence similarity matrix for characterizing the time similarity between the i-th node and the j-th node; (3) for the i-th node data, in the case where the crime occurrence time of the j-th node data is earlier than that of the i-th node data, using the sequence similarity matrix as the spatio-temporal feature corresponding to the i-th node data; (4) in the case where the crime occurrence time of the j-th node data is equal to that of the i-th node data, using the result of the sum of squares of the inter-node similarity matrix and the sequence similarity matrix as the spatio-temporal feature corresponding to the i-th node data. By repeating the above steps (1) - (4), the spatio-temporal feature corresponding to each node data can be calculated and obtained.

[0022] Specifically, the i-th node data in step (2) is the source node data, which may or may not be the same as the content included in the j-th node data. Correspondingly, the inter-node similarity matrix between the i-th node data and the j-th node data is expressed as:

[0023]

[0024] where is the spatio-temporal structure diagram, d ij is the distance between the i-th node data and the j-th node data, and η is the attenuation coefficient, which is used to control the distribution of similarity scores. The larger the attenuation coefficient, the flatter the function curve, and the smaller it is, the narrower and sharper the function curve.

[0025] And the sequence similarity matrix between the i-th node data and the j-th node data is expressed as:

[0026]

[0027] where, represents the DTW distance between the feature sequence of the i-th node data from time 0 to t and the feature sequence of the j-th node data from time 0 to t', and k is used to control the distribution of time dynamic similarity scores.

[0028] Steps (3) and (4) give the calculation method of the spatio-temporal feature corresponding to the i-th node data at time t, which can be specifically expressed as: where, That is, the original spatial similarity matrix is added with the identity matrix, which means adding the connection from the node itself to itself, and then the matrix normalization process is performed to obtain where each element is the sum of the corresponding row, which is used for normalization.

[0029] By using the inter-sequence similarity matrix when t′ < t the historical information from time 0 to time t at all times can be obtained.

[0030] Step 120: Input multiple spatio-temporal features into a gated recurrent unit for processing to obtain a corresponding set of multiple node hidden state data.

[0031] Here, the gated recurrent unit (GRU) is an improved recurrent neural network (RNN) structure. It captures long-term dependencies in time series by introducing an "update gate" and a "reset gate", while reducing the complexity of long-short-term memory (LSTM).

[0032] Exemplarily, use the GRU gated recurrent unit to process multiple spatio-temporal features at time t to obtain a set of multiple hidden states H t Specifically, at time t, there are t + 1 units set in the GRU gated recurrent unit, and the values recorded by it are input into the dynamic multi-level neural network in the last unit (the (t + 1)-th unit) to complete the convolution operation to obtain the feature processing result X t ', and take the feature processing result X t ' as the input of the GRU gated recurrent unit, continue to learn the time series features, and combine the hidden state H t-1 output by the GRU in the previous unit, and finally output the hidden state H t . It should be noted that by changing the number of GRU gated recurrent units, the output value of the GRU gated recurrent unit can be changed.

[0033] Step 130: Input the multiple node hidden state data into a pre-trained multi-object spatio-temporal prediction decoder to obtain an initial prediction result.

[0034] Here, step 130 specifically includes: for each node hidden state data, obtain the first learning parameter and the second learning parameter in the pre-trained multi-object spatio-temporal prediction decoder; perform a multiplication process on the first learning parameter and the node hidden state data to obtain a product result; perform an addition process on the product result and the second learning parameter to obtain the corresponding initial prediction result.

[0035] It should be noted that the pre-trained multi-objective spatio-temporal prediction decoder is obtained by iteratively training the initial MLP neural network using a historical dataset. A Multilayer Perceptron (MLP) is a feedforward artificial neural network model that can map multiple input datasets to a single output dataset. Specifically, during the training process, the pre-trained multi-objective spatio-temporal prediction decoder is obtained by minimizing the mean absolute error between the true value and the predicted value.

[0036] In one possible implementation, the initial prediction result is expressed as: where W is the first learning parameter and b is the second learning parameter.

[0037] Step 140: Use the node embedding function to perform arithmetic processing on multiple node hidden state data respectively to obtain corresponding multiple node state data.

[0038] Here, one node embedding function corresponds to one node hidden state data; the node embedding function is obtained in the following way: perform biased random walk sampling on each node data to obtain the corresponding neighbor node set; based on the neighbor node set corresponding to each node data, obtain the corresponding node embedding function.

[0039] Here, the state of a node is easily affected by the states of its neighbors. Use the bias term α pq (t) and the transition probability to control the sampling process. For example, given the source node v i and the current node v j , then the probability of accessing the next node v k can be calculated by the product of the bias term α pq (t) and the transition probability .

[0040] In one possible implementation, the calculation formula of the bias term α pq (t) is expressed as: where p controls the probability of re-visiting a traversed node, q distinguishes whether the node reaches the optimal prediction time accurately and in a timely manner during the search process, is the optimal prediction time.

[0041] Here, after obtaining the neighbor node set corresponding to each node data, use this set to obtain the corresponding node embedding function. Specifically, the expression for obtaining each node embedding function satisfies:

[0042]

[0043] where ft (·) is the node embedding function, N(v i ) is the node data v obtained by biased random walk sampling i The neighbor node set, Pr(N(v i )|f t (v i )) is in f t (v i ) under the condition that N(v i ) occurs, V is a collection of multiple node data.

[0044] After obtaining each node embedding function, the node embedding function is used to perform calculations on multiple node hidden state data respectively to obtain multiple node state data corresponding to each other, further including: for each node hidden state data, a deep walking algorithm is used to perform calculations on the node hidden state data and the corresponding node embedding function to obtain the corresponding node state data.

[0045] Here, the deep walk algorithm regards nodes as "words" and the node sequences obtained by random walk sampling as "sentences". By inputting these sequences into the word vector model, each node is embedded with contextual information in the graph structure to capture the background information of the node in the graph structure, such as the neighborhood structure information of the node, to obtain the corresponding node status data. This method is equivalent to format conversion of the original node data, which is conducive to subsequent data analysis and calculation.

[0046] Step 150: Determine action decision variables based on multiple node state data and a deep Q network using an ε-greedy strategy.

[0047] Here, step 150 specifically includes: for each node state data, inputting the node state data into a deep Q network that applies an ε-greedy strategy, and calculating a corresponding decision action value based on an updated parameter and the node state data.

[0048] Specifically, the deep Q network using the ε-greedy strategy consists of a fully connected layer and a rectified linear unit, which randomly selects an action with probability ε and selects an action with probability 1-ε such that ω T Q(s t ,a,ω;θ) the largest action value as the corresponding decision action value to avoid the "over-utilization" problem in multi-objective reinforcement learning. In the process of deep Q network calculation, the node state data s needs to be used t , weight value ω, update parameter θ.

[0049] Among them, for different node state data, the corresponding update parameter θ is also different. The update parameter θ is obtained through multiple training iterations. Specifically, the solution formula for the update parameter corresponding to each node state data satisfies:

[0050] L(θ) = (1 - λ)L A (θ) + λL B (θ);

[0051]

[0052] Among them, L(θ) is the final loss function, L A (·) is the main loss function, used to predict the reward value, L B (θ) is the auxiliary loss function, used to optimize the problem that the loss function L A (·) is not smooth, λ is the trade-off parameter between L A (·) and L B (θ), and its value range is 0 to 1. During training, the value of the parameter λ will gradually increase to 1. θ is the update parameter, E[·] is used to calculate the expectation, y is the target value, s t is the node state data, a t is the decision action value, ω is the weight value, and its value range is 0 to 1, which decays from 1 to 0 during the training process. Q is the decay value, is the mean square processing, |·| is to take the absolute value, T is the maximum value of the observation time window, γ is the decay coefficient, r t is the corresponding accuracy reward value at time t, and its calculation formula is: Among them, ρ is the trade-off parameter, argmax(·) is used to return the maximized independent variable of the function expression, ω is the set of multiple weight values at time t, Q(s t+1 , a, ω′; θ k ) is the set of multiple network Q values at time t + 1, ω' is the set of multiple weight values at time t + 1, a is the set of multiple action decision variables, θ k is the update parameter corresponding to the k-th iteration of the model.

[0053] Furthermore, the action decision variable is expressed as:

[0054]

[0055] Among them, a t is the set of multiple action decision variables at time t, ω T is the inversion of the set of multiple weight values, Q(s t, a, ω; θ) is a set of multiple network Q-values at time t, and the value range of ε is from 0 to 1, which decays from 1 to 0 during the training process.

[0056] Exemplarily, for node v i , means that time t is the optimal prediction time Conversely, if t ∈ [0, T - 1], then time T - 1 is the optimal prediction time.

[0057] Through step 150, each node state data is calculated to obtain an action decision variable a corresponding to each node state data t .

[0058] Step 160: Perform Hadamard product calculation on the initial prediction result and the action decision variable to obtain the final prediction result.

[0059] Exemplarily, the solution formula for the final prediction result satisfies:

[0060]

[0061] Where, is the final prediction result, is the initial prediction result.

[0062] Through steps 110 - 160, the crime probability prediction within the current time period can be calculated. In one possible implementation, this method further includes: updating the content included in multiple node data. Specifically, it is to re-obtain multiple node data for the next cycle and perform calculations using steps 110 - 160.

[0063] Here, node update is performed through the formula , that is, when updating the node, the node feature X at time t′ ∈ (0, t) t′ is multiplied by the weight , then multiplied by the corresponding similarity matrix set A t,t′ , then the results of all time steps are summed, and finally calculated by the activation function sigmoid. Here, the weight parameter is related to the number of hidden layer units h and will be optimized during model training to help the prediction model of the method of the present invention learn useful feature representations.

[0064] Figure 2 is an application example diagram of the real-time prediction method for urban crime events based on multi-objective reinforcement learning provided by the embodiments of the present invention. As Figure 2As shown, the input is the spatio-temporal sequence data of urban public safety. This data first enters the enhanced spatio-temporal predictor, where it is extracted using the encoder to obtain hidden data (i.e., node hidden state data). This hidden data is input into the decoder to generate a preliminary prediction result, and the hidden data is used as the input data for the next layer and is fed into the reinforcement learning state generator. Using the node embedding function therein, the node state data is calculated. This node state data is used as the input value for the optimal moment decision of reinforcement learning. Using the deep Q-network, the decision action value (or action decision variable) is calculated. The Hadamard product result of this value and the preliminary prediction result is used as the final optimized prediction result. Here, in the encoder unit, the dynamic adjustment of the number of encoder units composed of multi-level graph convolution and GRU gating is used to focus on different parts of the time series by the model method.

[0065] Aiming at the problem that the data mining ability of the crime prediction method is insufficient, resulting in low crime warning accuracy and unable to be well applied to real life, the present invention provides a real-time prediction method for urban crime events based on multi-objective reinforcement learning. This method obtains node data and uses the node data to extract the spatio-temporal interaction features between nodes. The hidden data information in the spatio-temporal interaction features is deeply mined through the gated recurrent unit, and the mined hidden data information is used for preliminary prediction. Subsequently, combined with the deep Q-network applying the ε-greedy strategy, the action decision variable for determining whether to detect is determined. Combining the action decision variable and the preliminary prediction result, the final prediction result is obtained. During the whole calculation process, the accuracy and timeliness of the prediction can be balanced, the best prediction timing can be dynamically adjusted, and the real-time performance and accuracy of crime event prediction can be effectively improved, providing strong data support and decision-making basis for urban public safety management.

[0066] To verify the accuracy of the real-time prediction method for urban crime events based on multi-objective reinforcement learning provided by the present invention, a simulation software is now used for comparative simulation. Here, the verification dataset is selected as the large-scale open-source dataset publicly available in the real world: the emergency situation dataset of the EMS New York Fire Department. The dataset contains 145 regions based on zip codes, with a sampling rate of 1 hour. The recording time range of the data is from January 1, 2011 to November 30, 2011. In addition, the evaluation metrics commonly used in the regression task are used to evaluate the performance of all models. The evaluation metrics include the mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), hypervolume (HV) metric, and spacing (S) metric. The specific simulation results are shown in Table 1 and Table 2. Among them, Table 1 is the performance metrics of different models on the public safety dataset, and Table 2 is the multi-objective optimization performance evaluation result.

[0067] Table 1

[0068]

[0069] Table 2

[0070]

[0071] As can be seen from Table 1, the present invention can accurately achieve the prediction of urban public safety data on the New York Fire Department emergency situation dataset, and the model of the present invention can obtain the best performance in all evaluation indicators. Moreover, the HV value corresponding to the present invention in Table 2 shows that the present invention has better convergence, which also indicates the proximity of the solution set to the actual Pareto front, indicating that the solution set has a wide coverage in the objective space. And the S value corresponding to the present invention shows that the prediction results given by the technology of the present invention are evenly distributed. This means that the technology of the present invention provides multiple optimal prediction times and provides a powerful and general technical invention framework for various actual application scenarios, effectively achieving the timely and reliable multi-objective prediction effect within a series of prediction times.

[0072] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A real-time prediction method for urban crime events based on multi-objective reinforcement learning, characterized in that: include: Extract features from the acquired multiple node data to obtain multiple space-time features that correspond to each other, where each node data includes information on the time of crime, the location of crime, and the type of crime; Inputting the multiple spatiotemporal features into a gated recurrent unit for processing to obtain multiple node hidden state data corresponding to each other; Inputting the plurality of node hidden state data into a pre-trained multi-objective spatiotemporal prediction decoder to obtain an initial prediction result; Using a node embedding function to perform calculations on the plurality of node hidden state data respectively to obtain a plurality of node state data corresponding to each other; Determining action decision variables based on the plurality of node state data and a deep Q network applying an ε-greedy strategy; A Hadamard product is calculated on the initial prediction result and the action decision variable to obtain a final prediction result.

2. The real-time prediction method for urban crime events based on multi-objective reinforcement learning according to claim 1 is characterized in that: The feature extraction is performed on the acquired multiple node data to obtain multiple one-to-one corresponding spatiotemporal features, including: Obtaining the i-th node data and the j-th node data from the plurality of node data, wherein the crime occurrence time of the j-th node data is earlier than or equal to the crime occurrence time of the i-th node data, i and j are both positive integers, and i=j or i≠j; Using the i-th node data and the j-th node data, establish an inter-node similarity matrix for characterizing the distance similarity between the i-th node and the j-th node, and an inter-sequence similarity matrix for characterizing the time similarity between the i-th node and the j-th node; For the i-th node data, when the crime occurrence time of the j-th node data is earlier than the crime occurrence time of the i-th node data, the inter-sequence similarity matrix is ​​used as the spatiotemporal feature corresponding to the i-th node data; When the crime occurrence time of the j-th node data is equal to the crime occurrence time of the ith node data, the result of square sum processing of the inter-node similarity matrix and the inter-sequence similarity matrix is ​​used as the spatiotemporal feature corresponding to the ith node data.

3. The real-time prediction method for urban crime events based on multi-objective reinforcement learning according to claim 1 is characterized in that: The step of inputting the plurality of node hidden state data into a pre-trained multi-objective spatiotemporal prediction decoder to obtain an initial prediction result includes: For each node hidden state data, obtaining a first learning parameter and a second learning parameter in the pre-trained multi-objective spatiotemporal prediction decoder; Multiplying the first learning parameter and the hidden state data of the node to obtain a product result; The multiplication result and the second learning parameter are added to obtain a corresponding initial prediction result.

4. The real-time prediction method for urban crime events based on multi-objective reinforcement learning according to claim 3 is characterized in that: The pre-trained multi-target spatiotemporal prediction decoder is obtained by iteratively training the initial MLP neural network using a historical data set.

5. The real-time prediction method for urban crime events based on multi-objective reinforcement learning according to claim 1 is characterized in that: A node embedding function corresponds to a node hidden state data; the node embedding function is obtained by: Perform biased random walk sampling on each node data to obtain the corresponding neighbor node set; Based on the neighbor node set corresponding to each node data, the corresponding node embedding function is obtained.

6. The real-time prediction method for urban crime events based on multi-objective reinforcement learning according to claim 5 is characterized in that: The expression for evaluating each node embedding function satisfies: Among them, f t (·) is the node embedding function, N(v i ) is the node data v obtained by biased random walk sampling i The neighbor node set, Pr(N(v i )|f t (v i )) is in f t (v i ) under the condition that N(v i ) occurs, V is the set of the multiple node data.

7. The real-time prediction method for urban crime events based on multi-objective reinforcement learning according to claim 5 is characterized in that: The node embedding function is used to perform calculations on the plurality of node hidden state data to obtain a plurality of node state data corresponding to each other, including: For each node hidden state data, the deep walk algorithm is used to perform calculations on the node hidden state data and the corresponding node embedding function to obtain the corresponding node state data.

8. The method for real-time prediction of urban crime events based on multi-objective reinforcement learning according to claim 1 is characterized in that: The action decision variables are determined based on multiple node state data and a deep Q network using an ε-greedy strategy, including: For each node state data, the node state data is input into the deep Q network applying the ε-greedy strategy, and the corresponding decision action value is calculated based on the updated parameters and the node state data.

9. The method for real-time prediction of urban crime events based on multi-objective reinforcement learning according to claim 8 is characterized in that: The solution formula for the update parameter corresponding to each node state data satisfies: L(θ)=(1-λ)L A (θ)+λL B (i); Among them, L(θ) is the final loss function, L A (·) is the main loss function used to predict the reward value, L B (θ) is the auxiliary loss function used to optimize the loss function L A (·) Non-smooth problem, λ is L A (·) and L B (θ), the value range is 0 to 1, θ is the update parameter, E[·] is used to calculate the expectation, y is the target value, s t is the node status data, a t is the decision action value, ω is the weight value, the value range is 0 to 1, and it decays from 1 to 0 during the training process. Q is the decay value. is the mean square processing, |·| is to obtain the absolute value, T is the maximum value of the observation time window, γ is the attenuation coefficient, r t is the corresponding accuracy reward value at time t, argmax(·) is used to return the maximized independent variable of the function expression, ω is the set of multiple weight values ​​at time t, Q(s t+1 ,a,ω′;θ k ) is a set of multiple network Q values ​​at time t+1, ω' is a set of multiple weight values ​​at time t+1, a is a set of multiple action decision variables, θ k is the update parameter corresponding to the kth iteration of the model.

10. The method for real-time prediction of urban crime events based on multi-objective reinforcement learning according to claim 1 is characterized in that: The solution formula of the final prediction result satisfies: in, is the final prediction result, is the initial prediction result, a t is a set of multiple action decision variables at time t, ω T is the inversion of a set of multiple weight values, Q(s t ,a,ω;θ) is the set of multiple network Q values ​​at time t, and the value range of ε is 0~1, which decays from 1 to 0 during the training process.