Alarm condition data analysis method based on graph auto-encoder
Through a deep learning method based on graph autoencoder and combined with the cross-modal dynamic fusion mechanism, a deep fusion network is built, which solves the problem of multimodal correlation in police data analysis, and achieves more accurate police data prediction and model stability improvement.
Patent Information
- Application Number
- CN202510376349.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-08-19
AI Technical Summary
The existing police data analysis methods fail to make full use of the correlation between multimodal data, resulting in inaccurate prediction effects, and traditional methods fail to deeply explore the correlation between modals, resulting in incomplete and accurate model training.
The deep learning method based on graph autoencoder is adopted, and by building a deep fusion network, combining a cross-modal dynamic fusion mechanism, efficient analysis and deep mining of police data are achieved, and information from multiple modal data types is used to adaptively adjust the fusion weight to realize the organic integration of multimodal data.
It improves the accuracy of police data prediction and the generalization ability of the model, can capture key information more accurately, and improves the accuracy and stability of the prediction.
Smart Images

Figure CN120508822A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of police data analysis, and in particular to a police data analysis method based on a graph autoencoder. Background Art
[0002] With the advancement of smart cities, police data has experienced explosive growth. Its multi-source, heterogeneous nature and the nonlinear correlations between multiple factors have posed significant challenges to data processing. However, current police data analysis cannot simply yield accurate results based on a single modality. For example, certain locations have extremely high crime rates during specific periods of time, and some items involved in crimes have a strong correlation with the methods used. Therefore, simply analyzing a single piece of police data, such as the time of occurrence or the items involved, is unlikely to yield accurate results. Therefore, fully leveraging the data across all modalities has become a pressing challenge in police data analysis.
[0003] Among existing inventions for analyzing police data, the mainstream approaches fall into two main categories: one is modal fusion, which simply concatenates police data at the data level without performing structural correlation; the other relies on manual feature engineering or simply extracts features from the data using shallow neural networks. While these inventions can achieve better predictions than raw data, they also suffer from at least two problems:
[0004] (1) Existing inventions mostly rely on single-modal data and ignore the correlation between multiple modalities. However, police data naturally have structural characteristics, which leads to the neglect of potential correlations and thus reduces the predictive effect of police data.
[0005] (2) Although some inventions use information from different modalities, they simply align or connect information from different modalities without deeply exploring the correlation between modalities, which leads to incomplete and inaccurate model training.
[0006] In view of this, simply using multiple modalities and simply aligning or connecting information from different modalities obviously cannot ensure that the police data can be fully utilized. Therefore, for intelligent street patrol systems, it is very necessary to design a method that can fully utilize police data analysis and obtain better police data prediction results to predict the results of police data processing. The present invention proposes a police data analysis method based on a graph autoencoder, which realizes the processing of police data by constructing a deeply integrated network structure, thereby realizing the prediction of the processing results of the given police data. Summary of the Invention
[0007] The purpose of the present invention is to provide a police data analysis method based on a graph autoencoder to achieve efficient analysis and deep mining of police data through deep learning and other related technologies, breaking through the limitations of single-modal analysis in traditional police data processing, achieving accurate prediction and classification of police data, and making up for the defect of insufficient information fusion in traditional police data processing.
[0008] In order to achieve the above objectives, the technical solution adopted by the present invention includes the following steps:
[0009] Step S1: Obtain historical police data and construct a training dataset;
[0010] Step S2: Building a deep learning network by training the data set in an artificial neural network;
[0011] Step S3: Build a fused autoencoder model with the help of a deep learning network;
[0012] Step S4: constructing a police information data analysis model based on the fused autoencoder model;
[0013] Step S5: Use the police situation data analysis model to predict the police situation data processing results.
[0014] The step S1 acquires historical police data and constructs a training data set; specifically includes the following steps:
[0015] Step S11: Collect historical police information data, including case type, location, time of occurrence, persons involved, items involved, and modus operandi; the case types include criminal cases, public security cases, traffic cases, and other cases; the location includes specific addresses in the standard address database, such as residential areas, streets, public places, etc.; the time of occurrence includes daytime, nighttime, weekdays, holidays, etc.; the persons involved include suspects, victims, and witnesses; the items involved include various items used to commit crimes, such as drugs, guns, stolen money, etc.; the modus operandi includes fraud, robbery, etc.;
[0016] Step S12: Clean the collected historical alarm data, including processing missing values, abnormal values, erroneous values, and deduplication of duplicate data, and integrate the cleaned alarm data;
[0017] Step S13: Convert the format of the fused police information data using different conversion methods according to their data types;
[0018] Step S14: Convert the fused data in the converted format into vector form and construct a specific edge set, further construct a training data set, and divide the training data set into a training set and a test set in a ratio of 8:2 according to time stratified sampling. The training set is used to train the model, and the test set is used to test the model performance.
[0019] Step S2 builds a deep learning network by training the data set in an artificial neural network; specifically includes the following steps:
[0020] Step S21: setting the structure of the artificial neural network and initializing it;
[0021] Step S22: Building an initial deep learning network by setting a back propagation algorithm on the initialized artificial neural network structure;
[0022] Step S23: Setting an optimizer on the initial deep learning network and training data therein to obtain a deep learning network.
[0023] The step S3 constructs a fused autoencoder model with the help of a deep learning network; specifically includes the following steps:
[0024] Step S31: constructing a symmetric graph autoencoder model with the help of the autoencoder model and the graph autoencoder model;
[0025] Step S32: Train the symmetric graph autoencoder model to obtain a fused autoencoder model.
[0026] The step S4 constructs a police data analysis model based on the fused autoencoder model, specifically including the following steps:
[0027] Step S41: Introduce a cross-modal dynamic fusion mechanism to further fuse the feature information of the autoencoder and the graph autoencoder. The specific steps are as follows:
[0028] Step S411: using linear combination operation to obtain initial fusion information;
[0029] Step S412: using a message passing operation to process the initial fusion information to obtain local structure enhancement information;
[0030] Step S413: introducing an autocorrelation learning mechanism to obtain a global structure enhancement matrix;
[0031] Step S414: obtaining the final fusion feature information through skip connection;
[0032] Step S42: Construct the final police data analysis model based on the fused feature information. The specific steps are as follows:
[0033] Step S421: obtaining a soft assignment matrix by calculating the similarity between the sample and the cluster center;
[0034] Step S422: generating a target distribution by processing the soft allocation matrix;
[0035] Step S423: Calculate the KL divergence formula between the soft allocation matrix and the target distribution;
[0036] Step S43: constructing a total loss function of the model to obtain a police data analysis model;
[0037] Step S44: Model determination. An iteration threshold is set in advance. When the number of model iterations is greater than the iteration threshold, the model training is completed and the test data set is input into the trained model to test the model performance. Otherwise, training continues.
[0038] The step S5 uses the police data analysis model to predict the police data processing results. This is based on the police data analysis model constructed in step S4. By inputting the police data to be analyzed, the model outputs the probability distribution of the case type to which the police data belongs, and selects the case type with the highest probability as the prediction result.
[0039] Compared with traditional methods of police information processing, the present invention has the following beneficial effects and advantages:
[0040] (1) This paper combines deep learning networks and autoencoder models to achieve complementary fusion of attribute information and structural information by constructing a symmetric graph autoencoder model. The designed hybrid reconstruction loss function can more effectively guide model training and improve the generalization ability and stability of the model.
[0041] (2) The present invention introduces a cross-modal dynamic fusion mechanism that can simultaneously integrate information of multiple data types, such as case type, location, time of occurrence, persons involved, items involved, and modus operandi, so as to more accurately capture key information in police data and thus improve the accuracy of prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a step diagram of a police data analysis method based on a graph autoencoder according to the present invention.
[0043] Figure 2 This is a diagram of the implementation process of a police data analysis method based on a graph autoencoder according to the present invention. DETAILED DESCRIPTION
[0044] Example:
[0045] like Figure 1 As shown, the technical solution of the present invention includes five steps: S1 obtains historical police data and constructs a training data set, S2 builds a deep learning network by training the data set in an artificial neural network, S3 uses the deep learning network to build a fused autoencoder model, S4 builds a police data analysis model based on the fused autoencoder model, and S5 uses the police data analysis model to predict the police data processing results.
[0046] Step S1 acquires historical alarm data and constructs a training dataset: the original alarm data is obtained by collecting historical alarm data, performing data cleaning and fusion, and then converting the original flight data format and vector conversion to obtain alarm data that can be used for data processing. The final training dataset is constructed by constructing a specific edge set on the data.
[0047] The step S2 builds a deep learning network by training the data set in the artificial neural network: building and initializing the artificial neural network, constructing an initial deep learning network with the help of the training data set, and finally adopting an optimization strategy to obtain the deep learning network;
[0048] The step S3 constructs a fused autoencoder model with the help of a deep learning network: first, a symmetric graph autoencoder model is constructed with the help of the autoencoder model and the graph autoencoder model, and then the symmetric graph autoencoder model is trained to obtain a fused autoencoder model;
[0049] Step S4 constructs a police data analysis model based on the fused autoencoder model: first, the feature information of the autoencoder and the graph autoencoder is further integrated through linear combination operations, message passing operations, and the introduction of an autocorrelation learning mechanism. Then, the fused information is used to calculate the soft assignment matrix and generate the target distribution. Then, the loss function of the model is defined and initialized to optimize the model. Finally, the model performance is tested using a test set.
[0050] The step S5 uses the police situation data analysis model to predict the police situation data processing result: the flight data of the aircraft to be located is input into the tested police situation data analysis model, and the case type to which the police situation data belongs is output.
[0051] Step S1 acquires historical police data and constructs a training data set; it includes the following steps:
[0052] Step S11: Collect historical police information data, including case type, location, time of occurrence, persons involved, items involved, and modus operandi; the case types include criminal cases, public security cases, traffic cases, and other cases; the location includes specific addresses in the standard address database, such as residential areas, streets, public places, etc.; the time of occurrence includes daytime, nighttime, weekdays, holidays, etc.; the persons involved include suspects, victims, and witnesses; the items involved include various items used to commit crimes, such as drugs, guns, stolen money, etc.; the modus operandi includes fraud, robbery, etc.;
[0053] Step S12 performs data cleaning on the collected historical police data: including using a hash algorithm to identify and merge duplicate records; using a hierarchical marking strategy to mark missing values in the police data, that is, marking key fields and non-key fields with different contents; correcting and matching case types and modus operandi by creating mapping tables and regular expressions; using address resolution to complete and use a forward maximum matching algorithm to match the standard address library to standardize the location data, such as completing "XX Community" to "XX Street XX Community"; using the time resolution library's processing method to unify the time format of the occurrence time, the format is It can be year, month, day, hour, minute, and second; establish a synonym database and unify the names of items involved in the case and the modus operandi through bidirectional mapping, such as "heroin" and "drugs", "drugs" and "heroin / methamphetamine / morphine", "snatch" and "robbery"; at the same time, establish classification standards to concretize the vague descriptions in the police data; unify the data by setting unique constraints and using check constraints to constrain the database integrity method, such as the case type can only be a predefined value, and the occurrence time must be province / city / district / street / district; and perform data fusion on the unified processed data through multidimensional aggregation methods;
[0054] Step S13 converts the format of the fused data: case types and modus operandi are converted into data formats using label coding (e.g., criminal case = 1, public security case = 2; fraud = 1, robbery = 2); the location of the case is converted using the Word2Vec word embedding method; the persons involved in the case are converted using the label coding form after name standardization; the data of the items involved in the case are converted and label-coded using the classification standard established by classification standardization (e.g., drugs = 1, guns = 2); the time of occurrence is converted by extracting time features; in particular, all extracted unique case types, locations, times of occurrence, persons involved, items involved, and modus operandi are used as nodes;
[0055] Step S14: Convert the fused data in the converted format into vector form and construct a specific edge set, and further construct a training data set. The training data set is divided into a training set and a test set in a ratio of 8:2 by time stratified sampling. The training is used to train the model, and the test set is used to test the model performance. In particular, the constructed specific edge sets are case type to location, case type to time of occurrence, case type to persons involved, case type to items involved, modus operandi to persons involved, modus operandi to location, items involved to modus operandi, and items involved to location; the vector form is: [case type, location, time of occurrence, persons involved, items involved, modus operandi].
[0056] Step S2 builds a deep learning network by training the data set in an artificial neural network, including the following steps:
[0057] Step S21: Set the structures of the artificial neural network autoencoder and the graph autoencoder and use He initialization to initialize the encoding part of the autoencoder structure and the graph autoencoder structure, and use Xavier initialization to initialize the decoding part of the graph autoencoder structure; the He initialization formula and Xavier initialization formula are as follows:
[0058] (1)
[0059] Where W is the weight matrix, N represents the normal distribution, is the number of neurons, b is the bias term;
[0060] (2)
[0061] Where W is the weight matrix, U represents uniform distribution, is the input dimension, is the output dimension;
[0062] Step S22: Build an initial deep learning network by setting the back propagation algorithm on the initialized artificial neural network structure, wherein the mean square error is used to calculate the loss and the chain rule is used to calculate the gradient; the network output and loss in the back propagation algorithm are calculated as follows:
[0063] (3)
[0064] (4)
[0065] Where, is the activation function, is the input data, N is the number of samples, Output for the network;
[0066] Step S23: A deep learning network is obtained by training the dataset in the initial deep learning network and optimizing the network using the Adam optimizer with set hyperparameters. The relevant formulas for the autoencoder and graph autoencoder are expressed as follows:
[0067] (5)
[0068] Where, is the attribute matrix, is the encoding part of the autoencoder;
[0069] (6)
[0070] Where, is the latent embedding of the autoencoder, It is the decoding part of the autoencoder;
[0071] (7)
[0072] (8)
[0073] Where, is a nonlinear activation function, such as or , are the learnable weights of the graph autoencoder, is the original adjacency matrix, is the corresponding degree matrix, and ; is the unit matrix, indicating that each node is connected to itself, and N is the number of data samples;
[0074] (9)
[0075] Where, is the latent embedding of the graph autoencoder, are the learnable weights of the decoding part of the graph autoencoder.
[0076] The step S3 constructs a fused autoencoder model with the help of a deep learning network, including the following steps:
[0077] Step S31 uses the autoencoder model and the graph autoencoder model to construct a symmetric graph autoencoder model, so that the attribute matrix and the adjacency matrix can be reconstructed simultaneously; the formulas for the encoding part and the decoding part are expressed as follows:
[0078] (10)
[0079] (11)
[0080] Where, and Representative Layer encoder and The learnable parameters of the layer decoder, is a nonlinear activation function, such as or .,and is the set of data nodes;
[0081] Step S32 trains the symmetric graph autoencoder model to obtain a fused autoencoder model: the fused autoencoder model is constructed by designing and minimizing the hybrid reconstruction loss function between the weighted attribute matrix and the adjacency matrix. The hybrid loss formula is as follows:
[0082] (12)
[0083] Where, are predefined hyperparameters, Reconstruction loss for weighted attributes, Reconstruction loss for the adjacency matrix;
[0084] (13)
[0085] Where, is the reconstructed weighted attribute matrix;
[0086] (14)
[0087] Where, is the reconstructed adjacency matrix.
[0088] The step S4 constructs a police data analysis model based on the fused autoencoder model, including the following steps:
[0089] Step S41: Introduce a cross-modal dynamic fusion mechanism to further fuse the feature information of the autoencoder and the graph autoencoder, specifically including the following steps:
[0090] Step S411: Use linear combination operation to obtain initial fusion information; it is expressed as follows:
[0091] (15)
[0092] Where, is a learnable parameter and is initialized to 0.5;
[0093] Step S412: Process the initial fusion information by message passing to obtain local structure enhancement information; this is expressed as follows:
[0094] (16)
[0095] Where, For local enhancement ;
[0096] Step S413: Introduce the autocorrelation learning mechanism to obtain the global structure enhancement matrix. First, normalize the autocorrelation matrix, and then use the global correlation between samples to reconstruct the local enhanced potential embedding; it is expressed as follows:
[0097] (17)
[0098] (18)
[0099] Where S is the reconstruction coefficient;
[0100] Step S414: Obtain the final fusion feature information of the autoencoder and the graph autoencoder through skip connection, which is expressed as follows:
[0101] (19)
[0102] Where, is the jump ratio parameter and is initialized to 0;
[0103] Step S42 constructs a police data analysis model based on the fused feature information, specifically including the following steps:
[0104] Step S421: Obtain the soft assignment matrix by calculating the similarity between the sample and the cluster center. Using Student's t-distribution as the kernel, calculate the similarity between each sample in the fused embedding space and the pre-calculated cluster center. It is expressed as follows:
[0105] (20)
[0106] Where, represents the i-th sample, represents the jth sample centroid;
[0107] Step S422: Generate target distribution by processing the soft allocation matrix, which is expressed as follows:
[0108] (twenty one)
[0109] Where, Represents the element of the target distribution, the probability that the i-th sample belongs to the j-th sample centroid;
[0110] Step S423: Calculate the KL divergence formula between the soft assignment matrix and the target distribution to ensure that the autoencoder, graph autoencoder, and fused autoencoder model are trained in the same framework; expressed as follows:
[0111] (twenty two)
[0112] Where, and denote the soft assignment distributions of the autoencoder and graph autoencoder respectively;
[0113] Step S43 constructs the model total loss function to train the police data analysis model: the total loss function mainly consists of three parts: the reconstruction loss of the autoencoder, the reconstruction loss of the fused autoencoder model, and the KL divergence formula between the three soft assignment matrices and the target distribution in step S423; it is expressed as follows:
[0114] (twenty three)
[0115] Step S44 completes the training and tests the model performance: Model judgment, an iteration threshold is set in advance, when the number of iterations of the model is greater than the iteration threshold, the model training is completed and the test data set is input into the trained model to test the model performance; otherwise, training continues.
[0116] Step S5 uses the police situation data analysis model to predict the police situation data processing results: Based on the police situation data analysis model constructed in step S4, by inputting the police situation data to be analyzed, the model outputs the probability distribution of the case types to which the police situation data belongs, and selects the case type with the highest probability as the prediction result; it is expressed as follows:
[0117] (twenty four)
[0118] Where, is the output of the jth neuron, Indicates the total number of case types.
[0119] like Figure 2 As shown, the specific implementation process of the police data analysis method based on the graph autoencoder of the present invention is as follows:
[0120] Step 1: First, collect historical data; then clean and fuse the data; then convert the data format; finally, convert the data into vector form and construct a graph dataset.
[0121] Step 2: Set up and initialize the artificial neural network; then set up the backpropagation algorithm to build the initial deep learning network; finally, train the initial deep network to obtain the deep learning network.
[0122] Step 3: Use the deep learning network to build a symmetric graph autoencoder model; then input the data into the symmetric autoencoder model for training to obtain a fused autoencoder model.
[0123] Step 4: First, introduce the cross-modal dynamic fusion mechanism to fuse feature information, then build a police data analysis model through the fused feature information, then construct the model total loss function to train the police data analysis model, and finally train and test the model.
[0124] Step 5: Use the police data analysis model to predict the police data processing results.
[0125] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A police data analysis method based on graph autoencoder, characterized in that: The steps include: Step S1: Obtain historical police data and construct a training dataset; Step S2: Building a deep learning network by training the data set in an artificial neural network; Step S3: Build a fused autoencoder model with the help of a deep learning network; Step S4: constructing a police information data analysis model based on the fused autoencoder model; Step S5: Use the police situation data analysis model to predict the police situation data processing results.
2. The method for analyzing police information based on a graph autoencoder according to claim 1, characterized in that: The step S1 acquires historical police data and constructs a training data set; specifically includes the following steps: Step S11: Collect historical police information data, including case type, location, time of occurrence, persons involved, items involved, and modus operandi; the case types include criminal cases, public security cases, traffic cases, and other cases; the location includes specific addresses in the standard address database, such as residential areas, streets, public places, etc.; the time of occurrence includes daytime, nighttime, weekdays, holidays, etc.; the persons involved include suspects, victims, and witnesses; the items involved include various items used to commit crimes, such as drugs, guns, stolen money, etc.; the modus operandi includes fraud, robbery, etc.; Step S12: Clean the collected historical alarm data, including processing missing values, abnormal values, erroneous values, and deduplication of duplicate data, and integrate the cleaned alarm data; Step S13: Convert the format of the fused police information data using different conversion methods according to their data types; Step S14: Convert the fused data in the converted format into vector form and construct a specific edge set, further construct a training data set, and divide the training data set into a training set and a test set in a ratio of 7:
3. The training set is used to train the model, and the test set is used to test the model performance.
3. The method for analyzing police information data based on a graph autoencoder according to claim 1, characterized in that: Step S2 builds a deep learning network by training the data set in an artificial neural network, specifically including the following steps: Step S21: setting the structure of the artificial neural network and initializing it; Step S22: Building an initial deep learning network by setting a back propagation algorithm on the initialized artificial neural network structure; Step S23: Setting an optimizer on the initial deep learning network and training data therein to obtain a deep learning network.
4. The method for analyzing police information data based on a graph autoencoder according to claim 1, characterized in that: The step S3 constructs a fused autoencoder model with the help of a deep learning network; specifically includes the following steps: Step S31: constructing a symmetric graph autoencoder model with the help of the autoencoder model and the graph autoencoder model; Step S32: Train the symmetric graph autoencoder model to obtain a fused autoencoder model.
5. The method for analyzing police information data based on a graph autoencoder according to claim 1, characterized in that: The step S4 constructs a police data analysis model based on the fused autoencoder model, specifically including the following steps: Step S41: Introduce a cross-modal dynamic fusion mechanism to further fuse the feature information of the autoencoder and the graph autoencoder. The specific steps are as follows: Step S411: using linear combination operation to obtain initial fusion information; Step S412: using a message passing operation to process the initial fusion information to obtain local structure enhancement information; Step S413: introducing an autocorrelation learning mechanism to obtain a global structure enhancement matrix; Step S414: obtaining the final fusion feature information through skip connection; Step S42: Construct the final police information analysis model through the fused feature information; the specific steps are as follows: Step S421: obtaining a soft assignment matrix by calculating the similarity between the sample and the cluster center; Step S422: generating a target distribution by processing the soft allocation matrix; Step S423: Calculate the KL divergence formula between the soft allocation matrix and the target distribution; Step S43: constructing a total loss function of the model to obtain a police data analysis model; Step S44: Model determination. An iteration threshold is set in advance. When the number of model iterations is greater than the iteration threshold, the model training is completed and the test data set is input into the trained model to test the model performance. Otherwise, training continues.
6. The method for analyzing police information data based on a graph autoencoder according to claim 1, characterized in that: The step S5 uses the police data analysis model to predict the police data processing results. This is based on the police data analysis model constructed in step S4. By inputting the police data to be analyzed, the model outputs the probability distribution of the case type to which the police data belongs, and selects the case type with the highest probability as the prediction result.