A method for detecting public safety emergencies based on big data
By constructing an emergency event detection model including spatiotemporal graph association network, global semantic aggregation and multi-scale spatiotemporal correlation hybrid attention module, the problem of insufficient processing of multimodal data association and spatiotemporal characteristics in the prior art is solved, and efficient and accurate detection of public safety emergencies is achieved.
Patent Information
- Application Number
- CN202410995815.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-07-24
AI Technical Summary
The prior art is difficult to effectively detect and identify public safety emergencies, especially when dealing with multimodal data and spatiotemporal characteristics, and lacks considerations for multimodal data association and information fusion of different data sources.
A sudden event detection model is built that includes a spatiotemporal graph association network module, a global semantic aggregation module and a multi-scale spatiotemporal correlation hybrid attention module. Through these modules, deep semantic information and spatiotemporal characteristics of text and image data are extracted, and comprehensive analysis is carried out to improve the accuracy of event detection.
It significantly improves the accuracy and efficiency of emergency detection, can accurately identify public emergencies, and combines global semantic aggregation and space-time graph association network modules to form comprehensive and accurate comprehensive features, and optimizes the model training process to ensure the accuracy of detection results.
Smart Images

Figure CN118940164B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of big data processing technology and public safety emergency event detection technology, and in particular to a public safety emergency event detection method based on big data. Background Art
[0002] With the popularity of various social media channels and the development of big data technology, people can obtain a large amount of public safety-related data through different channels, such as disaster events, emergencies, social events, etc. Large-scale public safety-related data is usually unstructured, including text data (such as social media messages, news reports) and multimedia data (such as images, videos); therefore, it is not easy to accurately detect and identify public safety emergencies from these massive data.
[0003] At present, the detection of emergencies is mostly aimed at one modal data, but the same emergency usually includes data of multiple modalities such as text, images, and videos, and data of different modalities may contain different information and can provide complementary information, while the existing technology does not take into account the correlation between multimodal data. At the same time, how to effectively integrate information from different data sources and conduct comprehensive analysis to improve the accuracy of event detection and identification is also an important challenge. Secondly, public safety events usually have spatiotemporal characteristics, that is, the distribution and evolution of events in time and space; how to model spatiotemporal correlation in data and how to analyze events at different scales is a key issue. Summary of the invention
[0004] To solve the above problems, the present invention provides a method for public safety emergency detection based on big data, including constructing and training an emergency detection model, inputting the public data to be detected into the trained emergency detection model, and outputting the detection results; the emergency detection model includes a spatiotemporal graph association network module, a global semantic aggregation module, and a multi-scale spatiotemporal association hybrid attention module;
[0005] The training process of the emergency detection model includes the following steps:
[0006] S1. Collecting and preprocessing multimodal data of public emergencies, wherein the multimodal data of public emergencies includes multiple groups of emergency data, each group of emergency data includes text modal data and image modal data corresponding to the same emergency;
[0007] S2. For the pre-processed emergency event data, its text modal data is sent to the global semantic aggregation module to obtain the global semantic information representation;
[0008] S3. Sending the processed image modality data corresponding to the text modality data in step S2 to the spatiotemporal graph association network module to obtain a spatiotemporal graph representation;
[0009] S4. Input the global semantic information representation and the graph spatiotemporal representation into the multi-scale spatiotemporal correlation hybrid attention module to obtain comprehensive features;
[0010] S5. Pass the comprehensive features through the linear layer to obtain the detection results, calculate the loss according to the spatiotemporal comprehensive cross entropy loss function and back-propagate the training model until the maximum number of iterations is reached and save the model parameters.
[0011] Beneficial effects of the present invention:
[0012] The public safety emergency detection method based on big data of the present invention significantly improves the accuracy and efficiency of emergency detection. By constructing an emergency detection model including a spatiotemporal graph association network module, a global semantic aggregation module and a multi-scale spatiotemporal association hybrid attention module, the present invention can make full use of the rich information of multimodal data to achieve accurate identification of public emergencies. The global semantic aggregation module effectively extracts the deep semantic information of the text modality, while the spatiotemporal graph association network module captures the spatial and temporal dependencies in the image modality. The multi-scale spatiotemporal association hybrid attention module further integrates these two types of information to form a more comprehensive and accurate comprehensive feature, thereby significantly improving the detection performance of the model. In addition, the present invention further ensures the accuracy of the detection results by optimizing model training through the spatiotemporal comprehensive cross entropy loss function. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a flow chart of the method of the present invention;
[0014] Figure 2 This is a training framework diagram of the emergency detection model of the present invention. DETAILED DESCRIPTION
[0015] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0016] The present invention provides a method for public safety emergency detection based on big data, including constructing and training an emergency detection model, inputting public data to be detected into the trained emergency detection model, and outputting detection results; the emergency detection model includes a spatiotemporal graph association network module, a global semantic aggregation module, and a multi-scale spatiotemporal association hybrid attention module.
[0017] like Figure 1 , Figure 2 As shown in Figure 1, the training process of the emergency detection model includes the following steps:
[0018] S1. Collecting and preprocessing multimodal data of public emergencies, wherein the multimodal data of public emergencies includes multiple groups of emergency event data, and each group of emergency event data includes text modal data and image modal data corresponding to the same emergency event.
[0019] Specifically, step S1 includes cleaning, denoising and standardizing the multimodal data of public emergencies.
[0020] S2. For the pre-processed emergency event data, its text modal data is sent to the global semantic aggregation module to obtain the global semantic information representation;
[0021] Specifically, the global semantic aggregation module includes a pre-trained BERT model, a position encoding module, a word segmentation module, a fully connected layer, a ReLU activation function layer, a Dropout layer and an output layer; step S2 sends the text modality data into the global semantic aggregation module to obtain a global semantic information representation, including:
[0022] S21. Use the pre-trained BERT model to encode the text modality data into word vectors to obtain a word vector sequence;
[0023] S22. Send the word vector sequence to the position encoding module to obtain the position information vector;
[0024] S23. Use the word segmentation module to perform word segmentation processing on the text modal data to obtain a word segmentation sequence;
[0025] S24. Perform part-of-speech tagging and entity recognition on the word segmentation sequence to obtain a part-of-speech sequence and an entity type sequence, and convert the part-of-speech sequence and the entity type sequence into word embedding vector representations to obtain a part-of-speech vector and an entity type vector;
[0026] S25. concatenate the word vector sequence, the part-of-speech vector, the entity type vector, and the position information vector to obtain a concatenated vector;
[0027] S26. Input the concatenated vector into the fully connected layer to obtain the connection vector, use the ReLU activation function layer to process the connection vector to obtain the correction vector, and input the correction vector into the Dropout layer and the output layer to obtain the global semantic information representation.
[0028] Specifically, the position encoding module in step S22 includes:
[0029]
[0030] In the formula, x i represents the i-th word vector in the word vector sequence, where PE(·) represents the position encoding operation, Re represents the real part operation, d represents the total length of the word vector sequence, α represents the random initial value, and θ i+1 =10000 -2(i+1) / d is the calculation parameter; q [2i:2i+1] It represents the query vector after the word vector corresponding to the [2i:2i+1]th token is integrated with the position information [2i:2i+1]. Indicates the key vector after the word vector corresponding to the [2i:2i+1]th token integrates the position information [2i:2i+1]; [2i:2i+1] means grouping the query and key vectors in pairs, using the slice writing method in the array. Specifically, when calculating the position encoding of each word vector, the elements at a fixed position in the word vector sequence must be calculated, that is, i ranges from 0 to d / 2-1. When i=0, q [2i:2i+1] The value is the query vector after the word vector corresponding to the 0th token is integrated with the position information 0. The value of is the key vector after the word vector corresponding to the 0th token integrates the position information 0; when i = 1, q [2i:2i+1] The value is the query vector after the word vector corresponding to the second token is integrated with the position information 2. At this time, it is the query vector of the second token and the key vector multiplied, and so on. The position encoding method of the present invention introduces the cosine function cos on the basis of the rotation encoding RoPE, so that a periodic change related to the position is introduced in the position encoding (PE) module to enhance the model's ability to capture the position information of words in the sequence. At the same time, the use of the cosine function helps to reduce the dimension of the data and simplify subsequent calculations and analysis. In addition, log(1-α) / 2 is used to adjust the position information of the position encoding module, thereby improving the performance of the model. The parameter α can control the strength and range of the position encoding. If α is relatively small, then log(1-α) / 2 will be relatively large, so that the strength of the position encoding will be relatively high; conversely, if α is relatively large, then log(1-α) / 2 will be relatively small, so that the strength of the position encoding will be relatively low.
[0031] Specifically, the calculation rules in step S25 include:
[0032]
[0033] In the formula, X w Represents a word vector sequence, X p represents the part-of-speech vector, X e represents the entity type vector, Represents the concatenation operation of vectors, PE(X w) represents the position information vector, X represents the concatenation vector, X i represents the i-th element of the splicing vector, n represents the length of the splicing vector; α i Indicates the corresponding X i , ω1 is the weight matrix of the fully connected layer, b1 is the bias vector, ReLU represents the rectified linear unit activation function, and Dp(·) represents the Dropout layer used to randomly set the output of some neurons to zero; ω2 is the weight matrix of the output layer, b2 is the bias vector of the output layer, and G is the final global semantic information representation.
[0034] S3. Sending the preprocessed image modality data corresponding to the text modality data in step S2 to the spatiotemporal graph association network module to obtain a spatiotemporal graph representation;
[0035] Specifically, in step S3, the image modality data is fed into the spatiotemporal graph association network module to obtain a spatiotemporal graph representation, including:
[0036] S31. For the image modality data {I1, I2, ..., I M} to extract features and obtain image feature data {F1, F2, ..., F M}; Among them, I i represents the i=1,2,…,Mth image in the image modality data, F i represents the feature vector corresponding to the i-th image, and M represents the number of images in the image modality data;
[0037] S32. Calculate the spatial similarity between images based on the image feature data;
[0038] S33. Extract the timestamp data {t1, t2, ..., t M}, calculate the time relationship between images based on the timestamp data; t i Represents the timestamp of the i-th image;
[0039] S34. Calculate the spatiotemporal relationship between images based on spatial similarity and temporal relationship, and construct a spatiotemporal relationship matrix based on the spatiotemporal relationship;
[0040] S35. After passing the spatiotemporal relationship matrix through multiple layers of graph convolution attention layers, the spatiotemporal representation of each node (i.e., each image in the image modality data) is obtained;
[0041] S36. Select the maximum spatiotemporal representation to perform maximum pooling, reshape the maximum pooling result into a one-dimensional tensor, and use the one-dimensional tensor as the graph spatiotemporal representation.
[0042] Specifically, the calculation formula of spatial similarity is:
[0043]
[0044] The calculation formula of time relationship is:
[0045]
[0046] The calculation formula of the time-space relationship is:
[0047] A ij =δS ij +(1-δ)T ij
[0048] In the formula, S ij represents the spatial similarity between the i-th image and the j-th image, T ij represents the time relationship between the i-th image and the j-th image, A ij It represents the spatiotemporal relationship between the i-th image and the j-th image; τ is a time scale parameter, and δ is a parameter that balances the spatial and temporal relationship.
[0049] Specifically, each graph convolution attention layer includes a graph spatiotemporal convolution and an attention mechanism; in step S35, after multiple layers of graph convolution attention layers, the spatiotemporal representation of each node is obtained, and the calculation rule is:
[0050] H (l+1) =σ(AH (l)′ W (l) )
[0051]
[0052] In the formula, H (l)′ Represents the input feature matrix of the l-th layer of graph spatiotemporal convolution, and its initialization input feature matrix is H (0)′ =[F1,F,...,F N ] T , A represents the space-time relationship matrix, W (l) H represents the weight matrix of the l-th layer of spatiotemporal convolution; (l+1) Represents the output feature matrix of the l-th layer of graph spatiotemporal convolution; Denotes the matrix H (l+1) The feature representation of the i-th node in ij represents the attention weight, Denotes the matrix H (l+1) The feature representation of the i-th node in is updated by the attention mechanism; a is the learnable parameter in the attention mechanism, || represents the connection operation of the vector, and σ represents the activation function. The output of the last layer of graph convolutional attention layer in Represents the spatiotemporal representation of the i-th node.
[0053] Specifically, in step S35, the spatiotemporal representations of all nodes are summarized to obtain the graph spatiotemporal representation, and the calculation rule is:
[0054]
[0055] In the formula, H global To represent the spatiotemporal representation of a graph, the method uses maximum pooling and then reduces or reshapes it into a one-dimensional tensor as input for subsequent steps.
[0056] S4. Input the global semantic information representation and spatiotemporal representation together into the multi-scale spatiotemporal correlation hybrid attention module to obtain comprehensive features.
[0057] Specifically, the processing of the multi-scale spatiotemporal correlation hybrid attention module in step S4 includes:
[0058]
[0059] O = ReLU(W f ·(β·A s +(1-β)·A c +S)+b f )
[0060] Among them, G represents the global semantic information representation, S represents the graph spatiotemporal representation, and A S represents the self-attention matrix, They represent the query weight matrix, key weight matrix, and value weight matrix in the self-attention mechanism respectively. is the scaling factor; A C represents the cross-attention matrix, Respectively represent the query weight matrix, key weight matrix, and value weight matrix in the cross-attention mechanism; O represents the comprehensive feature, β represents the attention weight mixing coefficient, and W f represents the fusion weight matrix, b f ReLU represents the rectified linear unit activation function.
[0061] S5. Pass the comprehensive features through the linear layer to obtain the detection results, calculate the loss according to the spatiotemporal integrated cross entropy loss function (STICE-Loss) and back-propagate the training model until the maximum number of iterations is reached and save the model parameters.
[0062] Specifically, the calculation rules of the spatiotemporal integrated cross entropy loss function (STICE-Loss) in step S5 include:
[0063]
[0064] Among them, the model's prediction result is P, the true label is Y, P i,jrepresents the model's predicted probability of the jth category for the i-th sample, Y i,j represents the true label of the jth category of the ith sample; N is the number of samples, C is the number of categories, γ is the sample category weight, which is used to balance the importance of positive and negative samples, and is the ratio of positive samples. Log is the natural logarithm function. The categories include: natural disasters, accidental disasters, public health events, and social security events. Each category includes its specific categories, and each category label corresponds to a unique digital identifier.
[0065] In the present invention, unless otherwise clearly stipulated and limited, the terms such as "installation", "setting", "connection", "fixation" and "rotation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral one; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. Unless otherwise clearly defined, ordinary technicians in this field can understand the specific meanings of the above terms in the present invention according to the specific circumstances.
[0066] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for detecting public safety emergencies based on big data, characterized in that: Construct and train an emergency detection model, input the public data to be detected into the trained emergency detection model, and output the detection results; the emergency detection model includes a spatiotemporal graph association network module, a global semantic aggregation module, and a multi-scale spatiotemporal association hybrid attention module; The training process of the emergency detection model includes the following steps: S1. Collecting and preprocessing multimodal data of public emergencies, wherein the multimodal data of public emergencies includes multiple groups of emergency data, each group of emergency data includes text modal data and image modal data corresponding to the same emergency; S2. For the pre-processed emergency event data, its text modal data is sent to the global semantic aggregation module to obtain the global semantic information representation; The global semantic aggregation module includes a pre-trained BERT model, a position encoding module, a word segmentation module, a fully connected layer, a ReLU activation function layer, a Dropout layer and an output layer; step S2 sends the text modality data to the global semantic aggregation module to obtain a global semantic information representation, including: S21. Use the pre-trained BERT model to encode the text modality data into word vectors to obtain a word vector sequence; S22. Send the word vector sequence to the position encoding module to obtain the position information vector; S23. Use the word segmentation module to perform word segmentation processing on the text modal data to obtain a word segmentation sequence; S24. Perform part-of-speech tagging and entity recognition on the word segmentation sequence to obtain a part-of-speech sequence and an entity type sequence, and convert the part-of-speech sequence and the entity type sequence into word embedding vector representations to obtain a part-of-speech vector and an entity type vector; S25. concatenate the word vector sequence, the part-of-speech vector, the entity type vector, and the position information vector to obtain a concatenated vector; S26. Input the concatenated vector into the fully connected layer to obtain a connection vector, use the ReLU activation function layer to process the connection vector to obtain a correction vector, and input the correction vector into the Dropout layer and the output layer to obtain a global semantic information representation; S3. Sending the processed image modality data corresponding to the text modality data in step S2 to the spatiotemporal graph association network module to obtain a spatiotemporal graph representation; In step S3, the image modality data is fed into the spatiotemporal graph association network module to obtain a spatiotemporal graph representation, including: S31. For the image modality data {I1, I2, ..., I M } to extract features and obtain image feature data {F1, F2, ..., F M }; Among them, I i represents the i=1,2,…,Mth image in the image modality data, F i represents the feature vector corresponding to the i-th image, and M represents the number of images in the image modality data; S32. Calculate the spatial similarity between images based on the image feature data; S33. Extract the timestamp data {t1, t2, ..., t M }, calculate the time relationship between images based on the timestamp data; t i Represents the timestamp of the i-th image; S34. Calculate the spatiotemporal relationship between images based on spatial similarity and temporal relationship, and construct a spatiotemporal relationship matrix based on the spatiotemporal relationship; S35. After passing the spatiotemporal relationship matrix through multiple layers of graph convolution attention layers, the spatiotemporal representation of each node is obtained; S36. Select the maximum spatiotemporal representation to perform maximum pooling, reshape the maximum pooling result into a one-dimensional tensor, and use the one-dimensional tensor as the graph spatiotemporal representation; S4. Input the global semantic information representation and the graph spatiotemporal representation together into the multi-scale spatiotemporal correlation hybrid attention module to obtain comprehensive features; Step S4: The processing of the multi-scale spatiotemporal correlation hybrid attention module includes: Among them, G represents the global semantic information representation, S represents the graph spatiotemporal representation, and A S represents the self-attention matrix, They represent the query weight matrix, key weight matrix, and value weight matrix in the self-attention mechanism respectively. is the scaling factor; A C represents the cross-attention matrix, They represent the query weight matrix, key weight matrix, and value weight matrix in the cross-attention mechanism respectively; O represents the comprehensive feature, β represents the attention weight mixing coefficient, represents the fusion weight matrix, represents the bias vector, ReLU represents the rectified linear unit activation function; S5. Obtain the detection result by passing the comprehensive features through the linear layer, calculate the loss according to the spatiotemporal comprehensive cross entropy loss function and back-propagate the training model until the maximum number of iterations is reached and save the model parameters.
2. A method for detecting public safety emergencies based on big data according to claim 1, characterized in that: The preprocessing of step S1 includes cleaning, denoising and standardizing the multimodal data of public emergencies.
3. A method for detecting public safety emergencies based on big data according to claim 1, characterized in that: The position encoding module in step S22 is expressed as: In the formula, x i represents the i-th word vector in the word vector sequence, where PE(·) represents the position encoding operation, Re represents the real part operation, d represents the total length of the word vector sequence, α represents the random initial value, and θ i+1 =10000 -2(i+1) / d ,q [2i:2i+1] It represents the query vector after the word vector corresponding to the [2i:2i+1]th token is integrated with the position information [2i:2i+1]. Indicates the key vector after the word vector corresponding to the [2i:2i+1]th token is integrated with the position information [2i:2i+1].
4. A method for detecting public safety emergencies based on big data according to claim 1, characterized in that: The calculation rules in step S25 include: In the formula, X w Represents a word vector sequence, X p represents the part-of-speech vector, X e Represents the entity type vector, PE(X w ) represents the position information vector, X represents the concatenation vector, X i represents the i-th element of the splicing vector, n represents the length of the splicing vector; α i Indicates the corresponding X i , ω1 is the weight matrix of the fully connected layer, b1 is the bias vector, ReLU represents the rectified linear unit activation function, and Dp(·) represents the Dropout layer; ω2 is the weight matrix of the output layer, b2 is the bias vector of the output layer, and G represents the global semantic information representation.
5. A method for detecting public safety emergencies based on big data according to claim 1, characterized in that: The calculation formula of spatial similarity is: The calculation formula of time relationship is: The calculation formula of the time-space relationship is: A ij =δS ij +(1-δ)T ij In the formula, S ij represents the spatial similarity between the i-th image and the j-th image, T ij represents the time relationship between the i-th image and the j-th image, A ij It represents the spatiotemporal relationship between the i-th image and the j-th image; τ is a time scale parameter, and δ is a parameter that balances the spatial and temporal relationship.
6. A method for detecting public safety emergencies based on big data according to claim 1, characterized in that: Each graph convolution attention layer includes a graph spatiotemporal convolution and an attention mechanism. In step S35, after multiple layers of graph convolution attention layers, the spatiotemporal representation of each node is obtained, and the calculation rule is: H (l+1) =σ(AH (l)′ W (l) ) In the formula, H (l)′ represents the input feature matrix of the l-th layer of graph spatiotemporal convolution; A represents the spatiotemporal relationship matrix, W (l) H represents the weight matrix of the l-th layer of spatiotemporal convolution; (l+1) Represents the output feature matrix of the l-th layer of graph spatiotemporal convolution; Denotes the matrix H (l +1) The feature representation of the i-th node in ij represents the attention weight, Denotes the matrix H (l+1) The feature representation of the i-th node after being updated by the attention mechanism; a is the learnable parameter in the attention mechanism, || represents the connection operation of the vector, and σ represents the activation function.
7. A method for detecting public safety emergencies based on big data according to claim 1, characterized in that: The calculation rules of the spatiotemporal integrated cross entropy loss function (STICE-Loss) in step S5 include: Among them, the model's prediction result is P, the true label is Y, P i,j represents the model's predicted probability of the jth category for the i-th sample, Y i,j represents the true label of the jth category of the ith sample; N is the number of samples, C is the number of categories, γ is the sample category weight, which is used to balance the importance of positive and negative samples, and its value is the ratio of positive samples. log is the natural logarithm function.
Citation Information
Patent Citations
Method for constructing a semantic representation model of service resources
CN113128237A
Argument extraction method for multi-view encoder by using topological dependency relationship
CN113222119A