A method, device and equipment for detecting violent incidents

By preprocessing the tag data set into tag prompt text, and splicing the text to be detected with the tag prompt text, feature extraction and reconstruction are performed, and brute force event detection is performed as input to the binary decoder, the problem of low detection accuracy and efficiency of brute force event in the prior art is solved, and more efficient and accurate detection effects are achieved.

CN115470348BActive Publication Date: 2025-06-20GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211097380.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-08
Publication Date
2025-06-20
Estimated Expiration
2042-09-08

AI Technical Summary

Technical Problem

The prior art has problems of inaccuracy and efficiency in violent event detection, especially in the context semantic recognition, overlap, typos, triggerless word classification and incremental learning of violent trigger words are difficult to effectively solve.

Method used

By obtaining the tag data set for preprocessing, the tag prompt text is obtained, and the text to be detected is spliced ​​with the tag prompt text to obtain the reconstructed input text. Then the reconstructed input text is encoded, the text and label representation sequence is extracted, feature extraction and reconstruction are integrated, and decoded as input of the binary decoder to output the brute force event detection result.

Benefits of technology

It improves the accuracy and efficiency of violent event detection, avoids the error transmission problem caused by the pattern of first extracting violent trigger words and then sorting events, and makes full use of the knowledge and label semantic information of the pre-trained model, enhancing the accuracy and speed of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115470348B_ABST
    Figure CN115470348B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of natural language processing, and more specifically, to a violent event detection method, apparatus and device. The method includes: obtaining a label data set; preprocessing the label data set to obtain label prompt text; splicing the text to be detected with the label prompt text to obtain a reconstructed input text; encoding the reconstructed input text to obtain a first text encoding sequence, and extracting a text representation sequence and a label representation sequence of the reconstructed input text from the first text encoding sequence; extracting features from the text representation sequence and the label representation sequence to obtain a label feature sequence; reconstructing a binary reconstructed input sequence by using the label feature sequence and the label representation sequence, and inputting the binary reconstructed input sequence into a binary decoder for decoding processing, and the binary decoder outputs a violent event detection result. The present invention improves the accuracy and efficiency of violent event detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and more specifically, to a method, device, and equipment for detecting violent events. Background Art

[0002] Nowadays, on major social networking platforms, a huge amount of information is being released and shared in real time every day, and a large amount of violent content exists among them. Violent events have a great impact on social stability. If violent events in social networks can be detected in a timely manner, the authorities can respond more effectively to real-time violent events and formulate various preventive policies according to geographical regions and event types.

[0003] Currently, for the task of detecting violent events based on text, usually, first, by using the BERT pre-trained model, the input text to be detected is converted into a sequence of vector representations, and then the sequence of vector representations is input into a neural network structure. The neural network extracts the trigger words of violent events of event types to achieve the detection of violent events.

[0004] However, the above method needs to overcome many difficulties, such as the recognition of the context semantics of violent trigger words, the overlap of violent trigger words, the misspelling of violent trigger words, the absence of violent trigger words or new violent trigger words, etc. That is, the method based on trigger word extraction needs to have capabilities such as error correction, classification without trigger words, and incremental learning, which makes the detection efficiency low. In addition, sometimes one type of violent event occurs along with the occurrence of another type of violent event, that is, there is a problem of detecting overlapping violent events, which also makes the method based on trigger word extraction have the defect of low detection accuracy. Summary of the Invention

[0005] The present invention provides a method, device, and equipment for detecting violent events to overcome the defects of low accuracy and efficiency in detecting violent events in the prior art.

[0006] To solve the above technical problems, the technical solution of the present invention is as follows:

[0007] In a first aspect, the present invention proposes a method for detecting violent events, including the following steps:

[0008] S1: Obtain a label data set; the label data set contains violent event labels.

[0009] S2: Preprocess the label data set to obtain label hint texts.

[0010] S3: Concatenate the text to be detected with the label hint texts to obtain a reconstructed input text.

[0011] S4: Perform encoding processing on the reconstructed input text to obtain a first text encoding sequence, and extract the text representation sequence and the label representation sequence of the reconstructed input text from the first text encoding sequence.

[0012] S5: Perform feature extraction on the text representation sequence and the label representation sequence to obtain a label feature sequence.

[0013] S6: Use the label feature sequence and the label representation sequence to reconstruct the input encoding of the binary decoder, input the input encoding into the binary decoder for decoding processing, and the binary decoder outputs the violent event detection result.

[0014] In a second aspect, the present invention proposes a violent event detection device, including:

[0015] An acquisition module, configured to acquire a label data set; the label data set contains violent event labels.

[0016] A preprocessing module, configured to preprocess the label data set to obtain a label prompt text.

[0017] A splicing module, configured to splice the text to be detected and the label prompt text to obtain a reconstructed input text.

[0018] An encoding module, configured to perform encoding processing on the reconstructed input text to obtain a first text encoding sequence.

[0019] An extraction module, configured to extract the text representation sequence and the label representation sequence of the reconstructed input text from the first text encoding sequence.

[0020] A feature extraction module, configured to perform feature extraction on the text representation sequence and the label representation sequence to obtain a label feature sequence.

[0021] A reconstruction module, configured to reconstruct the input encoding of the binary decoder by using the label feature sequence and the label representation sequence.

[0022] A detection module, configured to input the input encoding into the binary decoder for decoding processing, and the binary decoder outputs the violent event detection result.

[0023] In a third aspect, the present invention proposes an electronic device including a memory and a processor, and a computer program capable of running on the processor is stored on the memory. When the computer program is executed by the processor, the method described in the first aspect is implemented.

[0024] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows: The present invention preprocesses the violent event tags in the tag data set to obtain tag hint texts, then splices the text to be detected with the tag hint texts to obtain a reconstructed input text, and then extracts the feature information of the reconstructed input text, and reconstructs and fuses the feature information with the tag representation sequence as the input of the binary decoder to implement violent event detection, avoiding the problem of error transmission caused by the mode of first extracting violent trigger words and then classifying violent events. Moreover, the text to be detected and the violent event tags are jointly input into the binary decoder pre-training model in the form of tag hint texts, which can make full use of the knowledge of the pre-training model, endow the pre-training model with tag semantic information, co-occurrence information between tags, and interaction information between tags and texts, and improve the accuracy and efficiency of violent event detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a flowchart of the violent event detection method in Embodiment 1.

[0026] Figure 2 It is a schematic diagram of the violent event detection method in Embodiment 2.

[0027] Figure 3 It is a tag correlation attention map in Embodiment 2.

[0028] Figure 4 It is an architecture diagram of the violent event detection method in Embodiment 3. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0030] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0031] Embodiment 1

[0032] Please refer to Figure 1 , this embodiment proposes a violent event detection method, including the following steps:

[0033] S1: Obtain a tag data set; the tag data set contains violent event tags.

[0034] S2: Preprocess the tag data set to obtain tag hint texts.

[0035] S3: Splice the text to be detected with the tag hint texts to obtain a reconstructed input text.

[0036] S4: Encode the reconstructed input text to obtain a first text encoding sequence, and extract the text representation sequence and the label representation sequence of the reconstructed input text from the first text encoding sequence.

[0037] S5: Extract features from the text representation sequence and the label representation sequence to obtain a label feature sequence.

[0038] S6: Use the label feature sequence and the label representation sequence to reconstruct the input encoding of the binary decoder, input the input encoding into the binary decoder for decoding, and the binary decoder outputs the violent event detection result.

[0039] In the specific implementation process, the violent event label in the label dataset is preprocessed to obtain a label prompt text, then the text to be detected is concatenated with the label prompt text to obtain a reconstructed input text, and then the feature information of the reconstructed input text is extracted, and the feature information is reconstructed and fused with the label representation sequence as the input of the binary decoder to implement violent event detection, avoiding the problem of error propagation caused by the mode of first extracting violent trigger words and then classifying violent events, and jointly inputting the text to be detected and the violent event label in the form of label prompt text into the binary decoder pre-training model, which can make full use of the pre-training model knowledge, and endow the pre-training model with label semantic information, co-occurrence information between labels, and interaction information between labels and text, improving the accuracy and efficiency of violent event detection.

[0040] Embodiment 2

[0041] Refer to Figure 2 , this embodiment proposes a violent event detection method, including the following steps:

[0042] S1: Obtain a label dataset; the label dataset contains violent event labels.

[0043] In this embodiment, the label dataset uses the Spanish violent event detection dataset DAVINCIS. By counting the types of violent events, the label dataset contains a total of five violent event labels: accident, homicide, non-violent event, robbery, and kidnapping, and the violent event label set L = {l1, l2,..., l5} is obtained.

[0044] S2: Preprocess the label dataset to obtain a label prompt text. The specific steps include:

[0045] S2.1: Textualize the violent event labels in the label dataset to obtain a label text sequence.

[0046] S2.2: Reconstruct the label text sequence into a label prompt text of a natural language question.

[0047] In this embodiment, first, the violent event label set L = {l1, l2, …, l5} is texturized to obtain the label text sequence Y = {y1, y2, …, y5}, that is, the five labels of accident, homicide, non-violent event, robbery, and kidnapping are respectively represented by "Accident", "Asesinato", "Paz", "Robo", and "Secuestro". Then, the label text sequence Y is reconstructed into the form of the label prompt text Z = {y1, p1, p2, …, p5, y5} of a natural language question, where P = {p1, p2, …, p k}, k represents the length of the label prompt text, which is part of the natural languageization of the violent event label set, and the label prompt text is obtained: "accidente asesinato paz robo o secuestro?".

[0048] For another example, the original text to be detected is "Zhang San killed Li Si.", and the labels include: homicide, robbery, and rape. The reconstructed label prompt text is "Zhang San killed Li Si. Does it belong to homicide, robbery, or rape?".

[0049] S3: Concatenate the text to be detected with the label prompt text to obtain the reconstructed input text.

[0050] In this embodiment, the text to be detected X = {x1, x2, …, x r} is concatenated with the label prompt text Z to form the reconstructed input text r represents the length of the original text to be detected. By constructing the label prompt text and supplementing the label prompt text to the text to be detected to complete the reconstruction, the reconstructed input text is obtained.

[0051] S4: Perform encoding processing on the reconstructed input text to obtain the first text encoding sequence, and extract the text representation sequence and label representation sequence of the reconstructed input text from the first text encoding sequence.

[0052] In this embodiment, the BERT model is used to perform encoding processing on the reconstructed input text, and the BERT model outputs the first text encoding sequence.

[0053] In this embodiment, the reconstructed input text is used as the input of the BERT model, and the BERT model outputs the corresponding first text encoding sequence, that is, the vector representation. Then, from the first text encoding sequence, according to the input position index, the text representation sequence and the label representation sequence

[0054] S5: Extract features from the text representation sequence and the label representation sequence to obtain a label feature sequence. The specific steps include:

[0055] S5.1: Input the text representation sequence into a long short-term memory network, and the long short-term memory network outputs a forward context vector and a backward context vector.

[0056] In this embodiment, by constructing a bidirectional long short-term memory network, each vector representation in the text representation sequence is sequentially input into the bidirectional long short-term memory network to obtain a context vector with context information and a backward context vector respectively. Their expressions are as follows:

[0057]

[0058]

[0059] where t represents the time.

[0060] S5.2: Use the forward context vector and the backward context vector to construct a forward context sequence and a backward context sequence respectively. Their expressions are as follows:

[0061]

[0062]

[0063] where n represents the length of the forward context sequence and the backward context sequence .

[0064] S5.3: Concatenate the forward context sequence and the backward context sequence to obtain a second text encoding sequence containing complete context information

[0065] S5.4: Fuse the features of the second text encoding sequence and the label representation sequence to obtain a label feature sequence.

[0066] In this embodiment, first calculate the dot product matrix D of the second text encoding sequence G and the label representation sequence H Y . Its expression is as follows:

[0067]

[0068] where For the label sequence H Y is the transpose, and the dot product matrix is the relationship matrix between labels and text, which integrates context information.

[0069] Since there are only 5 dataset labels in this embodiment, this embodiment strengthens the features of the dot product matrix D by constructing a convolutional layer with 5 convolutional kernels, specifically including: using ReLU as the activation function of the convolutional layer and using the strategy of max pooling operation to extract the label feature sequence a representing the features of each label, and its expression is as follows:

[0070] a = tanh(Φ(D))

[0071] where the function Φ(·) represents ReLU activation and max pooling operation, and tanh(Φ(D)) represents re-activating the feature vector obtained after ReLU activation and max pooling operation with the tanh function.

[0072] S6: Reconstruct the input encoding of the binary decoder using the label feature sequence and the label representation sequence, input the input encoding into the binary decoder for decoding processing, and the binary decoder outputs the violent event detection result.

[0073] In this embodiment, multiply the label feature sequence a and the label representation sequence H Y to obtain the sequence H' containing the interaction information between labels Y ; add the sequence H' containing the interaction information between labels Y to the label representation sequence to obtain the input encoding K of the binary decoder, and the expression is as follows:

[0074] H' Y = H Y × a

[0075] K = H Y + H' Y

[0076] In this embodiment, input the input encoding into the binary decoder for decoding processing, and the binary decoder outputs the violent event detection result, and its expression is as follows:

[0077]

[0078]

[0079] where contains the prediction results of 5 labels, K represents the binary reconstruction input sequence, and FC(·) represents the fully connected layer, is the detection result of the i-th violent event label (whether the text belongs to or does not belong to this violent event label), sigmoid(·) means converting the numerical range of the output of the fully connected layer into a probability using the sigmoid function and taking the maximum value of the prediction using argmax(·).

[0080] In this embodiment, the prediction result is decoded by a binary decoder into the probabilities of belonging to and not belonging to event i

[0081] In the specific implementation process, the sigmoid function is used to make each element value in the detection result vector of the violent event label obey the interval [0, 1], and set the prediction threshold β = 0.5. If there is an element in the detection result vector of the violent event label of the text to be detected satisfies then it is determined that the text to be detected belongs to violent event i

[0082] In this embodiment, the asymmetric loss L is used as the objective function for training the model BERT to alleviate the problem of data imbalance (including the few-shot problem), and its expression is as follows

[0083]

[0084] γ - >γ +

[0085] where L + represents the loss function of positive samples, and L - represents the loss function of negative samples is the offset probability, which means performing a hard threshold processing on very simple negative samples so that the model can discard negative samples with very low probabilities during training. γ represents the focusing function represents the predicted probability output by the model,, γ + is the positive focusing parameter, and γ - is the negative focusing parameter

[0086] Set γ - >γ + is intended to emphasize the training contribution of positive samples. Perform a hard threshold processing on very easy negative samples so that negative samples with very low probabilities can be discarded during training. So far, one model training is completed. After the model training converges, as Figure 3 shown, take the average of the attention masks of the last three layers of the pre-trained model BERT, and the pre-trained model BERT learns the relationship between labels

[0087] Embodiment 3

[0088] Please refer to Figure 4 , this embodiment proposes a violent event system based on label hint and binary decoding, including: an acquisition module for acquiring a label data set; the label data set contains violent event labels.

[0089] A preprocessing module for preprocessing the label data set to obtain label hint text.

[0090] A splicing module for splicing the text to be detected with the label hint text to obtain a reconstructed input text.

[0091] An encoding module for encoding the reconstructed input text to obtain a first text encoding sequence.

[0092] An extraction module for extracting the text representation sequence and label representation sequence of the reconstructed input text from the first text encoding sequence.

[0093] A feature extraction module for extracting features from the text representation sequence and the label representation sequence to obtain a label feature sequence.

[0094] A reconstruction module for reconstructing the input encoding of the binary decoder by using the label feature sequence and the label representation sequence.

[0095] A detection module for inputting the input encoding into a binary decoder for decoding processing, and the binary decoder outputs a violent event detection result.

[0096] In the specific implementation process, the violent event labels in the label data set are preprocessed to obtain label hint text, then the text to be detected is spliced with the label hint text to obtain a reconstructed input text, and then the feature information of the reconstructed input text is extracted, and the feature information is reconstructed and fused with the label representation sequence as the input of the binary decoder to implement violent event detection, avoiding the problem of error propagation caused by the mode of first extracting violent trigger words and then classifying violent events, and inputting the text to be detected and violent event labels jointly in the form of label hint text into the binary decoder pre-training model, which can make full use of the pre-training model knowledge, and endow the pre-training model with label semantic information, co-occurrence information between labels, and interaction information between labels and text, improving the accuracy and efficiency of violent event detection.

[0097] The terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0098] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or alterations can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for detecting violent incidents, characterized in that, Including: S1: Obtain a label dataset; the label dataset contains violent event labels; S2: Preprocess the label dataset to obtain label prompt text; S3: Concatenate the text to be detected with the label prompt text to obtain a reconstructed input text; S4: Encode the reconstructed input text to obtain a first text encoding sequence, and extract the text representation sequence and label representation sequence of the reconstructed input text from the first text encoding sequence; S5: Extract features from the text representation sequence and the label representation sequence to obtain a label feature sequence, including: S5.1: Input the text representation sequence into a long short-term memory network, and the long short-term memory network outputs a forward context vector and a backward context vector; S5.2: Construct a forward context sequence and a backward context sequence using the forward context vector and the backward context vector respectively; S5.3: Concatenate the forward context sequence and the backward context sequence to obtain a second text encoding sequence containing context information; S5.4: Fuse the features of the second text encoding sequence and the label representation sequence to obtain a label feature sequence; S6: Reconstruct the input encoding of the binary decoder using the label feature sequence and the label representation sequence, input the input encoding into the binary decoder for decoding, and the binary decoder outputs a violent event detection result, and its expression is as follows: , ,..., } ) Among them, , ,..., } contains n the prediction results of n labels, where is a positive integer, represents the binary reconstruction input sequence, i is the detection result of the th violent event label, sigmoid represents converting the numerical range of the output of the fully connected layer into probability using and taking the maximum value of the prediction.

2. The method for detecting violent incidents according to claim 1, characterized in that, The specific steps of S2 include: S2.1: Textualize the violent event labels in the label dataset to obtain a label text sequence; S2.2: Reconstruct the label text sequence into label prompt text of a natural language question.

3. The method for detecting violent incidents according to claim 1, characterized in that, In S4, use a trained BERT model to encode the reconstructed input text, and the BERT model outputs a first text encoding sequence.

4. The method for detecting violent incidents according to claim 3, characterized in that, The objective function of the BERT model L is expressed as follows: Among them, represents the loss function of the positive sample, represents the loss function of the negative sample, is the offset probability, represents the focusing function, represents the predicted probability output by the model, is the positive focusing parameter, is the negative focusing parameter.

5. The method for detecting violent incidents according to claim 1, characterized in that, In S5.4, the specific steps include: Perform a dot product operation on the second text encoding sequence and the label representation sequence to obtain the relationship matrix between the labels and the text , and its expression is as follows: Among them, represents a second text coding sequence, is the transpose of the tag sequence ; Using a convolutional neural network to process the relationship matrix between the labels and the text to perform feature learning and obtain a label feature sequence representing the features of each label , and its expression is as follows: Among them, the function represents the ReLU activation and the max pooling operation, indicating that the feature vector obtained after the ReLU activation and the max pooling operation is re-activated using the function.

6. The method for detecting violent incidents according to claim 1, characterized in that, In S6, reconstruct a binary reconstruction input sequence using the label feature sequence and the label representation sequence, specifically including: Multiply the label feature sequence and the label representation sequence to obtain a sequence containing interaction information between labels, and its expression is as follows: Among them, is a sequence containing interaction information between tags, is a tag representation sequence, is a tag feature sequence; Add the sequence containing interaction information between labels and the label representation sequence to obtain a binary reconstruction input sequence, and its expression is as follows: Among them, is a binary reconstruction input sequence.

7. A violent incident detection device, applied to the violent incident detection method according to any one of claims 1 to 6, characterized in that, Including: An acquisition module for acquiring a label dataset; the label dataset contains violent event labels; A preprocessing module for preprocessing the label dataset to obtain label prompt text; A concatenation module for concatenating the text to be detected with the label prompt text to obtain a reconstructed input text; An encoding module for encoding the reconstructed input text to obtain a first text encoding sequence; An extraction module for extracting the text representation sequence and the label representation sequence of the reconstructed input text from the first text encoding sequence; A feature extraction module for extracting features from the text representation sequence and the label representation sequence to obtain a label feature sequence; A reconstruction module, configured to reconstruct an input encoding of a binary decoder by using the tag feature sequence and the tag representation sequence; A detection module, configured to input the input encoding into the binary decoder for decoding processing, and the binary decoder outputs a violent event detection result.

8. An electronic device, characterized in that, It includes a memory and a processor, and a computer program is stored on the memory and can run on the processor. When the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Video decoder with enhanced cabac decoding

    CN103959782A

  • Abstraction of text summarizaton

    US20190362020A1