Knowledge enhancement abnormal behavior detection method and device fusing rules and action fragments

By integrating rule knowledge and action clip examples, cross-modal matching and timing behavior analysis are carried out, which solves the problems of insufficient accuracy and poor adaptability of abnormal behavior detection in the prior art, and achieves high accuracy and robust abnormal behavior detection effects.

CN120032289APending Publication Date: 2025-05-23GOSUNCN TECH GRP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510006464.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy in video action recognition and abnormal behavior detection, poor adaptability to emerging abnormal behavior patterns, insufficient interpretability, and insufficient cross-modal fusion.

Method used

By fusing rule knowledge and action clip examples, vector features of real-time video, rules and action clips are extracted, cross-modal matching is performed, vector feature timing sequences with rules and action clip enhancement are generated, and timing behavior analysis is performed to achieve abnormal behavior detection.

Benefits of technology

It significantly improves the accuracy and robustness of abnormal behavior detection, enhances the adaptability and generalization capabilities of the system, improves the interpretability of the model, realizes deep cross-modal fusion, and improves the real-time performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032289A_ABST
    Figure CN120032289A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video monitoring and abnormal behavior detection, in particular to a knowledge enhancement abnormal behavior detection method and device fusing rules and action fragments, and the method comprises the steps: obtaining real-time video data; obtaining rule knowledge base data; action fragment knowledge base data are obtained, wherein the action fragment knowledge base data comprise action positive examples and interference examples; extracting real-time video vector features based on the real-time video data; extracting rule vector features and action vector features; performing cross-modal matching to obtain a matching result; based on a matching result, generating a vector feature time sequence with rules and action fragment enhancement; based on the vector feature time sequence with rules and action fragment enhancement, time sequence behavior analysis is carried out, and an abnormal behavior detection result is obtained; and outputting an abnormal behavior detection result, and through a knowledge enhancement technology fusing rules and action fragments, the method can more accurately identify abnormal behaviors in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video surveillance and abnormal behavior detection, and in particular to a method and device for abnormal behavior detection enhanced by knowledge of fusion rules and action clips. Background Art

[0002] In the field of video action recognition and behavior analysis, existing technologies have made significant progress, especially in multimodal data processing and behavior recognition algorithms. However, current technical solutions still have some limitations and cannot fully meet the high requirements for abnormal behavior detection in practical applications.

[0003] For example, the video action recognition method based on the multimodal large model CLIP can improve the performance and model versatility of video action recognition by extracting language and video information, using word vector models and image block segmentation technology, and combining temporal difference attention modules. This method achieves the recognition of video actions in complex scenes by fine-tuning the image encoder and text encoder of the CLIP model and calculating the similarity score between video features and category encoding features. However, this method mainly focuses on general action recognition and has limited ability to detect abnormal behavior.

[0004] Another technology provides a behavior recognition method that enhances the readability and interpretability of the behavior recognition model by building a scene graph and combining text, image, and optical flow recognition results, thereby improving the robustness of behavior recognition. This method uses sample videos to train the behavior recognition model and achieves accurate recognition of behaviors in image sequences. However, this method mainly targets predefined behavior categories and is insufficient for detecting unknown abnormal behaviors.

[0005] In the field of abnormal human behavior detection, some researchers have divided video scenes into multiple areas and marked the action sets that can be implemented in each area to obtain abnormal behavior judgment rules. This method can determine whether the human body has abnormal behavior by analyzing the movement trajectory and action sequence of the human body, and identify specific abnormal behaviors, thereby improving the recognition rate. However, this method relies too much on predefined areas and action sets, and is difficult to adapt to complex and changing actual scenarios.

[0006] Please refer to Figure 7The prior art closest to the present invention is the multimodal video analysis method disclosed in CN116385939A. The method includes real-time acquisition of information to be analyzed and extraction of information features, judging the type of information features, inputting a text feature encoder for feature extraction if it is text information, and inputting a picture feature encoder for feature extraction if it is picture information, and then performing similarity calculation and matching the video database content. This method can realize cross-modal retrieval of video data, but it has limitations in abnormal behavior detection.

[0007] The above-mentioned prior art solutions have the following main problems:

[0008] 1. Lack of specialized processing for abnormal behaviors makes it difficult to accurately identify abnormal behaviors in complex scenarios.

[0009] 2. It relies heavily on predefined behavior categories and has difficulty adapting to emerging abnormal behavior patterns.

[0010] 3. Insufficient interpretability, making it difficult to provide clear explanations for abnormal behavior detection results.

[0011] 4. Limited ability to update knowledge, making it difficult to adapt to ever-changing abnormal behavior patterns.

[0012] 5. Cross-modal fusion is not deep enough, making it difficult to fully utilize information from multiple modalities. Summary of the invention

[0013] The present invention aims to solve the above technical problems and provide a method and device for abnormal behavior detection by integrating rule knowledge and action fragments. The method significantly improves the accuracy and robustness of abnormal behavior detection by integrating rule knowledge and action fragment samples. At the same time, its dynamic update mechanism ensures that the system can continuously adapt to new abnormal behavior patterns.

[0014] The invention provides a method and device for enhancing abnormal behavior detection by protecting fusion rules and action fragments, including:

[0015] The acquisition steps include:

[0016] Get real-time video data;

[0017] Acquire rule knowledge base data, wherein the rule knowledge base data includes forward rules and reverse rules;

[0018] Acquire action segment knowledge base data, wherein the action segment knowledge base data includes action positive samples and interference samples;

[0019] Processing steps include:

[0020] Based on the real-time video data, extracting real-time video vector features;

[0021] Based on the rule knowledge base data and the action fragment knowledge base data, respectively extracting rule vector features and action vector features;

[0022] Perform cross-modal matching according to the real-time video vector feature, the rule vector feature and the action vector feature to obtain a matching result;

[0023] Based on the matching results, generating a vector feature time series sequence with rule and action segment enhancement;

[0024] Output steps include:

[0025] Based on the vector feature time series sequence with rule and action fragment enhancement, time series behavior analysis is performed to obtain abnormal behavior detection results;

[0026] The abnormal behavior detection result is output.

[0027] Preferably, the acquiring of real-time video data specifically includes:

[0028] Extract one frame of image from real-time video every second;

[0029] Every nine frames of images are combined to form a combined image.

[0030] Preferably, the extracting of real-time video vector features specifically includes:

[0031] Inputting the combined image into a Transformer-based image encoder;

[0032] The real-time video vector features are extracted through the image encoder.

[0033] Preferably, the extracting rule vector features specifically includes:

[0034] Inputting the text rules in the rule knowledge base data into a Transformer-based text encoder;

[0035] The rule vector feature is extracted by the text encoder.

[0036] Preferably, the extracting motion vector features specifically includes:

[0037] Inputting the action clip images in the action clip knowledge base data into a Transformer-based image encoder;

[0038] The motion vector feature is extracted by the image encoder.

[0039] Preferably, the cross-modal matching specifically includes:

[0040] Calculating the cosine similarity between the real-time video vector feature and the regular vector feature;

[0041] Calculating the cosine similarity between the real-time video vector feature and the action vector feature;

[0042] The cosine similarity is screened according to a preset threshold to obtain a matching result.

[0043] Preferably, the generating of the vector feature time series sequence with rule and action segment enhancement specifically comprises:

[0044] Selecting the forward rule vector feature and the reverse rule vector feature with the highest similarity to the real-time video vector feature;

[0045] Selecting the action positive sample vector feature and the interference sample vector feature that have the highest similarity to the real-time video vector feature;

[0046] The real-time video vector feature, the forward rule vector feature, the reverse rule vector feature, the action positive sample vector feature and the interference sample vector feature are combined to form an enhanced vector feature;

[0047] Organizing a plurality of the enhanced vector features in chronological order to form a temporal sequence of vector features enhanced with rules and action segments.

[0048] Preferably, the performing of the temporal behavior analysis specifically includes:

[0049] Inputting the vector feature time series sequence with rule and action segment enhancement into a Transformer-based feature encoder;

[0050] Encoding the vector feature time series sequence by the feature encoder;

[0051] Input the encoded features into the multi-layer perceptron head module;

[0052] The abnormal behavior detection result is obtained through the multi-layer perceptron head module.

[0053] As a preferred embodiment, the knowledge base updating step is also included:

[0054] Receive new rule definitions or action snippet samples;

[0055] Adding the new rule definition to the rule knowledge base, or adding the new action fragment sample to the action fragment knowledge base;

[0056] The rule knowledge base data or the action fragment knowledge base data is updated.

[0057] The device for detecting abnormal behaviors enhanced by the knowledge of fusion rules and action fragments of the method includes:

[0058] Get modules for:

[0059] Get real-time video data;

[0060] Acquire rule knowledge base data, wherein the rule knowledge base data includes forward rules and reverse rules;

[0061] Acquire action segment knowledge base data, wherein the action segment knowledge base data includes action positive samples and interference samples;

[0062] Processing modules for:

[0063] Based on the real-time video data, extracting real-time video vector features;

[0064] Based on the rule knowledge base data and the action fragment knowledge base data, respectively extracting rule vector features and action vector features;

[0065] Perform cross-modal matching according to the real-time video vector feature, the rule vector feature and the action vector feature to obtain a matching result;

[0066] Based on the matching results, generating a vector feature time series sequence with rule and action segment enhancement;

[0067] Output modules for:

[0068] Based on the vector feature time series sequence with rule and action fragment enhancement, time series behavior analysis is performed to obtain abnormal behavior detection results;

[0069] The abnormal behavior detection result is output.

[0070] The beneficial effects of the present invention are:

[0071] (1) Improved accuracy of abnormal behavior detection. By integrating rules and action fragments into knowledge enhancement technology, this method can more accurately identify abnormal behaviors in complex scenarios.

[0072] (2) Enhanced system adaptability and generalization capabilities. The dynamic knowledge base update mechanism enables the system to continuously learn and adapt to new abnormal behavior patterns.

[0073] (3) Improved the interpretability of the model. By introducing the rule knowledge base and action fragment knowledge base, this method provides a clear explanation basis for abnormal behavior detection results.

[0074] (4) Deep cross-modal fusion is achieved. This method deeply integrates rule knowledge, action examples and real-time video data, making full use of multimodal information.

[0075] (5) Improved real-time performance of the system. Through efficient feature extraction and matching algorithms, this method can quickly detect abnormal behaviors in real-time video streams.

[0076] (6) Enhanced system robustness. By combining rule knowledge and action examples, this method can better handle noise and incomplete data.

[0077] (7) It provides a flexible abnormal behavior definition mechanism. System administrators can easily add or modify rules and action samples to adapt to different application scenarios.

[0078] In summary, the method and device provided by the present invention effectively solve the problems existing in the existing abnormal behavior detection technology through innovative knowledge fusion and dynamic update mechanism, and provide an efficient, accurate and explainable abnormal behavior detection solution for practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 The figure is a complete flow chart of the method of the present invention.

[0080] Figure 2 A schematic diagram of constructing the rule knowledge base of the present invention.

[0081] Figure 3 The schematic diagram of constructing the action segment knowledge base of the present invention is shown in FIG.

[0082] Figure 4 It is a schematic diagram of real-time video vector feature extraction of the present invention.

[0083] Figure 5 Schematic diagram of cross-modal vector feature matching of the present invention.

[0084] Figure 6 It is a schematic diagram of cross-modal knowledge base retrieval of the present invention.

[0085] Figure 7 The following is a working principle diagram for comparison documents. DETAILED DESCRIPTION

[0086] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in each embodiment of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0087] Please refer to Figure 1-6The present invention provides a method and device for detecting abnormal behavior by integrating rules and action fragments. The method mainly includes an acquisition step, a processing step and an output step.

[0088] In the acquisition step, real-time video data is first acquired. These real-time video data usually come from surveillance cameras installed in specific areas, such as shopping malls, subway stations, or public squares. At the same time, this method also acquires rule knowledge base data and action fragment knowledge base data. The rule knowledge base data includes forward rules and reverse rules, while the action fragment knowledge base data includes action positive samples and interference samples.

[0089] The positive rules in the rule knowledge base define behavioral features that may be considered abnormal. For example, for detecting fighting behavior, a positive rule may be two or more people approaching quickly and making aggressive movements. The reverse rule defines behaviors that are easily misjudged as abnormal but are actually normal. For example, physical contact between athletes in a game may be misjudged as fighting, so it can be used as a reverse rule. Taking fighting as an example, the positive rules clearly point out what kind of behavior can be considered abnormal, such as physical attack: any party attacks the other party by using fists, kicks or other body parts, resulting in or likely to result in physical injury. In contrast, the reverse rule aims to clarify situations that are prone to misjudgment, such as animal behavior: fighting between animals does not involve human behavior.

[0090] The positive examples in the action fragment knowledge base are typical action sequences of known abnormal behaviors. For example, fighting behavior may include action fragments such as punching and kicking. Interference examples are normal action sequences that are easily confused with abnormal behaviors. For example, a warm hug may be similar to the action of fighting, but it is a normal behavior.

[0091] These forward and reverse rules are represented by text. In the process of building the rule knowledge base, the defined text rules are extracted with text vector features (text feature embedding) through the Transformer-based text encoder. These text vector features are combined with forward or reverse tags to form rule vector features (rule feature embedding). In the rule knowledge base, each algorithm contains multiple forward and reverse rule vector features. Among them, the forward rule is marked as 0 and the reverse rule is marked as 1.

[0092] By adjusting the forward and reverse rules of the corresponding algorithm in the knowledge base, the accuracy of the algorithm can be improved, or the algorithm definition can be modified. Similarly, new rules can be redefined and stored in the database to implement algorithm definition and identification of new abnormal behaviors.

[0093] In the processing step, the method first extracts real-time video vector features based on real-time video data. Preferably, a Transformer-based image encoder is used to accomplish this task. The Transformer model performs well in processing sequence data and can effectively capture the spatiotemporal relationship between video frames.

[0094] Next, the method extracts rule vector features and action vector features based on the rule knowledge base data and the action segment knowledge base data, respectively. For the rule vector features, a text encoder based on Transformer is used in one embodiment of the present invention to process the textual rules. For the action vector features, an image encoder that is the same as that used to process real-time video data is used.

[0095] Please refer to Figure 2 , build an action fragment knowledge base, refer to action samples and interference samples, improve the accuracy of abnormal behavior recognition, and enhance the applicability and migration ability in diverse scenarios. The action fragment knowledge base contains action samples and interference samples that represent abnormal behaviors.

[0096] These action clips are represented by images composed of multiple video frames representing key actions. In the process of building the action clip knowledge base, the combined image is passed through the Transformer-based image encoder to extract image vector features (image feature embedding). These image vector features are combined with action or interference tags to form action vector features (action feature embedding). In the action clip knowledge base, each algorithm contains action vector features corresponding to multiple action positive samples and interference samples. Among them, the action positive sample is marked as 0, and the interference sample is marked as 1.

[0097] By adjusting the positive examples and interference examples of the corresponding algorithm in the knowledge base, the accuracy of the algorithm can be improved, or the algorithm definition can be modified. Similarly, new action fragments can be re-entered to realize the algorithm definition and identification of new abnormal behaviors.

[0098] Please refer to Figure 3 , build an action fragment knowledge base, refer to action samples and interference samples, improve the accuracy of abnormal behavior recognition, and enhance the applicability and migration ability in diverse scenarios. The action fragment knowledge base contains action samples and interference samples that represent abnormal behaviors.

[0099] These action clips are represented by images composed of multiple video frames representing key actions. In the process of building the action clip knowledge base, the combined image is passed through the Transformer-based image encoder to extract image vector features (image feature embedding). These image vector features are combined with action or interference tags to form action vector features (action feature embedding). In the action clip knowledge base, each algorithm contains action vector features corresponding to multiple action positive samples and interference samples. Among them, the action positive sample is marked as 0, and the interference sample is marked as 1.

[0100] By adjusting the positive examples and interference examples of the corresponding algorithm in the knowledge base, the accuracy of the algorithm can be improved, or the algorithm definition can be modified. Similarly, new action fragments can be re-entered to realize the algorithm definition and identification of new abnormal behaviors.

[0101] Subsequently, the method performs cross-modal matching based on the real-time video vector features, rule vector features, and action vector features to obtain matching results. This step usually uses cosine similarity to calculate the similarity between different feature vectors. Preferably, a similarity threshold is set, such as 0.7 or 0.8, and only features whose similarity exceeds the threshold are retained. The selection of this threshold requires a trade-off between accuracy and recall, and the optimal value can be determined through experimental data.

[0102] Please refer to Figure 4 , the real-time video is extracted at one frame per second, and every 9 frames are combined to form a composite image. The composite image is extracted with the image encoder based on Transformer (image feature embedding). In this way, the real-time video will convert multiple image vector features with time sequence information to participate in abnormal behavior detection.

[0103] Based on the matching results, this method generates a time-series sequence of vector features with rule and action segment enhancement. This sequence combines real-time video features, the most similar rule features, and action segment features, thereby injecting prior knowledge on the basis of the original video features and enhancing the ability to characterize abnormal behaviors.

[0104] Please refer to Figure 5 , the rule vector feature / action vector feature removes the mark bit, and calculates the cosine similarity with the real-time video vector feature to obtain the similarity. At the same time, the current marking situation is determined by the value of the mark bit. That is, the rule matching result outputs the similarity and the forward rule or reverse rule mark; the action matching result outputs the similarity and the action positive sample or interference sample mark.

[0105] Please refer to Figure 6 The method of the present invention also includes cross-modal knowledge base retrieval, specifically including rule knowledge base retrieval and fragment knowledge base retrieval.

[0106] Rule knowledge base retrieval: traverse the rule knowledge base, perform cross-modal vector feature matching between the real-time video vector feature and each rule vector feature of the abnormal behavior type to be detected to obtain similarity. If it is greater than or equal to the set threshold, it will be retained, otherwise it will be discarded. Finally, take a forward rule vector feature and a reverse rule vector feature with the highest similarity to the current real-time video.

[0107] Action clip knowledge base retrieval: traverse the action clip knowledge base, perform cross-modal vector feature matching between the real-time video vector feature and each action vector feature of the abnormal behavior type to be detected to obtain similarity, retain if it is greater than or equal to the set threshold, otherwise discard. Finally, take the action positive sample vector feature and the interference sample vector feature with the highest similarity to the current real-time video.

[0108] In the output step, this method performs temporal behavior analysis based on the temporal sequence of vector features enhanced with rules and action fragments to obtain abnormal behavior detection results. This step usually uses another Transformer model as a feature encoder, followed by a multi-layer perceptron (MLP) head to complete the final classification task. Finally, this method outputs the abnormal behavior detection result, which can be a binary classification result (abnormal / normal) or a multi-classification result (normal, fighting, robbery, etc.).

[0109] A specific embodiment of the method of the present invention is that when acquiring real-time video data, one frame of image is extracted from the real-time video per second. This frame extraction frequency can well balance computational efficiency and information retention in most cases. Every nine frames of image are combined to form a combined image. The selection of nine frames as the combination unit is based on an empirical value, which can capture a sufficiently long time span to cover the characteristics of most abnormal behaviors without causing information redundancy or excessive computational burden.

[0110] When extracting real-time video vector features, the method inputs the combined image into a Transformer-based image encoder. This encoder can effectively process image sequences and capture the spatiotemporal relationship between frames. Through the image encoder, the method extracts real-time video vector features, which contain key spatiotemporal information in the video, laying the foundation for subsequent abnormal behavior detection. In the method of the present invention, extracting rule vector features is a key step. Specifically, the method inputs the text rules in the rule knowledge base data into a Transformer-based text encoder. This encoder is particularly suitable for processing natural language text and can capture semantic information and contextual relationships in the rules. Through the text encoder, the method extracts rule vector features, which are the representation of rule knowledge in the vector space, providing a basis for subsequent cross-modal matching.

[0111] Preferably, in one embodiment of the present invention, the text encoder adopts a BERT (Bidirectional Encoder Representations from Transformers) model or a variant thereof. The BERT model performs well in natural language processing tasks and can generate high-quality text representations. For regular text, the BERT model can well understand the semantic information therein, and even complex conditional statements or professional terms can be accurately encoded.

[0112] When extracting action vector features, this method adopts the same strategy as that used to process real-time video data. Specifically, the action clip images in the action clip knowledge base data are input into the Transformer-based image encoder. This consistency not only simplifies the system architecture, but also ensures that the real-time video features and action clip features are represented in the same feature space, which is beneficial for subsequent similarity calculation and matching.

[0113] The method of the present invention adopts an efficient and effective strategy when performing cross-modal matching. First, the cosine similarity between the real-time video vector feature and the regular vector feature is calculated. At the same time, the cosine similarity between the real-time video vector feature and the action vector feature is also calculated. Cosine similarity is an effective indicator to measure the degree of similarity between the directions of two vectors, and its value range is between -1 and 1, 1 means exactly the same, 0 means orthogonal, and -1 means opposite directions.

[0114] Preferably, in one embodiment of the present invention, a similarity threshold is set, such as 0.75. Only features whose similarity exceeds this threshold will be retained for subsequent processing. The selection of this threshold requires a trade-off between precision and recall. If the threshold is set too high, some potential abnormal behaviors may be missed; if it is set too low, noise may be introduced and the false alarm rate may increase. Through a large amount of experimental data, the value of 0.75 can achieve a good balance in most scenarios.

[0115] This method adopts an innovative fusion strategy when generating a time series of vector features with rule and action segment enhancement. First, the forward rule vector features and reverse rule vector features with the highest similarity to the real-time video vector features are selected. This ensures that the selected rules are the most relevant and can effectively enhance or suppress specific behavioral features.

[0116] At the same time, this method also selects the action positive sample vector features and interference sample vector features with the highest similarity to the real-time video vector features. These action samples provide more fine-grained behavior pattern information and can capture subtle action details that may be overlooked by the rules.

[0117] In a preferred embodiment of the present invention, real-time video vector features, forward rule vector features, reverse rule vector features, action positive sample vector features and interference sample vector features are merged to form enhanced vector features. This merging operation can use simple splicing or more complex fusion methods such as attention mechanisms or gating units. Regardless of the method used, this fusion can inject rich prior knowledge while retaining the original video information, thereby enhancing the ability to characterize abnormal behavior.

[0118] Finally, this method organizes multiple enhanced vector features in chronological order to form a temporal sequence of vector features with rule and action segment enhancements. This sequence not only contains the temporal information of the original video, but also integrates relevant rule knowledge and action patterns, providing a rich and comprehensive input for subsequent temporal behavior analysis.

[0119] In this way, the method of the present invention effectively combines rule knowledge and action samples with real-time video data, greatly enhancing the accuracy and robustness of abnormal behavior detection. This fusion can not only capture obvious abnormal behaviors, but also identify those subtle and easily overlooked abnormal patterns, thus playing an important role in practical applications. In the method of the present invention, temporal behavior analysis is a crucial step, which directly affects the final abnormal behavior detection results. Specifically, the method inputs the vector feature temporal sequence with rule and action fragment enhancement into a Transformer-based feature encoder. The selection of this encoder is based on its excellent performance in processing sequence data, especially in capturing long-distance dependencies.

[0120] Preferably, in one embodiment of the present invention, the feature encoder uses a multi-head self-attention mechanism. This mechanism allows the model to focus on different positions in the sequence at the same time, so as to better understand complex temporal patterns. For example, when detecting fighting behavior, the model can focus on the actions of multiple characters at the same time, and how these actions change over time.

[0121] After encoding the vector feature time series through the feature encoder, this method inputs the encoded features into the multi-layer perceptron head module. This module usually contains several fully connected layers, each followed by a nonlinear activation function such as ReLU. The role of the multi-layer perceptron is to further extract high-level features and ultimately map these features to abnormal behavior categories.

[0122] In a preferred embodiment of the present invention, the last layer of the multi-layer perceptron head module uses a softmax activation function to output the probability distribution of each abnormal behavior category. This design can not only give the most likely abnormal behavior type, but also provide a confidence estimate, providing more information for subsequent decision-making.

[0123] This method obtains abnormal behavior detection results through the multi-layer perceptron head module. This result can be a binary classification (normal / abnormal) or a multi-classification result, such as normal, fighting, robbery, loitering and other specific abnormal behavior types. The specific classification method can be adjusted according to the needs of the actual application scenario.

[0124] It is worth noting that the method of the present invention also includes an important knowledge base update step, which greatly improves the adaptability and scalability of the system. First, the method can receive new rule definitions or action fragment samples. This new knowledge may come from the definition of experts or may be automatically extracted from new data through machine learning algorithms.

[0125] After receiving new knowledge, the method will add new rule definitions to the rule knowledge base, or add new action fragment samples to the action fragment knowledge base. This process may involve knowledge verification and integration to ensure that the newly added knowledge is consistent with the existing knowledge base.

[0126] Preferably, in one embodiment of the present invention, a conflict check is performed when new knowledge is added. For example, if a new rule conflicts with an existing rule, the system will automatically mark it and remind the administrator to conduct a manual review. This mechanism can effectively prevent inconsistent or erroneous information from appearing in the knowledge base.

[0127] Updating the rule knowledge base data or action fragment knowledge base data is the key to maintaining the advancement of the method. By continuously updating and expanding the knowledge base, the method can adapt to new abnormal behavior patterns and improve the ability to handle unknown situations. This dynamic update mechanism enables the method of the present invention to remain efficient and accurate in long-term operation.

[0128] Finally, the present invention also provides a knowledge-enhanced abnormal behavior detection device corresponding to the above method by integrating rules and action segments. The device includes an acquisition module, a processing module and an output module, which correspond to the acquisition step, the processing step and the output step in the method respectively.

[0129] The acquisition module is responsible for acquiring real-time video data, rule knowledge base data, and action segment knowledge base data. This module may contain multiple submodules, such as video acquisition submodule, knowledge base interface submodule, etc. The processing module is the core of the entire device, responsible for extracting various features, performing cross-modal matching, generating enhanced vector feature time series, and other complex operations. The output module is responsible for the final time series behavior analysis and result output.

[0130] Through this modular design, the device of the present invention has good scalability and maintainability. Each module can be upgraded or replaced independently without affecting the operation of the overall system. In addition, this design is also easy to implement on different hardware platforms, such as deploying a computationally intensive processing module on a GPU server, and deploying a lightweight acquisition module and output module on an edge device.

[0131] In general, the method and device provided by the present invention significantly improve the accuracy and robustness of abnormal behavior detection by integrating rule knowledge and action fragment samples. At the same time, its dynamic update mechanism ensures that the system can continuously adapt to new abnormal behavior patterns, which has important value in practical applications.

[0132] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. The knowledge-enhanced abnormal behavior detection method integrating rules and action fragments is characterized by: include: The acquisition steps include: Get real-time video data; Acquire rule knowledge base data, wherein the rule knowledge base data includes forward rules and reverse rules; acquire action segment knowledge base data, wherein the action segment knowledge base data includes action positive samples and interference samples; Processing steps include: Based on the real-time video data, extracting real-time video vector features; Based on the rule knowledge base data and the action fragment knowledge base data, respectively extracting rule vector features and action vector features; Perform cross-modal matching according to the real-time video vector feature, the rule vector feature and the action vector feature to obtain a matching result; Based on the matching results, generating a vector feature time series sequence with rule and action segment enhancement; Output steps include: Based on the vector feature time series sequence with rule and action fragment enhancement, time series behavior analysis is performed to obtain abnormal behavior detection results; The abnormal behavior detection result is output.

2. The method according to claim 1, characterized in that The obtaining of real-time video data specifically includes: Extract one frame of image from real-time video every second; Every nine frames of images are combined to form a combined image.

3. The method according to claim 1, characterized in that The extracting of real-time video vector features specifically includes: Inputting the combined image into a Transformer-based image encoder; The real-time video vector features are extracted through the image encoder.

4. The method according to claim 1, characterized in that: The extraction rule vector feature specifically includes: Inputting the text rules in the rule knowledge base data into a Transformer-based text encoder; The rule vector feature is extracted by the text encoder.

5. The method according to claim 1, characterized in that The extracting of motion vector features specifically includes: Inputting the action clip images in the action clip knowledge base data into a Transformer-based image encoder; The motion vector feature is extracted by the image encoder.

6. The method according to claim 1, characterized in that The cross-modal matching specifically includes: Calculate the cosine similarity between the real-time video vector feature and the rule vector feature; calculate the cosine similarity between the real-time video vector feature and the action vector feature; filter the cosine similarity according to a preset threshold to obtain a matching result.

7. The method according to claim 1, characterized in that The generating of the vector feature time series sequence with rule and action segment enhancement specifically includes: Selecting the forward rule vector feature and the reverse rule vector feature with the highest similarity to the real-time video vector feature; Selecting the action positive sample vector feature and the interference sample vector feature that have the highest similarity to the real-time video vector feature; The real-time video vector feature, the forward rule vector feature, the reverse rule vector feature, the action positive sample vector feature and the interference sample vector feature are combined to form an enhanced vector feature; Organizing a plurality of the enhanced vector features in chronological order to form a temporal sequence of vector features enhanced with rules and action segments.

8. The method according to claim 1, characterized in that The timing behavior analysis specifically includes: Inputting the vector feature time series sequence with rule and action segment enhancement into a Transformer-based feature encoder; Encoding the vector feature time series sequence by the feature encoder; Input the encoded features into the multi-layer perceptron head module; The abnormal behavior detection result is obtained through the multi-layer perceptron head module.

9. The method according to claim 1, characterized in that: It also includes a knowledge base updating step: receiving new rule definitions or action fragment samples; Adding the new rule definition to the rule knowledge base, or adding the new action fragment sample to the action fragment knowledge base; The rule knowledge base data or the action fragment knowledge base data is updated.

10. A knowledge-enhanced abnormal behavior detection device for implementing the method according to any one of claims 1 to 9, characterized in that: include: Get modules for: Get real-time video data; Acquire rule knowledge base data, wherein the rule knowledge base data includes forward rules and reverse rules; acquire action segment knowledge base data, wherein the action segment knowledge base data includes action positive samples and interference samples; Processing modules for: Based on the real-time video data, extracting real-time video vector features; Based on the rule knowledge base data and the action fragment knowledge base data, respectively extracting rule vector features and action vector features; Perform cross-modal matching according to the real-time video vector feature, the rule vector feature and the action vector feature to obtain a matching result; Based on the matching results, generating a vector feature time series sequence with rule and action segment enhancement; Output modules for: Based on the vector feature time series sequence with rule and action fragment enhancement, time series behavior analysis is performed to obtain abnormal behavior detection results; The abnormal behavior detection result is output.

Citation Information

Patent Citations

  • Video analysis method based on multiple modes

    CN116385939A