Weak supervision abnormal behavior detection method based on multi-scale spatial-temporal characteristics

Through the improved C3D network model, combined with multi-scale space and time modules, the accuracy and efficiency of abnormal behavior detection are enhanced, and the problem of insufficient utilization of multi-scale spatiotemporal features in the prior art is solved, thereby achieving higher inter-class distinction and shorter training time.

CN120339895APending Publication Date: 2025-07-18THE INST OF AUTOMATION HEILONGJIANG ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510206130.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing anomaly behavior detection model fails to fully utilize the multi-scale spatiotemporal characteristics of the target object in the video data, resulting in inadequate detection accuracy and efficiency.

Method used

The improved C3D network model is adopted, combining multi-scale spatial modules and multi-scale time modules to capture more detailed spatial and temporal features of the input data, and through multi-instance learning and transfer learning techniques, the inter-class distinction and training efficiency of the model are enhanced.

Benefits of technology

It improves the accuracy and efficiency of abnormal behavior detection, solves the problem of blurred behavior boundaries in different monitoring scenarios, and achieves higher inter-class distinction and shorter training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339895A_ABST
    Figure CN120339895A_ABST
Patent Text Reader

Abstract

The invention discloses a weak supervision abnormal behavior detection method based on multi-scale spatial-temporal characteristics, and relates to the technical field of abnormal behavior detection, and the detection method comprises the steps: dividing a normal video clip set and an abnormal video clip set, and inputting the normal video clip set and the abnormal video clip set into an improved C3D network model; a multi-scale space module and a multi-scale time module are added to capture more detailed spatial and temporal characteristics of input data; outputting intra-class and inter-class abnormal scores of instances in the two sets through a comparison model, and carrying out regression sorting; in combination with a transfer learning technology, related weights of a pre-trained C3D model are loaded, parameters of the same module part of the network and the model are migrated to the algorithm, video-level weak marking is performed on a training sample, one-to-one mapping of positive and negative packets and labels is realized by adopting the improved C3D network model, and the algorithm is simple and convenient to operate. The inter-class distinction degree between the abnormal behaviors and the normal behaviors is increased, and the problems that the boundaries of the normal behaviors and the abnormal behaviors are fuzzy and uncertain in different monitoring scenes are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of abnormal behavior detection, and particularly relates to a method for weakly supervised abnormal behavior detection based on multi-scale spatio-temporal features. Background Technique

[0002] As one of the core technologies of intelligent video surveillance systems, higher requirements will be put forward for the detection accuracy and efficiency of abnormal behavior detection technology. When analyzing and processing specific abnormal behaviors in videos, it is found that the motion durations of different target objects are different, and the styles and sizes of these objects also vary with the camera position and viewing angle. However, existing abnormal behavior detection models pay more attention to the mining of high-level semantic information, have a low utilization rate of shallow feature information, and do not fully utilize the multi-scale spatio-temporal features of target objects in video data, thus affecting the accuracy and efficiency of abnormal behavior detection algorithms. Summary of the Invention

[0003] To solve the problems mentioned in the above background technique; the purpose of the present invention is to provide a method for weakly supervised abnormal behavior detection based on multi-scale spatio-temporal features.

[0004] A method for weakly supervised abnormal behavior detection based on multi-scale spatio-temporal features of the present invention has the following detection method: First, divide the normal video segment set and the abnormal video segment set and input them into an improved C3D network model, and capture more detailed spatio-temporal features of the input data by adding a multi-scale spatial module and a multi-scale temporal module; then, combine the multi-instance learning strategy and the Ranking loss function, and iterate multiple times to train and optimize the network model. By comparing the abnormal score sizes within and between classes of instances in the two sets output by the model, perform regression ranking; combine the transfer learning technology, load the relevant weights of the pre-trained C3D model, and transfer the parameters of the same module part of the network to the proposed algorithm to further shorten the training time of the network, and finally realize abnormal behavior detection.

[0005] Preferably, the detection method first introduces a multi-scale spatial module based on a 3D convolutional neural network, adds partial shallow feature information on the basis of the original C3D network, and incorporates not only low-level feature representations but also high-level semantic information during the last layer of classification, enabling the model to capture more spatial features at different scales. Secondly, a multi-scale temporal module is added, and 3D spatio-temporal convolutional networks with multiple different temporal scales are respectively used to obtain short-, medium-, and long-term video features and then perform fusion processing to mine the temporal features of the video frame sequence. Then, the normal video clip set and the abnormal video clip set are input into the network model, and the abnormal scores of their respective categories are regressed and sorted. It is ensured that the abnormal scores of abnormal behaviors must be higher than those of normal behaviors, and vice versa, further increasing the inter-class distance between the two. At the same time, a smoothing term and a sparsity term are added to the loss function. Finally, based on existing deep learning theoretical methods, the proposed algorithm uses transfer learning technology to transfer the parameters of the same module part of the original C3D network to the corresponding part of the proposed network, further reducing the training time required for the model.

[0006] Preferably, the multi-scale spatial module fuses different shallow feature maps into the last classification layer by adding 2 branches, enabling it to utilize not only high-level semantic information but also shallow surface feature information during the last layer of classification, and improving the overall performance of the model by fully considering features at different hierarchical stages.

[0007] Preferably, the multi-scale temporal module focuses on the detailed information of different temporal scales of the video to learn the dependency relationships of temporal features between a wider range of video frames, obtaining a more detailed end-to-end network model, and thus realizing abnormal behavior detection.

[0008] Preferably, the multi-scale temporal module consists of 3D convolutional kernels and pooling layers with different temporal depths, and its internal structure uses three convolutional kernels with different temporal lengths to extract short-, medium-, and long-term temporal information of the input video data respectively.

[0009] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0010] Weak video-level labels are applied to the training samples, and an improved C3D network model is used to achieve a one-to-one mapping between positive and negative bags and labels, increasing the inter-class discrimination between abnormal behaviors and normal behaviors, and solving problems such as the blurred and uncertain boundaries between normal behaviors and abnormal behaviors in different monitoring scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] For ease of explanation, the present invention is described in detail by the following specific embodiments and accompanying drawings.

[0012] Figure 1 is a flowchart of the anomaly detection algorithm;

[0013] Figure 2 Schematic diagram of the multi-scale space module;

[0014] Figure 3 Schematic diagram of the multi-scale time module. Detailed implementation manners

[0015] To make the objectives, technical solutions and advantages of the present invention clearer and more explicit, the present invention will be described below through specific embodiments shown in the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. The structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change in the ratio relationship or adjustment of the size, without affecting the effects that the present invention can produce and the objectives that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concept of the present invention.

[0016] Here, it should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less related to the present invention are omitted.

[0017] Such as Figure 1As shown in the figure, the following technical solutions are adopted in this specific implementation manner: In order to ensure that the network model can improve the feature representation ability in the spatial dimension and the temporal modeling ability in the temporal dimension at the same time, a weakly supervised abnormal behavior detection algorithm based on multi-scale spatio-temporal feature modeling is proposed. Based on the 3D convolutional neural network, this algorithm first introduces a multi-scale spatial module, adds some shallow feature information on the basis of the original C3D network, and incorporates not only low-level feature representations but also high-level semantic information during the classification of the last layer, enabling the model to capture more spatial features at different scales. Secondly, a multi-scale temporal module is added. 3D spatio-temporal convolutional networks with multiple different temporal scales are used to obtain short-term, medium-term, and long-term video features respectively and then perform fusion processing to mine the temporal features of the video frame sequence. Then, the normal video clip set and the abnormal video clip set are input into the network model, and the abnormal scores of the two categories are regressed and sorted. It is followed that the abnormal score of abnormal behavior must be higher than that of normal behavior, and vice versa, further increasing the inter-class distance between the two. At the same time, a smoothing term and a sparsity term are added to the loss function, which helps the model converge quickly and avoid overfitting. Finally, based on the existing deep learning theory methods, the proposed algorithm uses transfer learning technology to transfer the parameters of the same module part of the original C3D network to the corresponding part of the proposed network, further reducing the training time required for the model and improving the detection efficiency. The overall framework of the algorithm is shown in Figure 1 the figure.

[0018] First, divide the normal video clip set and the abnormal video clip set and input them into the improved C3D network model. By adding a multi-scale spatial module and a multi-scale temporal module, more detailed spatio-temporal features of the input data are captured. Then, combined with the multi-instance learning strategy and the Ranking loss function, the network model is trained and optimized iteratively for multiple times. By comparing the intra-class and inter-class abnormal score sizes of the instances in the two sets output by the model, regression sorting is performed. Combining transfer learning technology, load the relevant weights of the pre-trained C3D model, and transfer the parameters of the same module part of this network to the proposed algorithm, further shortening the training time of the network, and finally realizing abnormal behavior detection.

[0019] 1. Multi-scale spatial module:

[0020] By simultaneously utilizing the low-level surface information and high-level semantic information of the image, the clarity of the image can be gradually restored, and the detection accuracy and robustness of the model can be improved. Each layer of the original C3D structure only learns higher-level feature information from the feature maps output by the previous layer of the network as the final classification features, without fully combining the low-level feature information extracted earlier in the network. During the entire training process of the C3D model, the feature maps output by each layer of the network are only passed forward once, reducing the reuse rate of the feature maps and affecting the learning efficiency of the model. To address these issues, a multi-scale spatial module is proposed. By adding two branches to fuse different shallow feature maps into the last classification layer, it enables the use of not only high-level semantic information but also shallow surface feature information during the last layer classification, thereby improving the overall performance of the model by fully considering the features at different hierarchical stages. The specific implementation details of the multi-scale spatial module are as Figure 2 shown.

[0021] Based on the original C3D network structure, after the convolutional layer, a 1×1 convolutional network is used for processing. The number of convolutional kernels in this convolutional layer is set to be less than the number of channels of its input feature map to reduce the number of channels of its output feature map, thereby greatly reducing the number of parameters in the fully connected layer at the back of the network; then, an average pooling layer is used for pooling operation after the 1×1 convolutional layer. Since the convolutional operation in the spatio-temporal dimension after the convolutional layer is still in the initial stage and not a large number of convolutional operations have been performed, the feature information about the spatio-temporal dimension is still very rich. Using the max pooling operation will cause a large amount of key feature information to be lost. Therefore, the average pooling operation is used after the convolutional layer to retain more detailed information.

[0022] 2. Multi-scale temporal module:

[0023] In this embodiment, a multi-scale temporal module is added to the basic framework of the C3D structure of the model. By paying attention to the detailed information at different time scales of the video to learn the dependence relationship of the temporal features between a wider range of video frames, a more detailed end-to-end network model is obtained, thereby realizing the detection of abnormal behaviors. The schematic diagram of the multi-scale temporal module is as Figure 3 shown.

[0024] The main framework of the multi-scale time module is mainly composed of 3D convolutional kernels and pooling layers with different time depths. Its internal structure uses three convolutional kernels with different temporal lengths, which are used to extract short, medium, and long temporal information of the input video data respectively. Specifically, the feature vector input to the network model is represented by X, and three intermediate feature vectors {C1, C2, C3} are generated after three 3D convolutional operations with variable temporal lengths. Since these feature vectors {C1, C2, C3} are obtained by convolutional operations using convolutional kernels with different time depths, they have different time depths, but their spatial sizes are the same. To ensure that the output channel number of the model matches the input channel number of the next layer of the network, the feature vectors extracted by the three convolutional layers are concatenated to obtain the output feature X'. Then, a network layer with a convolutional kernel of 1*1*1 is used to perform the channel conversion operation. After each convolutional operation, batch normalization (BN) algorithm is used for normalization processing to make the output value of the model relatively stable. Finally, the ReLU activation function is used. By introducing non-linear factors, the model can fit any function and avoid overfitting problems. During the training process, the network model uses the constructed loss function to iteratively update the model parameters multiple times, and finally obtains a network model with multi-scale spatio-temporal features.

[0025] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention.

[0026] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for weakly supervised abnormal behavior detection based on multi-scale spatio-temporal features, characterized in that: Its detection method is as follows: first, divide the normal video clip set and the abnormal video clip set into the improved C3D network model, and capture more detailed spatiotemporal features of the input data by adding multi-scale spatial modules and multi-scale temporal modules; then, combine the multi-instance learning strategy and Ranking loss function, iterate multiple training to optimize the network model; by comparing the anomaly scores within and between instances in the two sets of model outputs, perform regression sorting; combine the transfer learning technology, load the relevant weights of the pre-trained C3D model, and migrate the parameters of the same module part of the network and the model to the proposed algorithm, further shorten the training time of the network, and finally realize abnormal behavior detection.

2. The method for weakly supervised abnormal behavior detection based on multi-scale spatio-temporal features according to claim 1, characterized in that: The detection method is based on a 3D convolutional neural network. First, a multi-scale spatial module is introduced. On the basis of the original C3D network, some shallow feature information is added. In the last layer of classification, not only low-level feature representations but also high-level semantic information are incorporated, so that the model captures more spatial features of different scales. Secondly, a multi-scale temporal module is added. Multiple 3D spatiotemporal convolutional networks of different time scales are used to obtain short-, medium- and long-term video features respectively, and then fusion processing is performed to mine the temporal features of the video frame sequence. Then, a set of normal video clips and a set of abnormal video clips are input into the network model, and the abnormal scores of the categories to which the two belong are regressed and sorted. The abnormal score of abnormal behavior must be higher than that of normal behavior, and vice versa, so as to further increase the inter-class distance between the two. At the same time, smoothing terms and sparse terms are added to the loss function. Finally, based on the existing deep learning theoretical methods, the proposed algorithm uses transfer learning technology to migrate the parameters of the original C3D network and the same module of the algorithm to the corresponding part of the proposed network, so as to further reduce the training time required for the model.

3. The method for weakly supervised abnormal behavior detection based on multi-scale spatio-temporal features according to claim 2, wherein: The multi-scale spatial module adds two branches to fuse different shallow feature maps to the last classification layer, so that when classifying the last layer, it not only utilizes high-level semantic information but also adds shallow surface feature information, thereby improving the overall performance of the model by fully considering the features of different levels.

4. The method for weakly supervised abnormal behavior detection based on multi-scale spatio-temporal features according to claim 2, wherein: The multi-scale time module obtains a more detailed end-to-end network model by focusing on the detailed information of different time scales of the video to learn the dependencies of temporal features between a wider range of video frames, thereby realizing abnormal behavior detection.

5. The method for weakly supervised abnormal behavior detection based on multi-scale spatio-temporal features according to claim 4, characterized in that: The multi-scale time module consists of 3D convolution kernels and pooling layers of different time depths. Its internal structure uses three convolution kernels of different time series lengths, which are used to extract short, medium and long time series information of input video data respectively.