Method and System for Analyzing Abnormal Actions in Examination Rooms Based on Feature Fusion
By introducing feature fusion and Moformer structures into the action recognition model, combining the pre-training-fine-tuning paradigm and weighted cross-entropy loss function, the problem of limited ability to extract actions continuous information and difficult data set annotation in the prior art is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411821878.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-11
AI Technical Summary
The existing deep learning-based action recognition method has limited ability to extract continuous information of actions, and it is difficult to label video data sets, insufficient data volume leads to poor recognition effect and poor generalization ability.
A method of abnormal motion analysis of examination room based on feature fusion was designed, continuous and discrete features were extracted through the ConFrames model, and feature fusion and further extraction were used using the Moformer structure, combining pre-training-fine-tuning paradigm and weighted cross-entropy loss function to improve the recognition accuracy and robustness of the model.
The ability of the action recognition model to extract continuous information of the action is improved, the recognition accuracy and robustness of the model are enhanced, and the dependence on the annotation and data volume of the video dataset is reduced.
Smart Images

Figure CN119649303B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of action recognition, and particularly to a method and system for analyzing abnormal actions in an examination room based on feature fusion. Background Art
[0002] Action Recognition is an important research field in computer vision, which mainly focuses on enabling a computer to understand actions in a video or image sequence through algorithms. This involves identifying, classifying, and localizing human behaviors in a video. Action recognition technology has a wide range of applications in many fields, such as intelligent monitoring systems, human-computer interaction, virtual reality, sports analysis, medical health, etc.
[0003] Before the rise of deep learning, action recognition mainly relied on manually designed features to describe actions in a video. Manually designed features refer to a series of feature descriptors carefully designed by researchers according to task requirements and prior knowledge, which can capture key information in the video, such as the appearance, shape, texture of objects, and their motion characteristics. In recent years, deep learning methods have made remarkable progress in the field of video action recognition, improving the performance and practicality of the system.
[0004] In previous years, a deep learning model, the SlowFast network, which fuses long-term features and short-term features in a video, was proposed in the field of action recognition. By introducing a multi-path mechanism to enhance the model's temporal perception ability, the SlowFast network can effectively capture action features at different time scales in a video, thus achieving remarkable results in the field of action recognition.
[0005] On the other hand, due to its advantages in processing sequential data, the Transformer model has been applied to video processing tasks such as frame synthesis, action recognition, and video retrieval. These tasks usually require processing spatio-temporal dimensional information of a video, and the Transformer model can well adapt to such tasks.
[0006] Although the deep learning-based action recognition method has achieved remarkable results, there are still some deficiencies: the model mainly performs dependent calculations on the spatial and temporal dimensions of video data respectively, and both time and space are calculated based on sampled discrete frames, while actions are continuous. The above approach has limited ability to extract continuous information of actions and ignores global information; it is difficult to annotate video datasets, and insufficient data volume will lead to poor recognition effects and poor generalization ability of the model, etc. Summary of the Invention
[0007] Although previous deep learning-based action recognition methods have achieved remarkable results, there are still some deficiencies. For example, the model has limited ability to extract continuous information of actions, and insufficient video data volume will lead to poor actual recognition effect of the model. To address the deficiencies of the existing technology, the present invention provides a method and system for analyzing abnormal actions in an examination room based on feature fusion; the action recognition model of the present invention is also based on deep learning methods, and better results have been achieved in terms of recognition accuracy and robustness.
[0008] On the one hand, a method for analyzing abnormal actions in an examination room based on feature fusion is provided, including:
[0009] Obtain the surveillance video of the examination room to be analyzed;
[0010] Input the surveillance video of the examination room to be analyzed into the trained abnormal action analysis model of the examination room to obtain the analysis result of abnormal actions in the examination room; wherein, the trained abnormal action analysis model of the examination room is used to extract primary features from the surveillance video of the examination room to be analyzed, extract environmental features and action features from the primary features respectively, fuse the environmental features and action features to obtain fused features, extract features from the fused features to obtain final features, and classify and recognize the final features to obtain the analysis result of abnormal actions in the examination room.
[0011] On the other hand, a system for analyzing abnormal actions in an examination room based on feature fusion is provided, including:
[0012] An acquisition module, which is configured to: obtain the surveillance video of the examination room to be analyzed;
[0013] An analysis module, which is configured to: input the surveillance video of the examination room to be analyzed into the trained abnormal action analysis model of the examination room to obtain the analysis result of abnormal actions in the examination room; wherein, the trained abnormal action analysis model of the examination room is used to extract primary features from the surveillance video of the examination room to be analyzed, extract environmental features and action features from the primary features respectively, fuse the environmental features and action features to obtain fused features, extract features from the fused features to obtain final features, and classify and recognize the final features to obtain the analysis result of abnormal actions in the examination room.
[0014] The above technical solution has the following advantages or beneficial effects:
[0015] 1. The present invention innovatively designs a brand-new feature extraction neural network ConFrames, and establishes two channels to collect continuous features and discrete features respectively. The continuous features are used to represent action features, and the discrete features are used to represent environmental features, and then they are fused through a feature fusion module. Next, the fused features are passed through Moformer again to obtain richer video spatio-temporal feature information.
[0016] 2. The present invention innovatively designs a lightweight Transformer structure, Moformer, which has three variants, namely Moformer_E, MoAformer_A, and Moformer, for environmental feature extraction, action feature extraction, and further feature processing respectively. They are both composed of two repeated structures, a corresponding block and a convolutional feed-forward network, but there are differences in details.
[0017] 3. The present invention innovatively designs a feature fusion module, FuseFrames, to fuse environmental features into action features. Since the feature dimensions of the two branches are different, the FuseFrames module first uses convolution to align the dimensions, and then uses a channel attention mechanism to help the model better capture key information, thereby enhancing the feature expression ability and improving the model robustness.
[0018] 4. The present invention incorporates context awareness into the self-attention mechanism. A dual-branch structure is adopted, and context-aware weights are fused in the local branch to aggregate high-frequency local information. For the global branch, ordinary self-attention is used, but K and V are downsampled to reduce FLOPs, which helps the model capture low-frequency global information. Then, the outputs of the local branch and the global branch are fused.
[0019] 5. The present invention uses a pre-training - fine-tuning paradigm. First, it is trained using a public dataset model, and then it is trained using an internal dataset. The parameters with better effects in the network are fine-tuned, achieving good recognition performance and reducing the impact of difficulties in video dataset annotation and insufficient data volume on the model.
[0020] 6. The present invention uses a weighted cross-entropy loss function to replace the ordinary cross-entropy loss function, and changes the weight of the loss function according to the action occurrence frequency in the examination room, improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0022] Figure 1 It is a flowchart of the method for the first embodiment;
[0023] Figure 2 It is the overall architecture of the abnormal action analysis model in the examination room for the first embodiment;
[0024] Figure 3 It is the internal structure diagram of Moformer_E for the first embodiment;
[0025] Figure 4The structural diagram of the MoEBlock of Moformer_E in Embodiment 1;
[0026] Figure 5 The structural diagram of the ConvFFN of Moformer_E in Embodiment 1;
[0027] Figure 6 The internal structural diagram of Moformer_A in Embodiment 1;
[0028] Figure 7 The structural diagram of the MoABlock of Moformer_A in Embodiment 1;
[0029] Figure 8 The structural diagram of the ConvFFN of Moformer_A in Embodiment 1;
[0030] Figure 9 The structural diagram of the feature fusion module in Embodiment 1;
[0031] Figure 10 The structural diagram of Moformer in Embodiment 1;
[0032] Figure 11 The structural diagram of the MoBlock in Embodiment 1. Detailed implementation manners
[0033] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0034] In terms of action recognition, the present invention has many advantages. Compared with the previous action recognition model SlowFast, the SlowFast model captures semantic information in the video through the slow channel, while the fast channel captures motion details. However, it involves multiple hyperparameters (such as frame rate, channel number ratio, etc.) and needs to be carefully adjusted to achieve the best performance. Moreover, the fast channel and the slow channel lack global information. In recent years, the self-attention mechanism has performed excellently in tasks such as video classification, object detection, and behavior recognition. The self-attention mechanism can effectively capture the long-range dependencies between video frames, whether in terms of time or space, helping the model understand dynamic changes and event relationships, allowing the model to adaptively focus on important regions or frames in the video without a fixed convolutional window, which can better handle the diversity of different scenarios and actions. Although self-attention can capture global information, it may ignore some local details, which may lead to information loss in some tasks. How to combine the advantages of the SlowFast model and the self-attention mechanism is an urgent problem to be solved, and the present invention provides a certain solution to this problem. The present invention designs a Transformer structure, Moformer, to replace the channels in SlowFast, then changes the sampling method to better adapt to the examination room environment, and uses the self-attention mechanism to process action features. Subsequently, the present invention designs a feature fusion module, FuseFrames, to effectively fuse the two types of features. The fused features pass through the Moformer structure again to extract features more fully.
[0035] The quality and diversity of the dataset directly affect the learning effect and generalization ability of the model. A high-quality dataset can help the model capture more features and patterns, thereby improving performance. However, this may be difficult to achieve in this task. Because in the examination room monitoring scenario, the action amplitude is small, there are many people, and the frequency of abnormal actions is relatively small, so it is extremely difficult to annotate, and the examination room environment is relatively homogeneous, so it is difficult to ensure the diversity of the dataset. The pre-training-fine-tuning paradigm is an effective model training strategy, especially in video understanding and other deep learning tasks. The model is trained on a large-scale dataset to learn general features and knowledge to improve the performance and generalization ability of the model. Then, the pre-trained model is further trained on the examination room monitoring dataset of the present invention to adjust the model parameters to better meet the requirements of specific tasks. The data volume requirement in this stage is small, and better results can be achieved with limited data.
[0036] Embodiment 1
[0037] As Figure 1 shown, this embodiment provides a method for analyzing abnormal actions in the examination room based on feature fusion, including:
[0038] S101: Obtain the examination room monitoring video to be analyzed;
[0039] S102: Input the examination room surveillance video to be analyzed into the trained analysis model for abnormal actions in the examination room to obtain the analysis result of abnormal actions in the examination room. Among them, the trained analysis model for abnormal actions in the examination room is used to extract primary features from the examination room surveillance video to be analyzed, separately extract environmental features and action features from the primary features, fuse the environmental features and action features to obtain fused features, extract features from the fused features to obtain final features, and perform classification and recognition on the final features to obtain the analysis result of abnormal actions in the examination room.
[0040] Further, in S101: Obtain the examination room surveillance video to be analyzed and collect it through a camera.
[0041] Further, the training process of the trained analysis model for abnormal actions in the examination room includes:
[0042] Construct a training set, where the training set is the examination room surveillance video with known abnormal action labels;
[0043] Input the training set into the analysis model for abnormal actions in the examination room to train the model. When the total loss function value of the model no longer decreases, or the number of iterations exceeds the set number of times, stop training to obtain the trained analysis model for abnormal actions in the examination room.
[0044] It should be understood that weighted cross-entropy is an improved version of the cross-entropy loss function, which is mainly used to solve the problem of class imbalance. In the environment of examination room surveillance, the number of samples of various action categories in the dataset is often unbalanced. If not processed, the model may tend to predict those categories with a larger number of samples, such as students doing questions, resulting in poor prediction performance for minority categories, such as picking up things, drinking water, etc. Therefore, the weighted cross-entropy loss function is adopted in the prediction stage. Here, the formula for weighted cross-entropy can be written as:
[0045] L w (p,q) = -∑ i w i q i logp i (3 - 8)
[0046] Here, w i is the weight of the i-th class and is determined based on the class frequency. In the present invention, it is set according to the reciprocal of the class frequency, and the formula can be expressed as:
[0047]
[0048] where N is the total number of samples, and n i is the number of samples of the i-th class.
[0049] Due to the difficulty in annotating video datasets, the cost of the examination hall audio dataset with labels is very high, and the oral English levels of candidates vary greatly. This makes it difficult to train a model with good recognition effect simply using the examination hall dataset. To overcome this problem, the present invention adopts a pre-training - fine-tuning paradigm, and the specific steps are as follows:
[0050] (1) Use the public dataset AVA (V2.2) to pre-train the ConFrames model, and the total duration of the video data is about 100h.
[0051] (2) Use the manually annotated internal examination hall monitoring dataset to fine-tune some parameters of the ConFrames model, and the total duration of the video data is about 1h.
[0052] For the selection of the parameters to be fine-tuned, the present invention fine-tunes the parameters of the channel self-attention calculation module of the feature fusion module. This is because in the examination hall scenario, the proportions of the environment and human actions are relatively fixed. Therefore, the present invention mainly fine-tunes the parameters of this module to achieve better recognition effect on the internal dataset.
[0053] Furthermore, as Figure 2 shown, the trained abnormal action analysis model for the examination hall includes:
[0054] A backbone module, the input end of which is used to input the examination hall monitoring video to be analyzed;
[0055] The output end of the backbone module is respectively connected to the input end of the first branch and the input end of the second branch. The first branch includes: a first downsampling layer and a Moformer_E module connected in sequence. The input end of the first downsampling layer is the input end of the first branch; the output end of the Moformer E module is the output end of the first branch;
[0056] The second branch includes: a Moformer_A module. The input end of the Moformer_A module is the input end of the second branch, and the output end of the Moformer_A module is the output end of the second branch;
[0057] The output ends of the first branch and the second branch are both connected to the input end of the feature fusion module. The output end of the feature fusion module is connected to the input end of the Moformer module, and the output end of the Moformer module outputs the abnormal action classification and recognition result.
[0058] It should be understood that the structure of the action recognition model ConFrames proposed by the present invention is as Figure 2As shown, the input of the model is the position tags of students and teachers and the clipped abnormal action video frames. First, the input video frame images pass through a convolutional backbone as the first layer of input processing. The backbone consists of three convolutional layers with stride values of 2, 2, and 1 respectively. Then it is divided into two groups and sent to two channels respectively, which are the environment channel and the action channel. The environment channel first uses a downsampling module to reduce the number of images, and then enters the Moformer_E module to extract environmental features. At the same time, the action channel uses the Moformer_A module to extract action features. Then, the extracted action features and environmental features are fused, and after fusion, they are input into the Moformer module again for further feature extraction, and finally the output result is used for classification and recognition.
[0059] Further, the backbone module includes: a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in sequence; the backbone module is used to extract primary features from the examination room monitoring video to be analyzed.
[0060] Further, as Figure 3 shown, the Moformer_E module includes:
[0061] A first MoEBlock module, a first ConvFFN module, a second MoEBlock module, and a second ConvFFN module connected in sequence.
[0062] Further, the Moformer_E module is used to extract environmental features from the primary features.
[0063] Further, as Figure 4 shown, the internal structures of the first MoEBlock module and the second MoEBlock module are the same. The first MoEBlock module includes:
[0064] A first normalization module, and the input end of the first normalization module is the input end of the first MoEBlock module;
[0065] The output end of the first normalization module is respectively connected to the input end of the first local branch and the input end of the second global branch;
[0066] The first local branch includes: a first fully-connected layer, the input end of the first fully-connected layer is connected to the output end of the first normalization module, and the output end of the first fully-connected layer is respectively connected to the input ends of a first depthwise separable convolution layer, a second depthwise separable convolution layer, and a third depthwise separable convolution layer; the output ends of the second depthwise separable convolution layer and the third depthwise separable convolution layer are connected to the input end of a first multiplier; the output end of the first multiplier is connected to the input end of a second fully-connected layer, the output end of the second fully-connected layer is connected to the input end of a first activation function layer, the output ends of the first depthwise separable convolution layer and the first activation function layer are both connected to the input end of a second multiplier, and the output end of the second multiplier is the output end of the first local branch;
[0067] The second global branch includes: a first self-attention mechanism layer, the input end of the first self-attention mechanism layer is connected to the output end of the first normalization module, the output end of the first self-attention mechanism layer is connected to the input end of a third fully-connected layer, and the output end of the third fully-connected layer is respectively connected to the input ends of a second downsampling layer, a third downsampling layer, and a third multiplier; the output end of the second downsampling layer is connected to the input end of a fourth multiplier; the output end of the third downsampling layer is connected to the input end of the third multiplier; the output end of the third multiplier is connected to the input end of a second activation function layer; the output end of the second activation function layer is connected to the input end of the fourth multiplier, and the output end of the fourth multiplier is the output end of the second global branch;
[0068] The output ends of the first local branch and the second global branch are both connected to the input end of a first connection unit, the output end of the first connection unit is connected to the input end of a fourth fully-connected layer; the output end of the fourth fully-connected layer is connected to the input end of a first adder, and the input end of the first adder is further connected to the input end of the first normalization module; the output end of the first adder is the output end of the first MoEBlock module.
[0069] Furthermore, the self-attention input transformation layer of the second global branch has the following formula expression:
[0070]
[0071] where X is the input vector representation, W q is the query weight matrix, Q gl is the generated query vector, W k is the key weight matrix, K gl is the generated key vector, W v is the value weight matrix, V gl is the generated value vector matrix.
[0072] First, downsample K and V, and then perform the standard attention calculation process on Q, K, and V to extract low-frequency global information. The extracted information is as follows:
[0073] X gl = Softmax(Q gl · Pool(K gl )) · Pool(V gl ) (3-2)
[0074] where Pool represents the downsampling operation.
[0075] The pattern of the second global branch effectively reduces the number of FLOPs required for attention and also obtains a global receptive field. However, although it can effectively capture low-frequency global information, it has insufficient ability to process high-frequency local information.
[0076] Furthermore, the first local branch, whose formula expression is as follows:
[0077] In the first local branch, first apply a linear transformation to obtain Q, K, and V:
[0078] Q, K, V = FC(X0) (3-3)
[0079] where X0 represents the initial input of the first local branch, and FC represents the fully connected layer;
[0080] After performing the linear transformation, first perform a local feature aggregation process with shared weights on V. Then, based on the processed V and Q, K, perform context-aware local enhancement. For V, use depthwise separable convolution (DWconv) to aggregate local information. The weights of DWconv are globally shared.
[0081] After integrating the local information of V with the shared weights, combine the joint Q and K to generate context-aware weights. Use two DWconv to separately aggregate the local information of Q and K.
[0082] Then, calculate the Hadamard product of Q and K, and perform a series of transformations on the result to obtain context-aware weights. Finally, use the generated weights to enhance the local features. The calculation process is as follows:
[0083]
[0084] where Softmax means using softmax as the activation function so that each element of the output is between [0,1].
[0085] Next, the first connection unit fuses the output of the first local branch and the output of the first global branch. The first connection unit connects the two outputs in the channel dimension, and then processes the data of the connected channel size using the fourth fully-connected layer.
[0086] Using the self-attention mechanism may reduce the efficiency of the model. Therefore, a lightweight transformer model needs to be used to improve the computational efficiency while maintaining the recognition effect. Therefore, the present invention adopts a dual-branch structure, which is divided into a first local branch and a first global branch. In the first local branch, context-aware weights are fused to aggregate high-frequency local information. For the first global branch, ordinary self-attention is used, but K and V are downsampled to reduce FLOPs, which helps the model capture low-frequency global information. Then, the outputs of the first local branch and the first global branch are fused.
[0087] The environmental features reflect the spatial features of the examination room. Extracting environmental features does not require time information. Therefore, it is not necessary to utilize all the input video frames. Downsampling can be performed first. In the past, some ideas were to select one frame every few frames. However, in the scenario of examination room monitoring, the environment is almost unchanged throughout the video. So, even only one frame can be taken, and then it can be used as an image for feature extraction. The performance of image feature extraction is better than that of video. Therefore, the present invention designs a lightweight transformer structure for environmental feature extraction, named Moformer_E. The structure of Moformer_E is as Figure 3 shown. Moformer_E consists of two repeated structures, and each structure is composed of a MoEBlock block and a convolutional feed-forward network (ConvFFN). Next, the structures of the MoEBlock block and the convolutional feed-forward network (ConvFFN) will be introduced respectively. First is the MoEBlock block, and its structure is as Figure 4 shown. Each block is composed of a local branch and a global branch. In the figure, the left branch is the local branch, and the right branch is the global branch. The present invention adds a local branch, which enables the model of the present invention to achieve high performance. It includes some standard attention operations.
[0088] Furthermore, as Figure 5 shown, the internal structures of the first ConvFFN module and the second ConvFFN module are the same. The first ConvFFN module includes:
[0089] A second layer normalization module, a fifth fully-connected layer, a third activation function layer, a fourth depthwise separable convolutional layer, a sixth fully-connected layer, and a second adder connected in sequence;
[0090] The input end of the second normalization module is the input end of the first ConvFFN module. The input end of the second normalization module is connected to the input end of the fifth depthwise separable convolutional layer. The output end of the fifth depthwise separable convolutional layer is connected to the input end of the third normalization module. The output end of the third normalization module is connected to the input end of the seventh fully connected layer. The output end of the seventh fully connected layer is connected to the input end of the second adder. The output end of the second adder is the output end of the first ConvFFN module.
[0091] ConvFFN refers to a module that combines the characteristics of convolutional layers and feedforward neural networks to enhance the model's ability to capture local and global features simultaneously. ConvFFN uses depthwise convolution (DWconv) after GELU activation, which enables ConvFFN to aggregate local information. With the help of DWconv, downsampling is directly performed in ConvFFN.
[0092] Furthermore, the working process of the first ConvFFN module includes:
[0093] The input data first passes through batch normalization and a fully connected layer, and then uses the GELU activation function for non-linear transformation. Then, after GELU activation, depthwise separable convolution (DWconv) is adopted to enable ConvFFN to aggregate local information. Again, it passes through a fully connected layer (FC Layer) for linear transformation.
[0094] Another line represents a skip connection, which is used to solve the problem of vanishing gradients in deep neural networks. In the skip connection, first, feature extraction is performed through a depthwise separable convolution (DW convolution). Then, batch normalization (BN) is performed for regularization. Then, the data enters a fully connected layer (FC Layer) for linear transformation. Among them, DW convolution and the fully connected layer are used to downsample the input and increase the dimension respectively. Finally, the results of the sixth fully connected layer and the seventh fully connected layer are added together to obtain the output of the first ConvFFN module.
[0095] Furthermore, as Figure 6 shown, the Moformer_A module includes:
[0096] The first MoABlock module, the third ConvFFN module, the second MoABlock module, and the fourth ConvFFN module connected in sequence.
[0097] Action features reflect the action changes of the people in the video. Compared with only sampling a small number of frames for environmental features, action features retain consecutive input frames, which can maximize the retention of action information. In the scenario of exam monitoring, the actions of people have the following characteristics: 1. The action amplitude is small; 2. The action duration varies; 3. The size of the people making the actions is relatively small. Therefore, the effect of using traditional methods to extract features is average. The present invention innovatively designs a novel action feature extraction module, named Moformer_A, whose structure is as Figure 6 shown. Moformer_A is also composed of two repeated structures, and each structure consists of a MoABlock block and a convolutional feed-forward network (ConvFFN). Next, the MoABlock block and ConvFFN will be introduced.
[0098] The structure of the MoABlock is as Figure 7 shown. Compared with the MoEBlock, the MoABlock only retains the local branch and abandons the global branch. This is because action features focus more on local features and do not require combining global information. Moreover, more video frames are input for extracting action features, and reducing one branch can also improve the running efficiency.
[0099] Furthermore, as Figure 7 shown, the internal structures of the first MoABlock module and the second MoABlock module are the same. The first MoABlock module includes:
[0100] The eighth fully connected layer, whose input end is the input end of the first MoABlock module; the output end of the eighth fully connected layer is respectively connected to the input ends of the first three-dimensional convolutional layer, the second three-dimensional convolutional layer, and the third three-dimensional convolutional layer;
[0101] The output ends of the second three-dimensional convolutional layer and the third three-dimensional convolutional layer are both connected to the input end of the fifth multiplier. The output end of the fifth multiplier is connected to the input end of the ninth fully connected layer. The output end of the ninth fully connected layer is connected to the input end of the Swish layer. Swish is an activation function in deep learning. Substantially, it is the product of the input x and its sigmoid value, where sigmoid(x) is a common S-shaped function that can map the input value to between 0 and 1. The output end of the Swish layer is connected to the input end of the tenth fully connected layer. The output end of the tenth fully connected layer is connected to the input end of the fourth activation function layer;
[0102] The output ends of the first three-dimensional convolutional layer and the fourth activation function layer are both connected to the input end of the sixth multiplier. The output end of the sixth multiplier is connected to the input end of the eleventh fully-connected layer. The output end of the eleventh fully-connected layer is connected to the input end of the third adder. The output end of the third adder is the output end of the first MoABlock module.
[0103] Furthermore, the first MoABlock module includes:
[0104] The input data also applies linear transformation to obtain Q, K, and V. For Q and K, first, feature extraction is performed through a 3D convolutional layer; then, the data enters a fully-connected layer for linear transformation; then, the Swish activation function is used for non-linear transformation; then, it passes through another fully-connected layer for linear transformation, and the softmax function is used to convert the result into a probability distribution; while V will be multiplied by it after passing through a 3D convolutional module to obtain a new feature representation; finally, a fully-connected layer is used to process the data with the connected channel size to obtain the final output.
[0105] Furthermore, the calculation process of the first MoABlock module is as follows:
[0106]
[0107] Among them, d is the dimension of the query (Q), key (K), and value (V). When d is very large, the calculated attention scores may become too large, resulting in numerical instability. To solve this problem, the present invention divides the attention scores by sqrt(d), that is, scales the attention scores. Compared with the previous method, this method introduces stronger non-linearity. Specifically, previously, in the context-aware weight generation process of local self-attention, the only non-linear operator was Softmax. However, in this module, in addition to Softmax, Swish is also executed. Stronger non-linearity will obtain higher-quality context-aware weights.
[0108] Then the structure of ConvFFN is different from that of the environment extraction module, as Figure 8 shown. First, the depthwise separable convolution in it is changed to 3D convolution, and secondly, all the operations in the skip connection are deleted to speed up the inference speed.
[0109] Furthermore, as Figure 8 shown, the internal structures of the third ConvFFN module and the fourth ConvFFN module are the same. The third ConvFFN module includes:
[0110] The fourth layer normalization module, the twelfth fully connected layer, the fifth activation function layer, the fourth three-dimensional convolutional layer, the thirteenth fully connected layer, and the fourth adder connected in sequence;
[0111] The input end of the fourth adder is also connected to the input end of the fourth layer normalization module;
[0112] The input end of the fourth layer normalization module is the input end of the third ConvFFN module;
[0113] The output end of the fourth adder is the output end of the third ConvFFN module.
[0114] Furthermore, the third ConvFFN module is used to first pass the input data through batch normalization and a fully connected layer, and then use the GELU activation function for non-linear transformation; then, after GELU activation, three-dimensional convolution (3dconv) is adopted to enable ConvFFN to aggregate spatio-temporal information; again, a linear transformation is performed through a fully connected layer (FC Layer);
[0115] Another line represents a skip connection. The skip connection of the third ConvFFN module does not require additional operations, and directly adds the input data and the result of the thirteenth fully connected layer to obtain the output of the third ConvFFN module.
[0116] Furthermore, as Figure 9 shown, the feature fusion module includes:
[0117] The fourth convolutional layer, and the input value of the fourth convolutional layer is the environmental feature;
[0118] The output end of the fourth convolutional layer is connected to the input end of the splicing unit, and the action feature is also input to the input end of the splicing unit;
[0119] The output end of the splicing unit is connected to the input end of the global average pooling layer;
[0120] The output end of the global average pooling layer is connected to the input end of the fourteenth fully connected layer;
[0121] The output end of the fourteenth fully connected layer is connected to the input end of the seventh multiplier;
[0122] The output end of the seventh multiplier is the output end of the feature fusion module.
[0123] Furthermore, after feature extraction through two branches, the feature fusion module respectively obtains an output with a dimension of and an output with a dimension of ; uses a 1×1 convolutional layer to expand the number of channels with a dimension of 1 to Then, the outputs of the two adjusted branches are spliced in the horizontal dimension;
[0124] Next, the SE module will be applied to adjust the fused features. The SE module includes a Squeeze operation layer, a fully connected network, and a feature mapping layer connected in sequence.
[0125] First, the Squeeze operation is performed. The Squeeze operation performs global average pooling on the feature map, compressing the three-dimensional spatial information into a one-dimensional vector. The Squeeze operation compresses all the spatial information in the feature map into a channel-level descriptor:
[0126] Z i = f gap (x i ) (3 - 5)
[0127] where x i is the feature map of the i-th channel, Z i is the global information of this channel, and f gap represents the global average pooling operation.
[0128] A fully connected network is used to generate the weights for each channel. This fully connected network contains a dimensionality reduction layer and a dimensionality increase layer. The former is used to reduce the number of parameters, and the latter restores the number of channels to the original. In this way, the model can learn the interdependencies between different channels and generate a weight vector s accordingly. Each element in this vector corresponds to an importance score for a channel:
[0129] s = W2(ReLU(W1(Z)) (3 - 6)
[0130] The ReLU activation function is used, and W1 and W2 are the weight matrices of the fully connected layers.
[0131] Finally, the generated weight vector s is applied to the original feature map through element-wise multiplication:
[0132] x i ′ = s i · x i (3 - 7)
[0133] Furthermore, as Figure 10 shown, the Moformer module includes:
[0134] The first MoBlock module, the fifth ConvFFN module, the second MoBlock module, and the sixth ConvFFN module connected in sequence.
[0135] Furthermore, as Figure 11As shown, the internal structures of the first MoBlock module and the second MoBlock module are the same. The first MoBlock module includes:
[0136] The fifth layer normalization module, and the input end of the fifth layer normalization module is the input end of the first MoBlock module;
[0137] The output end of the fifth layer normalization module is respectively connected to the input end of the second local branch and the input end of the second global branch;
[0138] The second local branch includes: the fifteenth fully connected layer, the input end of the fifteenth fully connected layer is connected to the output end of the fifth layer normalization module, and the output end of the fifteenth fully connected layer is respectively connected to the input ends of the sixth depthwise separable convolution layer, the seventh depthwise separable convolution layer, and the eighth depthwise separable convolution layer; the output ends of the seventh depthwise separable convolution layer and the eighth depthwise separable convolution layer are connected to the input end of the eighth multiplier; the output end of the eighth multiplier is connected to the input end of the sixteenth fully connected layer, the output end of the sixteenth fully connected layer is connected to the input end of the Swish layer, the output end of the Swish layer is connected to the input end of the seventeenth fully connected layer, the output end of the seventeenth fully connected layer is connected to the input end of the sixth activation function layer, the output ends of the sixth depthwise separable convolution layer and the sixth activation function layer are both connected to the input end of the ninth multiplier, and the output end of the ninth multiplier is the output end of the second local branch;
[0139] The second global branch includes: the second self-attention mechanism layer, the input end of the second self-attention mechanism layer is connected to the output end of the fifth layer normalization module, and the output end of the second self-attention mechanism layer is connected to the input end of the eighteenth fully connected layer, the output end of the eighteenth fully connected layer is respectively connected to the input ends of the fourth downsampling layer, the fifth downsampling layer, and the tenth multiplier; the output end of the fourth downsampling layer is connected to the input end of the eleventh multiplier; the output end of the fifth downsampling layer is connected to the input end of the tenth multiplier; the output end of the tenth multiplier is connected to the input end of the seventh activation function layer; the output end of the seventh activation function layer is connected to the input end of the eleventh multiplier, and the output end of the eleventh multiplier is the output end of the second global branch;
[0140] The output ends of the second local branch and the second global branch are both connected to the input end of the second connection unit, and the output end of the second connection unit is connected to the input end of the nineteenth fully connected layer; the output end of the nineteenth fully connected layer is connected to the input end of the fifth adder, and the input end of the fifth adder is also connected to the input end of the fifth layer normalization module; the output end of the fifth adder is the output end of the first MoBlock module.
[0141] Further, the calculation process of the first MoBlock module is as follows:
[0142]
[0143] Among them, X lo is the output of the second local branch, and X gl is the output of the second global branch. The calculation of the second global branch is exactly the same as that of the first global branch. X out is the final output of the MoBlock module.
[0144] Further, the internal structures of the fifth ConvFFN module and the sixth ConvFFN module are the same as those of the third ConvFFN module.
[0145] After feature fusion is completed, further processing is required to enhance the expression ability of the features. The fused features are input into Moformer. Moformer has many similarities with the previous Moformer_A and Moformer_E, only differing in details. It is also composed of two repeated structures, as Figure 10 shown.
[0146] Among them, the structure of ConvFFN is exactly the same as that of the action extraction module. Compared with MoEBlock, MoBlock also changes the depthwise separable convolution to 3D convolution. At the same time, on the local branch, referring to the structure of MoABlock, a Swish activation function and a fully connected layer are added to obtain stronger non-linearity, as Figure 11 shown.
[0147] Moformer combines some characteristics of Moformer_E and MoAfomer_A. It can not only combine global and local perception, but also combine temporal perception, further enhancing the expressiveness of the features.
[0148] Embodiment 2
[0149] This embodiment provides an in-examination-room abnormal action analysis system based on feature fusion, including:
[0150] An acquisition module configured to acquire an in-examination-room surveillance video to be analyzed;
[0151] An analysis module, which is configured to: input an examination room monitoring video to be analyzed into a trained analysis model for abnormal actions in the examination room to obtain an analysis result for abnormal actions in the examination room; wherein, the trained analysis model for abnormal actions in the examination room is used to extract primary features from the examination room monitoring video to be analyzed, respectively extract environmental features and action features from the primary features, fuse the environmental features and action features to obtain fused features, extract features from the fused features to obtain final features, and classify and identify the final features to obtain an analysis result for abnormal actions in the examination room.
[0152] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. The abnormal motion analysis method in the examination room based on feature fusion is characterized by: include: Obtain the examination room surveillance video to be analyzed; Input the examination room monitoring video to be analyzed into the trained examination room abnormal action analysis model to obtain the examination room abnormal action analysis result; wherein the trained examination room abnormal action analysis model is used to extract primary features from the examination room monitoring video to be analyzed, extract environmental features and action features from the primary features respectively, fuse the environmental features and action features to obtain fused features, extract features from the fused features to obtain final features, classify and identify the final features to obtain the examination room abnormal action analysis result, and the trained examination room abnormal action analysis model includes: A backbone module, the input end of which is used to input the examination room surveillance video to be analyzed; The output end of the backbone module is connected to the input end of the first branch and the input end of the second branch respectively, the first branch comprises: a first downsampling layer and a Moformer_E module connected in sequence, the input end of the first downsampling layer is the input end of the first branch; the output end of the Moformer_E module is the output end of the first branch; The second branch includes: a Moformer_A module, an input end of the Moformer_A module is an input end of the second branch, and an output end of the Moformer_A module is an output end of the second branch; The output end of the first branch and the output end of the second branch are both connected to the input end of the feature fusion module, the output end of the feature fusion module is connected to the input end of the Moformer module, and the output end of the Moformer module outputs the abnormal action classification and recognition result; among them, Moformer_E is used for environmental feature extraction, MoAformer_A is used for action feature extraction, and Moformer is used for further feature processing.
2. The method for analyzing abnormal movements in an examination room based on feature fusion as claimed in claim 1, characterized in that: The Moformer_E module includes: A first MoEBlock module, a first ConvFFN module, a second MoEBlock module, and a second ConvFFN module connected in sequence; The internal structures of the first MoEBlock module and the second MoEBlock module are consistent. The first MoEBlock module includes: a first-layer normalization module, the input end of the first-layer normalization module is the input end of the first MoEBlock module; the output end of the first-layer normalization module is respectively connected to the input end of the first local branch and the input end of the second global branch; The first local branch includes: a first fully connected layer, the input end of the first fully connected layer is connected to the output end of the first layer normalization module, the output end of the first fully connected layer is respectively connected to the input end of the first depth-separable convolution layer, the input end of the second depth-separable convolution layer and the input end of the third depth-separable convolution layer; the output end of the second depth-separable convolution layer and the output end of the third depth-separable convolution layer are connected to the input end of the first multiplier; the output end of the first multiplier is connected to the input end of the second fully connected layer, the output end of the second fully connected layer is connected to the input end of the first activation function layer, the output end of the first depth-separable convolution layer and the output end of the first activation function layer are both connected to the input end of the second multiplier, and the output end of the second multiplier is the output end of the first local branch; The second global branch includes: a first self-attention mechanism layer, wherein the input end of the first self-attention mechanism layer is connected to the output end of the first layer normalization module, the output end of the first self-attention mechanism layer is connected to the input end of the third fully connected layer, and the output end of the third fully connected layer is respectively connected to the input end of the second downsampling layer, the input end of the third downsampling layer and the input end of the third multiplier; the output end of the second downsampling layer is connected to the input end of the fourth multiplier; the output end of the third downsampling layer is connected to the input end of the third multiplier; the output end of the third multiplier is connected to the input end of the second activation function layer; the output end of the second activation function layer is connected to the input end of the fourth multiplier, and the output end of the fourth multiplier is the output end of the second global branch; The output end of the first local branch and the output end of the second global branch are both connected to the input end of the first connection unit, and the output end of the first connection unit is connected to the input end of the fourth fully connected layer; the output end of the fourth fully connected layer is connected to the input end of the first adder, and the input end of the first adder is also connected to the input end of the first layer normalization module; the output end of the first adder is the output end of the first MoEBlock module.
3. The method for analyzing abnormal movements in an examination room based on feature fusion as claimed in claim 2, characterized in that: The internal structures of the first ConvFFN module and the second ConvFFN module are consistent. The first ConvFFN module includes: A second normalization module, a fifth fully connected layer, a third activation function layer, a fourth depthwise separable convolutional layer, a sixth fully connected layer, and a second adder connected in sequence; The input end of the second-layer normalization module is the input end of the first ConvFFN module, the input end of the second-layer normalization module is connected to the input end of the fifth depth-separable convolutional layer, the output end of the fifth depth-separable convolutional layer is connected to the input end of the third-layer normalization module, the output end of the third-layer normalization module is connected to the input end of the seventh fully connected layer, the output end of the seventh fully connected layer is connected to the input end of the second adder, and the output end of the second adder is the output end of the first ConvFFN module; The working process of the first ConvFFN module includes: the input data first passes through batch normalization and a fully connected layer, and then uses the GELU activation function for nonlinear transformation; then, after the GELU activation, a depthwise separable convolution is used to make the ConvFFN aggregate local information; again, it passes through a fully connected layer for linear transformation; The other line represents the skip connection, which is used to solve the gradient vanishing problem in deep neural networks. In the skip connection, feature extraction is first performed through a depthwise separable convolution. Then, it is regularized through batch normalization. Then, the data enters a fully connected layer for linear transformation. The depthwise separable convolution and the fully connected layer are used to downsample and increase the dimension of the input, respectively. Finally, the results of the sixth and seventh fully connected layers are added together to obtain the output of the first ConvFFN module.
4. The method for analyzing abnormal movements in an examination room based on feature fusion as claimed in claim 1, characterized in that: The Moformer_A module includes: A first MoABlock module, a third ConvFFN module, a second MoABlock module and a fourth ConvFFN module connected in sequence; The internal structures of the first MoABlock module and the second MoABlock module are consistent. The first MoABlock module includes: An eighth fully connected layer, wherein the input end of the eighth fully connected layer is the input end of the first MoABlock module; the output end of the eighth fully connected layer is respectively connected to the input end of the first three-dimensional convolutional layer, the input end of the second three-dimensional convolutional layer, and the input end of the third three-dimensional convolutional layer; The output end of the second three-dimensional convolutional layer and the output end of the third three-dimensional convolutional layer are both connected to the input end of the fifth multiplier, the output end of the fifth multiplier is connected to the input end of the ninth fully connected layer, the output end of the ninth fully connected layer is connected to the input end of the Swish layer, the output end of the Swish layer is connected to the input end of the tenth fully connected layer, and the output end of the tenth fully connected layer is connected to the input end of the fourth activation function layer; The output end of the first three-dimensional convolutional layer and the output end of the fourth activation function layer are both connected to the input end of the sixth multiplier, the output end of the sixth multiplier is connected to the input end of the eleventh fully connected layer, the output end of the eleventh fully connected layer is connected to the input end of the third adder, and the output end of the third adder is the output end of the first MoABlock module; The first MoABlock module includes: the input data is also linearly transformed to obtain Q, K and V. For Q and K, first, a 3D convolution layer is used for feature extraction; then, the data enters a fully connected layer for linear transformation; then, the Swish activation function is used for nonlinear transformation; it is linearly transformed again through a fully connected layer, and the result is converted into a probability distribution using a softmax function; and V will pass through a 3D convolution module and then be multiplied with it to obtain a new feature representation; finally, a fully connected layer is used to process the data of the connected channel size to obtain the final output.
5. The method for analyzing abnormal movements in an examination room based on feature fusion as claimed in claim 4, characterized in that: The internal structures of the third ConvFFN module and the fourth ConvFFN module are consistent. The third ConvFFN module includes: A fourth normalization module, a twelfth fully connected layer, a fifth activation function layer, a fourth three-dimensional convolutional layer, a thirteenth fully connected layer and a fourth adder connected in sequence; The input end of the fourth adder is also connected to the input end of the fourth layer normalization module; The input of the fourth layer normalization module is the input of the third ConvFFN module; The output end of the fourth adder is the output end of the third ConvFFN module; The third ConvFFN module is used for inputting data, which first passes through batch normalization and a fully connected layer, and then uses a GELU activation function for nonlinear transformation; then, after GELU activation, a three-dimensional convolution is used to enable ConvFFN to aggregate spatiotemporal information; again, a linear transformation is performed through a fully connected layer FC Layer; another line represents a skip connection, and the skip connection of the third ConvFFN module does not require additional operations, and directly adds the input data and the result of the thirteenth fully connected layer to obtain the output of the third ConvFFN module.
6. The method for analyzing abnormal movements in an examination room based on feature fusion as claimed in claim 1, characterized in that feature fusion After the feature extraction of the two branches, the module obtains a dimension of and a dimension of Output: Use a 1×1 convolutional layer to expand the number of channels of dimension 1 to ; Then the outputs of the two adjusted branches are spliced in the horizontal dimension; Next, the SE module is applied to adjust the fused features. The SE module includes a squeeze operation layer, a fully connected network, and a feature mapping layer connected in sequence. First, the Squeeze operation is performed. The Squeeze operation performs global average pooling on the feature map to compress the three-dimensional spatial information into a one-dimensional vector. The Squeeze operation compresses all the spatial information in the feature map into a channel-level descriptor: (3-5) in, is the feature map of the ith channel, is the global information of the channel, represents the global average pooling operation; The weight of each channel is generated through a fully connected network to generate a weight vector s, in which each element corresponds to the importance score of a channel: (3-6) Using the ReLU activation function, and is the weight matrix of the fully connected layer; Finally, the generated weight vector Applied to the original feature map by element-by-element multiplication: (3-7)。 7. The method for analyzing abnormal movements in an examination room based on feature fusion as claimed in claim 1, characterized in that: The Moformer module includes: a first MoBlock module, a fifth ConvFFN module, a second MoBlock module and a sixth ConvFFN module connected in sequence; The internal structures of the first MoBlock module and the second MoBlock module are consistent. The first MoBlock module includes: The fifth layer normalization module, the input end of the fifth layer normalization module is the input end of the first MoBlock module; The output end of the fifth layer normalization module is connected to the input end of the second local branch and the input end of the second global branch respectively; The second local branch includes: a fifteenth fully connected layer, the input end of the fifteenth fully connected layer is connected to the output end of the fifth normalization module, the output end of the fifteenth fully connected layer is respectively connected to the input end of the sixth depth-separable convolutional layer, the input end of the seventh depth-separable convolutional layer and the input end of the eighth depth-separable convolutional layer; the output end of the seventh depth-separable convolutional layer and the output end of the eighth depth-separable convolutional layer are connected to the input end of the eighth multiplier; the output end of the eighth multiplier is connected to the input end of the sixteenth fully connected layer, the output end of the sixteenth fully connected layer is connected to the input end of the Swish layer, the output end of the Swish layer is connected to the input end of the seventeenth fully connected layer, the output end of the seventeenth fully connected layer is connected to the input end of the sixth activation function layer, the output end of the sixth depth-separable convolutional layer and the output end of the sixth activation function layer are both connected to the input end of the ninth multiplier, and the output end of the ninth multiplier is the output end of the second local branch; The second global branch includes: a second self-attention mechanism layer, the input end of the second self-attention mechanism layer is connected to the output end of the fifth normalization module, the output end of the second self-attention mechanism layer is connected to the input end of the eighteenth fully connected layer, the output end of the eighteenth fully connected layer is respectively connected to the input end of the fourth downsampling layer, the input end of the fifth downsampling layer and the input end of the tenth multiplier; the output end of the fourth downsampling layer is connected to the input end of the eleventh multiplier; the output end of the fifth downsampling layer is connected to the input end of the tenth multiplier; the output end of the tenth multiplier is connected to the input end of the seventh activation function layer; the output end of the seventh activation function layer is connected to the input end of the eleventh multiplier, and the output end of the eleventh multiplier is the output end of the second global branch; The output end of the second local branch and the output end of the second global branch are both connected to the input end of the second connection unit, and the output end of the second connection unit is connected to the input end of the nineteenth fully connected layer; the output end of the nineteenth fully connected layer is connected to the input end of the fifth adder, and the input end of the fifth adder is also connected to the input end of the fifth layer normalization module; the output end of the fifth adder is the output end of the first MoBlock module.
8. The method for analyzing abnormal movements in an examination room based on feature fusion as claimed in claim 7, characterized in that: The Moformer module includes: A first MoBlock module, a fifth ConvFFN module, a second MoBlock module, and a sixth ConvFFN module connected in sequence; The calculation process of the first MoBlock module is as follows: (3-8) in, is the output of the second local branch, is the output of the second global branch, which is calculated exactly the same as the first global branch. It is the final output of the MoBlock module.
9. The abnormal motion analysis system in the examination room based on feature fusion is characterized by: include: An acquisition module is configured to: acquire the examination room surveillance video to be analyzed; The analysis module is configured to: input the examination room monitoring video to be analyzed into the trained examination room abnormal action analysis model to obtain the examination room abnormal action analysis result; wherein the trained examination room abnormal action analysis model is used to extract primary features from the examination room monitoring video to be analyzed, extract environmental features and action features from the primary features respectively, fuse the environmental features and the action features to obtain fused features, extract features from the fused features to obtain final features, classify and identify the final features to obtain the examination room abnormal action analysis result, and the trained examination room abnormal action analysis model includes: A backbone module, the input end of which is used to input the examination room surveillance video to be analyzed; The output end of the backbone module is connected to the input end of the first branch and the input end of the second branch respectively, the first branch comprises: a first downsampling layer and a Moformer_E module connected in sequence, the input end of the first downsampling layer is the input end of the first branch; the output end of the Moformer_E module is the output end of the first branch; The second branch includes: a Moformer_A module, an input end of the Moformer_A module is an input end of the second branch, and an output end of the Moformer_A module is an output end of the second branch; The output end of the first branch and the output end of the second branch are both connected to the input end of the feature fusion module, the output end of the feature fusion module is connected to the input end of the Moformer module, and the output end of the Moformer module outputs the abnormal action classification and recognition result; among them, Moformer_E is used for environmental feature extraction, MoAformer_A is used for action feature extraction, and Moformer is used for further feature processing.
Citation Information
Patent Citations
Examination room abnormal behavior recognition method based on action recognition
CN114639166A
Examinee examination room abnormal behavior analysis method, system and device and storage medium
CN115880647A