A method for monitoring and identifying abnormal behaviors of personnel and an early warning system

By constructing a convolutional neural network model, the Frobenius norm and receptive field ratio of the feature matrix are used to give the feature map weight, and combined with the global attention layer, the problem of inaccurate recognition of dangerous objects and human bodies in the existing technology is solved, and the accuracy of abnormal behavior recognition is improved.

CN119810753BActive Publication Date: 2025-07-22ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510251698.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-22
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Existing image recognition models are difficult to accurately identify the differences between dangerous objects and human bodies, resulting in low accuracy in identifying abnormal behaviors.

Method used

Convolutional neural network model is constructed, including convolutional layer 1, channel multi-scale difference attention layer and convolutional layer 2. By calculating the Frobenius norm and receptive field ratio of the feature matrix, the detection model is obtained by training.

Benefits of technology

The model's attention to channels with large differences in multiple scales is improved, the ability to distinguish dangerous objects from the human body is enhanced, and the accuracy of abnormal behavior recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810753B_ABST
    Figure CN119810753B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of monitoring abnormal behaviors of personnel, and particularly to a method for monitoring and identifying abnormal behaviors of personnel and an early warning system. The method includes: constructing a convolutional neural network model, which includes a first convolutional layer, a channel multi-scale difference attention layer, and a second convolutional layer; in the channel multi-scale difference attention layer, using a convolutional kernel to perform convolution on each first feature map output by the first convolutional layer to obtain two second feature maps, performing pooling on each second feature map to obtain a feature matrix, calculating the Frobenius norm of the two feature matrices, and using the Frobenius norm as the weight of the first feature map; performing weighted connection on multiple first feature maps and inputting them into the second convolutional layer; training the convolutional neural network model to obtain a detection model, inputting the acquired personnel image into the detection model to detect abnormal behaviors of personnel, and if abnormal behaviors of personnel are detected, an early warning is issued. The present invention improves the accuracy of the neural network model in identifying abnormal behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of monitoring of abnormal behaviors of personnel, and particularly to a method for monitoring and identifying abnormal behaviors of personnel and an early warning system. Background Art

[0002] Traditional monitoring in public areas usually involves security personnel patrolling and inspecting corresponding areas, as well as inspecting passing pedestrians, vehicles, and carried items to ensure the safety of public areas. When the flow of people is large, it requires a large amount of human resources. With the development of Internet technology, video monitoring technology has been applied in the monitoring of abnormal behaviors of personnel. By detecting the behavior actions, expressions of personnel, or whether they carry dangerous items, it is determined whether the corresponding personnel show abnormal behaviors so as to take timely measures to prevent dangerous behaviors.

[0003] The Chinese patent application document with the publication number CN110826522A discloses a method, system, storage medium, and monitoring device for monitoring abnormal behaviors of the human body. The method includes: obtaining a video data set uploaded to a human behavior database; using octave convolution to update a preset I3D model to obtain a target I3D model; training and testing the video data set through the target I3D model to obtain a behavior prediction model; inputting the video information of the personnel flow in the monitoring area into the behavior prediction model to obtain the abnormal information of the personnel with abnormal behaviors in the monitoring area, and generating early warning and warning information according to the types of the abnormal behaviors.

[0004] However, in the actual detection process, due to the small sizes of dangerous items, such as knives, clubs, etc., the gap between dangerous items and the human body is large, the number of pixel points occupied by dangerous items in the image is small and the features are relatively blurred. Existing image recognition models are prone to jointly recognize dangerous items and the human body as the human body, thus recognizing abnormal behaviors as normal behaviors, resulting in low accuracy of the model in recognizing abnormal behaviors. Summary of the Invention

[0005] To solve the problem that it is not easy to identify dangerous items with small sizes, the present invention provides a method for monitoring and identifying abnormal behaviors of personnel and an early warning system.

[0006] In a first aspect, the present invention provides a method for monitoring and identifying abnormal behaviors of personnel, adopting the following technical solution:

[0007] Construct a convolutional neural network model, where the convolutional neural network model includes a first convolutional layer, a channel multi-scale difference attention layer, and a second convolutional layer;

[0008] In the channel multi-scale difference attention layer, a convolutional kernel is used to convolve each first feature map output by the first convolutional layer to obtain two second feature maps. Pooling is performed on each second feature map to obtain a feature matrix, and the Frobenius norm of the two feature matrices is calculated. The Frobenius norm is used as the weight of the first feature map. The weighted connection of multiple first feature maps is input into the second convolutional layer;

[0009] The convolutional neural network model is trained to obtain a detection model. The obtained personnel image is input into the detection model to detect the abnormal behavior of the personnel. If abnormal behavior of the personnel is detected, a warning is issued.

[0010] By assigning corresponding weights to the first feature maps through the channel multi-scale difference attention layer, different first feature maps can be distinguished, improving the model's attention to channels with large multi-scale differences and enhancing the accuracy of the model in identifying abnormal behaviors.

[0011] Preferably, the convolutional neural network model further includes an input layer and a multi-scale convolutional layer. The input layer inputs the three-channel image of the personnel image into the multi-scale convolutional layer. Multiple convolutional kernels are used to convolve the three-channel image to obtain multiple initial feature maps. The receptive field of the initial feature maps in the personnel area of the personnel image is calculated, and the weight of the corresponding initial feature map is calculated according to the ratio of the receptive field to the personnel area. The weighted connection of multiple initial feature maps is input into the first convolutional layer.

[0012] By calculating the weights of each initial feature map through the receptive field corresponding to the initial feature map, the greater the weight, the more important the corresponding feature map. Compared with the traditional fixed-weight connection method, the accuracy of identifying discriminative features (hazardous items) is improved.

[0013] Preferably, the weight expression of the initial feature map is:

[0014] ;

[0015] In the formula, represents the weight of the initial feature map, avg represents the average function, represents the receptive field of the initial feature map on the personnel image, is the personnel area with abnormal behavior on the personnel image.

[0016] Calculating the weights of the corresponding initial feature maps through the receptive field facilitates understanding the importance of the initial feature maps, thereby improving the accuracy of identifying hazardous substances.

[0017] Preferably, the method for training the convolutional neural network model is:

[0018] A training set is obtained. The training set includes multiple personnel images, and corresponding labels are set for each personnel image. The labels include normal and abnormal;

[0019] Use cross - entropy loss as the loss function, and input the training set into the convolutional neural network model for training;

[0020] Stop training in response to the number of training times reaching the set maximum number of training times or the loss of the convolutional neural network model being less than the preset loss threshold.

[0021] By training the convolutional neural network model, a detection model is obtained, and the detection model can identify abnormal behaviors of personnel.

[0022] Preferably, the expression of the loss function is:

[0023] ;

[0024] In the formula, represents the prediction loss of the i - th personnel image, r represents the label of the i - th personnel image, and its value is 0 or 1, and p represents the probability that the convolutional neural network model predicts as abnormal.

[0025] By determining the loss function, the accuracy of training the convolutional neural network model to identify abnormal behaviors is improved.

[0026] Preferably, global average pooling or global max pooling is performed on each second feature map.

[0027] Preferably, the convolutional neural network model further includes a global attention layer, a fully - connected layer, and an output layer.

[0028] In a second aspect, the present invention provides a personnel abnormal behavior warning system, adopting the following technical solution:

[0029] A personnel abnormal behavior warning system, including a processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above - mentioned personnel abnormal behavior monitoring and recognition method is implemented.

[0030] The beneficial effect is: generating a computer program for the above - mentioned personnel abnormal behavior monitoring and recognition method and storing it in the memory to be loaded and executed by the processor. Thus, a system is made according to the memory and the processor, which is convenient to use.

[0031] The present invention has the following technical effects:

[0032] By assigning corresponding weights to the first feature map through the channel multi-scale difference attention layer, different first feature maps can be distinguished, improving the model's attention to channels with large multi-scale differences, giving greater weights to the first feature maps that are beneficial for dividing small objects and humans, enabling the trained model to improve its attention to small objects while maintaining a certain attention to human recognition, thereby avoiding the phenomenon of misidentifying dangerous objects as humans and enhancing the accuracy of the model's abnormal behavior recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] By referring to the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of example and not limitation, and the same or corresponding reference numerals represent the same or corresponding parts.

[0034] Figure 1 It is a flowchart of a method for monitoring and recognizing abnormal behaviors of personnel according to an embodiment of the present invention.

[0035] Figure 2 It is a flowchart of the operation of the multi-scale convolutional layer according to an embodiment of the present invention.

[0036] Figure 3 It is a flowchart of the operation of the channel multi-scale difference attention layer according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0038] It should be understood that when the claims, specifications, and drawings of the present invention use terms such as "first" and "second", they are only used to distinguish different objects and not to describe a specific order. The terms "including" and "comprising" used in the specifications and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0039] An embodiment of the present invention discloses a method for monitoring and recognizing abnormal behaviors of personnel, referring to Figure 1 , including the following steps, specifically as follows:

[0040] S1: Construct a convolutional neural network model.

[0041] The convolutional neural network model includes an input layer, a multi-scale convolutional layer, a first convolutional layer, a channel multi-scale difference attention layer, a second convolutional layer, a global attention layer, a fully connected layer, and an output layer.

[0042] S11: The input layer inputs the three-channel images corresponding to the personnel image into the multi-scale convolutional layer. Among them, the three channels are the R, G, and B channels, that is, the R channel corresponds to one image, the G channel corresponds to one image, and the B channel corresponds to one image. The images input by the input layer can be regarded as an input tensor. Exemplarily, an input tensor is: , represents the input tensor, and H, W, and C respectively represent the height, width, and number of channels of the input tensor.

[0043] S12: The multi-scale convolutional layer uses multiple convolutional kernels to perform convolution on the three-channel images to obtain multiple initial feature maps, calculates the receptive field of the initial feature maps in the personnel area of the personnel image, calculates the weights of the corresponding initial feature maps according to the ratio of the receptive field to the personnel area, and performs weighted connection on the multiple initial feature maps and inputs them into the first convolutional layer.

[0044] Exemplarily, as shown in Figure 2 , the size of the three-channel image is 32×32. Use a 3×3 convolutional kernel to perform convolution on the input three-channel image, the convolution stride is 1, and the paddings are 1, 2, and 3 respectively (the purpose of padding is to ensure that the size of the output initial feature map is the same as the size of the input three-channel image). After convolution, an initial feature map with a size of 32×32 is obtained . Similarly, use convolutional kernel to perform convolution on the input tensor to obtain an initial feature map with a size of 32×32 . Use convolutional kernel to perform convolution on the input tensor to obtain an initial feature map with a size of 32×32 , and the corresponding tensor is: .

[0045] For the initial feature map , calculate the receptive field of the initial feature map in the personnel area of the personnel image, and calculate the ratio of the receptive field to the personnel area. The expression is: , represents the receptive field of the initial feature map on the personnel image, is the personnel area with abnormal behavior on the personnel image.

[0046] In the historical personnel images with abnormal behavior, each personnel image corresponds to a ratio. Take the average of the multiple ratios as the weight of the initial feature map . The weight expression of the initial feature map is:

[0047] ;

[0048] In the formula, represents the weight of the initial feature map, avg represents the average function, represents the receptive field of the initial feature map on the personnel image, is the personnel area with abnormal behavior on the personnel image. Thus, it can be seen that the initial feature map , and each correspond to a weight, and the weights are normalized to obtain , and .

[0049] Multiple initial feature maps are weighted and connected to obtain a weighted initial feature map. The expression for weighted connection is: , is the weighted feature map after weighted connection, is the weight corresponding to the initial feature map , is the weight corresponding to the initial feature map , is the weight corresponding to the initial feature map . The weights of each initial feature map are calculated through the receptive field corresponding to the initial feature map. The larger the weight, the more important the information contained in the corresponding feature map. Compared with the traditional method of connecting with fixed weights, the accuracy of the convolutional neural network model in identifying discriminative features (hazardous items) is improved.

[0050] S13: The convolutional layer 1 performs convolution processing on the weighted feature map output by the multi-scale convolutional layer to obtain the first feature map.

[0051] Exemplarily, the convolutional layer 1 contains 32 convolutional kernels of size 3×3, with a stride of 1 and a padding of 1. The convolutional layer 1 outputs 32 first feature maps with a size of 32×32.

[0052] S14: The channel multi-scale difference attention layer uses convolutional kernels to perform convolution on each first feature map output by the convolutional layer 1 to obtain two second feature maps, performs pooling on each second feature map to obtain a feature matrix, calculates the Frobenius norm of the two feature matrices, and uses the Frobenius norm as the weight of the first feature map; the weighted connection of multiple first feature maps is input into the convolutional layer 2. The pooling method adopts global average pooling or global max pooling.

[0053] Exemplarily, combined with Figure 3As shown, a second feature map p is obtained by convolving the first feature map g with a 3×3 convolutional kernel, and a second feature map q is obtained by convolving the first feature map g with a 7×7 convolutional kernel. A feature matrix A is obtained by pooling the second feature map p, and a feature matrix B is obtained by pooling the second feature map q. The Frobenius norm between the feature matrix A and the feature matrix B is calculated. The expression of the Frobenius norm is:

[0054] ;

[0055] In the formula, D represents the Frobenius norm between the feature matrix A and the feature matrix B, represents the element in the x-th row and y-th column of the feature matrix A, represents the element in the x-th row and y-th column of the feature matrix B, and m and n respectively represent the number of rows and columns of the feature matrix. The Frobenius norm is used as the weight of the first feature map b. Multiple first feature maps are weighted and connected to obtain a weighted first feature map. The way of weighted connection of the first feature map is the same as that of the initial feature map, which will not be elaborated here.

[0056] Since there are large volume differences between dangerous items (such as knives and spikes) and the human body in abnormal human behaviors, specifically manifested as: the characteristics of dangerous items are mainly reflected in small-scale features, and the characteristics of humans are mainly reflected in large-scale features; for dangerous items, the smaller the scale difference between dangerous items and humans, the easier it is for the model to identify dangerous items and the human body as one and misclassify abnormal behaviors as normal behaviors. The larger the scale difference between dangerous items and humans, the easier it is for the model to accurately distinguish and identify dangerous items and the human body. Therefore, by assigning corresponding weights to the first feature map through the channel multi-scale difference attention layer, different first feature maps can be distinguished, the attention of the model to channels with large multi-scale differences can be improved, and a larger weight is given to the first feature map that is beneficial to distinguishing small objects and the human body, so as to distinguish dangerous items from humans.

[0057] S15: The second convolutional layer performs convolution processing on the weighted first feature map output by the channel multi-scale difference attention layer to obtain a third feature map.

[0058] Exemplarily, the second convolutional layer contains 64 convolutional kernels of size with a stride of 1 and a padding of 1. The second convolutional layer outputs 64 third feature maps with a size of 32×32.

[0059] S16: The third feature map is input into the global attention layer. The global attention layer performs convolution on all the third feature maps input by the second convolutional layer to obtain query vectors and key vectors at different positions, and calculates the similarity between any two positions in the same third feature map.

[0060] The expression for similar figures is:

[0061] ;

[0062] Wherein, represents the similarity between position i and position j in the third feature map, f represents the dimension of the query vector or key vector, represents the value of the query vector at position i in the l dimension, represents the value of the key vector at position j in the l dimension. The softmax function is used to normalize the similarities of all positions to obtain the attention weight matrix. The attention weight matrix is multiplied by the third feature map to obtain the weighted third feature map, and the weighted third feature maps are concatenated to obtain the output feature map.

[0063] Exemplarily, the number of output feature maps is 64, and the size is 32×32.

[0064] S17: Input the output feature map into the fully connected layer, and output the behavior category of the person in the person image through the output layer. The behavior category is normal or abnormal.

[0065] Exemplarily, input 64 feature maps into the fully connected layer with 128 output neurons. The output size of the fully connected layer is 1×1×128, and the output layer outputs that the behavior category of the person in the person image is abnormal.

[0066] S2: Train the convolutional neural network model to obtain the detection model.

[0067] The method for training the convolutional neural network model is: Obtain the training set. The training set includes multiple person images, and set corresponding labels for each person image. The labels include normal and abnormal; Use cross-entropy loss as the loss function. The expression of the loss function is:

[0068] ;

[0069] Wherein, represents the prediction loss of the i-th person image, r represents the label of the i-th person image, and its value is 0 or 1. 1 represents abnormal, 0 represents normal, and p represents the probability that the convolutional neural network model predicts as abnormal. Input the training set into the convolutional neural network model for training; Stop training in response to the training times reaching the set maximum training times or the loss of the convolutional neural network model being less than the preset loss threshold.

[0070] S3: Input the obtained person image into the detection model to detect the abnormal behavior of the person. If the abnormal behavior of the person is detected, an early warning is issued.

[0071] Obtain the image of a person in real time and input it into a detection model to detect the abnormal behavior of the person. If an abnormal behavior of the person is detected, a warning is issued to prompt the corresponding security personnel to pay attention to the abnormal person in time and prevent the abnormal person from causing harm to the surrounding people.

[0072] An embodiment of the present invention also discloses a warning system for abnormal behavior of a person, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a method for monitoring and identifying abnormal behavior of a person according to the present invention is implemented.

[0073] The above system further includes a communication bus, a communication interface and other components well known to those skilled in the art. Their settings and functions are known in the art, so they will not be described in detail here.

[0074] In the present invention, the aforementioned memory can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device or device. For example, a computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high bandwidth memory (HBM), a hybrid memory cube (HMC), etc., or any other medium that can be used to store the required information and can be accessed by an application program, a module or both. Any such computer storage medium can be part of the device or accessible or connectable to the device.

[0075] Although this specification has shown and described multiple embodiments of the present invention, it is obvious to those skilled in the art that such embodiments are provided only by way of example. Those skilled in the art will think of many changes, alterations and alternative ways without departing from the spirit and concept of the present invention. It should be understood that various alternative solutions to the embodiments of the present invention described herein can be adopted in the process of practicing the present invention.

[0076] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for monitoring and identifying abnormal behaviors of personnel, characterized in that, Including the steps: Construct a convolutional neural network model, which includes a first convolutional layer, a channel multi-scale difference attention layer, and a second convolutional layer; the convolutional neural network model also includes a global attention layer, a fully connected layer, and an output layer; In the channel multi-scale difference attention layer, use a convolutional kernel to perform convolution on each first feature map output by the first convolutional layer to obtain two second feature maps, perform pooling on each second feature map to obtain a feature matrix, calculate the Frobenius norm of the two feature matrices, and use the Frobenius norm as the weight of the first feature map; perform weighted connection on multiple first feature maps and input them into the second convolutional layer; the second convolutional layer performs convolution processing on the weighted first feature maps output by the channel multi-scale difference attention layer to obtain a third feature map; input the third feature map into the global attention layer, and the global attention layer performs convolution on the third feature maps input by the second convolutional layer to obtain query vectors and key vectors at different positions, and calculate the similarity between any two positions in the same third feature map; The expression for similar figures is: ; In the formula, represents the similarity between position i and position j in the third feature map, f represents the dimension of the query vector or the key vector, represents the value of the query vector at position i on the l dimension, represents the value of the key vector at position j on the l dimension; the softmax function is used to normalize the similarities of all positions to obtain the attention weight matrix, the attention weight matrix is multiplied by the third feature map to obtain the weighted third feature map, and the weighted third feature maps are concatenated to obtain the output feature map; the output feature map is input into the fully connected layer, and the behavior category of the person in the person image is output through the output layer, and the behavior category is normal or abnormal; Train the convolutional neural network model to obtain a detection model, input the obtained personnel image into the detection model to detect the abnormal behavior of the personnel, and if the abnormal behavior of the personnel is detected, issue a warning; The convolutional neural network model also includes an input layer and a multi-scale convolutional layer. The input layer inputs the three-channel image of the personnel image into the multi-scale convolutional layer, uses multiple convolutional kernels to perform convolution on the three-channel image to obtain multiple initial feature maps, calculates the receptive field of the initial feature maps in the personnel area of the personnel image, calculates the weights of the corresponding initial feature maps according to the ratio of the receptive field to the personnel area, and performs weighted connection on the multiple initial feature maps and inputs them into the first convolutional layer.

2. The method for monitoring and identifying abnormal behaviors of personnel according to claim 1, wherein, The weight expression of the initial feature map is as follows: ; In the formula, represents the weight of the initial feature map, and avg represents the average function. represents the receptive field of the initial feature map on the person image. is the person area with abnormal behavior on the person image.

3. The personnel abnormal behavior monitoring and recognition method according to claim 1, characterized in that The method for training the convolutional neural network model is: Obtain a training set, which includes multiple personnel images, and set corresponding labels for each personnel image, and the labels include normal and abnormal; Use cross-entropy loss as the loss function, and input the training set into the convolutional neural network model for training; Stop training in response to the training times reaching the set maximum training times or the loss of the convolutional neural network model being less than the preset loss threshold.

4. The method for monitoring and identifying abnormal behaviors of personnel according to claim 3, characterized in that, The expression of the loss function is: ; Wherein, represents the prediction loss of the i-th personnel image, r represents the label of the i-th personnel image, and its value is 0 or 1, and p represents the probability that the convolutional neural network model predicts as abnormal.

5. The method for monitoring and identifying abnormal behaviors of personnel according to claim 1, wherein, Perform global average pooling or global max pooling on each second feature map.

6. A personnel abnormal behavior warning system, characterized in that, Including: A processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, it implements a method for monitoring and recognizing abnormal behavior of personnel according to any one of claims 1-5.

Citation Information

Patent Citations

  • Human body abnormal behavior monitoring method and system, storage medium and monitoring equipment

    CN110826522A

  • Neural network image classification method based on gating local channel attention

    CN114419361A

  • Escalator pedestrian abnormal behavior recognition method, device and equipment and storage medium

    CN118587760A

  • Multi-modal image fusion method and device based on extended residual attention network

    CN119131043A