Abnormal scene recognition methods, devices, electronic equipment and storage media
By initially screening pedestrian limb kinetic energy at the smart camera end and generating samples using posture transfer, combined with a CNN-LSTM network, the problems of low accuracy and high hardware consumption in abnormal scene recognition are solved, achieving efficient and real-time abnormal behavior recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-01
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies have low accuracy in identifying abnormal scenes, especially in public gathering places and densely populated areas. When abnormal behavior recognition models are transferred to real-world scenarios, their generalization and accuracy decrease. Furthermore, existing algorithms have high computational and hardware consumption requirements, making it difficult to meet real-time requirements.
The system performs initial screening at the smart camera end, uses multi-scale directional relative gradients to determine the kinetic energy of pedestrian end limbs, and sends the data to the pose estimation network for secondary screening if it exceeds a threshold. It also generates realistic abnormal behavior samples through pose transfer and combines them with a CNN-LSTM spatiotemporal network for identification.
It improves the accuracy and real-time performance of abnormal scene recognition, reduces the number of videos to be judged in the cloud, reduces hardware consumption, expands the number of samples, and improves the generalization ability of the model.
Smart Images

Figure CN115424191B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to an abnormal scene recognition method, apparatus, electronic device, and storage medium. Background Technology
[0002] In response to public safety and security issues, abnormal behavior analysis systems can be used in public gathering places and densely populated areas to intelligently analyze and judge various abnormal behaviors of individuals and groups in key areas, effectively assisting back-end personnel in judging and handling abnormal situations that occur in the monitoring scene.
[0003] Collecting samples related to abnormal behavior is quite difficult. A large portion of the samples are obtained by crawling video clips from the Internet. However, video clips are very different from real video surveillance scenes. Models trained using these data samples face significant challenges when transferred to real-world scenarios, and their generalization and accuracy will decrease significantly. Summary of the Invention
[0004] This invention provides an abnormal scene recognition method, device, electronic device, and storage medium to solve the technical problem of low accuracy in abnormal scene recognition in the prior art.
[0005] This invention provides a method for identifying abnormal scenes, comprising:
[0006] Determine the kinetic energy of the extremities of the target pedestrian in the target image;
[0007] If the kinetic energy of the distal limb exceeds the target threshold, the target image is sent to the pose estimation network for secondary screening.
[0008] Optionally, determining the extremity kinetic energy of the target pedestrian in the target image includes:
[0009] Extract the target pedestrian from the target image;
[0010] Track the target object;
[0011] Divide the target region into N blocks of different scales based on the longest side, and add N blocks to the target edge;
[0012] Determine the multi-scale directional relative gradient of each block within the target region; the multi-scale directional relative gradient of the block is used to characterize the block's kinetic energy.
[0013] The end-limb kinetic energy of the target pedestrian is determined based on the multi-scale directional relative gradient of each block.
[0014] Optionally, sending the target image to the pose estimation network for secondary screening includes:
[0015] Feature extraction is performed on the target image to obtain a feature map;
[0016] A pose map of anomalous behavior sequences is generated using a pose estimation network;
[0017] A secondary screening is performed based on the feature map and the pose map.
[0018] Optionally, the secondary screening based on the feature map and the pose map includes:
[0019] The pose map and feature map are sent to the decoder for upsampling to obtain the target person image after pose transfer;
[0020] Abnormal behaviors in the target person image are identified by a pose discriminator.
[0021] The present invention also provides an abnormal scene recognition device, comprising:
[0022] The determination module is used to determine the kinetic energy of the extremities of a target pedestrian in the target image;
[0023] The sending module is used to send the target image to the attitude estimation network for secondary screening when the kinetic energy of the end limb exceeds the target threshold.
[0024] Optionally, the determining module includes a first extraction submodule, a tracking submodule, a segmentation submodule, a first determining submodule, and a second determining submodule;
[0025] The first extraction submodule is used to extract the target pedestrian in the target image;
[0026] The tracking submodule is used to track the target object;
[0027] The segmentation submodule is used to divide the target region into N blocks of different scales based on the longest side, and to add N blocks to the target edge;
[0028] The first determining submodule is used to determine the multi-scale directional relative gradient of each block within the target region; the multi-scale directional relative gradient of the block is used to characterize the block's kinetic energy;
[0029] The second determining submodule is used to determine the end-limb kinetic energy of the target pedestrian based on the multi-scale directional relative gradient of each block.
[0030] Optionally, the sending module includes a second extraction submodule, a generation submodule, and a screening submodule:
[0031] The second extraction submodule is used to extract features from the target image to obtain a feature map;
[0032] The generation submodule is used to generate a pose graph of the abnormal behavior sequence through a pose estimation network;
[0033] The screening submodule is used to perform secondary screening based on the feature map and the pose map.
[0034] Optionally, the screening submodule includes an upsampling unit and an identification unit;
[0035] The upsampling unit is used to send the pose map and feature map to the decoder for upsampling to obtain the target person image after pose transfer.
[0036] The recognition unit is used to identify abnormal behaviors in the target person image through a pose discriminator.
[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the above-described abnormal scene recognition methods.
[0038] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described abnormal scene recognition methods.
[0039] The abnormal scene recognition method, device, electronic device and storage medium provided by the present invention use the regional kinetic energy algorithm of the smart camera to provide initial screening of videos, reduce the number of videos to be judged in the cloud, and improve detection efficiency and hardware utilization. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating the abnormal scene recognition method provided by the present invention;
[0042] Figure 2 This is one of the schematic diagrams illustrating the abnormal scene recognition principle provided by the present invention;
[0043] Figure 3 This is the second schematic diagram of the abnormal scene recognition principle provided by the present invention;
[0044] Figure 4 This is a schematic diagram of the abnormal scene recognition device provided by the present invention;
[0045] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0046] To address public safety and security concerns, anomaly analysis systems can be used in public gathering places and densely populated areas to intelligently analyze and assess various abnormal behaviors of individuals and groups in key areas. This effectively assists back-end personnel in judging and handling abnormal situations that occur in monitored scenarios. Fight detection is both a challenge and a key aspect of anomaly event analysis.
[0047] Currently, there are relatively few datasets related to abnormal behavior recognition. Most algorithms are trained on general behavior recognition datasets such as KTH, Kinetics, and UCF-101, which suffer from low image clarity, small sample sizes, and a lack of samples from real-world video surveillance scenarios. This results in low accuracy for abnormal behavior recognition in physical security scenarios. Current mainstream behavior recognition algorithms include dual-stream neural networks (DNNs) that simultaneously extract RGB semantic information and optical flow for fusion, 3D convolutional neural networks, and spatiotemporal networks combining CNNs and LSTMs. Deep learning requires a large sample size, and current datasets are insufficient for engineering-level algorithm applications. Therefore, how to efficiently and economically obtain high-quality samples is a hot research area.
[0048] Meanwhile, in city-level surveillance scenarios, due to the deep network layers and the need to combine LSTM context learning with network learning, simultaneous pose recognition monitoring of multiple video streams requires extremely high hardware computing power and cannot meet real-time requirements. Therefore, how to meet the real-time requirements and low hardware consumption of application scenarios based on algorithm construction is also a key focus of algorithm implementation research.
[0049] Collecting samples related to abnormal behavior is challenging, with a large portion obtained by crawling video clips from the internet. These videos differ significantly from real-world video surveillance scenarios, making it difficult to transfer models trained on such data to real-world environments, resulting in a noticeable decrease in generalization and accuracy. This application proposes a pose transfer-based solution that can rapidly generate a large number of samples for specific actions, reducing the time and cost of data collection and processing. Existing sample generation schemes based on generative adversarial networks produce samples of poor quality, often including incomplete or redundant body parts, and are frequently blurry and lack fine-grained detail. Existing behavior recognition methods rely on optical flow or 3D convolution operations, which are computationally intensive and demanding on high-performance computing equipment, making them unsuitable for deployment in edge AI inference. They also fail to meet the requirements for real-time recognition of abnormal behavior.
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0051] Figure 1 This is a flowchart illustrating the abnormal scene recognition method provided by the present invention, as shown below. Figure 1 As shown in the figure, this application provides a method for identifying abnormal scenes. The method includes:
[0052] Step 101: Determine the kinetic energy of the end limbs of the target pedestrian in the target image.
[0053] Specifically, Figure 2 This is one of the schematic diagrams illustrating the abnormal scene recognition principle provided by the present invention, such as... Figure 2 As shown, the first step is to use a smart camera to perform a preliminary screening to determine the kinetic energy of the blocks.
[0054] For high-resolution videos of abnormal fighting behavior (falling to the ground, fighting, etc.) captured by video surveillance, sparse sampling is performed to extract key short segments. Then, each segment is uniformly sampled and the resulting frames are combined into a candidate video sequence of abnormal behavior.
[0055] This paper introduces target block kinetic energy index detection at the smart camera end. It uses a multi-scale method to distinguish the kinetic energy of the target within a block by analyzing the pixel fluctuation range of the block within and adjacent images. The method is as follows:
[0056] 1. First, an object detection algorithm is used to extract the bounding box, and then the Intersection over Union (IOU) tracking algorithm is used to track the target pedestrian. This yields time-series images of consecutive frames for the same pedestrian ID.
[0057] 2. Divide the target region into five scale blocks of s = (3, 5, 7, 10) based on the longest side. Add the same number of blocks to the target edge. For example, if the target region size is 100*200, then the block sizes are 70, 40, 30, and 20 respectively. Add s blocks of the corresponding size to the target edge for each of these sizes.
[0058] 3. Convert the target image to an HSV image and use average pooling to average the values of each block. Then, based on the time series, stack blocks of different scales from the time series. Calculate the gradients in the surrounding eight directions for each block within the target region. Divide the overall target gradient by the directional gradients to obtain the multi-scale relative directional gradient G. s d The relative gradient of a block is defined as the block kinetic energy.
[0059] 4. Weight each scale using a normal distribution. The weighting ratio of the system is determined by an adjustment coefficient n. When n is adjusted to the point where the scale influence is greatest at s=7, the kinetic energy of the target's extremities can be determined. The calculation formula is as follows:
[0060]
[0061] After determining the kinetic energy of the end limbs of the target pedestrian in the target image, a threshold is set. When the limb kinetic energy exceeds the threshold, it is judged as a suspected video clip of a fight.
[0062] Step 102: If the kinetic energy of the distal limb exceeds the target threshold, the target image is sent to the pose estimation network for secondary screening.
[0063] Specifically, when the kinetic energy of the end limb exceeds the target threshold, the target image in the suspected video clip is sent to the pose estimation network in the cloud for secondary screening.
[0064] Figure 3 This is the second schematic diagram of the abnormal scene recognition principle provided by the present invention, as shown below. Figure 3 As shown, the pose estimation map of the abnormal behavior sequence generated by the pose estimation network is used as the template for the target pose sequence to be transferred. For a new human target, pose estimation is also generated, and then the human pose image, human image, and abnormal behavior pose template are input into the trained pose transfer adversarial generative network to obtain the abnormal behavior image sequence of the target human and synthesize a video clip as the abnormal behavior video sample data of the new target human.
[0065] Pose estimation is performed on a human image using OpenPose, estimating the pose of 18 keypoints and generating corresponding data sample sets of the human body and its keypoints. An adversarial generative network (GAN) is constructed based on the current image, the target human body image, the current human pose, and the target pose image, enabling the generation of realistic human images from the pose map. The GAN includes an encoder and a decoder module. The encoder consists of a lightweight convolutional network and a pose transfer module, used for generating image feature heatmaps and asymptotic pose transformation. The decoder includes an upsampled deconvolutional network to generate the target pose image. The discriminator includes a texture scoring discriminator and a pose scoring discriminator to determine the consistency between the generated target pose image and the actual pose. The discriminator consists of downsampled convolutional layers and residual convolutional modules.
[0066] Posture transfer is defined here as a gradual process, treating human posture as a set of states, and progressively transferring the original posture to the target posture through a series of intermediate states. This includes the following sub-steps:
[0067] First, feature maps P0 are extracted from the original image using the MobileNetV3 lightweight convolutional network. Then, the initial pose and the final pose images are dimensionally superimposed to obtain S0.
[0068] These are fed into the pose transfer module, which uses a two-layer convolutional network. The pose map is then transformed into weights by an activation function, which are multiplied by the feature map to obtain a new feature map guided by the pose. The new feature map is then concatenated with the pose map obtained from the pose transfer module to obtain a new pose map.
[0069] Multiple pose transfer modules continuously update the pose map and feature map, and send the final pose map and feature map to the decoder module for upsampling to obtain the target person image after pose transfer.
[0070] Finally, the generated target pose image and the original target pose image are concatenated and input into the texture discriminator. The concatenated target pose map and the generated target pose image are then fed into the pose discriminator for scoring. The network's overall loss function is defined as the discriminator's loss plus an L1 regularization term.
[0071] To increase the diversity of sample scenes, we selected key public places such as squares, train stations, subway stations, and campuses. Image fusion methods were used to blend the generated target person with the background, employing algorithms such as Poisson image fusion and multi-focus image fusion. Furthermore, to address the diversity of the person's physical appearance, a style transfer GAN network was used to modify the texture of clothing or implement clothing transformations.
[0072] For abnormal behavior recognition, a CNN-LSTM spatiotemporal network is used, combining the powerful spatial feature extraction capabilities of convolutional networks with the excellent temporal information representation capabilities of LSTM networks. The training video frame size is adjusted to 224×224 resolution. For video segment samples, the sequence images are fed into a MobileNetV3-Small network, which combines the depthwise separable convolutions of MobileNetV1, the inverse residual structure with linear bottlenecks of MobileNetV2, and the lightweight attention model based on squeeze and excitation structures of MnasNet. The feature vectors obtained from each sequence frame are input into LSTM units, with the number of LSTM units set to 256. The final prediction result is used as the abnormal behavior classification result. For segments extracted within a certain time period, the mode of the classification results is taken as the classification result for the current behavior.
[0073] The system utilizes regional kinetic energy algorithms in smart cameras to provide initial video screening. This reduces the number of videos requiring cloud-based judgment, improving detection efficiency and hardware utilization. A pose transfer-based approach generates new anomalous behavior sample data, addressing the challenge of acquiring relevant data, rapidly expanding the sample size, and reducing the probability of poor generalization due to model overfitting. Furthermore, customized behavioral pose design increases sample diversity, making it more closely resemble changes in complex real-world scenarios and expanding the scope of behavioral research.
[0074] Figure 4 This is a schematic diagram of the abnormal scene recognition device provided by the present invention, as shown below. Figure 4 As shown, this application embodiment provides an abnormal scene recognition device, which includes a determining module 401 and a sending module 402, wherein:
[0075] The determination module 401 is used to determine the end-limb kinetic energy of the target pedestrian in the target image; the sending module 402 is used to send the target image to the pose estimation network for secondary screening when the end-limb kinetic energy exceeds the target threshold.
[0076] Optionally, the determining module includes a first extraction submodule, a tracking submodule, a segmentation submodule, a first determining submodule, and a second determining submodule;
[0077] The first extraction submodule is used to extract the target pedestrian in the target image;
[0078] The tracking submodule is used to track the target object;
[0079] The segmentation submodule is used to divide the target region into N blocks of different scales based on the longest side, and to add N blocks to the target edge;
[0080] The first determining submodule is used to determine the multi-scale directional relative gradient of each block within the target region; the multi-scale directional relative gradient of the block is used to characterize the block's kinetic energy;
[0081] The second determining submodule is used to determine the end-limb kinetic energy of the target pedestrian based on the multi-scale directional relative gradient of each block.
[0082] Optionally, the sending module includes a second extraction submodule, a generation submodule, and a screening submodule:
[0083] The second extraction submodule is used to extract features from the target image to obtain a feature map;
[0084] The generation submodule is used to generate a pose graph of the abnormal behavior sequence through a pose estimation network;
[0085] The screening submodule is used to perform secondary screening based on the feature map and the pose map.
[0086] Optionally, the screening submodule includes an upsampling unit and an identification unit;
[0087] The upsampling unit is used to send the pose map and feature map to the decoder for upsampling to obtain the target person image after pose transfer.
[0088] The recognition unit is used to identify abnormal behaviors in the target person image through a pose discriminator.
[0089] The abnormal scene recognition device provided in this application embodiment can be used to execute the methods described in the above-mentioned corresponding embodiments. The specific steps of executing the methods described in the above-mentioned corresponding embodiments by the device provided in this embodiment are the same as those in the above-mentioned corresponding embodiments, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail.
[0090] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute an abnormal scene recognition method, which includes:
[0091] Determine the kinetic energy of the extremities of the target pedestrian in the target image;
[0092] If the kinetic energy of the distal limb exceeds the target threshold, the target image is sent to the pose estimation network for secondary screening.
[0093] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0094] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the abnormal scene recognition method provided by the above methods, the method comprising:
[0095] Determine the kinetic energy of the extremities of the target pedestrian in the target image;
[0096] If the kinetic energy of the distal limb exceeds the target threshold, the target image is sent to the pose estimation network for secondary screening.
[0097] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the aforementioned abnormal scene identification methods, the method comprising:
[0098] Determine the kinetic energy of the extremities of the target pedestrian in the target image;
[0099] If the kinetic energy of the distal limb exceeds the target threshold, the target image is sent to the pose estimation network for secondary screening.
[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An abnormal scene recognition method, characterized in that, include: Determine the kinetic energy of the extremities of the target pedestrian in the target image; If the kinetic energy of the extremity exceeds the target threshold, the target image is sent to the pose estimation network for secondary screening. Determining the extremity kinetic energy of the target pedestrian in the target image includes: Extract the target pedestrian from the target image; The target pedestrians are tracked to obtain time-series images of the same pedestrian ID; Divide the target region into N blocks of different scales based on the longest side, and add N blocks to the target edge; Determine the multi-scale relative gradient of each block within the target region, wherein the multi-scale relative gradient of the block is used to characterize the block's kinetic energy, specifically including: The target image is converted into an HSV image, and average pooling is used to extract the average pixel value for each block in the HSV image. Based on the time series of the time series image, the average pixel value of each block is superimposed in the time dimension. The multi-scale directional gradient is calculated for each superimposed block in the target region. The multi-scale directional gradient of each block is divided by the overall target gradient to obtain the multi-scale relative gradient of each block. The end-limb kinetic energy of the target pedestrian is determined based on the multi-scale directional relative gradient of each block, specifically including: The relative gradients at different scales are weighted using a normal distribution function; the weighted relative gradients at each scale are divided by the total number of scales to obtain the calculation results for each scale; the calculation results for each scale are summed to obtain the end-limb kinetic energy of the target pedestrian.
2. The abnormal scene identification method according to claim 1, characterized in that, The step of sending the target image to the pose estimation network for secondary screening includes: Feature extraction is performed on the target image to obtain a feature map; A pose map of anomalous behavior sequences is generated using a pose estimation network; A secondary screening is performed based on the feature map and the pose map.
3. The abnormal scene recognition method according to claim 2, characterized in that, The secondary screening based on the feature map and the pose map includes: The pose map and feature map are sent to the decoder for upsampling to obtain the target person image after pose transfer; Abnormal behaviors in the target person image are identified by a pose discriminator.
4. An abnormal scene recognition device, characterized in that, include: The determination module is used to determine the kinetic energy of the extremities of a target pedestrian in the target image; The sending module is used to send the target image to the pose estimation network for secondary screening when the kinetic energy of the end limb exceeds the target threshold; The determining module includes a first extraction submodule, a tracking submodule, a segmentation submodule, a first determining submodule, and a second determining submodule; The first extraction submodule is used to extract the target pedestrian in the target image; The tracking submodule is used to track the target pedestrian and obtain a time-series image of the same pedestrian ID; The segmentation submodule is used to divide the target region into N blocks of different scales based on the longest side, and to add N blocks to the target edge; The first determining submodule is used to determine the multi-scale relative gradient of each block within the target region, wherein the multi-scale relative gradient of the block is used to characterize the block's kinetic energy, specifically including: The target image is converted into an HSV image, and average pooling is used to extract the average pixel value for each block in the HSV image. Based on the time series of the time series image, the average pixel value of each block is superimposed in the time dimension. The multi-scale directional gradient is calculated for each superimposed block in the target region. The multi-scale directional gradient of each block is divided by the overall target gradient to obtain the multi-scale relative gradient of each block. The second determining submodule is used to determine the end-limb kinetic energy of the target pedestrian based on the multi-scale directional relative gradient of each block, specifically including: The relative gradients at different scales are weighted using a normal distribution function; the weighted relative gradients at each scale are divided by the total number of scales to obtain the calculation results for each scale; the calculation results for each scale are summed to obtain the end-limb kinetic energy of the target pedestrian.
5. The abnormal scene recognition device according to claim 4, characterized in that, The sending module includes a second extraction submodule, a generation submodule, and a screening submodule: The second extraction submodule is used to extract features from the target image to obtain a feature map; The generation submodule is used to generate a pose graph of the abnormal behavior sequence through a pose estimation network; The screening submodule is used to perform secondary screening based on the feature map and the pose map.
6. The abnormal scene recognition device according to claim 5, characterized in that, The screening submodule includes an upsampling unit and an identification unit; The upsampling unit is used to send the pose map and feature map to the decoder for upsampling to obtain the target person image after pose transfer. The recognition unit is used to identify abnormal behaviors in the target person image through a pose discriminator.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the abnormal scene recognition method as described in any one of claims 1 to 3.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the abnormal scene recognition method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Behavior analysis method based on human body key point time sequence, storage medium and equipment
CN111079536A
Subway door-punching behavior detection method and system based on neural network
CN111985453A
Method, device and equipment for recognizing action and storage medium
CN112784765A