Community-oriented abnormal behavior identification method and related equipment

Through feature enhancement processing network and adaptive grayscale conversion network, the monitoring images are processed, which solves the accuracy of abnormal behavior recognition in community monitoring and realizes efficient automatic recognition of abnormal behavior.

CN120279476APending Publication Date: 2025-07-08CHINA TOWER CO LTD GUANGXI ZHUANG AUTONOMOUS REGION BRANCH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510242381.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, abnormal behavior recognition in community monitoring is easily affected by subjective factors of manual observation and is difficult to accurately identify.

Method used

The feature enhancement processing network is used to perform feature enhancement processing on the monitoring image, and the low-level visual features and high-level semantic features are fused through the double-feed feature fusion module. The image feature recognition is improved by combining the adaptive grayscale conversion network, and the image feature recognition model is finally input to identify the abnormal behavior recognition model.

Benefits of technology

It improves the accuracy of identification of abnormal behaviors in the monitoring image, can accurately represent key behavior characteristics, reduce manual intervention, and improve recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279476A_ABST
    Figure CN120279476A_ABST
Patent Text Reader

Abstract

The invention discloses a community-oriented abnormal behavior identification method, and aims to solve the problem that abnormal behaviors in monitoring are difficult to accurately identify in the prior art. The method comprises the following steps: acquiring a monitoring image of a target area in a target community; inputting the monitoring image into a feature enhancement processing network for feature enhancement processing to obtain an enhanced monitoring image feature; the feature enhancement processing network comprises a double-fed feature fusion module, and the double-fed feature fusion module is used for carrying out bidirectional interactive fusion on low-level visual features and high-level semantic features of the monitoring image through two gating mechanisms to obtain enhanced monitoring image features; and inputting the enhanced monitoring image features into the trained abnormal behavior recognition model for abnormal behavior recognition, and outputting a behavior classification result and a confidence score. The invention further discloses a community-oriented abnormal behavior recognition device, electronic equipment, a computer readable storage medium and a computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular to an abnormal behavior recognition method and related devices for communities. Background Art

[0002] Currently, most communities usually use traditional monitoring combined with manual viewing of videos to identify abnormal behaviors in the community. However, manual observation is easily affected by the subjective factors of the observers and may be difficult to accurately identify abnormal behaviors in the monitoring.

[0003] Based on this, how to accurately identify abnormal behaviors in the monitoring is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0004] Embodiments of this application provide an abnormal behavior recognition method for communities to solve the problem in the prior art that it is difficult to accurately identify abnormal behaviors in the monitoring.

[0005] Embodiments of this application also provide an abnormal behavior recognition device for communities, an electronic device, a computer-readable storage medium, and a computer program product.

[0006] Embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, this application provides an abnormal behavior recognition method for communities, including:

[0008] Obtain the monitoring images of the target area in the target community;

[0009] Input the monitoring images into a feature enhancement processing network for feature enhancement processing to obtain enhanced monitoring image features;

[0010] The feature enhancement processing network includes a double-feed feature fusion module, and the double-feed feature fusion module is used to perform bidirectional interactive fusion on the low-level visual features and high-level semantic features of the monitoring images through two gating mechanisms to obtain enhanced monitoring image features;

[0011] The low-level visual features include the color information, texture information, and edge information of the monitoring images; the high-level semantic features include the active target objects in the monitoring images and the motion trajectory information of the active target objects;

[0012] Among them, performing bidirectional interactive fusion on the low-level visual features and high-level semantic features of the monitoring images through two gating mechanisms includes:

[0013] Input the high-level semantic features into the first gating mechanism of the two gating mechanisms to perform weighted fusion on the low-level visual features;

[0014] Input the low-level visual features into the second gating mechanism among the two gating mechanisms to perform weighted fusion on the high-level semantic features;

[0015] Fuse the low-level visual features and high-level semantic features after weighted fusion by the two gating mechanisms through feature concatenation to obtain enhanced surveillance image features;

[0016] Input the enhanced surveillance image features into the trained abnormal behavior recognition model for abnormal behavior recognition, and output the behavior classification result and confidence score.

[0017] Optionally, the feature enhancement processing network includes an encoder and a dual-feed feature fusion decoder, and the dual-feed feature fusion decoder includes a dual-feed feature fusion module;

[0018] The encoder is used to obtain the low-level visual features of the surveillance image;

[0019] The dual-feed feature fusion decoder is used to obtain the high-level semantic features of the surveillance image.

[0020] Optionally, the encoder includes four convolutional layers, and among the four convolutional layers, there are three first convolutional layers with a convolution kernel of 3×3 and one second convolutional layer with a convolution kernel of 1×1.

[0021] Optionally, the dual-feed feature fusion decoder includes four dual-feed feature fusion modules and five convolutional layers. Among the five convolutional layers, there is one third convolutional layer with a convolution kernel of 1×1, one fourth convolutional layer with a convolution kernel of 4×2, and three fifth convolutional layers with a convolution kernel of 4×1.

[0022] Optionally, before inputting the surveillance image into the feature enhancement processing network for feature enhancement processing, the method further includes:

[0023] Perform adaptive gray conversion on the surveillance image through an adaptive gray conversion network to obtain the sharpened gray image features of the surveillance image; the adaptive gray conversion network includes a second encoder, a bottleneck layer, and a second decoder; the image feature recognition degree of the sharpened gray image features is greater than or equal to a preset image feature recognition degree threshold;

[0024] The second encoder is used to map the surveillance image to a high-level feature space to obtain the high-level feature space representation of the surveillance image;

[0025] The bottleneck layer is used to restore the gray features of the surveillance image based on the high-level feature space representation;

[0026] The second decoder is used to map the gray features restored by the bottleneck layer to the target recognition level that can be extracted by the abnormal behavior recognition model to obtain the sharpened gray image features.

[0027] Optionally, the bottleneck layer includes four multi-receptive field residual aggregation modules and one channel-spatial attention module; among them, the multi-receptive field residual aggregation module is used to obtain the image features of the surveillance image in different receptive fields;

[0028] The channel-spatial attention module is used to perform channel attention and spatial attention on the image features of the surveillance image in different receptive fields respectively, so as to recover the gray-scale features of the surveillance image from the image features of the surveillance image in different receptive fields.

[0029] Optionally, restoring the gray-scale features of the surveillance image based on the high-level feature space representation includes:

[0030] Performing channel number compression processing on the high-level feature space representation to obtain the compressed features of the high-level feature space representation;

[0031] According to the compressed features, obtain the image features of the surveillance image in different receptive fields through the multi-receptive field residual aggregation module in the bottleneck layer;

[0032] According to the image features of the surveillance image in different receptive fields, perform weighted processing in the channel dimension and the spatial dimension respectively through the channel-spatial attention module in the bottleneck layer to restore the gray-scale features of the surveillance image.

[0033] Optionally, the multi-receptive field residual aggregation module includes four cascaded dilated convolutional layers, and the dilation rate of each dilated convolutional layer is different.

[0034] Optionally, the second encoder includes three sixth convolutional layers with a stride of 2 and a convolution kernel of 3×3.

[0035] Optionally, the second decoder includes three transposed convolutional layers with a stride of 2 and a convolution kernel of 3×3.

[0036] Optionally, the method further includes:

[0037] When the behavior classification result is abnormal and / or the confidence score is less than the preset confidence score threshold, start the corresponding disposal measures according to the behavior classification result and the confidence score.

[0038] In a second aspect, the present application provides a community-oriented abnormal behavior recognition device, including a surveillance image acquisition module, a feature enhancement processing module, and an abnormal behavior recognition module, wherein:

[0039] The surveillance image acquisition module is used to acquire the surveillance image of the target area in the target community;

[0040] The feature enhancement processing module is used to input the surveillance image into the feature enhancement processing network for feature enhancement processing to obtain the enhanced surveillance image features;

[0041] The feature enhancement processing network includes a double-feed feature fusion module, which is used to perform two-way interactive fusion on the low-level visual features and high-level semantic features of the surveillance image through two gating mechanisms to obtain enhanced surveillance image features;

[0042] The low-level visual features include the color information, texture information, and edge information of the surveillance image; the high-level semantic features include the moving target objects in the surveillance image and the motion trajectory information of the moving target objects;

[0043] Among them, the double-feed feature fusion module is specifically used for:

[0044] Input the high-level semantic features into the first gating mechanism of the two gating mechanisms to perform weighted fusion on the low-level visual features;

[0045] Input the low-level visual features into the second gating mechanism of the two gating mechanisms to perform weighted fusion on the high-level semantic features;

[0046] Fuse the low-level visual features and high-level semantic features after weighted fusion by the two gating mechanisms through feature splicing to obtain enhanced surveillance image features;

[0047] The abnormal behavior recognition module is used to input the enhanced surveillance image features into the trained abnormal behavior recognition model for abnormal behavior recognition, and output the behavior classification result and confidence score.

[0048] Optionally, the feature enhancement processing network includes an encoder and a double-feed feature fusion decoder, and the double-feed feature fusion decoder includes a double-feed feature fusion module;

[0049] The encoder is used to obtain the low-level visual features of the surveillance image;

[0050] The double-feed feature fusion decoder is used to obtain the high-level semantic features of the surveillance image.

[0051] Optionally, the encoder includes four convolutional layers, and among the four convolutional layers, there are three first convolutional layers with a convolutional kernel of 3×3 and one second convolutional layer with a convolutional kernel of 1×1.

[0052] Optionally, the double-feed feature fusion decoder includes four double-feed feature fusion modules and five convolutional layers. Among the five convolutional layers, there is one third convolutional layer with a convolutional kernel of 1×1, one fourth convolutional layer with a convolutional kernel of 4×2, and three fifth convolutional layers with a convolutional kernel of 4×1.

[0053] Optionally, before inputting the surveillance image into the feature enhancement processing network for feature enhancement processing, the method further includes:

[0054] The surveillance image is adaptively grayscale-converted through an adaptive grayscale conversion network to obtain the sharpened grayscale image features of the surveillance image; the adaptive grayscale conversion network includes a second encoder, a bottleneck layer, and a second decoder; the image feature recognition degree of the sharpened grayscale image features is greater than or equal to a preset image feature recognition degree threshold value;

[0055] The second encoder is used to map the surveillance image to a high-level feature space to obtain a high-level feature space representation of the surveillance image;

[0056] The bottleneck layer is used to restore the grayscale features of the surveillance image based on the high-level feature space representation;

[0057] The second decoder is used to map the grayscale features restored by the bottleneck layer to the target recognition level that can be extracted by the abnormal behavior recognition model to obtain the sharpened grayscale image features.

[0058] Optionally, the bottleneck layer includes four multi-receptive field residual aggregation modules and one channel-spatial attention module; among them, the multi-receptive field residual aggregation module is used to obtain the image features of the surveillance image in different receptive fields;

[0059] The channel-spatial attention module is used to perform channel attention and spatial attention on the image features of the surveillance image in different receptive fields respectively, so as to restore the grayscale features of the surveillance image from the image features of the surveillance image in different receptive fields.

[0060] Optionally, the bottleneck layer is specifically used for:

[0061] Performing channel number compression processing on the high-level feature space representation to obtain a compressed feature of the high-level feature space representation;

[0062] According to the compressed feature, the multi-receptive field residual aggregation module in the bottleneck layer is used to obtain the image features of the surveillance image in different receptive fields;

[0063] According to the image features of the surveillance image in different receptive fields, the channel-spatial attention module in the bottleneck layer performs weighted processing in the channel dimension and the spatial dimension respectively to restore the grayscale features of the surveillance image.

[0064] Optionally, the multi-receptive field residual aggregation module includes four cascaded dilated convolutional layers, and the dilation rate of each dilated convolutional layer is different.

[0065] Optionally, the second encoder includes three sixth convolutional layers with a stride of 2 and a convolution kernel of 3×3.

[0066] Optionally, the second decoder includes three transposed convolutional layers with a stride of 2 and a convolution kernel of 3×3.

[0067] Optionally, the device is further used for:

[0068] When the behavior classification result is abnormal and / or the confidence score is less than the preset confidence score threshold, according to the behavior classification result and the confidence score, initiate the corresponding disposal measures corresponding to the classification result and the confidence score.

[0069] In a third aspect, the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, it implements the steps of the above-mentioned community-oriented abnormal behavior recognition method.

[0070] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above-mentioned community-oriented abnormal behavior recognition method.

[0071] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned community-oriented abnormal behavior recognition method.

[0072] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:

[0073] By using the method provided in the embodiments of the present application, before performing abnormal behavior recognition through an abnormal behavior recognition model, the monitored image can be first input into a feature enhancement processing network for feature enhancement processing. Since the feature enhancement processing network includes a dual-feed feature fusion module, and the dual-feed feature fusion module can perform bidirectional interactive fusion on the low-level visual features and high-level semantic features of the monitored image through two gating mechanisms, so as to ensure that the enhanced monitored image features can accurately represent the key behavior features related to abnormal behavior in the monitored image, and further enable the abnormal behavior recognition model to accurately recognize the abnormal behavior in the monitoring based on the enhanced monitored image features. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0075] Figure 1a It is a schematic flowchart of the implementation of a community-oriented abnormal behavior recognition method provided by an embodiment of the present application;

[0076] Figure 1b It is a schematic network structure diagram of a dual-feed feature fusion module provided by an embodiment of the present application;

[0077] Figure 1cSchematic diagram of the network structure of a feature enhancement processing network provided by an embodiment of the present application;

[0078] Figure 2a Schematic diagram of the network structure of an adaptive grayscale conversion network provided by an embodiment of the present application;

[0079] Figure 2b Schematic diagram of the implementation process of a method for recovering the grayscale features of a surveillance image based on high-level feature space representation provided by an embodiment of the present application;

[0080] Figure 3 Schematic diagram of the specific structure of an abnormal behavior recognition device for communities provided by an embodiment of the present application;

[0081] Figure 4 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0082] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0083] The following will detail the technical solutions provided by each embodiment of the present application in conjunction with the drawings.

[0084] Embodiment 1

[0085] To solve the problem in the prior art that it is difficult to accurately identify abnormal behaviors in surveillance, an embodiment of the present application provides an abnormal behavior recognition method for communities.

[0086] The execution subject of this method can be various types of computing devices, or it can be an application program or application (Application, APP) installed on a computing device. The computing device can be, for example, a user terminal such as a mobile phone, a tablet computer, a smart wearable device, etc., or a server, etc.

[0087] For ease of description, in this embodiment of the present application, the execution subject of this method is taken as an example of a server to introduce this method. Those skilled in the art can understand that taking this server as an example to introduce the method in this embodiment of the present application is only an illustrative explanation and does not limit the protection scope of the corresponding claims of this solution.

[0088] Specifically, the implementation process of this method provided by this embodiment of the present application is as Figure 1a shown and includes the following steps:

[0089] Step 11: Obtain the surveillance images of the target area within the target community.

[0090] The target community refers to the community where abnormal behavior detection is to be carried out.

[0091] The target area includes community public areas or public surveillance areas within the target community where people's activities are relatively frequent and security issues are likely to be involved. For example, the target area may include, but is not limited to, the entrances and exits, corridors, parking lots, garbage disposal points, etc. of the target community, which are community public areas with frequent people's activities within the target community.

[0092] In the embodiments of the present application, the surveillance images of the target area within the target community can be obtained through the existing surveillance cameras within the target community.

[0093] Step 12: Input the surveillance images into the feature enhancement processing network for feature enhancement processing to obtain the enhanced surveillance image features.

[0094] The feature enhancement processing network includes a double-feed feature fusion module. The double-feed feature fusion module is used to perform two-way interactive fusion on the low-level visual features and high-level semantic features of the surveillance images through two gating mechanisms to obtain the enhanced surveillance image features.

[0095] The low-level visual features include the color information, texture information, and edge information of the surveillance images.

[0096] The high-level semantic features include the moving target objects in the surveillance images and the motion trajectory information of the moving target objects.

[0097] As Figure 1b shown, it is a schematic diagram of the network structure of a double-feed feature fusion module provided by the embodiments of the present application. As can be seen from Figure 1b , the double-feed feature fusion module includes a low-level visual feature input unit, a high-level semantic feature input unit, a gating mechanism unit, a multiplication operation unit, an addition operation unit, and a channel dimension splicing unit. Among them, the low-level visual feature input unit is used to receive the low-level visual features of the surveillance images, such as the color information, texture information, and edge information of the surveillance images, etc.

[0098] The high-level semantic feature input unit is used to receive the high-level semantic features of the surveillance images, such as the moving target objects in the surveillance images and the motion trajectory information of the moving target objects, etc.

[0099] The gating mechanism unit is used to perform weighted adjustment on the input features of the dual-feed feature fusion module (such as low-level visual features and high-level semantic features) according to their differences, so as to selectively retain or ignore some information in the input features of the dual-feed feature fusion module, thereby making the fusion of the input features of the dual-feed feature fusion module more accurate.

[0100] The multiplication operation unit is used to perform a multiplication operation on the low-level visual features and high-level semantic features of the monitoring image during the two-way interactive fusion process, so as to retain the important information in the low-level visual features and high-level semantic features of the monitoring image and avoid the loss of important information.

[0101] The addition operation unit is used to perform an addition operation on the low-level visual features and high-level semantic features of the monitoring image during the two-way interactive fusion process to ensure that the features after the two-way interactive fusion can simultaneously contain the low-level visual features and high-level semantic features and achieve an effective combination of the two.

[0102] The channel dimension splicing unit is used to splice the results of the two-way interactive fusion of the low-level visual features and high-level semantic features in the channel direction to form a multi-dimensional feature representation, ensuring the diversity and richness of the features of the two-way interactive fusion.

[0103] Optionally, the specific implementation process of the two-way interactive fusion of the low-level visual features and high-level semantic features of the monitoring image through two gating mechanisms is as follows:

[0104] (1) Input the high-level semantic features into the first gating mechanism of the two gating mechanisms to perform weighted fusion on the low-level visual features;

[0105] In the embodiment of the present application, inputting the high-level semantic features into the first gating mechanism of the two gating mechanisms to perform weighted fusion on the low-level visual features can highlight the key detail information of the monitoring image.

[0106] (2) Input the low-level visual features into the second gating mechanism of the two gating mechanisms to perform weighted fusion on the high-level semantic features;

[0107] In the embodiment of the present application, inputting the low-level visual features into the second gating mechanism of the two gating mechanisms to perform weighted fusion on the high-level semantic features can fuse the key detail information and global semantic information and achieve an effective combination of the two.

[0108] (3) Fuse the low-level visual features and high-level semantic features after being weighted and fused by the two gating mechanisms through feature splicing to obtain the enhanced monitoring image features.

[0109] It should be noted that the above steps (1) and (2) can be executed simultaneously or in sequence one after another.

[0110] In an alternative embodiment, before the two gating mechanisms perform bidirectional interactive fusion on the low-level visual features and high-level semantic features of the surveillance image, the dual-feed feature fusion module can also perform gating mechanism fusion learning through the two gating mechanisms based on the low-level visual features and high-level semantic features of the surveillance image in the following manner, so as to obtain the gating weights of the first gating mechanism and the gating weights of the second gating mechanism:

[0111] g en = α(Conv 1×1 (Y en ))

[0112] where g en represents the gating weight of the first gating mechanism; Conv 1×1 represents a convolutional layer with a kernel size of 1; Y en represents the low-level visual features; α represents a preset initial gating weight.

[0113] g de = α(Conv 1×1 (Y de ))

[0114] where g de represents the gating weight of the second gating mechanism; Y de represents the high-level semantic features.

[0115] After obtaining the gating weights of the first gating mechanism and the gating weights of the second gating mechanism, the low-level visual features and high-level semantic features of the surveillance image can be bidirectionally interactively fused in the following manner:

[0116]

[0117] where Y' en represents the low-level visual features after bidirectional interactive fusion; Y' de represents the high-level semantic features after bidirectional interactive fusion. represents the addition operation; ⊙ represents the multiplication operation.

[0118] Finally, Y' en and Y' de are fused by feature concatenation to obtain the enhanced surveillance image features.

[0119] Optionally, the feature enhancement processing network further includes an encoder and a dual-feed feature fusion decoder, and the dual-feed feature fusion decoder includes a dual-feed feature fusion module; the encoder is used to obtain low-level visual features of the monitoring image; the dual-feed feature fusion decoder is used to obtain high-level semantic features of the monitoring image.

[0120] Optionally, the encoder includes four convolutional layers, and among the four convolutional layers, there are three first convolutional layers with a convolutional kernel of 3×3 and one second convolutional layer with a convolutional kernel of 1×1.

[0121] Optionally, the dual-feed feature fusion decoder includes four dual-feed feature fusion modules and five convolutional layers. Among the five convolutional layers, there is one third convolutional layer with a convolutional kernel of 1×1, one fourth convolutional layer with a convolutional kernel of 4×2, and three fifth convolutional layers with a convolutional kernel of 4×1.

[0122] As Figure 1c shown, it is a schematic diagram of the network structure of a feature enhancement processing network provided by an embodiment of the present application. As can be seen from Figure 1c , the feature enhancement processing network includes two parts: an encoder and a dual-feed feature fusion decoder. Among them, the encoder includes four convolutional layers, and among the four convolutional layers, there are three first convolutional layers with a convolutional kernel of 3×3 and one second convolutional layer with a convolutional kernel of 1×1. The dual-feed feature fusion decoder includes four dual-feed feature fusion modules (that is, the dual-feed feature fusion modules 1 to 4 in Figure 1c ) and five convolutional layers. Among them, the network structures of the four dual-feed feature fusion modules are exactly the same ( Figure 1c shows the detailed network structure taking the dual-feed feature fusion module 1 as an example, and the detailed network structures of the remaining dual-feed feature fusion modules 2 to 3 are exactly the same as that of the dual-feed feature fusion module 1), and among the five convolutional layers, there is one third convolutional layer with a convolutional kernel of 1×1, one fourth convolutional layer with a convolutional kernel of 4×2, and three fifth convolutional layers with a convolutional kernel of 4×1.

[0123] It should be noted that Figure 1c the schematic diagram of the network structure of the feature enhancement processing network shown is only an exemplary illustration of an embodiment of the present application and does not impose any limitation on the embodiments of the present application.

[0124] Step 13, input the enhanced monitoring image features into the trained abnormal behavior recognition model for abnormal behavior recognition, and output the behavior classification result and the confidence score.

[0125] Optionally, the behavior classification result may include two cases: behavior abnormal and behavior normal.

[0126] In an alternative embodiment, when the behavior classification result is abnormal and / or the confidence score is less than a preset confidence score threshold, corresponding handling measures can be initiated based on the behavior classification result and the confidence score.

[0127] (1) Optionally, if the confidence score is greater than a preset value and the behavior classification result is abnormal behavior, a warning message can be sent to the security personnel closest to the location where the abnormal behavior occurred through push methods such as text messages or community security applications. The warning message can include the location where the abnormal behavior occurred, the appearance characteristics of the involved person, etc. For example, the content of the warning message can be as follows: A man wearing a black coat and a hat appeared next to the trash can in Area A of the community.

[0128] Optionally, when sending the warning message, the optimal route for the security personnel to the location where the abnormal behavior occurred can also be planned through the community security application. Moreover, the access control system at the entrance and exit of the target community can be remotely controlled to temporarily prevent abnormal personnel from entering and leaving, and the community broadcast can be activated to remind the surrounding residents to pay attention to the nearby situation.

[0129] (2) Optionally, if the confidence score is less than or equal to the threshold and the behavior classification result is abnormal behavior, a notification message can be sent to the security personnel closest to the location where the abnormal behavior occurred through push methods such as text messages or community security applications to notify the security personnel to go to the scene for verification.

[0130] (3) If the confidence score is less than or equal to the threshold and the behavior classification result is normal behavior, no corresponding handling measures will be taken temporarily, and supervision will continue.

[0131] Using the method provided in the embodiment of the present application, before identifying abnormal behavior through the abnormal behavior recognition model, the surveillance image can be first input into the feature enhancement processing network for feature enhancement processing. Since the feature enhancement processing network includes a dual-feed feature fusion module, and the dual-feed feature fusion module can perform two-way interactive fusion on the low-level visual features and high-level semantic features of the surveillance image through two gating mechanisms, so as to ensure that the enhanced surveillance image features can accurately represent the key behavior features related to abnormal behavior in the surveillance image, and further enable the abnormal behavior in the surveillance to be accurately identified based on the enhanced surveillance image features through the abnormal behavior recognition model.

[0132] Embodiment 2

[0133] In order to further improve the accuracy of the abnormal behavior recognition model in recognizing abnormal behaviors in surveillance images, in the embodiments of the present application, before inputting the surveillance image into the feature enhancement processing network for feature enhancement processing, the surveillance image can be further adaptively grayscale-converted through an adaptive grayscale conversion network to obtain a sharpened grayscale image feature of the surveillance image. Among them, the image feature recognition degree of the sharpened grayscale image feature is greater than or equal to the preset image feature recognition degree threshold.

[0134] Optionally, considering that usually, the color display space of the surveillance image is usually in the RGB space, and for the abnormal behavior recognition model, it is usually not easy to capture clear feature information in the surveillance image in this color space. Therefore, to avoid this problem, in the embodiments of the present application, before adaptively grayscale-converting the surveillance image through the adaptive grayscale conversion network, the surveillance image can be first color space-converted in the following manner:

[0135] C in = C gt (1 - M)

[0136] Among them, C in represents the surveillance image after color space conversion; C gt represents the surveillance image; M represents a binary mask; 1 represents the image area in the surveillance image that needs to be color space-converted.

[0137] As Figure 2a shown, it is a schematic diagram of the network structure of an adaptive grayscale conversion network provided by the embodiments of the present application. The adaptive grayscale conversion network includes a second encoder, a bottleneck layer, and a second decoder; the second encoder is used to map the surveillance image to a high-level feature space to obtain a high-level feature space representation of the surveillance image. The bottleneck layer is used to restore the grayscale feature of the surveillance image based on the high-level feature space representation. The second decoder is used to map the grayscale feature restored by the bottleneck layer to the target recognition level that can be extracted by the abnormal behavior recognition model to obtain a sharpened grayscale image feature.

[0138] Optionally, the second encoder includes three sixth convolutional layers with a stride of 2 and a convolutional kernel of 3×3.

[0139] Optionally, the second decoder includes three transposed convolutional layers with a stride of 2 and a convolutional kernel of 3×3.

[0140] Optionally, the bottleneck layer includes four multi-receptive field residual aggregation modules (i.e., Figure 2a the multi-receptive field residual aggregation modules 1 to 4 in

[0141] Among them, the channel-spatial attention module is used to perform channel attention and spatial attention on the image features of the surveillance image in different receptive fields respectively, so as to recover the grayscale features of the surveillance image from the image features of the surveillance image in different receptive fields.

[0142] Optionally, as Figure 2a shown, the multi-receptive field residual aggregation module provided by the embodiment of the present application may include four cascaded dilated convolutional layers and two convolutional layers, wherein the dilation rates of each dilated convolutional layer are different. And, as Figure 2a shown, one of the two convolutional layers is connected to the output of the second encoder and is used to receive the output result of the second encoder. The other convolutional layer is used to receive the result of concatenating the four cascaded dilated convolutional layers along the channel dimension.

[0143] In the embodiment of the present application, the stride and the convolutional kernel size of the two convolutional layers in the multi-receptive field residual aggregation module are not limited and can be set according to actual needs.

[0144] In the embodiment of the present application, by stacking multiple multi-receptive field residual aggregation modules, it can be ensured that the context information in the surveillance image can be repeatedly aggregated to the area that needs to perform color space transformation, which is beneficial to filling the surveillance image with a large color change area.

[0145] In the embodiment of the present application, each of the four cascaded dilated convolutional layers can collect the receptive field features with different color information, which are specifically expressed as follows:

[0146]

[0147] X r+1 represents the output representation of the (r + 1)-th dilated convolutional layer; X r represents the output representation of the previous dilated convolutional layer of the (r + 1)-th dilated convolutional layer; ReLU represents the rectified linear unit, which is used to further expand the receptive field of the adaptive grayscale conversion network and is beneficial to the diverse acquisition of information. represents the dilated convolutional layer with a dilation rate of 2 r .

[0148] Optionally, the dilation rates of the four cascaded dilated convolutional layers in the multi-receptive field residual aggregation module can be set to 2 0 , 2 1 , 2 2 , 2 3 .

[0149] Optionally, as Figure 2b shown, the bottleneck layer recovers the grayscale features of the surveillance image based on the high-level feature space representation, and specifically includes the following steps:

[0150] Step 21: Perform channel number compression processing based on the high-level feature space representation to obtain the compressed features of the high-level feature space representation.

[0151] In the embodiment of the present application, the channel number of the high-level feature space representation of the monitoring image can be compressed to one-fourth of the original through a convolutional layer, so as to obtain the compressed features of the high-level feature space representation.

[0152] Step 22: According to the compressed features, obtain the image features of the monitoring image in different receptive fields through the multi-receptive field residual aggregation module in the bottleneck layer.

[0153] In the embodiment of the present application, after the compressed features are processed through four cascaded dilated convolutional layers of the multi-receptive field residual aggregation module in the bottleneck layer, 4 features with different receptive fields can be obtained, denoted as X1 to X4. Then, X1 to X4 can be concatenated along the channel dimension, and the dimension concatenation result can be integrated into a comprehensive context feature X through a convolutional layer C , and this X C has the same dimension as the high-level feature space representation of the monitoring image. Finally, according to the high-level feature space representation of the monitoring image and the context feature X C , a residual connection is used to add the two to obtain the output of the multi-receptive field residual aggregation module, that is, the image features of the monitoring image in different receptive fields:

[0154]

[0155] where, represents the output of the multi-receptive field residual aggregation module; X represents the high-level feature space representation of the monitoring image.

[0156] Step 23: According to the image features of the monitoring image in different receptive fields, perform weighted processing in the channel dimension and the spatial dimension respectively through the channel-spatial attention module in the bottleneck layer to restore the gray-scale features of the monitoring image.

[0157] In the embodiment of the present application, after obtaining the image features of the monitoring image in different receptive fields, the channel-spatial attention module in the bottleneck layer can first perform weighted processing on the image features in the channel dimension, and the specific process is as follows:

[0158]

[0159] where, represents the result obtained by performing weighted processing on the image features of the monitoring image in different receptive fields in the channel dimension; ε represents the Sigmoid function; conv 4*1 represents a convolutional layer with a convolution kernel size of 4*1; Avgpool s represents average pooling in the spatial dimension;

[0160] After that, it is possible to further perform weighted processing on the in the spatial dimension, and the specific process is as follows:

[0161]

[0162] Among them, represents the result obtained by performing weighted processing on the in the spatial dimension, that is, the gray-scale feature of the surveillance image; Avgpool c represents average pooling according to the channel dimension.

[0163] By adopting the method provided in the embodiment of the present application, before inputting the surveillance image into the feature enhancement processing network for feature enhancement processing, the surveillance image can be further adaptively gray-scale converted through the adaptive gray-scale conversion network to obtain the clear gray-scale image feature of the surveillance image; then, the clear gray-scale image feature is input into the feature enhancement processing network for feature enhancement processing. On the one hand, since the image feature recognition degree of the clear gray-scale image feature is greater than or equal to the preset image feature recognition degree threshold, it is possible to ensure that the key behavior features in the surveillance image are more prominent and easier to be recognized, thereby improving the recognition accuracy of the subsequent abnormal behavior recognition model for abnormal behaviors; on the other hand, since the feature enhancement processing network includes a double-feed feature fusion module, and the double-feed feature fusion module can perform two-way interactive fusion on the low-level visual features and high-level semantic features of the surveillance image through two gating mechanisms, so as to ensure that the enhanced surveillance image features can accurately represent the key behavior features related to abnormal behaviors in the surveillance image, and further enable the abnormal behavior recognition model to accurately recognize the abnormal behaviors in the surveillance based on the enhanced surveillance image features.

[0164] Embodiment 3

[0165] To solve the problem in the prior art that it is difficult to accurately recognize abnormal behaviors in surveillance, the embodiment of the present application provides a device for recognizing abnormal behaviors for a community. The specific structural schematic diagram of the device is as Figure 3 shown, including a surveillance image acquisition module 31, a feature enhancement processing module 32, and an abnormal behavior recognition module 33. The functions of each module are as follows:

[0166] The surveillance image acquisition module 31 is used to acquire the surveillance image of the target area in the target community;

[0167] The feature enhancement processing module 32 is used to input the surveillance image into the feature enhancement processing network for feature enhancement processing to obtain the enhanced surveillance image features;

[0168] The feature enhancement processing network includes a double-feed feature fusion module, which is used to perform bidirectional interactive fusion on the low-level visual features and high-level semantic features of the surveillance image through two gating mechanisms to obtain enhanced surveillance image features;

[0169] The low-level visual features include the color information, texture information, and edge information of the surveillance image; the high-level semantic features include the active target objects in the surveillance image and the motion trajectory information of the active target objects;

[0170] Among them, the double-feed feature fusion module is specifically used for:

[0171] Input the high-level semantic features into the first gating mechanism of the two gating mechanisms to perform weighted fusion on the low-level visual features;

[0172] Input the low-level visual features into the second gating mechanism of the two gating mechanisms to perform weighted fusion on the high-level semantic features;

[0173] Fuse the low-level visual features and high-level semantic features after being weighted and fused by the two gating mechanisms through feature splicing to obtain enhanced surveillance image features;

[0174] The abnormal behavior recognition module 33 is used to input the enhanced surveillance image features into the trained abnormal behavior recognition model for abnormal behavior recognition, and output the behavior classification result and the confidence score.

[0175] Optionally, the feature enhancement processing network includes an encoder and a double-feed feature fusion decoder, and the double-feed feature fusion decoder includes a double-feed feature fusion module;

[0176] The encoder is used to obtain the low-level visual features of the surveillance image;

[0177] The double-feed feature fusion decoder is used to obtain the high-level semantic features of the surveillance image.

[0178] Optionally, the encoder includes four convolutional layers, and among the four convolutional layers, there are three first convolutional layers with a convolutional kernel of 3×3 and one second convolutional layer with a convolutional kernel of 1×1.

[0179] Optionally, the double-feed feature fusion decoder includes four double-feed feature fusion modules and five convolutional layers. Among the five convolutional layers, there is one third convolutional layer with a convolutional kernel of 1×1, one fourth convolutional layer with a convolutional kernel of 4×2, and three fifth convolutional layers with a convolutional kernel of 4×1.

[0180] Optionally, before inputting the surveillance image into the feature enhancement processing network for feature enhancement processing, the method further includes:

[0181] The surveillance image is adaptively grayscale-converted through an adaptive grayscale conversion network to obtain the sharpened grayscale image features of the surveillance image; the adaptive grayscale conversion network includes a second encoder, a bottleneck layer, and a second decoder; the image feature recognition degree of the sharpened grayscale image features is greater than or equal to a preset image feature recognition degree threshold;

[0182] The second encoder is used to map the surveillance image to a high-level feature space to obtain the high-level feature space representation of the surveillance image;

[0183] The bottleneck layer is used to restore the grayscale features of the surveillance image based on the high-level feature space representation;

[0184] The second decoder is used to map the grayscale features restored by the bottleneck layer to the target recognition level that can be extracted by the abnormal behavior recognition model to obtain the sharpened grayscale image features.

[0185] Optionally, the bottleneck layer includes four multi-receptive field residual aggregation modules and one channel-spatial attention module; among them, the multi-receptive field residual aggregation module is used to obtain the image features of the surveillance image in different receptive fields;

[0186] The channel-spatial attention module is used to perform channel attention and spatial attention on the image features of the surveillance image in different receptive fields respectively, so as to restore the grayscale features of the surveillance image from the image features of the surveillance image in different receptive fields.

[0187] Optionally, the bottleneck layer is specifically used for:

[0188] Performing channel number compression processing on the high-level feature space representation to obtain the compressed features of the high-level feature space representation;

[0189] According to the compressed features, the multi-receptive field residual aggregation module in the bottleneck layer is used to obtain the image features of the surveillance image in different receptive fields;

[0190] According to the image features of the surveillance image in different receptive fields, the channel-spatial attention module in the bottleneck layer performs weighted processing in the channel dimension and the spatial dimension respectively to restore the grayscale features of the surveillance image.

[0191] Optionally, the multi-receptive field residual aggregation module includes four cascaded dilated convolutional layers, and the dilation rate of each dilated convolutional layer is different.

[0192] Optionally, the second encoder includes three sixth convolutional layers with a stride of 2 and a convolution kernel of 3×3.

[0193] Optionally, the second decoder includes three transposed convolutional layers with a stride of 2 and a convolution kernel of 3×3.

[0194] Optionally, the device is further used for:

[0195] When the behavior classification result is abnormal and / or the confidence score is less than the preset confidence score threshold, corresponding handling measures are initiated according to the behavior classification result and the confidence score. By using the device provided in the embodiments of the present application, before abnormal behavior recognition is performed through the abnormal behavior recognition model, the monitoring image can be input into the feature enhancement processing network for feature enhancement processing. Since the feature enhancement processing network includes a dual-feed feature fusion module, and the dual-feed feature fusion module can perform two-way interactive fusion on the low-level visual features and high-level semantic features of the monitoring image through two gating mechanisms, so as to ensure that the enhanced monitoring image features can accurately represent the key behavior features related to abnormal behavior in the monitoring image, and further enable the abnormal behavior in the monitoring to be accurately recognized based on the enhanced monitoring image features through the abnormal behavior recognition model.

[0196] Embodiment 4

[0197] Figure 4 FIG. 4 is a schematic hardware structure diagram of an electronic device for implementing various embodiments of the present application. The electronic device may include a processor 401 and a memory 402 storing computer program instructions. Specifically, the processor 401 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0198] The memory 402 may include a mass storage for data or instructions. By way of example and not limitation, the memory 402 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In a suitable case, the memory 402 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 402 may be internal or external to the electronic device. In a particular embodiment, the memory 402 may be a non-volatile solid-state memory.

[0199] In one embodiment, the memory 402 may be a read only memory (ROM). In one embodiment, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory or a combination of two or more of these.

[0200] The processor 401 reads and executes the computer program instructions stored in the memory 402 to implement any one of the community-oriented abnormal behavior recognition methods in the above embodiments.

[0201] In one example, the electronic device may further include a communication interface 403 and a bus 410. Among them, as Figure 4 shown, the processor 401, the memory 402, and the communication interface 403 are connected through the bus 410 and complete communication with each other.

[0202] The communication interface 403 is mainly used to implement communication between each module, device, unit, and / or device in the embodiments of the present application.

[0203] The bus 410 includes hardware, software, or both, and couples the components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 410 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0204] In addition, in combination with the community-oriented abnormal behavior recognition method in the above embodiments, the embodiments of the present application can provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by the processor, any one of the community-oriented abnormal behavior recognition methods in the above embodiments is implemented.

[0205] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between the steps after understanding the spirit of the present application.

[0206] As described above, the above are only specific embodiments of the examples of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0207] Secondly, those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0208] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0209] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0210] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0211] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0212] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0213] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0214] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0215] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A community-oriented abnormal behavior recognition method, characterized in that, Including: Obtain the surveillance images of the target area within the target community; Input the surveillance images into a feature enhancement processing network for feature enhancement processing to obtain enhanced surveillance image features; The feature enhancement processing network includes a dual-feed feature fusion module, and the dual-feed feature fusion module is used to perform two-way interactive fusion on the low-level visual features and high-level semantic features of the surveillance images through two gating mechanisms to obtain the enhanced surveillance image features; The low-level visual features include the color information, texture information, and edge information of the surveillance images; the high-level semantic features include the active target objects in the surveillance images and the motion trajectory information of the active target objects; Among them, performing two-way interactive fusion on the low-level visual features and high-level semantic features of the surveillance images through two gating mechanisms includes: Input the high-level semantic features into the first gating mechanism of the two gating mechanisms to perform weighted fusion on the low-level visual features; Input the low-level visual features into the second gating mechanism of the two gating mechanisms to perform weighted fusion on the high-level semantic features; Fuse the low-level visual features and high-level semantic features after being weighted and fused by the two gating mechanisms through feature splicing to obtain the enhanced surveillance image features; Input the enhanced surveillance image features into a trained abnormal behavior recognition model for abnormal behavior recognition, and output the behavior classification result and confidence score.

2. The method according to claim 1, wherein The feature enhancement processing network includes an encoder and a dual-feed feature fusion decoder, and the dual-feed feature fusion decoder includes the dual-feed feature fusion module; The encoder is used to obtain the low-level visual features of the surveillance images; The dual-feed feature fusion decoder is used to obtain the high-level semantic features of the surveillance images.

3. The method according to claim 2, wherein The encoder includes four convolutional layers, and among the four convolutional layers, there are three first convolutional layers with a convolutional kernel of 3×3 and one second convolutional layer with a convolutional kernel of 1×1.

4. The method according to claim 2, wherein The dual-feed feature fusion decoder includes four dual-feed feature fusion modules and five convolutional layers, and among the five convolutional layers, there are one third convolutional layer with a convolutional kernel of 1×1, one fourth convolutional layer with a convolutional kernel of 4×2, and three fifth convolutional layers with a convolutional kernel of 4×1.

5. The method according to claim 1, characterized in that Before inputting the surveillance images into the feature enhancement processing network for feature enhancement processing, the method further includes: Perform adaptive grayscale conversion on the surveillance images through an adaptive grayscale conversion network to obtain the sharpened grayscale image features of the surveillance images; the adaptive grayscale conversion network includes a second encoder, a bottleneck layer, and a second decoder; the image feature recognition degree of the sharpened grayscale image features is greater than or equal to a preset image feature recognition degree threshold; The second encoder is used to map the surveillance images to a high-level feature space to obtain the high-level feature space representation of the surveillance images; The bottleneck layer is used to restore the grayscale features of the surveillance images based on the high-level feature space representation; The second decoder is used to map the grayscale feature after the bottleneck layer is restored to the target recognition level that can be extracted by the abnormal behavior recognition model, so as to obtain the clarified grayscale image feature.

6. The method according to claim 5, wherein The bottleneck layer includes four multi-receptive field residual aggregation modules and one channel-spatial attention module; wherein, the multi-receptive field residual aggregation module is used to obtain the image features of the monitoring image in different receptive fields; The channel-spatial attention module is used to perform channel attention and spatial attention on the image features of the monitoring image in different receptive fields respectively, so as to restore the grayscale feature of the monitoring image from the image features of the monitoring image in different receptive fields.

7. The method according to claim 5 or 6, characterized in that, Restoring the grayscale feature of the monitoring image based on the high-level feature space representation includes: Performing channel number compression processing on the high-level feature space representation to obtain the compressed feature of the high-level feature space representation; According to the compressed feature, obtaining the image features of the monitoring image in different receptive fields through the multi-receptive field residual aggregation module in the bottleneck layer; According to the image features of the monitoring image in different receptive fields, performing weighted processing in the channel dimension and the spatial dimension respectively through the channel-spatial attention module in the bottleneck layer to restore the grayscale feature of the monitoring image.

8. The method according to claim 5, 6 or 7, characterized in that The multi-receptive field residual aggregation module includes four cascaded dilated convolutional layers, and the dilation rate of each dilated convolutional layer is different.

9. The method according to claim 5, wherein The second encoder includes three sixth convolutional layers with a stride of 2 and a convolution kernel of 3×3.

10. The method according to claim 5, wherein The second decoder includes three transposed convolutional layers with a stride of 2 and a convolution kernel of 3×3.

11. The method according to claim 1, wherein The method further includes: When the behavior classification result is abnormal and / or the confidence score is less than the preset confidence score threshold, according to the behavior classification result and the confidence score, initiate the disposal measures corresponding to the classification result and the confidence score.

12. An abnormal behavior recognition device for a community, characterized in that, Including a monitoring image acquisition module, a feature enhancement processing module and an abnormal behavior recognition module, wherein: The monitoring image acquisition module is used to acquire the monitoring image of the target area in the target community; The feature enhancement processing module is used to input the monitoring image into the feature enhancement processing network for feature enhancement processing to obtain the enhanced monitoring image feature; The feature enhancement processing network includes a dual-feed feature fusion module, and the dual-feed feature fusion module is used to perform two-way interactive fusion on the low-level visual feature and the high-level semantic feature of the monitoring image through two gating mechanisms to obtain the enhanced monitoring image feature; The low-level visual feature includes the color information, texture information and edge information of the monitoring image; the high-level semantic feature includes the active target object in the monitoring image and the motion trajectory information of the active target object; Among them, the dual-feed feature fusion module is specifically used for: Inputting the high-level semantic feature into the first gating mechanism of the two gating mechanisms to perform weighted fusion on the low-level visual feature; Inputting the low-level visual feature into the second gating mechanism of the two gating mechanisms to perform weighted fusion on the high-level semantic feature; Fuse the low-level visual features and high-level semantic features after weighted fusion by the two gating mechanisms through feature splicing to obtain the enhanced surveillance image features; An abnormal behavior recognition module, configured to input the enhanced surveillance image features into a trained abnormal behavior recognition model for abnormal behavior recognition, and output a behavior classification result and a confidence score.

13. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the community-oriented abnormal behavior recognition method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the community-oriented abnormal behavior recognition method according to any one of claims 1 to 11 are implemented.

15. A computer program product, characterized in that, Including a computer program, and when the computer program is executed by a processor, the steps of the community-oriented abnormal behavior recognition method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Multi-target detection method based on convolutional neural network

    CN112906718A

  • Multispectral image semantic segmentation method based on asymmetric cross fusion

    CN114445442A

  • Target behavior recognition method and device based on audio-visual feature fusion and application

    CN114581749A

  • Intelligent video monitoring method and system based on multi-source data fusion

    CN119169536A

  • In-station anomaly analysis method and system based on multi-modal retrieval enhancement generation and readable medium

    CN119360119A