A fighting behavior detection method, device, electronic device and storage medium
By using area candidate modules and frame difference in the video surveillance system to build a network, the areas where fighting behaviors may occur are pre-selected and the probability of fighting behaviors are calculated, which solves the problem of large amount of combat behavior detection and low efficiency in the prior art, and realizes efficient fighting behavior recognition.
Patent Information
- Application Number
- CN202510031909.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-09
AI Technical Summary
In the prior art, combat behavior detection is computationally expensive and inefficient, and there is a lack of effective solutions.
By obtaining the sequence of continuous image frames in the video, detecting and tracking the human target object, using the area candidate module to preselect the areas where fighting behavior may occur, and constructing the network through the frame difference to calculate the probability of fighting behavior, obtaining the human target object where fighting behavior occurs.
The calculation amount of fighting behavior detection is reduced, the detection efficiency is improved, and real-time and efficient identification of strong timing fighting behaviors is achieved.
Smart Images

Figure CN119479078B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of visual behavior detection, and in particular to a fighting behavior detection method, device, electronic device and storage medium. Background Art
[0002] In the field of video surveillance and security, more and more industries or scenarios require the recognition of human behavior. According to the length of time required for recognition, it can be roughly divided into two categories. One category is weak temporal behavior, such as smoking, making phone calls, etc., which can be recognized in a single frame image; the other category is strong temporal behavior, such as fighting, which can only be recognized by analyzing the image frame sequence over a period of time. In addition to analyzing the two-dimensional features of the image, the recognition process also needs to analyze the features of the image time dimension. Therefore, the recognition of strong temporal behavior is more difficult and complicated.
[0003] There are two existing mainstream methods. One is to use the traditional optical flow algorithm and frame difference information for recognition. However, the traditional algorithm has low computational efficiency and cannot be accelerated using the currently popular intelligent computing units such as GPU and NPU. In actual use, it will consume a lot of CPU computing resources. The other is to add 3D convolution to the neural network, extracting and classifying the temporal features of the image sequence based on conventional convolution. Therefore, a certain number of image frames need to be cached before 3D convolution, which involves more image copy operations. In addition, the amount of calculation of 3D convolution is doubled compared to ordinary convolution. In video surveillance scenarios with high real-time requirements, these will put great pressure on the hardware storage and computing resources.
[0004] With regard to the problems of large amount of computation and low efficiency in related technologies for fighting behavior detection, no effective solution has been proposed so far. Summary of the invention
[0005] In this embodiment, a fighting behavior detection method, device, electronic device and storage medium are provided to solve the problem of large amount of calculation and low efficiency in fighting behavior detection in related technologies.
[0006] In a first aspect, a fighting behavior detection method is provided in this embodiment, including:
[0007] Obtain a sequence of continuous image frames in a video;
[0008] Detecting and tracking human target objects in the continuous image frame sequence to obtain a number of human target object frames to be detected;
[0009] Inputting the plurality of human target object frames to be detected into a preset region candidate module to obtain pre-selected regions;
[0010] Inputting each of the pre-selected regions into a preset frame difference construction network, so that the frame difference construction network calculates the probability of fighting behavior of each of the human target object frames to be detected in each of the pre-selected regions, and obtains a number of the human target objects that have engaged in fighting behavior;
[0011] The minimum bounding box of the human target objects where the fighting behavior occurs is calculated to obtain a fighting behavior detection result.
[0012] In some embodiments, the inputting of the plurality of human target object frames to be detected into a preset region candidate module to obtain the pre-selected regions includes:
[0013] The distances between the plurality of human target object frames to be detected are calculated by the preset region candidate module, and the region consisting of the human target object frames to be detected that meet the preset distance is determined as the pre-selected region.
[0014] In some embodiments, the preset frame difference construction network includes a primary frame difference construction network and a secondary frame difference construction network, and the inputting of each pre-selected area into the preset frame difference construction network so that the frame difference construction network calculates the probability of fighting behavior of each of the human target object frames to be detected in each pre-selected area, including:
[0015] Inputting each pre-selected area into a primary frame difference construction network, so that the primary frame difference construction network calculates the probability of fighting behavior occurring in each pre-selected area, and obtains each target pre-selected area;
[0016] The human target object frames to be detected in the target pre-selected areas are input into the secondary frame difference construction network, so that the secondary frame difference construction network calculates the probability of fighting behavior of the human target object frames to be detected in the target pre-selected areas.
[0017] In some embodiments, the inputting of each pre-selected area into a primary frame difference construction network so that the primary frame difference construction network calculates the probability of fighting behavior occurring in each pre-selected area to obtain each target pre-selected area includes:
[0018] Calculating the feature map difference of each adjacent moment of the pre-selected area through the first-level frame difference construction network to obtain a first-level frame difference feature map;
[0019] The probability of fighting behavior occurring in the primary frame difference feature map is calculated, and the pre-selected area that meets the preset required probability is identified as the target pre-selected area.
[0020] In some of the embodiments, calculating the feature map difference of each adjacent moment of the pre-selected area through the first-level frame difference construction network to obtain the first-level frame difference feature map includes:
[0021] Extracting the initial feature map of the preselected area through the preset first-level frame difference construction network;
[0022] Performing dimensionality reduction processing on the initial feature map to obtain a two-dimensional feature map;
[0023] The difference calculation is performed on the two-dimensional feature map at each adjacent moment to obtain the first-level frame difference feature map.
[0024] In some embodiments, the step of inputting each of the to-be-detected human target object frames within each of the target pre-selected regions into the secondary frame difference construction network so that the secondary frame difference construction network calculates the probability of fighting behavior of each of the to-be-detected human target object frames within each of the target pre-selected regions, comprises:
[0025] In the same target pre-selected area, extracting a feature map of each human target object frame to be detected in the target pre-selected area through the secondary frame difference construction network;
[0026] Calculating the feature map difference between the frames of the human target objects to be detected at adjacent moments, and obtaining a plurality of secondary frame difference feature maps of the human target objects to be detected at adjacent moments;
[0027] The plurality of secondary frame difference feature maps of the human target object frames to be detected are spliced to obtain a target frame difference feature map;
[0028] The probability of the target frame difference feature graph having a fighting behavior is calculated to obtain the human target object having the fighting behavior.
[0029] In some embodiments, the step of splicing the plurality of secondary frame difference feature maps of the human target object frames to be detected to obtain a target frame difference feature map includes:
[0030] Performing maximum pooling processing on the plurality of secondary frame difference feature maps of the human target object frames to be detected to obtain a plurality of enhanced secondary frame difference feature maps;
[0031] Performing dimensionality reduction processing on the enhanced secondary frame difference feature maps to obtain a plurality of one-dimensional feature vectors;
[0032] The plurality of one-dimensional feature vectors are concatenated to obtain the target frame difference feature map.
[0033] In a second aspect, a fighting behavior detection device is provided in this embodiment, comprising: an acquisition module, a preselected area detection module, a fighting behavior detection module, and a calculation module, wherein:
[0034] The acquisition module is used to acquire a continuous image frame sequence in the video; detect and track human target objects in the continuous image frame sequence to obtain a number of human target object frames to be detected;
[0035] The pre-selected region detection module is used to input the plurality of human target object frames to be detected into a preset region candidate module to obtain each pre-selected region;
[0036] The fighting behavior detection module is used to input the pre-selected areas into a preset frame difference construction network, so that the frame difference construction network calculates the probability of fighting behavior of each of the human target object frames to be detected in the pre-selected areas, and obtains a number of human target objects that have engaged in fighting behavior;
[0037] The calculation module is used to calculate the minimum bounding box where the human target objects that have engaged in fighting behaviors are located, and obtain a fighting behavior detection result.
[0038] In a third aspect, an electronic device is provided in this embodiment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the fighting behavior detection method described in the first aspect when executing the computer program.
[0039] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored, and when the program is executed by a processor, the fighting behavior detection method described in the first aspect is implemented.
[0040] Compared with the related art, the fighting behavior detection method provided in the present embodiment obtains a continuous image frame sequence in a video; detects and tracks human target objects in the continuous image frame sequence to obtain a number of human target object frames to be detected; inputs the number of human target object frames to be detected into a preset area candidate module to obtain each pre-selected area; inputs the each pre-selected area into a preset frame difference construction network, so that the frame difference construction network calculates the probability of fighting behavior of each of the human target object frames to be detected in the each pre-selected area to obtain a number of the human target objects that have engaged in fighting behavior; calculates the minimum external bounding box in which the several human target objects that have engaged in fighting behavior are located to obtain a fighting behavior detection result, thereby solving the problem of large calculation amount and low efficiency in fighting behavior detection, reducing the detection calculation amount, and improving the detection efficiency of fighting behavior.
[0041] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0043] Figure 1 is a hardware structure block diagram of a terminal of the fighting behavior detection method of this embodiment;
[0044] Figure 2 is a flow chart of the fighting behavior detection method of this embodiment;
[0045] Figure 3 is a calculation flow chart of the area pre-selection module in the fighting behavior detection method of this embodiment;
[0046] Figure 4 is a processing flow chart of constructing a network by first-level frame difference in the fighting behavior detection method of this embodiment;
[0047] Figure 5 is a processing flow chart of a secondary frame difference network construction method of the fighting behavior detection method of this embodiment;
[0048] Figure 6 is a flow chart of another fighting behavior detection method of the present embodiment;
[0049] Figure 7 It is a structural block diagram of the fighting behavior detection device of this embodiment. DETAILED DESCRIPTION
[0050] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0051] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0052] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 FIG. 1 is a hardware structure diagram of a terminal of the fighting behavior detection method of this embodiment. Figure 1 As shown, the terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 and a memory 104 for storing data, wherein the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the above terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.
[0053] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the fighting behavior detection method in the present embodiment, and the processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0054] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider of the terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.
[0055] In video surveillance and behavior recognition methods, behavior recognition can be roughly divided into weak temporal behaviors and strong temporal behaviors according to the length of time required for behavior recognition. For example, behaviors with weak temporal sequences, such as smoking and making phone calls, can be recognized in a frame of video image at a certain moment; but behaviors with strong temporal sequences, such as fighting, require analysis of several continuous image frame sequences within a certain period of time to recognize such strong temporal behaviors. In addition to analyzing the two-dimensional features of the image, it is also necessary to analyze the temporal features of the image. Therefore, the recognition of behaviors with strong temporal sequences is more difficult and complicated.
[0056] In this embodiment, a method for detecting fighting behavior is provided. Figure 2 is a flow chart of the fighting behavior detection method of this embodiment, such as Figure 2 As shown, the process includes the following steps:
[0057] Step S201, obtaining a continuous image frame sequence in a video; detecting and tracking human target objects in the continuous image frame sequence to obtain a number of human target object frames to be detected.
[0058] Specifically, in order to analyze strong temporal behaviors, it is necessary to analyze the temporal characteristics of the image, obtain a continuous image frame sequence from time t to time t+n from the video to be analyzed, and perform real-time detection and tracking of human targets in each frame of the image through target detection and target tracking, and preliminarily obtain a number of human target object frames to be detected, among which the target detection and tracking algorithm includes but is not limited to the detection and tracking algorithm widely used in the current image processing field, and this embodiment does not make specific limitations on this.
[0059] Step S202: input a plurality of human target object frames to be detected into a preset region candidate module to obtain pre-selected regions.
[0060] Specifically, the minimum bounding rectangle between each human target object frame to be detected is calculated through a preset region candidate module, and the region composed of human target object frames within the minimum bounding rectangle that meets the requirements is identified as a pre-selected region, and human target object frames that do not meet the requirements are excluded. Through the region candidate module, on the one hand, the region where fighting behavior may occur can be pre-selected in advance, and on the other hand, relatively free human targets in the image frame can be filtered out. These relatively free human targets have a low probability of fighting behavior and do not need to be calculated and analyzed in the subsequent process. By filtering out human target objects with low probability in advance, the calculation amount of subsequent modules can be saved and the speed of fighting behavior detection can be accelerated.
[0061] Step S203, inputting each pre-selected area into a preset frame difference construction network, so that the frame difference construction network calculates the probability of fighting behavior of each human target object frame to be detected in each pre-selected area, and obtains a number of human target objects that have engaged in fighting behavior.
[0062] Specifically, a convolutional neural network is used as the basic network, and the structure and related processing flow of frame difference construction are added to the network to obtain a frame difference construction network. The frame difference construction network retains the temporal correlation characteristics, and the temporal correlation features can be extracted using ordinary convolution. Among them, the basic network includes but is not limited to the currently widely used convolutional networks such as VGGNet, ResNet, etc., and this embodiment does not specifically limit this. The probability of fighting in each pre-selected area is calculated by the frame difference construction network. If the probability exceeds the specified threshold, it is considered that there is a fighting behavior in the pre-selected area, and the pre-selected area is identified as the target pre-selected area. Among them, the frame difference construction network calculates a pre-selected area each time, and the pre-selected area of each frame image from time t to (t+n) needs to be input into the network in sequence to construct a frame difference feature map of the pre-selected area from time t to (t+n). The frame difference feature map construction network constructs a frame difference feature map by performing operations such as difference calculation of the frame feature maps before and after the same area, feature map dimension compression, dimension conversion and splicing. The frame difference feature map can reflect the changes in the features before and after the image frame, suppress the background with weak feature changes, and is more conducive to the extraction and classification of temporal features. When the frame difference feature map construction network calculates that there is a fight in a pre-selected area, all the human target frames to be detected in the area are processed and calculated in turn, and then the probability of fighting for each human target object to be detected is determined. If the probability is greater than the preset threshold, it is considered that the human target object has a fight, and several human target objects with fighting behaviors are obtained. Similar to the target pre-selected area determination process, each time a human target object frame to be detected is calculated, the human target object frame to be detected in each frame image from time t to (t+n) needs to be input into the network in turn to construct the frame difference feature map of the human target frame to be detected from time t to (t+n).
[0063] Step S204, calculating the minimum bounding box of several human target objects that are engaged in fighting behavior, and obtaining a fighting behavior detection result.
[0064] Specifically, after detecting several human target objects that are fighting, the minimum bounding rectangle of all human target objects that are fighting is calculated to obtain a fighting behavior detection result. The area where the rectangle is located is the area where fighting may occur, and fighting exists in the area.
[0065] Through the above steps S201 to S204, a continuous image frame sequence in the video is obtained; the human target object in the continuous image frame sequence is detected and tracked to obtain a number of human target object frames to be detected; the number of human target object frames to be detected are input into a preset region candidate module to obtain each pre-selected region; each pre-selected region is input into a preset frame difference construction network, so that the frame difference construction network calculates the probability of fighting behavior of each human target object frame to be detected in each pre-selected region, and obtains a number of human target objects that have fought; the minimum external bounding box where a number of human target objects that have fought are located is calculated to obtain a fighting behavior detection result. Compared with the prior art of detecting fighting behavior by extracting temporal features through optical flow algorithm and 3D convolution, this embodiment pre-selects the area where fighting behavior may occur in advance through the region candidate module, and filters out the target frame where fighting behavior is unlikely to occur, thereby reducing the amount of calculation in the subsequent detection process. Then, by adding the frame difference construction module and processing flow to the convolutional neural network, the frame difference feature map is constructed by performing operations such as difference calculation of the feature maps of the previous and next frames in the same area, feature map dimension compression, dimension conversion and splicing, so as to suppress the background with weak feature changes, which is more conducive to the extraction and classification of temporal features. There is no need to use 3D convolution with higher computational complexity to retain the temporal nature of the feature map, and the use of ordinary convolution can be used to extract and classify temporal features, so as to realize real-time and efficient recognition of fighting behaviors with strong temporal nature, reduce the computational complexity of fighting behavior detection and improve detection efficiency.
[0066] In some of the embodiments, a plurality of human target object frames to be detected are input into a preset region candidate module to obtain pre-selected regions, including:
[0067] The distances between a number of human target object frames to be detected are calculated through a preset region candidate module, and a region consisting of human target object frames to be detected that meet the preset distance is determined as a pre-selected region.
[0068] Specifically, the area pre-selection module pre-screens large areas where fighting may occur, and excludes areas where fighting is unlikely to occur, so as to reduce the amount of calculation in the subsequent detection process. Figure 3 is a calculation flow chart of the area preselection module in the fighting behavior detection method of this embodiment, such as Figure 3 As shown, the calculation process includes:
[0069] Step S301, initializing a target set r_set to be selected and a target set r_unselect to be unselected;
[0070] Specifically, all human target object frames to be detected in continuous image frames are taken as the target set to be selected r_set, and an empty set r_unselect is set as the set for storing unselected human target objects each time the candidate area is searched, i.e., the unselected target set. In each round of searching for a new human target object frame to be detected, the human target object frames in the unselected target set will be cleared, and the human target object frames therein will be moved to the target set to be selected r_set. Before the detection starts, the target set to be selected r_set and the unselected target set r_unselect are initialized, and then the video image to be detected is loaded.
[0071] Step S302, determine whether r_set is empty, if so, execute step S313; otherwise, execute step S303;
[0072] Specifically, it is determined whether there are any human target object frames to be detected in the target set r_set to be selected. If not, it means that all human target object frames to be detected have been detected and the fighting behavior detection is finished.
[0073] Step S303: randomly select a human target object frame r to be detected from r_set n , as the reference human target object frame, the human target object frame to be detected r n As the preselected area R n ;
[0074] Specifically, take out any human target object frame r from the target set r_set n As the reference human target object frame, start detection and set the human target object frame to be detected r n The area is the pre-selected area R n The initial pre-selected area.
[0075] Step S304, determine whether r_set is empty, if so, execute step S313; otherwise, execute step S305;
[0076] Specifically, it is determined whether there are any human target object frames to be detected in the target set r_set to be selected. If not, it means that all human target object frames to be detected have been detected and the fighting behavior detection is finished.
[0077] Step S305: randomly select a human target object frame r to be detected from r_set m , as a comparison human target object box;
[0078] Specifically, take out any target box r from the target set r_set mAs a comparison human target object frame, and the reference human target object frame r n Make a comparison.
[0079] Step S306, calculate the current R n and r m The minimum enclosing rectangle R n_tmp ;
[0080] Specifically, calculate the pre-selected area R n The human target object frame r to be detected n With r m The minimum enclosing rectangle R n_tmp ; The minimum bounding rectangle includes the pre-selected area R n The human target object frame to be detected and the human target object frame r to be compared m .
[0081] Step S307, determine R n_tmp Whether the short side length of is less than the safety threshold, if so, execute step S308, otherwise, execute step S312;
[0082] Specifically, by judging R n_tmp The short side length of the human body object frame that may be involved in a fight is screened out by checking whether the short side length of the human body object frame meets the safety threshold, and the range where these human body object frames are located is included in the pre-selected area. Among them, the safety threshold can be configured by the user according to the actual scene, and this embodiment does not make a specific limitation on this.
[0083] Step S308, update R n =R n_tmp ;
[0084] Specifically, when R n_tmp When the short side length meets the safety threshold, it means that the human target object box r is compared. m With the preselected area R n There may be fighting in the composed area, and the pre-selected area R n The area is replaced by the comparison human target object box r m Area R n_tmp ; R n_tmp As pre-selected areas for the next round of comparison.
[0085] Step S309, determine whether r_set is empty, if so, execute step S310; otherwise, execute step S304;
[0086] Step S310: n Add to the selected pre-selected area set R_select and initialize the current R n ;
[0087] Specifically, when all the human target object frames to be detected in the target set r_set are detected, the currently obtained pre-selected area R n The area where the fight is located is added to the selected pre-selected area set R_select, that is, the pre-selected area where fighting may occur is obtained, and then the current pre-selected area R n Perform initialization processing in preparation for the next round of fighting behavior detection.
[0088] Step S311, put all human target object frames in r_unselect back into r_set; and return to execute step S302;
[0089] Specifically, all human target object frames in the unselected target set r_unselect are put back into the to-be-selected target set r_set, and the next round of fighting behavior detection is performed again.
[0090] Step S312: m Put it into the r_unselect set and execute step S304;
[0091] Step S313, end.
[0092] Through the above steps S301 to S313, the distance relationship between the human target object frames in the image frame is analyzed and calculated by the region candidate module. On the one hand, it is possible to obtain the area where fighting is likely to occur, and on the other hand, it also filters out the human targets in a free state in the image frame. These human targets do not need to be subsequently calculated and analyzed, thereby saving computing resources to a certain extent, reducing the amount of calculation in the subsequent detection process, and improving the detection efficiency.
[0093] In another embodiment, the preset frame difference construction network includes a primary frame difference construction network and a secondary frame difference construction network, and each pre-selected area is input into the preset frame difference construction network so that the frame difference construction network calculates the probability of fighting behavior of each human target object frame to be detected in each pre-selected area, including:
[0094] Each pre-selected area is input into the first-level frame difference construction network, so that the first-level frame difference construction network calculates the probability of fighting behavior occurring in each pre-selected area, and obtains each target pre-selected area; each of the human target object frames to be detected in each target pre-selected area is input into the second-level frame difference construction network, so that the second-level frame difference construction network calculates the probability of fighting behavior of each human target object frame to be detected in each target pre-selected area.
[0095] Specifically, this embodiment adopts a two-level frame difference construction network to extract and classify features of regions and human targets from the whole to the part, thereby improving the speed and efficiency of fighting behavior recognition. The preset frame difference construction network includes a primary frame difference construction network and a secondary frame difference construction network.
[0096] The first-level frame difference construction network uses a convolutional neural network as the basic network, adds the structure of frame difference construction and related processing procedures to the network, retains the temporal correlation characteristics so that ordinary convolution can perform feature extraction. The basic network includes but is not limited to the currently widely used convolutional networks such as VGGNet and ResNet. The first-level frame difference construction network calculates the probability of fighting in each pre-selected area in turn. If the probability exceeds the specified threshold, it is considered that there is fighting in the area, and the area is identified as the target pre-selected area. Then, all the human target object frames to be detected in each target pre-selected area are input into the second-level frame difference construction network. The second-level frame difference construction network processes and calculates all the human target object frames to be detected in the target pre-selected area in turn to determine the probability of fighting in each human target object frame to be detected. If the probability is greater than a certain threshold, it is considered that the human target object belongs to fighting.
[0097] In some embodiments, each pre-selected area is input into a primary frame difference construction network, so that the primary frame difference construction network calculates the probability of fighting behavior occurring in each pre-selected area, and obtains each target pre-selected area, including:
[0098] The network is constructed through the first-level frame difference to calculate the difference of the feature maps of the pre-selected area at each adjacent moment, and the first-level frame difference feature map is obtained; the probability of fighting behavior occurring in the first-level frame difference feature map is calculated, and the pre-selected area that meets the preset required probability is identified as the target pre-selected area.
[0099] Specifically, the first-level frame difference construction network calculates a pre-selected area each time, and the pre-selected area of each frame image from time t to (t+n) needs to be input into the network in sequence to construct the frame difference feature map of the pre-selected area from time t to (t+n). The probability of fighting in each pre-selected area is calculated in sequence based on the frame difference feature map. If the probability exceeds the specified threshold, it is considered that fighting exists in the area, and the area is identified as the target pre-selected area.
[0100] In another embodiment, the feature map difference of each adjacent time of the pre-selected area is calculated by constructing a network through a first-level frame difference to obtain a first-level frame difference feature map, including:
[0101] The network is constructed by using a preset first-level frame difference to extract the initial feature map of the preselected area; the initial feature map is subjected to dimensionality reduction processing to obtain a two-dimensional feature map; the difference between the two-dimensional feature maps at each adjacent moment is calculated to obtain a first-level frame difference feature map.
[0102] Specifically, by taking a pre-selected region R in R_select as an example, the processing flow of constructing a network using a first-level frame difference is described. Figure 4 : is a processing flow chart of the first-level frame difference construction network of the fighting behavior detection method of this embodiment, such as Figure 4 As shown in the figure, taking the time t and (t+1) as an example, each pre-selected area is input into the first-level frame difference construction network so that the first-level frame difference construction network calculates the probability of fighting in each pre-selected area. In the process of obtaining the pre-selected areas of each target, the area of the pre-selected area R at the time t and (t+1) is respectively input into the first-level frame difference construction network for feature extraction, and the feature maps at the time t and (t+1) are respectively obtained, and each set of feature maps contains multiple feature channels; then a single 1x1 convolution kernel and a convolution layer with 1 convolution kernel are used to extract features from the feature map. The 1x1 convolution kernel can ensure that the width and height of the feature map remain unchanged after the feature extraction of this layer, and the number of convolution kernels is 1, which can make the number of channels of the output feature map of this layer become 1, that is, the channel number dimension can be omitted, thereby achieving the purpose of feature dimensionality reduction. After passing through this layer, the two-dimensional feature maps at time t and (t+1) are obtained respectively; then, the difference calculation is performed on the two-dimensional feature maps at time t and (t+1), that is, the feature values of the corresponding positions of the two-dimensional feature maps are subtracted to obtain a two-dimensional frame difference feature map. Finally, each frame of the pre-selected area R from time t to (t+n) is processed in the above manner, and (n-1) two-dimensional frame difference feature maps are accumulated. The (n-1) two-dimensional frame difference feature maps are spliced into a three-dimensional frame difference feature map to obtain a first-level frame difference feature map, and the number of channels is (n-1). The first-level frame difference feature map is then sent to the subsequent network layer for further feature extraction and classification, and finally the probability of fighting in the pre-selected area R is output. If the probability exceeds the preset threshold, it is considered that a fight has occurred in the pre-selected area R. Among them, n and the threshold can be configured by the user according to the actual situation, and this embodiment does not make specific restrictions on this.
[0103] In some embodiments, each human target object frame to be detected in each target pre-selected area is input into a secondary frame difference construction network, so that the secondary frame difference construction network calculates the probability of fighting behavior of each human target object frame to be detected in each target pre-selected area, including:
[0104] In the same target pre-selected area, a network is constructed through secondary frame differences to extract feature maps of each human target object frame to be detected in the target pre-selected area; the feature map differences between adjacent moments of each human target object frame to be detected are calculated to obtain a number of secondary frame difference feature maps of each human target object to be detected at each adjacent moment; the several secondary frame difference feature maps of each human target object frame to be detected are spliced to obtain a target frame difference feature map; the probability of fighting behavior occurring in the target frame difference feature map is calculated to obtain the human target object that has undergone fighting behavior.
[0105] Specifically, after being processed by the first-level frame difference construction network, the target pre-selected area that is judged to have fighting behavior will further input each human target object frame in the target pre-selected area into the second-level frame difference construction network for processing in sequence, and more meticulously determine whether each human target object has fighting behavior. Each time the second-level frame difference construction network calculates a human target object frame, it is necessary to input the human target object frame of each frame image from time t to (t+n) into the network in sequence, and calculate the feature map difference between adjacent moments of each human target object frame, construct the frame difference feature map of the human target object frame from time t to (t+n), and splice the frame difference feature map to obtain the target frame difference feature map. The target frame difference feature map describes the transformation features of the human target object frame r from time t to (t+n). After feature extraction and classification of the subsequent convolutional layers in the network, the probability of the current human target object having fighting behavior is given. If the probability exceeds the preset threshold, the current human target object is judged to be fighting behavior, that is, the human target object having fighting behavior is obtained. Among them, n and the threshold value can be configured by the user according to actual conditions, and this embodiment does not make specific limitations on this.
[0106] In another embodiment, a plurality of secondary frame difference feature maps of each human target object frame to be detected are spliced to obtain a target frame difference feature map, including:
[0107] Performing maximum pooling processing on several secondary frame difference feature maps of each human target object frame to be detected to obtain several enhanced secondary frame difference feature maps; performing dimensionality reduction processing on the enhanced several secondary frame difference feature maps to obtain several one-dimensional feature vectors; and splicing several one-dimensional feature vectors to obtain a target frame difference feature map.
[0108] Specifically, taking a human target object frame r in a region R that has been determined to have fighting behavior as an example, the structure and processing flow of the secondary frame difference construction network are explained. Figure 5 : is a processing flow chart of the secondary frame difference construction network of the fighting behavior detection method of this embodiment, such as Figure 5As shown, in the process of inputting each human target object frame to be detected in each target pre-selected area into the secondary frame difference construction network so that the secondary frame difference construction network calculates the probability of fighting behavior of each human target object frame to be detected in each target pre-selected area, the human target object frame r from time t to (t+n) is input into the secondary frame difference construction network. Taking the human target frame r at time t and time (t+1) as an example, after several convolutional layers are used for feature extraction, feature maps at time t and time (t+1) are obtained respectively. The feature map contains multiple feature channels. The feature map at time t and time (t+1) is subjected to difference calculation, that is, the feature values at corresponding positions are subtracted to obtain a secondary frame difference feature map. The width, height, and number of channels of the frame difference feature map remain unchanged before the difference calculation, and still contain multiple feature channels. Then, the maximum pooling operation is performed on the secondary frame difference feature map to extract the maximum feature value of each feature channel in the secondary frame difference feature map, and obtain several enhanced secondary frame difference feature maps. The feature value can describe the most significant change of the human target object frame r between time t and time (t+1). After the maximum pooling, the width and height of the frame difference feature map are changed to 1, and the number of feature channels remains unchanged. Then, the 1×1×C frame difference feature map (C is the number of feature channels) after the maximum pooling is reshaped, and the 1×1×C frame difference feature map is transformed into a C×1 feature vector, thereby realizing feature dimension reduction. The time sequence information of the human target object is still preserved in the reduced feature vector. The above processing is performed on the human target object frame r from time t to (t+n), and (n-1) C×1 feature vectors can be accumulated. These feature vectors are concatenated in sequence to obtain a (n-1)×C×1 frame difference feature map with a channel number of 1. This feature map describes the transformation characteristics of the target r from time t to (t+n). After feature extraction and classification of the subsequent convolutional layers in the network, the probability of the current human target object fighting behavior is obtained. If the probability exceeds the preset threshold, the current human target object is judged as fighting behavior, and the human target object with fighting behavior is obtained.
[0109] Among them, the design ideas and purposes of the first-level frame difference construction network and the second-level frame difference construction network are the same, but there are certain differences in the specific structural design. The first-level frame difference construction network is mainly used to identify whether there is fighting behavior in the area. While paying attention to the temporal changes in the area, it is also necessary to pay attention to the feature information of the two-dimensional plane of the area. Therefore, the first-level frame difference construction network uses a single 1x1 convolution kernel to compress the feature map dimension without losing the two-dimensional plane features of the image. The second-level frame difference construction network is mainly used to detect the fighting behavior of each human target object frame. It no longer needs to pay attention to the two-dimensional plane information of the human target object, but only needs to pay attention to the changes in the realization of the human target object frame features. Therefore, the second-level frame difference construction network uses maximum pooling to highlight the changes in target features, and can also achieve feature dimensionality reduction.
[0110] This embodiment also provides a fighting behavior detection method. Figure 6 is a flow chart of another fighting behavior detection method of this embodiment, such as Figure 6 As shown, the process includes the following steps:
[0111] Step S601, obtaining a continuous image frame sequence in a video;
[0112] Step S602, detecting human target objects in a continuous image frame sequence to obtain a number of human target object frames to be detected;
[0113] Step S603, calculating the distances between a plurality of human target object frames to be detected by a preset region candidate module, and determining a region consisting of human target object frames to be detected that meet the preset distance as a pre-selected region;
[0114] Step S604, extracting an initial feature map of the preselected area through a preset first-level frame difference construction network; performing dimensionality reduction processing on the initial feature map to obtain a two-dimensional feature map; performing difference calculation on the two-dimensional feature maps at each adjacent moment to obtain a first-level frame difference feature map;
[0115] Step S605, calculating the probability of fighting behavior occurring in the primary frame difference feature map, and identifying the pre-selected area that meets the preset required probability as the target pre-selected area;
[0116] Step S606, in the same target pre-selected area, extracting a feature map of each human target object frame to be detected in each target pre-selected area by constructing a network through a secondary frame difference;
[0117] Step S607, calculating the feature map difference between adjacent moments of each human target object frame to be detected, and obtaining a plurality of secondary frame difference feature maps of each human target object to be detected at adjacent moments;
[0118] Step S608, performing maximum pooling processing on the plurality of secondary frame difference feature maps to obtain a plurality of enhanced secondary frame difference feature maps; performing dimensionality reduction processing on the plurality of enhanced secondary frame difference feature maps to obtain a plurality of one-dimensional feature vectors; and concatenating the plurality of one-dimensional feature vectors to obtain a target frame difference feature map;
[0119] Step S609, calculating the probability of a fighting behavior occurring in the target frame difference feature graph, and obtaining a human target object that has a fighting behavior;
[0120] Step S610, calculating the minimum bounding box of several human target objects that are engaged in fighting behavior, and obtaining a fighting behavior detection result.
[0121] Through the above steps S601 to S610, compared with the prior art of extracting temporal features through optical flow algorithm and 3D convolution to detect fighting behavior, this embodiment analyzes and calculates the distance relationship between human target object frames in the image frame through the region candidate module, pre-selects the pre-selected area where fighting behavior may occur, and filters out the human target object frames where fighting behavior is impossible, reducing the calculation amount of subsequent modules. Then, the target pre-selected area where fighting behavior may occur is calculated through the first-level frame difference construction network; the pre-selected area where fighting behavior is impossible is excluded, further reducing the calculation amount of the subsequent network. Then, the human target object in the target pre-selected area where fighting behavior may occur is calculated through the second-level frame difference construction network, thereby obtaining the human target object where fighting behavior occurs. Ordinary convolution can be used to extract features from the frame difference feature map without using 3D convolution with higher computational complexity, reducing the calculation amount of fighting behavior detection in the video; the two-level frame difference is used to construct the network, and feature extraction and classification are performed on the region and human target from the whole to the part, respectively, which improves the speed and recognition efficiency of fighting behavior recognition.
[0122] In the present embodiment, a fighting behavior detection device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and the descriptions thereof will not be repeated. The terms "module", "unit", "subunit" etc. used below can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware is also possible and contemplated.
[0123] Figure 7 : is a structural block diagram of the fighting behavior detection device of this embodiment, such as Figure 7 As shown, the device 70 includes: an acquisition module 71, a pre-selected area detection module 72, a fighting behavior detection module 73, and a calculation module 74, wherein:
[0124] The acquisition module 71 is used to acquire a continuous image frame sequence in the video; detect and track human target objects in the continuous image frame sequence to obtain a number of human target object frames to be detected;
[0125] The pre-selected region detection module 72 is used to input a plurality of human target object frames to be detected into a preset region candidate module to obtain each pre-selected region;
[0126] The fighting behavior detection module 73 is used to input each pre-selected area into a preset frame difference construction network, so that the frame difference construction network calculates the probability of fighting behavior of each human target object frame to be detected in each pre-selected area, and obtains a number of human target objects that have engaged in fighting behavior;
[0127] The calculation module 74 is used to calculate the minimum bounding box of several human target objects that are engaged in fighting behavior, and obtain the fighting behavior detection result.
[0128] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0129] In this embodiment, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0130] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0131] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0132] S1, obtain a continuous image frame sequence in the video;
[0133] S2, detecting and tracking human target objects in a continuous image frame sequence to obtain a number of human target object frames to be detected;
[0134] S3, inputting a number of human target object frames to be detected into a preset region candidate module to obtain pre-selected regions;
[0135] S4, inputting each pre-selected area into a preset frame difference construction network, so that the frame difference construction network calculates the probability of fighting behavior of each human target object frame to be detected in each pre-selected area, and obtains a number of human target objects that have engaged in fighting behavior;
[0136] S5, calculating the minimum bounding box of several human target objects that are engaged in fighting behavior, and obtaining a fighting behavior detection result.
[0137] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.
[0138] In addition, in combination with the fighting behavior detection method provided in the above embodiments, a storage medium may be provided in this embodiment to implement the fighting behavior detection method. The storage medium stores a computer program; when the computer program is executed by a processor, any one of the fighting behavior detection methods in the above embodiments is implemented.
[0139] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.
[0140] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0141] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.
[0142] The term "embodiment" in this application refers to a specific feature, structure or characteristic described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various locations in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is clearly or implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.
[0143] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0144] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.
Claims
1. A fighting behavior detection method, characterized in that: include: Obtain a sequence of continuous image frames in a video; Detecting and tracking human target objects in the continuous image frame sequence to obtain a number of human target object frames to be detected; Inputting the plurality of human target object frames to be detected into a preset region candidate module to obtain pre-selected regions; Inputting each of the pre-selected regions into a preset frame difference construction network, wherein the preset frame difference construction network includes a primary frame difference construction network and a secondary frame difference construction network; Calculating the feature map difference of each adjacent moment of the pre-selected area through the first-level frame difference construction network to obtain a first-level frame difference feature map; Calculating the probability of fighting behavior occurring in the primary frame difference feature map, and identifying the pre-selected area that meets the preset required probability as the target pre-selected area; Inputting each of the human target object frames to be detected in each of the target pre-selected areas into the secondary frame difference construction network, so that the secondary frame difference construction network calculates the probability of fighting behavior of each of the human target object frames to be detected in each of the target pre-selected areas, and obtains a number of human target objects that have engaged in fighting behavior; The minimum bounding box of the human target objects where the fighting behavior occurs is calculated to obtain a fighting behavior detection result.
2. The fighting behavior detection method according to claim 1, characterized in that: The step of inputting the plurality of human target object frames to be detected into a preset region candidate module to obtain each pre-selected region includes: The distances between the plurality of human target object frames to be detected are calculated by the preset region candidate module, and the region consisting of the human target object frames to be detected that meet the preset distance is determined as the pre-selected region.
3. The fighting behavior detection method according to claim 1, characterized in that: The step of calculating the feature map difference values of each adjacent moment of the pre-selected area through the first-level frame difference construction network to obtain the first-level frame difference feature map includes: Extracting the initial feature map of the preselected area through the preset first-level frame difference construction network; Performing dimensionality reduction processing on the initial feature map to obtain a two-dimensional feature map; The difference calculation is performed on the two-dimensional feature map at each adjacent moment to obtain the first-level frame difference feature map.
4. The fighting behavior detection method according to claim 1, characterized in that: The step of inputting each of the to-be-detected human target object frames in each of the target pre-selected regions into the secondary frame difference construction network so that the secondary frame difference construction network calculates the probability of fighting behavior of each of the to-be-detected human target object frames in each of the target pre-selected regions, comprises: In the same target pre-selected area, extracting a feature map of each human target object frame to be detected in the target pre-selected area through the secondary frame difference construction network; Calculating the feature map difference between the frames of the human target objects to be detected at adjacent moments, and obtaining a plurality of secondary frame difference feature maps of the human target objects to be detected at adjacent moments; The plurality of secondary frame difference feature maps of the human target object frames to be detected are spliced to obtain a target frame difference feature map; The probability of the target frame difference feature graph having a fighting behavior is calculated to obtain the human target object having the fighting behavior.
5. The fighting behavior detection method according to claim 4, characterized in that: The step of splicing the plurality of secondary frame difference feature maps of the human target object frames to be detected to obtain a target frame difference feature map comprises: Performing maximum pooling processing on the plurality of secondary frame difference feature maps of the human target object frames to be detected to obtain a plurality of enhanced secondary frame difference feature maps; Performing dimensionality reduction processing on the enhanced secondary frame difference feature maps to obtain a plurality of one-dimensional feature vectors; The plurality of one-dimensional feature vectors are concatenated to obtain the target frame difference feature map.
6. A fighting behavior detection device, characterized in that: include: An acquisition module, a preselected area detection module, a fighting behavior detection module, and a calculation module, wherein: The acquisition module is used to acquire a continuous image frame sequence in the video; detect and track human target objects in the continuous image frame sequence to obtain a number of human target object frames to be detected; The pre-selected region detection module is used to input the plurality of human target object frames to be detected into a preset region candidate module to obtain each pre-selected region; The fighting behavior detection module is used to input each pre-selected area into a preset frame difference construction network, wherein the preset frame difference construction network includes a primary frame difference construction network and a secondary frame difference construction network; calculate the feature map difference value of each adjacent moment of the pre-selected area through the primary frame difference construction network to obtain a primary frame difference feature map; calculate the probability of fighting behavior in the primary frame difference feature map, and identify the pre-selected area that meets the preset required probability as the target pre-selected area; input each of the human target object frames to be detected in each target pre-selected area into the secondary frame difference construction network, so that the secondary frame difference construction network calculates the probability of fighting behavior of each of the human target object frames to be detected in each target pre-selected area, and obtains a number of human target objects that have fought; The calculation module is used to calculate the minimum bounding box where the human target objects that have engaged in fighting behaviors are located, and obtain a fighting behavior detection result.
7. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the fighting behavior detection method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the fighting behavior detection method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Identification method and device for fighting behavior, storage medium and electronic device
CN111860430A
Taking identification method and system and storage medium thereof
CN117831119A