Method, apparatus, electronic device, and storage medium for identifying target behavior

By combining behavior identification methods of video and audio data, the violations caused by lack of monitoring in the control areas are solved, and high-accurate behavior identification and timely control are achieved.

CN115100558BActive Publication Date: 2025-08-05ANHUI IFLYTEK INTELLIGENT SYST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210521485.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-08-05
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

The lack of behavior monitoring in the control areas leads to frequent illegal and irregular behaviors, affecting production order.

Method used

By combining video data and audio data, a pre-trained target behavior recognition model is used to fuse video and audio information for behavior feature extraction to identify the behavior of the detection target.

Benefits of technology

It improves the accuracy of behavior recognition, can promptly detect violations, and ensure the effectiveness of control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100558B_ABST
    Figure CN115100558B_ABST
Patent Text Reader

Abstract

The present application proposes a method, device, electronic device and storage medium for identifying target behavior. The method includes obtaining target video data and target audio data, wherein the target video data and target audio data are respectively obtained by collecting video data and audio data of the detection area where the detection target is located. Based on the target video data and target audio data, it is possible to determine whether the detection target exhibits target behavior, thereby ensuring the management and control effect. This solution combines target video data and target audio data to identify target behavior from multiple modalities, effectively improving the accuracy of target behavior identification. The above solution is applied to the management and control of sand mining areas such as rivers and seas, which can timely discover whether sand mining ships are engaging in illegal mining and ensure the management and control effect of sand mining areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of behavior recognition technology, and in particular to a method, device, electronic device and storage medium for identifying a target behavior. Background Art

[0002] If there is a lack of behavioral control in the controlled areas, illegal and irregular behaviors are likely to occur, affecting the production order in the controlled areas. For example, in sand mining areas, if there is no supervision of sand mining, disorderly mining and illegal mining of sand and gravel may occur.

[0003] Behavior monitoring in controlled areas is the key to effective behavior control in controlled areas and preventing violations. Therefore, there is an urgent need for a technical solution that can monitor behavior in controlled areas to ensure the effectiveness of control in controlled areas. Summary of the Invention

[0004] Based on the above needs, this application proposes a method, device, electronic device and storage medium for identifying target behavior. The method provides a technical solution that can monitor the behavior of the controlled area to ensure the control effect of the controlled area.

[0005] The technical solutions proposed in this application are as follows:

[0006] In one aspect, the present invention provides a method for identifying a target behavior, comprising:

[0007] Acquire target video data and target audio data; wherein the target video data and the target audio data are respectively obtained by collecting video data and audio data of a detection area where the detection target is located;

[0008] Based on the target video data and the target audio data, it is determined whether the detection target performs a target behavior.

[0009] Furthermore, the determining whether the detection target performs a target behavior based on the target video data and the target audio data includes:

[0010] fusing the target video data and the target audio data to obtain fusion information;

[0011] extracting behavioral features of the detection target from the fusion information;

[0012] Based on the behavior characteristics, it is determined whether the detection target performs a target behavior.

[0013] Furthermore, the fusing of the target video data and the target audio data to obtain fusion information, extracting behavioral features of the detection target from the fusion information, and determining whether the detection target performs a target behavior based on the behavioral features, includes:

[0014] Inputting the target video data and the target audio data into a pre-trained target behavior recognition model, causing the target behavior recognition model to fuse the target video data and the target audio data to obtain fusion information, extracting the behavior features of the detection target from the fusion information, and obtaining a target behavior recognition result based on the behavior features;

[0015] The target behavior recognition result includes that the detection target performs the target behavior, or that the detection target does not perform the target behavior.

[0016] Furthermore, before inputting the target video data and the target audio data into the pre-trained target behavior recognition model, the method further includes:

[0017] Performing frame extraction processing on the target video data, and / or performing downsampling processing on the target audio data;

[0018] The target video data and the target audio data are time-synchronized data.

[0019] Furthermore, before obtaining the target video data and the target audio data, the method further includes:

[0020] determining whether the detection target has the conditions to perform the target behavior in the detection area;

[0021] If the detection target meets the conditions for performing the target behavior in the detection area, the step of acquiring the target video data and the target audio data is performed.

[0022] Furthermore, determining whether the detection target has the conditions for performing the target behavior in the detection area includes:

[0023] Get the length of time the detection target stays in the detection area;

[0024] If the detection target stays in the detection area for a predetermined period of time, it is determined that the detection target has the conditions to perform the target behavior in the detection area.

[0025] Furthermore, obtaining the duration of time the detection target stays in the detection area includes:

[0026] When it is determined that the detection target enters the detection area, the detection target is tracked from the surveillance video of the detection area to determine the duration for which the detection target appears in the surveillance video of the detection area;

[0027] The duration for which the detection target appears in the surveillance video of the detection area is determined as the duration for which the detection target stays in the detection area.

[0028] Furthermore, tracking the detection target from the surveillance video of the detection area includes:

[0029] Determine a bounding box area of the detection target in a first video frame of a surveillance video of the detection area;

[0030] Identifying a bounding box area of a target object from a second video frame of a surveillance video of the detection area; wherein the second video frame is a video frame subsequent to the first video frame, and the target object is an identification object of the same type as the detection target;

[0031] If the intersection-and-union ratio of the bounding box area of any target object in the second video frame and the bounding box area of the detection target in the first video frame of the surveillance video of the detection area is greater than a set threshold, the target object in the second video frame is determined to be the detection target.

[0032] Furthermore, the target video data includes monitoring video data of the sand mining area, the target audio data includes audio data collected from the sand mining area, and the target behavior includes sand mining behavior.

[0033] In another aspect, the present invention provides a target behavior recognition device, comprising:

[0034] An acquisition module, configured to acquire target video data and target audio data; wherein the target video data and the target audio data are respectively acquired by collecting video data and audio data of a detection area where the detection target is located;

[0035] A determination module is used to determine whether the detection target performs a target behavior based on the target video data and the target audio data.

[0036] In another aspect, the present invention provides an electronic device comprising:

[0037] memory and processor;

[0038] Wherein, the memory is used to store programs;

[0039] The processor is configured to implement the target behavior recognition method as described above by running the program in the memory.

[0040] On the other hand, the present invention provides a storage medium, comprising: a computer program stored on the storage medium, and when the computer program is executed by a processor, each step of the target behavior identification method described in any one of the above is implemented.

[0041] The target behavior identification method, device, electronic device and storage medium of the present application include obtaining target video data and target audio data, wherein the target video data and the target audio data are respectively obtained by collecting video data and audio data of the detection area where the detection target is located. Based on the target video data and the target audio data, it is possible to determine whether the detection target exhibits the target behavior, thereby ensuring the control effect.

[0042] Furthermore, this solution combines target video and audio data to identify target behavior from multiple modalities, effectively improving the accuracy of target behavior recognition. Applying this solution to the management and control of sand mining areas such as rivers and seas can promptly detect whether sand mining vessels are engaging in illegal mining, ensuring effective control of sand mining areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0044] Figure 1 This is a flow chart of a target behavior identification method provided by an embodiment of the present application;

[0045] Figure 2 This is a flowchart of determining whether a detection target performs a target behavior according to an embodiment of the present application;

[0046] Figure 3 This is a flowchart of determining whether a detection target meets the conditions for performing a target behavior in a detection area, provided by an embodiment of the present application;

[0047] Figure 4 This is a flow chart of obtaining the length of time a detection target stays in a detection area, provided in an embodiment of the present application;

[0048] Figure 5 This is a flow chart of tracking and detecting a target from a surveillance video provided by an embodiment of the present application;

[0049] Figure 6 is a schematic diagram of the union area of two bounding box areas provided in an embodiment of the present application;

[0050] Figure 7 is a schematic diagram of the intersection area of two bounding box areas provided in an embodiment of the present application;

[0051] Figure 8 This is a schematic diagram of the structure of a target behavior recognition device provided in an embodiment of the present application;

[0052] Figure 9 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The technical solution of the embodiment of the present application is applicable to application scenarios for identifying the target behavior of the detection target in the detection area, for example, it is applicable to application scenarios for identifying the illegal mining behavior of sand mining ships in the sand mining area. The technical solution of the embodiment of the present application, combined with the video data and audio data of the detection area where the detection target is located, performs target behavior recognition based on multiple modal data, effectively improving the accuracy of target behavior recognition. Applied to the management and control of the sand mining area, it can promptly detect whether the sand mining ship is engaging in illegal mining, thereby ensuring the management and control effect of the sand mining area.

[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0055] This embodiment proposes a method for identifying target behavior, see Figure 1 As shown, the method includes:

[0056] S101: Acquire target video data and target audio data.

[0057] The target video data is obtained by collecting video data of the detection area where the detection target is located; the target audio data is obtained by collecting audio data of the detection area where the detection target is located.

[0058] Specifically, the target video data and target audio data mentioned above refer to the video and audio data collected when the target is within the detection area. In this case, the target video data includes the image of the target, and the target audio data includes the audio signals generated by the target's behavior.

[0059] The detection target refers to the subject whose behavior needs to be identified, the detection area refers to the area where the behavior of the detection target is identified, and the target behavior refers to the behavior that needs to be controlled. For example, in the application scenario of identifying the illegal mining behavior of sand dredging ships in a sand mining area, the detection target refers to the sand dredging ship, the detection area refers to the sand mining area, and the target behavior refers to the sand mining behavior of the sand dredging ship. For another example, in the application scenario of identifying the illegal mining behavior of excavators in a soil mining area, the detection target refers to the excavator, the detection area refers to the soil mining area, and the target behavior refers to the soil mining behavior of the excavator.

[0060] Video data acquisition can be achieved by installing cameras in the detection area. The cameras are installed within the detection area, or at the periphery of the detection area, so that the video acquisition range of the cameras covers the entire detection area. It should be noted that the number of cameras required and the installation position of each camera can be determined based on the total size of the detection area and the video acquisition range of each camera, to ensure that there are no blind spots in the detection area where video data cannot be acquired, and that the video acquisition range of the cameras covers the entire detection area.

[0061] For example, the imaging device may be a camera. In order to facilitate nighttime control and realize nighttime shooting operations, the imaging device may also be a camera with night vision function, such as an infrared camera.

[0062] As another example, the camera device can adopt a fixed camera device that only shoots in a single direction, or a rotating camera device that can rotate and shoot at a certain angle. The video acquisition range of the rotating camera device is larger than the video acquisition range of the fixed camera device, so the use of a rotating camera device can reduce the number of camera devices installed and save installation costs. However, if a rotating camera device is used, since the rotating camera device is rotating to shoot, it is impossible to collect video data of the entire detection area at the same time, so it may not be possible to detect the detection target entering the detection area in time, and then it is impossible to identify in time whether the detection target performs the target behavior, affecting the management and control effect. Based on this, in order to be able to identify in time whether the detection target performs the target behavior and ensure the management and control effect, this embodiment preferably adopts a fixed camera device.

[0063] Audio data collection can be achieved by installing a recording device in the detection area. The recording device can be installed in the detection area, or even outside the detection area, so that the audio collection range of the recording device covers the entire detection area. It should be noted that the number of recording devices required and the installation location of each recording device can be determined based on the total size of the detection area and the audio collection range of each recording device to ensure that there are no blind spots in the detection area where audio data cannot be collected, and that the audio collection range of the audio collection device covers the entire detection area.

[0064] For example, the recording device can be a voice recorder. Moreover, the detection area is generally a relatively open outdoor area with a lot of ambient noise. In order to ensure the quality of sound collection and the accuracy of target behavior recognition, a recording device with noise reduction function can be used.

[0065] As another example, recording devices can be evenly arranged in the detection area according to the audio collection range of the recording devices to avoid blind spots where audio data cannot be collected. A larger number of recording devices can also be set in areas with high incidence of target behavior in the detection area to improve the quality of audio data collection.

[0066] In addition, if there is a special location in the detection area where video data collection and audio data collection can be performed simultaneously, then a camera and a recording device can be set up at this location at the same time; in order to reduce the amount of equipment installed, an audio and video collection device with both camera and recording functions can also be set up to obtain the target video data and target audio data collected by the audio and video collection device.

[0067] For example, in the application scenario of identifying the illegal mining behavior of sand mining ships in the sand mining area, the sand mining area is generally in the river channel, so the camera device can be installed on one side or both sides of the river channel, and the recording device can be set underwater to facilitate the collection of the sound of the sand mining ship when mining sand; in the application scenario of identifying the illegal mining behavior of excavators in the soil mining area, the soil mining area is generally in a deep pit for soil mining, so the camera device can be installed around the deep pit, and the recording device can be set in the deep pit to facilitate the collection of the sound of the excavator when taking soil. In other application scenarios for identifying the target behavior of the detection target in the detection area, those skilled in the art can combine the records of the installation positions of the camera device and the recording device in this embodiment and set the camera device and the recording device according to actual conditions, which will not be elaborated in this embodiment.

[0068] The camera device and the audio recording device can collect data synchronously, starting from the same time point to collect target video data and target audio data of the same duration; or they can collect data asynchronously, which is not limited in this embodiment.

[0069] S102: Determine whether the detection target performs a target behavior based on the target video data and the target audio data.

[0070] After acquiring the target video data and the target audio data, the target video data and the target audio data are analyzed and processed so as to identify the target behavior from the two modalities of video and audio, and then determine whether the detection target performs the target behavior.

[0071] Specifically, the target video data and target audio data can be fused. The target video data only contains image information, while the target audio data only contains sound information. By fusing the separate target video data and the separate target audio data, fused information containing both image and sound information is obtained. Feature extraction is performed based on the fused target video data and target audio data to determine whether the target is performing the target behavior.

[0072] In this embodiment, based on the target video data and target audio data, it is possible to determine whether the target behavior is being detected, thereby ensuring the control effect. Furthermore, this embodiment combines the target video data and target audio data to identify the target behavior from multiple modalities, effectively improving the accuracy of target behavior identification.

[0073] If the above scheme is applied to the application scenario of identifying illegal mining behavior of sand dredging ships in the sand dredging area, video data and audio data can be collected from the sand dredging area where the sand dredging ship is located to obtain target video data and target audio data. By analyzing the target video data and target audio data, it can be determined whether the sand dredging ship is engaging in illegal mining, thereby ensuring the management and control effect of the sand dredging area.

[0074] As an optional implementation, such as Figure 2 As shown, in another embodiment of the present application, the steps of the above embodiment are based on the target video data and the target audio data to determine whether the detection target performs the target behavior, which can be specifically achieved by the following steps:

[0075] S201: Fuse target video data and target audio data to obtain fusion information.

[0076] This fusion combines target video data containing only image information with target audio data containing only sound information, resulting in fused information containing both image and sound information. The image information contains the image features of the target's behavior, and the sound information contains the sound features of the target's behavior. Therefore, fused information containing both image and sound information also contains both audio and video features of the target's behavior.

[0077] S202: extracting behavioral features of the detection target from the fused information.

[0078] The aforementioned behavioral features are used to characterize the behavior of the detection target. Specifically, each fused information may contain multiple behavioral features of the detection target, and these multiple behavioral features need to be combined and analyzed to determine the specific behavior of the detection target. For example, in the application scenario of identifying illegal sand mining by sand dredging vessels in a sand mining area, the behavioral features of the detection target extracted from the fused information include three video features: turbid water around the sand dredging vessel, ripples in the water around the sand dredging vessel, and the sand dredging vessel's location in a sand-rich area; as well as two audio features: the sound of the sand dredging vessel's engine during sand mining and the sound of sand and gravel excavation. These five behavioral features need to be combined to make a comprehensive judgment to determine whether the sand dredging vessel is engaged in sand mining. If the analysis determines that the above five behavioral features are those of the sand dredging vessel's illegal sand mining, it can be determined that the sand dredging vessel is engaged in sand mining.

[0079] Exemplarily, the above-mentioned step of extracting the behavioral features of the detection target from the fused information can be implemented based on a pre-trained feature extraction model. The fused information corresponding to various possible behaviors of the detection target in the detection area is used as training samples, and behavioral features are annotated for each fused information as sample labels. The feature extraction model is trained until the output of the feature extraction model meets the training requirements, and the feature extraction model training is completed. The fused information obtained by fusing the target video data and the target audio data in this embodiment is input into the trained feature extraction model to obtain the behavioral features of the detection target.

[0080] For example, in an application scenario involving identifying illegal sand mining by dredging vessels within a sand mining area, training samples can include first-order fused information of the dredging vessel mining within the area, second-order fused information of the dredging vessel moving within the area, and third-order fused information of the dredging vessel stationary within the area. The first-order fused information can label behavioral features such as turbid water surrounding the dredging vessel, ripples in the water surrounding the dredging vessel, the sound of the dredging vessel's engine while mining, and the sound of sand and gravel excavation. The second-order fused information can label turbid water surrounding the dredging vessel, ripples in the water surrounding the dredging vessel, and the sound of the dredging vessel's engine while moving. The third-order fused information can label behavioral features such as clear water surrounding the dredging vessel, still water surrounding the dredging vessel, and the absence of engine sounds, to train the feature extraction model. After the feature extraction model is trained, the resulting fused information is input into the trained feature extraction model to obtain the behavioral features of the dredging vessel.

[0081] It should be noted that the above-mentioned fusion information and behavioral features corresponding to the various behaviors of sand mining ships are only used to assist in explaining the role of the behavioral features of the detection target and the training process of the feature extraction model, and do not form a specific limitation. When this solution is applied to other scenarios, such as the application scenario of identifying the illegal mining behavior of excavators in the soil mining area, technical personnel in this field can determine the various behaviors that the detection target may perform and the behavioral features corresponding to each behavior without consuming any creativity, and then select training samples and labels for model training to realize the extraction of the detection target behavior features. Therefore, this embodiment does not elaborate on them one by one.

[0082] S203: Determine whether the detection target performs the target behavior based on the behavior characteristics.

[0083] In this step, a comprehensive judgment is made as to whether the behavioral characteristics of the detection target extracted from the fused information are consistent with the behavioral characteristics of the pre-set target behavior; if the behavioral characteristics of the detection target extracted from the fused information are consistent with the behavioral characteristics of the target behavior, it indicates that the detection target is in a state of executing the target behavior; if the behavioral characteristics of the detection target extracted from the fused information are not consistent with the behavioral characteristics of the target behavior, it indicates that the detection target is not in a state of executing the target behavior.

[0084] For example, in an application scenario where we need to identify illegal sand mining by dredging vessels in a sand mining area, we can pre-set the behavior as long as the behavior features include the sound of sand excavation. If the behavior features extracted from the fused information include turbid water around the dredging vessel, ripples in the water around the dredging vessel, and the sound of sand excavation, then the behavior meets the characteristics of a sand dredging vessel, indicating that the dredging vessel is currently engaged in sand mining.

[0085] It should be noted that this is merely an example of extracting the behavioral characteristics of the detection target from the fusion information, and does not constitute any limitation. When applying this solution to any specific scenario, those skilled in the art can set the behavioral characteristics of the target behavior according to the actual situation, and this embodiment will not be repeated.

[0086] In this embodiment, the target behavior is identified from both audio features and video features. Compared with the target behavior identification method of only performing video detection to extract and detect target video features, and the target behavior identification method of only performing audio detection to extract and detect target audio features, the accuracy of target behavior identification can be effectively improved.

[0087] Furthermore, the steps of the above embodiment fuse the target video data and the target audio data to obtain fused information, extract the behavioral features of the detection target from the fused information, and determine whether the detection target performs the target behavior based on the behavioral features. This can be specifically achieved through the following steps:

[0088] The target video data and target audio data are input into a pre-trained target behavior recognition model, so that the target behavior recognition model fuses the target video data and the target audio data to obtain fusion information, extracts the behavior features of the detection target from the fusion information, and obtains the target behavior recognition result based on the behavior features.

[0089] A pre-trained target behavior recognition model can be used to determine whether the target has performed the target behavior. Specifically, the target video data and target audio data can be input into the pre-trained target behavior recognition model. The target behavior recognition model processes the target video data and target audio data, including fusing the target video data and target audio data, extracting behavioral features from the fused information, and obtaining a target behavior recognition result based on the behavioral features. The target behavior recognition result may include whether the target has performed the target behavior or whether the target has not performed the target behavior.

[0090] In this embodiment, target behavior recognition from both audio and video features can effectively improve the accuracy of target behavior recognition. Furthermore, using a pre-trained target behavior recognition model to identify and detect whether a target is performing a target behavior is not only simple to operate, but also, if the required number of training samples is met, the target behavior recognition model can achieve a high accuracy rate, further improving the accuracy of target behavior recognition.

[0091] Furthermore, before the steps of the above embodiment input the target video data and the target audio data into the pre-trained target behavior recognition model, the following steps are further included:

[0092] Perform frame extraction processing on the target video data and / or perform downsampling processing on the target audio data.

[0093] Specifically, the target video data and the target audio data are time-synchronized data, where the time-synchronized data means that the target video data and the target audio data have the same duration and are aligned according to the timeline.

[0094] If the camera device and the recording device collect data synchronously and collect target video data and target audio data of the same length starting from the same time point, then the target video data and the target audio data have the same length and are aligned, and the target video data can be frame-extracted and / or the target audio data can be downsampled.

[0095] If the camera and the audio recording device are not synchronously collected, the target video data and the target audio data need to be aligned at the same time. After alignment, the target video data and the target audio data of the same duration are intercepted from the aligned portion. It should be noted here that the target video data and the target audio data can be aligned and intercepted first, and then the target video data is subjected to frame extraction and / or the target audio data is subjected to downsampling. Alternatively, the target video data can be subjected to frame extraction and / or the target audio data is subjected to downsampling, and then the target video data and the target audio data are aligned and intercepted. However, the method of first performing frame extraction and / or downsampling on the target video data and then performing alignment and interception on the target video data and the target audio data requires a large amount of calculation. Therefore, in this embodiment, it is preferred to first perform alignment and interception on the target video data and then perform frame extraction and / or downsampling on the target audio data.

[0096] For example, video frames can be extracted at preset time intervals to achieve downsampling of the target video data. The preset time interval can be set according to actual conditions and is not limited in this embodiment. For example, a video frame can be extracted every 1 second. If the target video data is 30 seconds long, the target video data will still be 30 seconds long after the frame extraction process, containing a total of 30 video frames.

[0097] In another exemplary embodiment, the target audio data may be downsampled according to a preset standard and then normalized after the downsampling. The preset standard may be set according to actual conditions and is not limited in this embodiment. For example, the target audio data may be downsampled to 1024 times per second and normalized to between 0 and 1.

[0098] In this embodiment, the target video data and / or target audio data are downsampled before being input into the pre-trained target behavior recognition model to reduce the computational complexity of the target behavior recognition model and improve the computational speed of the target behavior recognition model.

[0099] As an optional implementation method, the target behavior recognition model can be trained through the following steps:

[0100] Input the training sample into the target behavior recognition model to obtain the output result of the target behavior recognition model; determine the loss value of the target behavior recognition model based on the output result and label of the target behavior recognition model; adjust the operation parameters of the target behavior recognition model based on the loss value; repeat the above process until the loss value of the target behavior recognition model is less than the set value.

[0101] Specifically, a large amount of video data of the detection target when the detection target exhibits target behavior in the detection area, audio data of the detection target when the detection target exhibits target behavior in the detection area, video data of the detection target when the detection target does not exhibit target behavior in the detection area, audio data of the detection target when the detection target does not exhibit target behavior in the detection area, video data of the detection area when the detection target does not appear in the detection area, and audio data of the detection area when the detection target does not appear in the detection area can be collected as sample data.

[0102] For example, in an application scenario for identifying illegal mining activities of sand dredging ships in a sand dredging area, a large amount of video data of the sand dredging ships when they engage in sand dredging activities in the sand dredging area, audio data of the sand dredging ships when they engage in sand dredging activities in the sand dredging area, video data of the sand dredging ships when they do not engage in sand dredging activities in the sand dredging area, audio data of the sand dredging ships when they do not engage in sand dredging activities in the sand dredging area, video data of the sand dredging area when the sand dredging ships do not appear in the sand dredging area, and audio data of the sand dredging area when the sand dredging ships do not appear in the sand dredging area can be collected as sample data.

[0103] Another exemplary application scenario for identifying illegal mining behavior by excavators in a soil mining area can collect a large amount of video data of the excavator when it is taking soil in the soil mining area, audio data of the excavator when it is taking soil in the soil mining area, video data of the excavator when it is not taking soil in the soil mining area, audio data of the excavator when it is not taking soil in the soil mining area, video data of the soil mining area when the excavator is not in the soil mining area, and audio data of the soil mining area when the excavator is not in the soil mining area as sample data.

[0104] If this solution is applied to other scenarios, those skilled in the art can determine the sample data without any creative effort, and this embodiment will not elaborate on it.

[0105] Then, the sample data is preprocessed to obtain training samples. The preprocessing steps include:

[0106] All video data and audio data are cropped according to a preset length, and video data and audio data that are shorter than the preset length are discarded. The preset length can be set according to actual conditions. For example, the preset length can be set to 30 seconds, which is not limited in this embodiment. The cropped video data is subjected to frame extraction processing, and the cropped audio data is subjected to downsampling processing to obtain training samples. The frame extraction process here is the same as the frame extraction process described in the above embodiment, and the downsampling process is the same as the downsampling process described in the above embodiment. Those skilled in the art can refer to the description of the above embodiment, and will not be described in detail here.

[0107] The screen sizes corresponding to the video data captured by different shooting devices may be different. In this embodiment, the video data is cropped or scaled so that the screen sizes corresponding to all the video data are the same.

[0108] It should be noted that, in this step, the video data may be first subjected to frame extraction processing, the audio data may be subjected to downsampling processing, and then all the video data and audio data may be cropped according to a preset length. This embodiment does not limit this.

[0109] It should also be noted that for audio data and video data collected at the same time in the same detection area, the video data and audio data should first be aligned according to the same timeline. When all the video data and audio data are cropped according to the preset length, only the aligned part of the video data and audio data should be retained.

[0110] After the above preprocessing steps are completed, the training samples are input into the target behavior recognition model to train the target behavior recognition model. The target behavior recognition model in this embodiment includes a backbone network and a fully connected layer, wherein the fully connected layer is connected to the output end of the backbone network. Exemplarily, a Swin Transformer can be used as the backbone network.

[0111] Specifically, the training samples are input into the backbone network of the target behavior recognition model. For example, if the video data has a screen size of 224×224, 3 channels, and a length of 30 seconds, then the video data input dimensions are 224×224×3×30. If the audio data has 1024 times per second and is 30 seconds long, then the audio data input dimensions are 1024×30.

[0112] The backbone network fuses the video and audio data from the training samples to obtain fused information. Features are extracted from the fused information corresponding to each set of video and audio data, and the features are normalized to obtain normalized feature data corresponding to each set of video and audio data. The size of the normalized feature data can be adjusted based on actual conditions. In this embodiment, the size of the normalized data is k × 1, where k is a hyperparameter and is set to 512.

[0113] The backbone network inputs the normalized feature data corresponding to each set of video and audio data into the fully connected layer. The fully connected layer integrates the normalized data and outputs the target behavior recognition results corresponding to each set of video and audio data as the output of the target behavior recognition model. In this embodiment, we only focus on the target behavior recognition results corresponding to the video and audio data of the target when the target behavior occurs in the detection area, the target behavior recognition results corresponding to the video and audio data of the target when the target does not occur in the detection area, and the target behavior recognition results corresponding to the video and audio data of the detection area when the target does not occur in the detection area. Therefore, the output size of the fully connected layer is set to 3×1.

[0114] Moreover, since this embodiment only focuses on the target behavior recognition results of the above three pairs of video data and audio data combinations, labels can be set only for the above three pairs of video data and audio data combinations. For example, in the training sample, when the target behavior occurs in the detection area, the video data and audio data of the target are labeled as 100, when the target behavior does not occur in the detection area, the video data and audio data of the target are labeled as 010, and when the target does not appear in the detection area, the video data and audio data of the detection area are labeled as 001.

[0115] Based on the output results and labels of the target behavior recognition model, the loss value of the target behavior recognition model is determined, and the operational parameters of the target behavior recognition model are adjusted with the goal of reducing the loss value; the above process is repeated until the loss value of the target behavior recognition model is less than the set value. The above adjustment of the operational parameters of the target behavior recognition model with the goal of reducing the loss value can be achieved by adjusting the loss function. For example, the loss function uses cross entropy loss.

[0116] The target video data and target audio data are input into a pre-trained target behavior recognition model, and the target behavior recognition result is determined based on the output of the target behavior recognition model. In this embodiment, since the output size of the fully connected layer is set to 3×1, the output of the target behavior recognition model consists of a matrix of three data points. In this embodiment, the subscript corresponding to the data point with the largest value in the matrix is taken as the output of the target behavior recognition model. When the output result is 0, it indicates that the detected target is performing the target behavior in the detection area. For example, if the output matrix is [0.80.10.1], then the data point with the largest value in the matrix is 0.8, and the subscript corresponding to 0.8, that is, the order of 0.8 in the matrix is 0, then the output of the target behavior recognition model is 0, indicating that the detected target is performing the target behavior in the detection area. For another example, if the output matrix is [0.30.90.1], then the data point with the largest value in the matrix is 0.9, and the subscript corresponding to 0.9, that is, the order of 0.9 in the matrix is 1, then the output of the target behavior recognition model is 1, indicating that the detected target is not performing the target behavior in the detection area.

[0117] In this embodiment, the target behavior recognition model is trained from two perspectives: audio features and video features, so as to further improve the accuracy of target behavior recognition.

[0118] As an optional implementation method, another embodiment of the present application discloses that in the steps of the above embodiment, if it is determined that the detection target performs the target behavior in the detection area, an early warning information is generated and the early warning information is sent to the user end, and / or the target video data and target audio data are sent to the user end.

[0119] Specifically, if it is determined that the detected target is performing the target behavior in the detection area, the corresponding management personnel of the detection area need to be notified. In this embodiment, an early warning message can be generated and sent to the user terminal to inform the management personnel that the detected target is performing the target behavior in the detection area. In addition, the target video data and target audio data can also be sent to the user terminal so that the management personnel can make further judgments based on the target video data and target audio data.

[0120] Exemplarily, the warning information is sent to the user terminal and the target video data and target audio are sent to the user terminal, which can be sent via mobile phone text messages, emails, WeChat messages, DingTalk messages, etc.

[0121] In this embodiment, when it is determined that the detection target is performing the target behavior in the detection area, an early warning is issued to the management personnel, which facilitates the management of the detection area by the detection personnel. This solution is applied to the application scenario of identifying the illegal mining behavior of sand mining ships in the sand mining area. When it is determined that the sand mining ship is performing the sand mining behavior in the sand mining area, an early warning is issued to the management personnel of the sand mining area, so that the management personnel can promptly discover the illegal mining behavior and strengthen the management of the sand mining area.

[0122] As an optional implementation, another embodiment of the present application discloses that before obtaining the target video data and the target audio data in the steps of the above embodiment, the following steps may also be included:

[0123] Determine whether the detection target has the conditions to perform the target behavior in the detection area; if the detection target has the conditions to perform the target behavior in the detection area, execute the steps of obtaining target video data and target audio data.

[0124] Specifically, in this implementation, before obtaining the target video data and target audio data and identifying the behavior of the detection target, it is first determined whether the detection target has the conditions to perform the target behavior in the detection area. When the detection target has the conditions to perform the target behavior in the detection area, the steps of obtaining the target video data and target audio data of the above embodiment are executed to identify the behavior of the detection target and then determine whether the detection target performs the target behavior.

[0125] For example, if the detection target appears in the detection area, or the detection target appears in the detection area and stays in the detection area for a preset time, it means that the detection target has the conditions to perform the target behavior in the detection area.

[0126] In this embodiment, when it is detected that the detection target has the conditions to perform the target behavior in the detection area, the steps of obtaining the target video data and the target audio data in the above embodiment are triggered to reduce the amount of calculation in the process of identifying the target behavior.

[0127] As an optional implementation, such as Figure 3 As shown, in another embodiment of the present application, the steps of the above embodiment are used to determine whether the detection target has the conditions to perform the target behavior in the detection area, which can be specifically achieved through the following steps:

[0128] S301: Obtain the duration that the detection target stays in the detection area.

[0129] Determine whether a detection target appears in the detection area. When a detection target appears in the detection area, start timing to determine the length of time the detection target remains in the detection area. In this embodiment, the method for obtaining the length of time the detection target remains in the detection area is not limited, as long as the method can accurately obtain the length of time the detection target remains in the detection area.

[0130] Exemplarily, whether a detection target appears in the detection area can be determined by radar detection. Specifically, the detection area can be determined as a radar detection area. When the detection target is detected to appear in the radar detection area, timing is started to determine the length of time the detection target stays in the radar detection area, and the length of time the detection target stays in the radar detection area is used as the length of time the detection target stays in the detection area. Whether a detection target appears in the detection area can also be determined by video detection. Specifically, a surveillance video of the detection area can be obtained. When the detection target is detected to appear in the surveillance video of the detection area, timing is started to determine the length of time the detection target stays in the surveillance video of the detection area, and the length of time the detection target stays in the surveillance video of the detection area is used as the length of time the detection target stays in the detection area.

[0131] The above-mentioned residence time refers to the time that the detection target appears in the detection area. After the detection target appears in the detection area, whether it is the time it remains stationary in the detection area or the time it moves in the detection area, it is considered the time that the detection target stays in the detection area.

[0132] S302: If the detection target stays in the detection area for a predetermined time, it is determined that the detection target has the conditions to perform the target behavior in the detection area.

[0133] If, after detection, it is determined that the detection target has stayed in the detection area for a preset time, it means that the detection target has the conditions to perform the target behavior in the detection area, and the steps of obtaining target video data and target audio data in the above embodiment can be executed.

[0134] In this embodiment, when the detection target stays in the detection area for a preset time, the steps of obtaining target video data and target audio data in the above embodiment are executed, and then based on the target video data and target audio data, it is determined whether the detection target performs the target behavior. Compared with the method of executing the steps of obtaining target video data and target audio data in the above embodiment when the detection target appears, this embodiment can further reduce the amount of calculation and improve the recognition speed of the target behavior.

[0135] For example, in an application scenario for identifying illegal mining activities of sand dredging ships in a sand dredging area, it is necessary to detect the length of time the sand dredging ship stays in the sand dredging area. If the length of time the sand dredging ship stays in the sand dredging area reaches a preset length of time, it is determined that the sand dredging ship has the conditions to perform sand dredging activities in the sand dredging area. Then, video data and audio data of the sand dredging area where the sand dredging ship is located are collected to obtain target video data and target audio data. By analyzing the target video data and target audio data, it is determined whether the sand dredging ship has engaged in illegal mining activities.

[0136] As an optional implementation, such as Figure 4 As shown, in another embodiment of the present application, the steps of the above embodiment for obtaining the duration of the detection target staying in the detection area may specifically include the following steps:

[0137] S401: When it is determined that a detection target enters a detection area, the detection target is tracked from a surveillance video of the detection area to determine a duration for which the detection target appears in the surveillance video of the detection area.

[0138] In this embodiment, feature extraction is performed on the surveillance video of the monitoring area to determine whether the detection target has entered the detection area. Exemplarily, the surveillance video of the detection area is used as a training sample, and the presence of the detection target in the detection area is used as a label to train the feature extraction model. The output result of the feature extraction model is obtained; the loss value of the feature extraction model is determined based on the output result and label of the feature extraction model; the operation parameters of the feature extraction model are adjusted based on the loss value; and the above process is repeated until the loss value of the feature extraction model is less than the set value. Among them, the feature extraction model can use the yolo neural network as the basic model.

[0139] The surveillance video of the detection area is input into the above-mentioned feature extraction model to obtain the feature extraction result of the feature extraction model; wherein the feature extraction result of the feature extraction model includes the detection target entering the detection area, or the detection target not entering the detection area.

[0140] For example, in an application scenario for identifying illegal sand mining by dredging vessels in a sand mining area, a feature extraction model is trained using surveillance videos of the sand mining area as training samples and the presence of a sand mining vessel in the sand mining area as a label. After the feature extraction model is trained, the surveillance video of the sand mining area is input into the feature extraction model to obtain feature extraction results. The feature extraction results of the feature extraction model include either a sand mining vessel entering the sand mining area or a sand mining vessel not entering the sand mining area.

[0141] When it is determined that the detection target enters the surveillance video of the detection area, that is, when it is determined that the detection target enters the detection area, the detection target begins to record the duration of its appearance in the surveillance video of the detection area, and the detection target's movement trajectory in the surveillance video of the detection area is tracked to determine whether the detection target disappears from the surveillance video of the detection area. If the detection target does not disappear from the surveillance video of the detection area, the detection target continues to record the duration of its appearance in the surveillance video of the detection area. If the detection target disappears from the surveillance video of the detection area, the detection target ceases to record the duration of its appearance in the surveillance video of the detection area.

[0142] S402: Determine the duration that the detection target appears in the surveillance video of the detection area as the duration that the detection target stays in the detection area.

[0143] The surveillance video of the detection area can reflect the actual situation in the detection area in real time. When the detection target is detected in the surveillance video of the detection area, it is determined that the detection target has entered the detection area; when the detection target disappears from the surveillance video of the detection area, it is determined that the detection target has left the detection area. Therefore, the duration that the detection target appears in the surveillance video of the detection area is equal to the duration that the detection target remains in the detection area.

[0144] In this embodiment, it is only necessary to install a device such as a camera that can collect surveillance video, so that the length of time the detection target stays in the detection area can be determined based on the surveillance video of the monitoring area. The installation cost is low, and the surveillance video of the detection area reflects the actual situation of the detection area in real time, and the monitoring is highly timely.

[0145] For example, in an application scenario for identifying illegal mining activities of sand dredging ships in a sand dredging area, when it is determined that the sand dredging ship has entered the sand dredging area, the sand dredging ship is tracked from the surveillance video of the sand dredging area, and then the length of time the sand dredging ship appears in the surveillance video of the sand dredging area is determined. The length of time the sand dredging ship appears in the surveillance video of the sand dredging area is determined as the length of time the sand dredging ship stays in the sand dredging area.

[0146] As an optional implementation, such as Figure 5 As shown, in another embodiment of the present application, the steps of the above embodiment for tracking the detection target from the surveillance video of the detection area may specifically include the following steps:

[0147] S501: Determine a bounding box area of a detection target in a first video frame of a surveillance video of a detection area.

[0148] The surveillance video of the detection area is composed of multiple video frames. In this embodiment, the bounding box area of the detection target in the first video frame is determined. Feature extraction can be performed on the video frame to obtain the bounding box area of the detection target in the first video frame by extracting the detection target in the first video frame.

[0149] Exemplarily, the above-mentioned step of extracting features from the video frame to obtain the bounding box area of the detection target in the first video frame can be obtained through a pre-trained bounding box labeling model. Specifically, the bounding box labeling model can be trained using the video frame of the detection target as a training sample and the bounding box position of the detection target as a label; the result output by the bounding box labeling model is obtained; the loss value of the bounding box labeling model is determined based on the output result and label of the bounding box labeling model; the operation parameters of the bounding box labeling model are adjusted based on the loss value; and the above process is repeated until the loss value of the bounding box labeling model is less than the set value. Among them, the bounding box labeling model can use the yolo neural network as the basic model.

[0150] After the model training is completed, the first video frame of the surveillance video in which the detection target is in the detection area is input into the bounding box marking model to obtain the output result of the bounding box marking model, and the bounding box area in the first video frame of the surveillance video in which the detection target is in the detection area is determined.

[0151] S502: Identify a bounding box area of the target object from a second video frame of the surveillance video of the detection area.

[0152] The second video frame is the next video frame of the first video frame.

[0153] The target object is an identification object of the same type as the detection target. For example, if the detection target is a sand mining ship, then the target object is also a sand mining ship; if the detection target is an excavator, then the target object is also an excavator.

[0154] Feature extraction can be performed on the second video frame to extract the target object in the second video frame to obtain a bounding box region of the target object in the second video frame. For example, the second video frame can be input into the bounding box labeling model of the above embodiment to determine a bounding box region in the second video frame of the surveillance video of the target object in the detection area.

[0155] S503: If the intersection-over-union ratio of the bounding box area of any target object in the second video frame and the bounding box area of the detection target in the first video frame of the surveillance video of the detection area is greater than a set threshold, the target object in the second video frame is determined to be the detection target.

[0156] A rectangular coordinate system is established on the plane of the surveillance video displaying the detection area to determine the vertex coordinates of the bounding box area of the detection target in the first video frame and the vertex coordinates of the bounding box area of any target object in the second video frame.

[0157] Based on the vertex coordinates of the bounding box area of the detection target in the first video frame and the vertex coordinates of the bounding box area of each target object in the second video frame, the intersection area and the union area of the bounding box area of the detection target in the first video frame and the bounding box area of each target object in the second video frame can be determined, and the ratio of the intersection area and the union area of the bounding box area of the detection target in the first video frame and the bounding box area of each target object in the second video frame is calculated. The ratio is compared with a set threshold. If there is a bounding box area of a target object in the second video frame, and the intersection-union ratio with the bounding box area of the detection target in the first video frame of the surveillance video of the detection area is greater than the set threshold, it indicates that the target object in the second video frame is the detection target.

[0158] If there are special circumstances where at least two IoU ratios are greater than the set threshold, the target object corresponding to the largest IoU ratio can be taken as the detection target; or manual identification can be performed; the feature similarity between the detection target and the target object can be further calculated, and the detection target can be determined in the second video frame based on the feature similarity.

[0159] If all the calculated intersection-over-union ratios are less than the set threshold, it means that the target objects in the second video frame are not the detection targets in the first video frame, and the detection targets have left the detection area.

[0160] like Figure 6 and Figure 7 As shown, the vertices of the bounding box area of the detection target in the first video frame are A1, B1, C1, and D1. The coordinates of A1 are (x11, y11), the coordinates of B1 are (x12, y12), the coordinates of C1 are (x13, y13), and the coordinates of D1 are (x14, y14). The vertices of the bounding box area of the target object are A2, B2, C2, and D2. The coordinates of A2 are (x21, y21), the coordinates of B2 are (x22, y22), the coordinates of C2 are (x23, y23), and the coordinates of D2 are (x24, y24). The intersection points of the bounding box area in the first video frame and the bounding box area of the target object are a1 and a2.

[0161] but Figure 6 The area S1 formed by the midpoints A1, B1, a1, B2, C2, D2, a2, and D1 is the union area of the bounding box area of the detection target in the first video frame and the bounding box area of the target object in the second video frame; Figure 7The area S2 formed by midpoints A2, a2, C1, and a1 is the intersection area of the bounding box area of the detection target in the first video frame and the bounding box area of the target object in the second video frame. It should be noted that if the bounding box area in the first video frame does not intersect with the bounding box area of the target object, then the union area of the bounding box area of the detection target in the first video frame and the bounding box area of the target object in the second video frame is the sum of the areas of the two bounding box areas, and the intersection area of the bounding box area of the detection target in the first video frame and the bounding box area of the target object in the second video frame is zero.

[0162] In this embodiment, as the monitoring video of the detection area is generated, the area of the bounding box region of the detection target in the first video frame in which the detection target exists is continuously calculated, and the intersection and union ratio of the area of the bounding box region of each target object in the second video frame is calculated to determine whether the detection target exists in the next video frame. If the detection target exists in the next video frame, the movement trajectory of the detection target can be generated, and then the detection target can be tracked based on the movement trajectory.

[0163] For example, in a scenario where there are multiple target objects in a video frame, in order to further improve the accuracy of tracking the detected target, another embodiment of the present application calculates the intersection-over-union ratio of the bounding box area of the detected target in the first video frame and the bounding box area of any target object in the second video frame, and calculates the feature similarity between the detected target in the first video frame and any target object in the second video frame, so as to determine the detected target among the multiple target objects in the second video frame and achieve tracking of the detected target.

[0164] Specifically, determine the feature similarity between the detection target in the first video frame of the surveillance video of the detection area and each target object in the second video frame of the surveillance video of the detection area; and determine the intersection-and-union ratio of the bounding box area of the detection target in the first video frame of the surveillance video of the detection area and the bounding box area of each target object in the second video frame of the surveillance video of the detection area; calculate the feature similarity between the detection target in the first video frame of the surveillance video of the detection area and each target object in the second video frame of the surveillance video of the detection area, and the weighted sum of the intersection-and-union ratios of the bounding box area of the detection target in the first video frame of the surveillance video of the detection area and the bounding box area of each target object in the second video frame of the surveillance video of the detection area; determine that the target object whose weighted sum is greater than a preset weighted threshold is the detection target.

[0165] The feature similarity calculation described above can be performed using a similarity calculation model. For example, the video frames of the detected targets are used as training samples, and the similarity between each pair of detected targets is used as a label to train the similarity calculation model. The first and second video frames are input into the trained similarity calculation model, and the similarity calculation model outputs the feature similarity between the monitored target and each target object.

[0166] The calculation method of the above-mentioned intersection-and-union ratio is the same as the calculation method of the intersection-and-union ratio of the bounding box area of the detected target in the first video frame and the bounding box area of each target object in the second video frame in the above embodiment. Those skilled in the art can refer to the records of the above embodiments and will not be repeated here.

[0167] When calculating the weighted sum of the feature similarity and the intersection-over-union ratio, the weights corresponding to the feature similarity and the intersection-over-union ratio can be set according to actual conditions, and this embodiment does not limit this.

[0168] In this embodiment, the target object whose weighted sum is greater than a preset weighted threshold is determined to be the detection target. If, in a special case, the weighted sum corresponding to at least two target objects is greater than the preset weighted threshold, the target object corresponding to the weighted sum with the largest value is selected as the detection target, or manual detection is performed. If the weighted sum corresponding to all target objects is less than the preset weighted threshold, it indicates that the target object in the second video frame is not the detection target in the first video frame and the detection target has left the detection area.

[0169] In this embodiment, the detection target is identified from multiple target objects in the second video frame based on feature similarity and intersection-over-union ratio, which can effectively improve the accuracy of detection target identification.

[0170] Corresponding to the above target behavior recognition method, the present application embodiment also discloses a target behavior recognition device, see Figure 8 As shown, the device includes:

[0171] An acquisition module 100 is configured to acquire target video data and target audio data; wherein the target video data and target audio data are acquired by respectively collecting video data and audio data of a detection area where the detection target is located;

[0172] The determination module 110 is configured to determine whether the detection target performs a target behavior based on the target video data and the target audio data.

[0173] In this embodiment, the acquisition module 100 acquires target video data and target audio data, wherein the target video data and target audio data are respectively obtained by collecting video data and audio data of the detection area where the detection target is located. The determination module 110 can determine whether the detection target exhibits target behavior based on the target video data and target audio data, thereby ensuring the control effect.

[0174] Furthermore, this solution combines target video and audio data to identify target behavior from multiple modalities, effectively improving the accuracy of target behavior recognition. Applying this solution to the management and control of sand mining areas such as rivers and seas can promptly detect whether sand mining vessels are engaging in illegal mining, ensuring effective control of sand mining areas.

[0175] As an optional implementation, another embodiment of the present application discloses that the determination module 110 in the above embodiment includes:

[0176] a fusion unit, configured to fuse target video data and target audio data to obtain fusion information;

[0177] An extraction unit, used to extract the behavioral features of the detection target from the fusion information;

[0178] The determination unit is used to determine whether the detection target performs the target behavior based on the behavior characteristics.

[0179] As an optional implementation, another embodiment of the present application discloses that the determination module 110 in the above embodiment, when fusing the target video data and the target audio data to obtain fusion information, extracting the behavior characteristics of the detection target from the fusion information, and determining whether the detection target performs the target behavior based on the behavior characteristics, is specifically configured to:

[0180] Inputting the target video data and the target audio data into a pre-trained target behavior recognition model, causing the target behavior recognition model to fuse the target video data and the target audio data to obtain fused information, extracting the behavioral features of the detection target from the fused information, and obtaining the target behavior recognition result based on the behavioral features;

[0181] The target behavior recognition result includes detecting that the target performs the target behavior, or detecting that the target does not perform the target behavior.

[0182] As an optional implementation, another embodiment of the present application discloses that the determining module 110 in the above embodiment further includes:

[0183] A pre-processing module, configured to perform frame extraction processing on the target video data and / or down-sample the target audio data before inputting the target video data and the target audio data into a pre-trained target behavior recognition model;

[0184] The target video data and the target audio data are time-synchronized data.

[0185] As an optional implementation, another embodiment of the present application discloses that the target behavior recognition device of the above embodiment further includes:

[0186] The condition determination module is also used to determine whether the detection target has the conditions to perform the target behavior in the detection area; if the detection target has the conditions to perform the target behavior in the detection area, the acquisition module 100 executes the steps of acquiring target video data and target audio data.

[0187] As an optional implementation, another embodiment of the present application discloses that the condition determination module in the above embodiment further includes:

[0188] An acquisition unit, used to obtain the length of time the detection target stays in the detection area;

[0189] The condition determination unit is used to determine that the detection target has the conditions to perform the target behavior in the detection area if the detection target stays in the detection area for a preset time.

[0190] As an optional implementation, another embodiment of the present application discloses that the acquisition unit of the above embodiment includes:

[0191] A tracking subunit is configured to, when it is determined that a detection target enters a detection area, track the detection target from the surveillance video of the detection area to determine the duration that the detection target appears in the surveillance video of the detection area;

[0192] The determination subunit is used to determine the duration that the detection target appears in the monitoring video of the detection area as the duration that the detection target stays in the detection area.

[0193] As an optional implementation, another embodiment of the present application discloses that the tracking subunit of the above embodiment, when tracking the detection target from the surveillance video of the detection area, is specifically used to:

[0194] Determine a bounding box area of the detection target in a first video frame of a surveillance video of the detection area;

[0195] Identifying a bounding box area of a target object from a second video frame of a surveillance video of the detection area, wherein the second video frame is a subsequent video frame of the first video frame, and the target object is an identification object of the same type as the detection target;

[0196] If the intersection-and-union ratio of the bounding box area of any target object in the second video frame and the bounding box area of the detection target in the first video frame of the surveillance video of the detection area is greater than the set threshold, the target object in the second video frame is determined to be the detection target.

[0197] As an optional implementation method, disclosed in another embodiment of the present application, the target video data of the above embodiment includes monitoring video data of the sand mining area, the target audio data of the above embodiment includes audio data collected from the sand mining area, and the target behavior of the above embodiment includes sand mining behavior.

[0198] Specifically, for the specific working contents of each unit of the above-mentioned target behavior recognition device, please refer to the contents of the above-mentioned method embodiment, which will not be repeated here.

[0199] Another embodiment of the present application further provides an electronic device, see Figure 9 As shown, the device includes:

[0200] Memory 200 and processor 210;

[0201] The memory 200 is connected to the processor 210 and is used to store programs;

[0202] The processor 210 is configured to implement the target behavior recognition method disclosed in any of the above embodiments by running the program stored in the memory 200 .

[0203] Specifically, the electronic device may further include: a bus, a communication interface 220 , an input device 230 and an output device 240 .

[0204] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are interconnected via a bus.

[0205] A bus may include a pathway that transfers information between components of a computer system.

[0206] Processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, or the like, or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component.

[0207] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.

[0208] The memory 200 stores a program for executing the technical solution of the present application, and may also store an operating system and other key services. Specifically, the program may include program code, and the program code includes computer operating instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, disk storage, flash, etc.

[0209] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.

[0210] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speakers, etc.

[0211] The communication interface 220 may include any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0212] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of the target behavior identification method provided in the above embodiment of the present application.

[0213] Another embodiment of the present application further provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the various steps of the target behavior identification method provided in any of the above embodiments.

[0214] Specifically, the specific working contents of each part of the above-mentioned electronic device, as well as the specific processing contents of the computer program on the above-mentioned storage medium when being executed by the processor, can be found in the contents of each embodiment of the above-mentioned target behavior identification method, and will not be repeated here.

[0215] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0216] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.

[0217] The steps in the methods of each embodiment of the present application can be adjusted in sequence, merged, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0218] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be merged, divided, and deleted according to actual needs.

[0219] In the several embodiments provided in this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or submodules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or module, which can be electrical, mechanical or other forms.

[0220] The modules or submodules described as separate components may or may not be physically separate, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules may be selected to achieve the purpose of this embodiment according to actual needs.

[0221] In addition, each functional module or submodule in each embodiment of the present application may be integrated into a processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into a single module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or software functional modules or submodules.

[0222] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0223] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, software units executed by a processor, or a combination of the two. The software units may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0224] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0225] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identifying a target behavior, characterized in that: include: Obtaining a duration for which a detection target stays in a detection area, wherein the duration for which the detection target stays in the detection area is obtained by tracking the detection target from a surveillance video of the detection area; The step of tracking the detection target from the surveillance video of the detection area includes: Determine an intersection-over-union ratio of a bounding box area of any target object in a second video frame of a surveillance video of the detection area to a bounding box area of the detection target in a first video frame of the surveillance video of the detection area; wherein the second video frame is a video frame subsequent to the first video frame, and the target object is an identification object of the same type as the detection target; If there are at least two intersection-over-union ratios greater than a set threshold, determining the feature similarity between the detection target and each target object, and determining the target object whose weighted sum of the feature similarity and the intersection-over-union ratio is greater than a preset weighted threshold as the detection target; If the detection target stays in the detection area for a predetermined time, it is determined that the detection target meets the conditions for performing the target behavior in the detection area, and target video data and target audio data are obtained; wherein the target video data and the target audio data are respectively obtained by collecting video data and audio data of the detection area where the detection target is located; Based on the target video data and the target audio data, it is determined whether the detection target performs a target behavior.

2. The method according to claim 1, characterized in that The determining, based on the target video data and the target audio data, whether the detection target performs a target behavior includes: fusing the target video data and the target audio data to obtain fusion information; extracting behavioral features of the detection target from the fusion information; Based on the behavior characteristics, it is determined whether the detection target performs a target behavior.

3. The method according to claim 2, characterized in that The fusing the target video data and the target audio data to obtain fusion information, extracting the behavior characteristics of the detection target from the fusion information, and determining whether the detection target performs a target behavior based on the behavior characteristics, includes: Inputting the target video data and the target audio data into a pre-trained target behavior recognition model, causing the target behavior recognition model to fuse the target video data and the target audio data to obtain fusion information, extracting the behavior features of the detection target from the fusion information, and obtaining a target behavior recognition result based on the behavior features; The target behavior recognition result includes that the detection target performs the target behavior, or that the detection target does not perform the target behavior.

4. The method according to claim 3, characterized in that Before inputting the target video data and the target audio data into the pre-trained target behavior recognition model, the method further includes: Performing frame extraction processing on the target video data, and / or performing downsampling processing on the target audio data; The target video data and the target audio data are time-synchronized data.

5. The method according to claim 1, wherein The obtaining of the duration of time the detection target stays in the detection area includes: When it is determined that the detection target enters the detection area, the detection target is tracked from the surveillance video of the detection area to determine the duration for which the detection target appears in the surveillance video of the detection area; The duration for which the detection target appears in the surveillance video of the detection area is determined as the duration for which the detection target stays in the detection area.

6. The method according to claim 5, characterized in that Tracking the detection target from the surveillance video of the detection area includes: If the intersection-and-union ratio of the bounding box area of any target object in the second video frame and the bounding box area of the detection target in the first video frame of the surveillance video of the detection area is greater than a set threshold, the target object in the second video frame is determined to be the detection target.

7. The method according to claim 1, characterized in that The target video data includes monitoring video data of a sand mining area, the target audio data includes audio data collected from the sand mining area, and the target behavior includes sand mining behavior.

8. A target behavior recognition device, characterized in that: include: An acquisition module is used to acquire a duration of time that a detection target stays in a detection area, wherein the duration of time that the detection target stays in the detection area is obtained by tracking the detection target from a surveillance video of the detection area; The step of tracking the detection target from the surveillance video of the detection area includes: Determine an intersection-over-union ratio of a bounding box area of any target object in a second video frame of a surveillance video of the detection area to a bounding box area of the detection target in a first video frame of the surveillance video of the detection area; wherein the second video frame is a video frame subsequent to the first video frame, and the target object is an identification object of the same type as the detection target; If there are at least two intersection-over-union ratios greater than a set threshold, determining the feature similarity between the detection target and each target object, and determining the target object whose weighted sum of the feature similarity and the intersection-over-union ratio is greater than a preset weighted threshold as the detection target; If the detection target stays in the detection area for a predetermined time, it is determined that the detection target meets the conditions for performing the target behavior in the detection area, and target video data and target audio data are obtained; wherein the target video data and the target audio data are respectively obtained by collecting video data and audio data of the detection area where the detection target is located; A determination module is used to determine whether the detection target performs a target behavior based on the target video data and the target audio data.

9. An electronic device, characterized in that: include: memory and processor; Wherein, the memory is used to store programs; The processor is configured to implement the target behavior recognition method according to any one of claims 1 to 7 by running the program in the memory.

10. A storage medium, characterized in that: include: The storage medium stores a computer program, and when the computer program is executed by the processor, the computer program implements the steps of the target behavior identification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Warning method and device for criminal activity, storage medium and server

    CN108351968A

  • Monitoring method and system of target vehicle and computer readable storage medium

    CN109800696A