Convolutional neural network training sample augmentation method for monitoring video small target detection

Through the combination of drone aerial photography and Unreal Engine, high-quality small target samples are generated and automatic labeling is realized, which solves the problem of scarce sample data and poor labeling accuracy in small-scale object detection tasks in surveillance videos, and significantly improves the detection accuracy and generalization capabilities of the model.

CN120147776APending Publication Date: 2025-06-13GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510204154.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Small and medium-sized object detection tasks in monitoring videos face the problems of scarce sample data and poor labeling accuracy, which affects the training accuracy and generalization ability of the object detection model.

Method used

Through aerial photography of a drone, a three-dimensional scene model of the park is generated, and imported into Unreal Engine for optimization and display, and decomposed into continuous frame images. Based on these independent image files, a large semantic segmentation model is used to dynamically generate masks to guide copy and paste, obtain high-quality and evenly distributed small target samples, and realize automatic labeling.

Benefits of technology

A high-quality and evenly distributed small target samples were successfully generated, which solved the problems of few small target samples, uneven distribution and low manual labeling accuracy, expanded the number of small target training, improved the labeling quality, and improved the detection accuracy and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147776A_ABST
    Figure CN120147776A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video monitoring, in particular to a convolutional neural network training sample augmentation method for monitoring video small target detection, and the method comprises the steps: collecting multi-angle image data through unmanned aerial vehicle aerial photography, processing the influence data, and generating a park three-dimensional scene model; importing the park three-dimensional scene model into an unreal engine for optimization and display, and decomposing the park three-dimensional scene model into continuous frame images to obtain independent image files; based on an independent image file, a large-scale semantic segmentation model is used for dynamically generating a mask to guide copying and pasting, a mask image is obtained and stored as a file with the same name of an original image, the method utilizes a virtual engine and a copying and pasting strategy based on the large-scale semantic segmentation model mask, and small target samples which are high in quality and uniform in distribution are successfully generated; according to the method, the Airsim is used for generating the labeling box of the virtual assets, the mask obtained by the segmentation model is used for fitting the labeling box of the targets, automatic labeling of the samples is successfully achieved, and the method can expand the training number and diversity of the small targets and improve the labeling quality of the small targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video surveillance, and in particular to a method for augmenting training samples of a convolutional neural network for small target detection in surveillance videos. Background Art

[0002] The task of small target detection in video surveillance is usually defined as identifying and locating small-sized targets (such as pedestrians, bicycles, cars) that only occupy a small pixel area in the surveillance video, which is of great significance for improving the intelligence level of the surveillance system and expanding its applications in scenarios such as traffic, security, and industrial surveillance.

[0003] In recent years, deep learning artificial intelligence technologies represented by networks such as Faster R-CNN, YOLO, and SSD have achieved great success in visual tasks such as target detection and classification recognition. However, the actual application performance of these networks depends on high-quality training sample data, and the number of small target samples in existing publicly available datasets is generally insufficient. For example, only 52.3% of the images in the COCO dataset contain small target samples and are unevenly distributed in the images. The lack of small target sample information will undoubtedly affect the training accuracy and generalization ability of the target detection model. On the other hand, due to the low resolution of surveillance videos, the loss of image details will greatly reduce the network's feature extraction ability for small targets. At the same time, small targets have very small pixel ratios, weak feature expressions, blurred boundaries, and are easily affected by the background, etc., which will often reduce the quality of their manual annotation. In the task of small target detection, the network is highly sensitive to the positioning information of small targets, and poor annotation quality will directly lead to a significant decline in model performance. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for augmenting training samples of a convolutional neural network for small target detection in surveillance videos, aiming to solve the problems of scarce sample data and poor annotation accuracy faced by the current small target detection task in surveillance videos.

[0005] To achieve the above purpose, the present invention provides a method for augmenting training samples of a convolutional neural network for small target detection in surveillance videos, including the following steps;

[0006] Collect multi-angle image data through drone aerial photography, and process the image data to generate a three-dimensional scene model of the park;

[0007] Import the three-dimensional scene model of the park into the Unreal Engine for optimization and display, and decompose it into consecutive frame images to obtain independent image files;

[0008] Based on the independent image files, use a large semantic segmentation model to dynamically generate a mask to guide copy and paste, obtain a mask image, and save it as a file with the same name as the original image.

[0009] Among them, the specific method of collecting multi-angle image data through UAV aerial photography and processing the image data to generate a 3D scene model of the park:

[0010] Collect multi-angle image data through UAV aerial photography to obtain image data;

[0011] Process the image data based on the oblique photography modeling technology to generate an initial 3D model of the park;

[0012] Perform a model repair operation on the initial 3D model of the park to obtain a 3D scene model of the park.

[0013] Among them, the model repair operation includes optimizing the geometric details of buildings, smoothing the terrain surface, and repairing void and defect areas caused by data collection limitations.

[0014] Among them, the specific method of importing the 3D scene model of the park into the Unreal Engine for optimization and display, and decomposing it into consecutive frame images to obtain independent image files:

[0015] Import the 3D scene model of the park into the Unreal Engine to simulate pedestrian activities in the park;

[0016] Use a behavior tree automatic logic to design a random walking pattern and combine an air wall to restrict the pedestrian activity area;

[0017] Record the dynamic video of the park scene, and use a function to obtain the spatial position of objects and the status data of object production to determine the position of the detection frame;

[0018] Decompose the dynamic video into consecutive frame images and save them as independent image files.

[0019] Among them, the large semantic segmentation model can distinguish road, sky, and building semantic categories, and the output segmentation mask not only has clear boundaries but also can reflect the actual semantic structure of the scene.

[0020] The method for augmenting training samples of a convolutional neural network for small target detection in surveillance videos of the present invention collects multi-angle image data through aerial photography by a drone, processes the image data to generate a three-dimensional scene model of a park; imports the three-dimensional scene model of the park into the Unreal Engine for optimization and display, and decomposes it into consecutive frame images to obtain independent image files; uses a large semantic segmentation model to dynamically generate a mask for guiding copy-paste based on the independent image files, obtains a mask image, and saves it as a file with the same name as the original image. This method uses the virtual engine (UE) and the copy-paste strategy based on the mask of the large semantic segmentation model (mask) to successfully generate high-quality and evenly distributed small target samples, solves the problems of few small target samples and uneven distribution, uses Airsim to generate annotation frames of virtual assets, and uses the mask obtained by the segmentation model to fit the annotation frames of the target, successfully realizing automatic annotation of samples, solving the problem of low accuracy of manual annotation. This method can expand the quantity and diversity of small target training and improve its annotation quality, solves the deficiencies such as scarce sample data and poor annotation accuracy faced by the current small target detection task in surveillance videos, and has broad market application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0022] Figure 1 It is a schematic structural diagram of a three-dimensional scene model of a park.

[0023] Figure 2 It is a schematic diagram of an air wall setting scheme.

[0024] Figure 3 It is a schematic diagram of data generated based on UE4 and Airsim.

[0025] Figure 4 It is a schematic diagram of the comparison effect between real data and UE synthetic data.

[0026] Figure 5 It is a flowchart of data augmentation by copy-paste guided by a Mask mask.

[0027] Figure 6 It is a schematic diagram of the inference result of adding UE synthetic data. Among them, (a) is without adding synthetic data, and (b) is with adding synthetic data.

[0028] Figure 7It is a schematic diagram of the visualized comparison result in the service area scenario. Among them, (a) is the Ground Truth, (b) is without data augmentation, and (c) is with data augmentation.

[0029] Figure 8 It is a schematic diagram of the visualized comparison result in the construction site scenario. Among them, (a) is the Ground Truth, (b) is without data augmentation, and (c) is with data augmentation.

[0030] Figure 9 It is a flowchart of the convolutional neural network training sample augmentation method for small target detection in surveillance videos provided by the present invention.

[0031] Figure 10 It is a flowchart of the convolutional neural network training sample augmentation method for small target detection in surveillance videos provided by the present invention.

[0032] Figure 11 It is a flowchart of the specific method for collecting multi-angle image data by drone aerial photography and processing the influence data to generate a 3D scene model of the park.

[0033] Figure 12 It is a flowchart of the specific method for importing the 3D scene model of the park into the Unreal Engine for optimization and display, and decomposing it into continuous frame images to obtain independent image files. Specific Embodiment

[0034] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0035] Please refer to Figures 1 to 12 , the present invention provides a convolutional neural network training sample augmentation method for small target detection in surveillance videos, including the following steps;

[0036] S1 Collect multi-angle image data by drone aerial photography and process the influence data to generate a 3D scene model of the park;

[0037] Specific Method:

[0038] S11 Collect multi-angle image data by drone aerial photography to obtain image data;

[0039] In the embodiment of the present invention, multi-angle image data is collected by drone aerial photography, covering most of the buildings, roads, and greening areas within the Academy of Sciences campus.

[0040] S12 processes the image data based on oblique photography modeling technology to generate an initial three-dimensional model of the park;

[0041] S13 performs a modeling operation on the initial three-dimensional model of the park to obtain a three-dimensional scene model of the park.

[0042] In the embodiment of the present invention, in order to further improve the accuracy and visual effect of the model, the initial 3D model of the park is modified, including optimizing the geometric details of the buildings, smoothing the terrain surface, and repairing the holes and defective areas caused by data collection limitations. Finally, a detailed and high-precision 3D scene model of the park is successfully constructed.

[0043] S2 imports the park three-dimensional scene model into Unreal Engine for optimization and display, and decomposes it into continuous frame images to obtain independent image files;

[0044] Specific method:

[0045] S21 imports the park 3D scene model into Unreal Engine to simulate pedestrian activities in the park;

[0046] In an embodiment of the present invention, the three-dimensional scene model of the park is imported into the Unreal Engine (UE4) for further optimization and display. With the powerful rendering capability of UE4, the three-dimensional model of the park shows realistic visual effects, and truly restores the proportions and spatial relationships of buildings and terrain. In order to enhance the immersiveness of the virtual scene, the character assets provided by UE4 are used to simulate the activities of pedestrians in the park. By adding virtual characters of different genders, ages and clothing styles, the scene is closer to real life.

[0047] S22 uses behavior tree automatic logic to design a random walking pattern, and combines air walls to restrict pedestrian activity areas;

[0048] In the embodiment of the present invention, the NPC behavior tree automatic logic is used to design a random walking mode for the virtual pedestrians, and the pedestrian activity area is restricted by combining the air wall (Invisible Walls), such as Figure 2 The behavior tree logic automatically generates paths, allowing pedestrians to move freely within the park according to the predetermined area. In addition, this paper further adjusts the rhythm and stop points of pedestrians through the path editing tool to simulate the daily activity scenarios of people in the park.

[0049] S23 records a dynamic video of the park scene, and uses functions to obtain the spatial position of the object and the status data of the object asset to determine the position of the detection frame;

[0050] In the embodiments of the present invention, in order to expand the function of data acquisition, the AirSim plugin is integrated into the virtual scene. The first-person view function of the AirSim drone is used as a fixed virtual camera to simulate the high-altitude surveillance camera in the real scene. Through the first-person view function of AirSim, the dynamic video of the park scene is recorded. The get_object_pose() function of Airsim is used to obtain the spatial position and other information of the object, and the client.getPersonState() function is used to obtain the status data of the person assets. Then, the position of the detection box is determined according to these data. In UE4, the DrawDebugBox() function is used to draw a cube box to represent the target detection box.

[0051] S24 Decompose the dynamic video into consecutive frame images and save them as independent image files.

[0052] In the embodiments of the present invention, the video is decomposed into consecutive frame images and saved as independent image files. In addition, the Airsim plugin is used to automatically project the detection box onto a two-dimensional plane to obtain the two-dimensional detection box coordinates of each frame of pedestrian assets, and an.xml file with asset position information is automatically exported to directly obtain the annotation file of the person assets. Finally, the required UE small target synthesis data is obtained, as Figure 3 shown.

[0053] Figure 4 is the comparison effect between the real data and the UE synthesis data of this study. Figure 4 (a) Real data of the Academy of Sciences park. From the visualization results, the synthesis data of this study ( Figure 4 (b)) is infinitely close to the real data in terms of both background and sample details. The method of this study starts from three-dimensional scene modeling, combines virtual pedestrian behavior simulation, dynamic video recording, and detection box generation to construct a complete virtual data generation system. Through this method, a large number of small target training samples can be obtained quickly, and high-quality annotation boxes can be directly obtained automatically, solving the problems of scarce small target samples and low annotation quality, and providing reliable data support for subsequent model training.

[0054] S3 Use a large semantic segmentation model to dynamically generate a mask to guide copy and paste based on the independent image file, obtain the mask image, and save it as a file with the same name as the original image.

[0055] In an embodiment of the present invention, a large semantic segmentation model SAM2 (SemanticAlignment Model 2) is used to dynamically generate masks to guide the copy and paste process. In traffic scenes, SAM2 can accurately distinguish semantic categories such as roads, sky and buildings, and the segmentation mask output by it not only has clear boundaries, but also accurately reflects the actual semantic structure of the scene. This high-quality mask provides a strong guide for target generation and ensures the physical rationality of the target pasting position. For example, when data enhancement is performed on a vehicle, the road area mask generated based on SAM2 can effectively limit the vehicle target to be pasted only in the road area, avoiding the physical unreasonable phenomenon (such as the vehicle is suspended in the air or overlapped with the building) that may be caused by random pasting in the traditional method. In addition, the mask dynamically generated based on SAM2 has a high scene adaptability. This feature enables the present method to be widely used in a variety of complex scenes, not just limited to a single road scene. In indoor object detection tasks, SAM2 can segment areas such as desktops, floors, walls, etc., thereby guiding the generation of targets in appropriate locations and avoiding the irrationality of target distribution. In natural scenes, SAM2 can identify trees, ground, or water areas to ensure the semantic consistency of pasted targets.

[0056] For scenes with complex backgrounds that are difficult to segment or have poor segmentation quality, interactive mask drawing can be implemented through OpenCV. You can draw a mask by mouse operation in the interface. Hold down the left button and drag to draw a black mask area. Right click to undo the operation. You can also adjust the brush size through the slider. After the mask is drawn, it will be automatically saved as a _mask.jpg file with the same name as the original image for subsequent use. Apply the generated mask image to all images in the specified folder in batches.

[0057] To better understand the present technical solution, the following embodiments are provided for further explanation:

[0058] The character assets provided by Unreal Engine (UE) were applied to the 3D map of the Institute of Materials, and reasonable walking paths and boundary conditions were set to simulate the behavior of people in the real environment. The skeletal meshes in the scene were detected through the Airsim simulation platform to generate corresponding images and annotation boxes. This method can quickly generate a large number of annotated detection samples, omitting the tedious process of manual annotation and greatly improving the efficiency of data generation.

[0059] To verify the effectiveness of UE synthetic data in small object detection, 5000 real data of the Academy of Sciences campus were used as the training result as the baseline, and different proportions of UE synthetic data were gradually added. Through this experiment, it aims to study the impact of the addition of synthetic data on the model performance. To ensure the fairness and objectivity of the experiment, 500 real data were used as the validation set to evaluate the performance of the model under different data combinations. The experimental results are shown in Table 1 below.

[0060] Table 1

[0061]

[0062]

[0063] The experimental results show that with the gradual addition of synthetic data, the detection accuracy has been significantly improved. However, the quantity of synthetic data is not the more the better. In some cases, continuously increasing the synthetic data will instead lead to a decline in the model performance. This phenomenon is closely related to the authenticity of the synthetic data. When the quality of the synthetic data is high and the similarity with the real data is strong, the contribution of the synthetic data to the model is more obvious; but if the authenticity of the synthetic data is insufficient, adding too much synthetic data may cause the model to overfit the data, thus affecting the generalization ability of the model on real data. Therefore, the quantity of synthetic data should be appropriately increased on the premise of ensuring quality. Among them, Figure 6 (a) represents the inference result without adding synthetic data, and (b) represents the result after adding UE synthetic data.

[0064] It further explores whether synthetic data can completely replace real data. In this experiment, while keeping the quantity of synthetic data unchanged, the proportion of real data was gradually reduced, and the changes in the model performance were observed under different training set compositions. Table 2 shows the experimental results. The experiment shows that although synthetic data can improve the performance of the model to a certain extent, it cannot completely replace real data. Especially when the difference between synthetic data and real data is large, reducing the proportion of real data will lead to a significant decline in the performance of the model in the real environment. Therefore, although synthetic data has high generation efficiency and annotation accuracy, it still needs to be combined with real data to ensure the robustness and generalization ability of the model.

[0065] Table 2

[0066]

[0067] The effectiveness of our data augmentation method was verified through ablation experiments. The Mask-guided copy-paste method overcomes the problems of boundary incoherence and semantic conflicts easily caused by traditional random paste methods by introducing physical rationality and semantic constraints during the data augmentation process. This method ensures that small objects in the augmented samples maintain good structural and semantic consistency, thus effectively improving the model's detection ability for small objects. On the other hand, the SAHI method optimizes the model's input by spatially slicing the image to increase the resolution. Different from directly making significant modifications to the model, the SAHI method can be seamlessly integrated into existing object detection frameworks without additional complex calculations. This augmentation method can significantly improve the model's inference ability without changing the network structure, especially showing outstanding effects in the detection of high-resolution small objects. By combining these two methods, this paper can not only utilize the Mask-guided augmentation strategy to provide semantic constraints and physical rationality but also optimize the model input through the SAHI method to improve the effectiveness and diversity of training samples.

[0068] In this experiment, the training results of 5000 pieces of real data from the service area were used as the baseline, and the YOLOv7 model was used for training. Table 3 shows the training results under different data augmentation strategies, where UE is adding UE synthetic data, and MCP is using mask-guided copy-paste. The experimental results show that after adding UE synthetic data, the model performance is significantly improved. When using the mask-guided copy-paste data augmentation method alone, the detection accuracy of the model is significantly improved. This indicates that the mask-guided copy-paste augmentation method can effectively maintain the structural and semantic consistency of the target, thus enhancing the model's detection ability for small objects. Similarly, when using the SAHI method alone, the detection accuracy of the model is also improved. SAHI optimizes the quality of the input data by spatially slicing the image and increasing the resolution, thus effectively improving the model's inference ability, especially performing excellently when dealing with high-resolution small objects. After combining these three methods, the detection accuracy of the model reaches the best level. Figure 7 and Figure 8 Shows the visualization results of model inference in the service area scenario and the construction site scenario after using this method. Figure 7 and Figure 8 In (a) is the ground truth label, (b) is without using data augmentation, and (c) is the inference result after adding the synthetic data of this patent.

[0069] Table 3

[0070]

[0071] By testing on different object detection networks, the generality of the combination of the two methods was verified. To comprehensively evaluate the applicability of the method, experiments were conducted on three different object detection frameworks in this paper, namely YOLOv7, YOLOv8, and RT-DETR (Real-Time Detection Transformer). These three networks represent different architectures in the current object detection field: YOLOv7 and YOLOv8, as classic single-stage detectors, have high real-time performance and accuracy; while RT-DETR, as a Transformer-based network, can better handle complex scenes and small object detection tasks. The experimental results are shown in Table 4.

[0072] Table 4

[0073]

[0074] The experimental results show that the model combined with the Mask-guided copy-paste data augmentation method and SAHI demonstrated good performance improvement on these three different network architectures. Whether in the YOLO series of models or in a more advanced Transformer-based detection framework like RT-DETR, the combination of the two methods can effectively improve the detection accuracy of the model, especially showing significant advantages in the recognition of small objects.

[0075] The above-disclosed is only the preferred embodiment of the convolutional neural network training sample augmentation method for small object detection in the monitoring video of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A convolutional neural network training sample augmentation method for small target detection in surveillance video, characterized in that: The following steps are involved: Collect multi-angle image data through drone aerial photography, process the impact data, and generate a three-dimensional scene model of the park; Importing the park three-dimensional scene model into Unreal Engine for optimization and display, and decomposing it into continuous frame images to obtain independent image files; Based on the independent image file, a large semantic segmentation model is used to dynamically generate a mask to guide copying and pasting, obtain a mask map, and save it as a file with the same name as the original image.

2. The method for augmenting convolutional neural network training samples for small target detection in surveillance video according to claim 1, characterized in that: The specific method of collecting multi-angle image data through drone aerial photography and processing the impact data to generate a three-dimensional scene model of the park is as follows: Collect multi-angle image data through drone aerial photography to obtain image data; Processing the image data based on oblique photography modeling technology to generate an initial three-dimensional model of the park; The initial three-dimensional model of the park is modified to obtain a three-dimensional scene model of the park.

3. The method for augmenting convolutional neural network training samples for small target detection in surveillance video according to claim 2, characterized in that: The modeling operations include optimizing the geometric details of buildings, smoothing the terrain surface, and repairing holes and defective areas caused by data acquisition limitations.

4. The method for augmenting convolutional neural network training samples for small target detection in surveillance video according to claim 1, characterized in that: The specific method of importing the park 3D scene model into Unreal Engine for optimization and display, and decomposing it into continuous frame images to obtain independent image files is as follows: Importing the three-dimensional scene model of the park into the Unreal Engine to simulate pedestrian activities in the park; Use behavior tree automatic logic to design random walking patterns, and combine air walls to restrict pedestrian activity areas; Record dynamic video of the park scene, and use functions to obtain the spatial position of objects and the status data of objects to determine the position of the detection frame; The dynamic video is decomposed into continuous frame images and saved as independent image files.

5. The method for augmenting convolutional neural network training samples for small target detection in surveillance video according to claim 1, characterized in that: The large-scale semantic segmentation model is able to distinguish road, sky and building semantic categories, and its output segmentation mask not only has clear boundaries but also reflects the actual semantic structure of the scene.

Citation Information

Cited By

  • Food material data generation method and device based on image processing, storage medium and electronic device

    CN121482526A