Personnel target detection network design method for open scene video monitoring
By simulating multi-view scenes and improving the detection model, the problem of missed detection due to small target size or unclear features in open scenes is solved, and high-precision detection of low-resolution small targets and extreme pose targets is achieved.
Patent Information
- Application Number
- CN202510241626.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-06
AI Technical Summary
Existing personnel target detection problems are missed due to the small target size or unclear characteristics in open scenarios.
Multiple cameras simulate multi-view personnel object detection scenarios, build training sets and test sets, and improve the detection model, including replacing IoU with NWD metric method, introducing initial receptive field enhancement module, deformable convolution and spatial enhanced attention module.
The improved YOLOv9 target detection network can improve the detection accuracy of low-resolution targets, enhance the adaptability to extreme attitude changes and shape transformations, and improve the identification and positioning capabilities of the monitoring system in complex scenarios.
Smart Images

Figure CN119942599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video surveillance, and in particular to a method for designing a personnel target detection network for open scene video surveillance. Background Art
[0002] Video object detection is of great value in many applications such as transportation, security, case detection, and sports event broadcasting. It is of great significance to improve the intelligence level of monitoring systems and expand their application scenarios. In recent years, deep learning artificial intelligence technology represented by the YOLO series of networks has effectively avoided the limitations of artificial feature design in traditional methods and achieved higher detection accuracy and faster reasoning speed by automatically learning effective feature representations from labeled data in object detection and classification and recognition visual tasks.
[0003] However, the actual monitoring scene, especially the open scene environment, is complex. The appearance features of people in the video at different times are affected by complex factors such as monitoring perspective, lighting conditions, image resolution, motion posture, background changes, target occlusion, and inter-frame loss. It is difficult to ensure the consistency and continuity of the image feature capture of the target person, resulting in the deep learning artificial intelligence technology still facing many challenges in the task of monitoring video person target detection in open scenes. For example, the human targets in the monitoring video usually have different resolution sizes. The traditional convolutional neural network relies on high-resolution input images to extract rich feature information, but the human targets in the low-resolution video may appear small and blurred. The loss of image details will greatly reduce the network's feature extraction ability for small human targets. The posture changes of people in the monitoring video under various actions such as walking, running, turning, jumping, squatting, etc. will cause significant differences in the shape and appearance features of the people, thereby increasing the difficulty of target detection, resulting in mismatch between the detection frame and the target, and inaccurate target positioning, missed detection or false detection. The human targets in the monitoring video are often blocked by other targets or similar targets, and the occlusion will directly affect the feature extraction and positioning accuracy of the target, and even lead to the inability to accurately identify the target. Summary of the invention
[0004] The purpose of the present invention is to provide a method for designing a human target detection network for open scene video surveillance, aiming to solve the problem of missed detection due to small target size or unclear features in existing human target detection.
[0005] To achieve the above object, the present invention provides a method for designing a human target detection network for open scene video surveillance, comprising the following steps:
[0006] Use multiple cameras to simulate multi-viewpoint human target detection scenarios and construct training and test sets.
[0007] Improve the detection model;
[0008] The detection model is trained and tested by using the training set and the test set to obtain a target detection model;
[0009] Input videos from different perspectives into the target detection model to generate result videos with detection box information.
[0010] Among them, in "Simulating multi-viewpoint human target detection scenarios through multiple cameras and building training sets and test sets", the following steps are included:
[0011] Set up multiple cameras for video acquisition;
[0012] The collected videos are divided into training sets and test sets.
[0013] Among them, in "setting up multiple cameras for video acquisition", the total video recording time of each camera is 35 minutes, and the scene picture is extracted every 20 frames.
[0014] Among them, "improving the detection model" includes the following steps:
[0015] Use the NWD measurement method to replace the IoU measurement method of the original detection model;
[0016] Introduce the receptive field enhancement module into the detection model;
[0017] Introducing deformable convolution into detection models;
[0018] Add the spatial enhanced attention module to the detection model.
[0019] Among them, the NWD measurement method is used to model the bounding box as a two-dimensional Gaussian distribution, and the normalized Wasserstein distance (NWD) is used to replace the traditional IoU to reduce the position sensitivity of low-resolution small targets; the initial receptive field enhancement module is based on the multi-branch structure of dilated convolution, captures multi-scale information through branches with different expansion rates, and fuses weighted features; the deformable convolution is used to dynamically adjust the convolution kernel sampling position to improve the detection ability of deformed targets; the spatial enhanced attention module restores occlusion features through channel and spatial mixing modules to enhance the target area response.
[0020] The invention discloses a method for designing a human target detection network for open scene video monitoring, comprising the following steps: simulating a multi-view human target detection scene through multiple cameras, constructing a training set and a test set; improving a detection model; training and testing the detection model through the training set and the test set to obtain a target detection model; inputting videos of different viewpoints into the target detection model to generate a result video with detection frame information. The improved YOLOv9 target detection network of the invention can improve the detection accuracy of low-resolution small targets, especially has significant advantages in human target detection in long-distance monitoring videos. This improvement enhances the recognition and positioning capabilities of the monitoring system in complex scenes, especially in key areas such as security and traffic monitoring, can improve the detection accuracy of extreme posture targets or targets with large shape changes, can enhance the adaptability of the network to extreme posture changes and shape transformations, effectively solves the shortcomings of traditional methods in dynamic scenes, can more accurately identify targets in different postures, action states or with large morphological changes, thereby improving the reliability and intelligent recognition capabilities of the system in actual monitoring and security protection applications, thereby solving the problem of missed detection due to small target size or unclear features in existing human target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0022] Figure 1 It is a flow chart of a method for designing a personnel target detection network for open scene video surveillance provided by the present invention.
[0023] Figure 2 This is a schematic diagram of image collection of people in open scenes.
[0024] FIG3( a ) is a schematic diagram of the improved YOLOv9 network structure; FIG3( b ) is a schematic diagram of the basic YOLOv9 network structure.
[0025] Figure 4 (a) is a schematic diagram of the IoU (Intersection over Union) indicator that is deviated to keep the scales of Box_A and Box_B always the same; Figure 4 (b) Schematic diagram of the IoU (Intersection over Union) metric that is deviated to keep the side length of Box_B half of Box_A.
[0026] Figure 5(a) Schematic diagram of the NWD (normalized Wasserstein distance) indicator that deviates to keep the scales of Box_A and Box_B always the same; Figure 5 (b) Schematic diagram of the NWD (normalized Wasserstein distance) metric that is deviated to keep the side length of Box_B half of Box_A.
[0027] Figure 6 It is a schematic diagram of the initial receptive field enhancement module introducing diverse receptive field information.
[0028] Figure 7 It is a schematic diagram of the receptive field enhancement module that introduces dilated convolution.
[0029] Figure 8 is a schematic diagram of the reparameterized convolution of the reference plan.
[0030] Fig. 9 It is a schematic diagram of optimizing the structure of the receptive field enhancement module.
[0031] Fig.10 (a) is a schematic diagram of a common convolution operation. Fig.10 (b), (c), and (d) are schematic diagrams of deformable convolution.
[0032] Fig.11 It is a schematic diagram of standard convolution and deformable convolution.
[0033] Fig.12 Schematic diagram of the spatial enhanced attention module introducing enhanced responses to unoccluded targets to compensate for the response loss of occluded targets.
[0034] Fig.13 It is the structural diagram of the improved CBLiner_RFE module (receptive field enhanced linear layer module) and RepNCSPELAN4_RFE module (receptive field enhanced cross-stage efficient feature aggregation layer module).
[0035] Fig.14 It is the structural diagram of the improved Conv_D module (deformable convolution layer module).
[0036] Fig.15 It is the structural diagram of the improved Conv_SEA module (spatial attention enhanced convolutional layer module) and RepNCSPELAN4_SEA module (spatial attention enhanced cross-stage efficient feature aggregation layer module).
[0037] Fig.16 It is a summary of the impact of different metrics of the YOLOv9 model on the test results of this video scene data.
[0038] Fig.17 and Fig.18This is a summary of the impact of introducing the receptive field enhancement module (RFEM) or deformable convolution (DCN) into the YOLOv9 model on the test results on this dataset.
[0039] Fig.19 This is a summary of the impact of the spatial enhanced attention module (SEAM) introduced by the YOLOv9 model on the test results on this dataset.
[0040] Fig. 20 and Fig.21 The left side of the arrow in the figure is a schematic diagram of the original YOLOv9 method and the inference result; the right side of the arrow is the inference result of the improved YOLOv9 method.
[0041] Fig. 22 This is a flowchart of simulating multi-viewpoint human target detection scenarios through multiple cameras and constructing training sets and test sets.
[0042] Fig.23 It is a flow chart for improving the detection model. DETAILED DESCRIPTION
[0043] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0044] See also Figures 1 to 23 The present invention provides a method for designing a human target detection network for open scene video surveillance, comprising the following steps:
[0045] S1 uses multiple cameras to simulate multi-viewpoint human target detection scenarios and construct training and test sets;
[0046] S11 sets up multiple cameras for video acquisition;
[0047] The total video recording time of each camera is 35 minutes, and scene pictures are extracted every 20 frames.
[0048] Specifically, four cameras are used to simulate multi-view target detection scenarios, such as Figure 2As shown in the figure. During the simulation, each participant not only needs to perform different postures such as standing, walking, and sitting, but also needs to keep the pixel changes of the monitoring camera during the execution process to ensure the diversity of data and the true restoration of the scene. Each camera collects data at different time periods of the day. The total video recording time of each camera is about 35 minutes, and the scene pictures are extracted every 20 frames. The video scene under each camera covers 2 to 7 identities in a certain period of time, and each identity has about 400 to 800 images.
[0049] S12 divides the collected video into a training set and a test set.
[0050] Specifically, the training set and test set are divided into 8:2, and more than 10,000 images are selected as the training set, and more than 2,600 images that try to include low-resolution small targets, images with large posture changes, and images of people with partial occlusion are selected as the test set.
[0051] S2 improves the detection model;
[0052] The NWD metric method is used to model the bounding box as a two-dimensional Gaussian distribution, and the normalized Wasserstein distance (NWD) is used to replace the traditional IoU to reduce the position sensitivity of low-resolution small targets; the initial receptive field enhancement module is based on the multi-branch structure of dilated convolution, captures multi-scale information through branches with different expansion rates, and fuses weighted features; the deformable convolution is used to dynamically adjust the convolution kernel sampling position to improve the detection ability of deformed targets; the spatial enhanced attention module restores occlusion features through channel and spatial mixing modules to enhance the target area response.
[0053] S21 uses the NWD metric method to replace the IoU of the detection model;
[0054] Specifically, the current detection model is based on the Intersection over Union (IoU) metric, which is very sensitive to the position deviation of small low-resolution objects. IoU will greatly reduce the detection performance when used in anchor-based detectors. Figure 4 The horizontal axis value represents the number of pixel deviations between the two bounding boxes A and B, and the vertical axis value represents the metric value. Figure 4 (a) Keep the scale of Box_A and Box_B the same. Figure 4 In (b), the side length of Box_B is kept half of Box_A, and the value in the upper right corner represents the side length of Box_B. It can be seen that as the target size becomes smaller, the curve drops faster. The sensitivity of IoU stems from the particularity that the position of the bounding box can only change discretely.
[0055] Figure 4The phenomenon leads to serious defects in label matching. For low-resolution small objects, even a slight position deviation will cause the anchor box label to be reversed, resulting in the characteristics of positive samples being very similar, and insufficient supervision information for training small object detection. Although the dynamic matching strategy can adaptively obtain the IoU threshold for assigning positive and negative samples based on the statistical characteristics of the target, the sensitivity of IoU makes it difficult to find a good threshold for low-resolution small object detection and provide high-quality positive and negative samples.
[0056] Therefore, we propose to use the NWD metric to measure the similarity of bounding boxes to replace the standard IoU. Specifically, the bounding boxes are first modeled as two-dimensional Gaussian distributions, and then the normalized Wasserstein distance new metric is used to calculate the similarity between them through their corresponding Gaussian distributions. Moreover, the proposed NWD metric can be easily embedded in any positive and negative sample matching process based on anchor box detectors to replace the traditional IoU.
[0057] The Wasserstein distance between two bounding boxes can be expressed as:
[0058]
[0059] and Represents the bounding box A = (cx a ,cy a ,w a ,h a ) and bounding box B = (cx b ,cy b ,w b ,h b ) modeled by a Gaussian distribution. Since is a distance metric. Therefore, it is normalized using its exponential form to obtain the normalized Wasserstein distance (NWD):
[0060]
[0061] When describing the weights of different pixels in the bounding box, the weight of the pixel in the center of the bounding box is the highest, and the weight of the pixel decreases gradually from the center to the edge. Figure 5 No longer appears in Figure 4 The NWD metric is used to measure the similarity of bounding boxes instead of the standard IoU in YOLOv9, which effectively reduces the sensitivity of IoU to the position deviation of low-resolution small targets.
[0062] S22 introduces the receptive field enhancement module into the detection model;
[0063] Specifically, the initial receptive field enhancement module (RFEM) introduces diverse receptive field information such as Figure 6As shown in the figure, a five-branch structure is adopted, which includes two architectural ideas of Inception and ResNet. First, it uses a 1×1 convolution layer to reduce the number of channels to one-fourth of the previous layer. Then, 1×k and k×1 (k=3 and 5) convolution layers are used to provide diversity in the receptive field. The feature maps of the four branches are spliced through another 1×1 convolution layer. In addition, a shortcut path is applied to retain the receptive field of the previous layer of the original path to avoid information loss. Regarding the consideration of the number of parameters, the improved RFEM makes a reasonable trade-off and uses the principle of dilated convolution, that is, by inserting holes (that is, areas where no convolution calculations are performed) between the elements in the convolution kernel to increase the receptive field of the convolution without increasing the amount of calculation or the number of parameters. It has significant advantages, especially in tasks that require expanding the receptive field without increasing the computational complexity, such as Figure 7 As shown in the figure. The structure uses three dilated convolution branches with different expansion rates to perform convolution operations to capture multi-scale information. In addition, different branches share weights to reduce the number of parameters. The only difference is their respective receptive fields. In addition, residual connections are used to prevent the problem of gradient explosion. Finally, the features of the four branches are averaged through a pooling layer to generate a feature map with a larger receptive field and richer context information to obtain the output feature layer. In summary, the RFEM module mainly consists of multi-branch and aggregate weighted layers based on dilated convolution. Multi-branch is used to capture multi-scale information and dependencies of different ranges. Aggregate weighted layers are used to collect information from different branches and weight the features of each branch.
[0064] For the planned receptive field enhancement module, the main reference is the planned reparameterized convolution, such as Figure 8 As shown in the description of its four ablation experiments, when the main network module is doing residual connection structure, if the residual connection structure is introduced again in the last sub-network module, it will have a negative impact on the performance of the model, resulting in a decline in the overall residual structure performance. That is, both the reparameterized convolution and the residual structure have the same connection structure, and the two identical connections conflict. According to the idea of P-RepResNet, the RFEM structure is optimized, and the residual identity connection structure is added in the reasonable place, such as Fig. 9 The method shown in Figure 1 is called P-RFEM.
[0065] After introducing RFEM and P-RFEM into YOLO, different receptive fields are provided for the model to detect human targets in extreme postures, which helps to detect human targets with large scale changes, thereby optimizing the detection accuracy of YOLOv9. The location where RFEM is introduced into the YOLOv9 network is shown in Figure 3(a), which is located in the CBLiner layer (linear layer) and RepNCSPELAN4 layer (cross-stage efficient feature aggregation layer) of the Backbone part.
[0066] S23 introduces deformable convolution into the detection model;
[0067] Specifically, the deformable convolution is a convolution where the position of the convolution is deformable, rather than performing convolution on the traditional N×N grid. The advantage of this is that the desired features can be extracted more accurately, while the traditional convolution can only extract the features of the rectangular frame. Fig.10 As shown, (a) is a normal convolution operation, (b), (c), and (d) are deformable convolutions, and (c) and (d) are special cases of (b), indicating that deformable convolution generalizes various transformations of scale, length, width, and rotation.
[0068] Fig.11 In the convolutional neural network, the fixed weights of the ordinary convolution kernel Conv result in the same network having the same receptive field size when processing different position areas of an image, which is unreasonable for deep convolutional neural networks that encode position information. Because different positions may correspond to objects of different scales or deformations, these layers need to be able to automatically adjust the scale or receptive field. The effect of target detection depends largely on the bounding box based on feature extraction, which is not the best method, especially for non-grid targets. DCN can better capture the spatial structure information of the target object, thereby improving the model's recognition performance for deformed targets. Deformable convolution introduces an offset vector Δp of the spatial position n , so that the convolution kernel can dynamically adjust the position and shape of its receptive area when performing feature extraction. The following formula is the operation formula of traditional convolution.
[0069]
[0070] 1p n Listed All positions in the deformable convolution DCN is to add a position offset vector Δp to the traditional convolution calculation formula n . Among them, {Δp n |n=1,2,…,N}, Then the computation of the deformable convolution becomes:
[0071]
[0072] Where y is the new pixel coordinate, Δp n is the offset, sampling is done at irregular and offset positions p n +Δp n Up, p 0 +p n +Δp n is the new coordinate, x(p 0 +p n +Δp n) means taking out the pixel value at that position, w(p n ) represents the convolution kernel relative to p 0 +p n +Δp n The weight value of the pixel at the position. Iteratively add the values w(p n )·x(p 0 +p n +Δp n ) outputs the resulting pixel coordinate value.
[0073] After DCN is introduced into YOLO, the sampling position of the convolution kernel is adjusted through the offset learned in each convolution operation, so that the problem of large changes in target morphology can be more effectively dealt with, thereby optimizing the detection accuracy of YOLOv9. The location of DCN introduced into the YOLOv9 network is shown in Figure 3(a), which is located in the Conv layer (convolution layer) of the Backbone part.
[0074] S24 adds the spatial enhanced attention module to the detection model.
[0075] Specifically, the spatial enhanced attention module (SEAM) exploits the relationship between feature maps to recover occluded features, such as Fig.12 As shown in the figure, the architecture of SEAM is on the left and the structure of CSMM (Channel and Spatial Mixing Module) is on the right. CSMM uses different blocks to process multi-scale features and uses deep separable convolution to learn the correlation between spatial dimensions and channels. Object occlusion can cause three problems: category alignment error, local aliasing, and feature loss. This module has three purposes: to achieve multi-scale object detection, emphasize the object area in the image, and weaken the background area accordingly. The first part of SEAM, CSMM, is the deep separable convolution with residual connection. Deep separable convolution is operated depth by depth, that is, convolution separated by channel. Although deep separable convolution can learn the importance of different channels and reduce the number of parameters, it ignores the information relationship between channels. To make up for this loss, the outputs of different depth convolutions are then combined through point-by-point (1x1) convolution. A two-layer fully connected network is then used to fuse the information of each channel, allowing the network to strengthen the connection between all channels. It is hoped that the model can make up for the above losses in occluded scenes by learning the relationship between occluded and unoccluded objects in the previous step. Next, the output score logits learned by the fully connected layer is processed by an exponential function to expand the value range from [0, 1] to [1, e]. This exponential normalization provides a monotonic mapping relationship, which makes the result more tolerant to position errors. Finally, the output of the SEAM module is multiplied by the original feature as attention, so that the model can handle target occlusion more effectively.
[0076] After SEAM is introduced into YOLO, the response loss of the occluded target is compensated by enhancing the response of the unoccluded target, solving the problem of feature loss caused by partial occlusion of people or partial overlap of people in multi-view surveillance video scenes, thereby optimizing the detection accuracy of YOLOv9. The location of SEAM introduced in the YOLOv9 network is shown in Figure 3(a), which is located in the Conv layer (convolution layer) of the Neck part and the RepNCSPELAN4 layer (cross-stage efficient feature aggregation layer).
[0077] S3 trains and tests the detection model using the training set and the test set to obtain a target detection model;
[0078] Specifically, the improved target detection network is used to train the data set. By continuously adjusting the model parameters and structure, the characteristics of the target category are learned, and ultimately the network can efficiently and accurately identify and locate human targets, achieving higher detection performance.
[0079] S4 inputs videos of different perspectives into the target detection model to generate result videos with detection box information.
[0080] Specifically, the data in the actual scene is inferred by improving the network training model. The result data after inference not only contains the category and location box of each detected target, but also the average detection accuracy in the data scenario, thereby providing accurate target positioning and classification information and reliable accuracy data support for subsequent applications.
[0081] Beneficial Effects
[0082] 1. The improved YOLOv9 target detection network of this method can improve the detection accuracy of low-resolution small targets, especially in the detection of human targets in long-distance surveillance videos. This improvement enhances the recognition and positioning capabilities of the monitoring system in complex scenarios, especially in key areas such as security and traffic monitoring, and helps to enhance the adaptability of the public safety system, which has important social value.
[0083] 2. The improved YOLOv9 target detection network by this method can improve the detection accuracy of targets with extreme postures or large shape changes, and can enhance the network's adaptability to extreme posture changes and shape transformations, effectively solving the shortcomings of traditional methods in dynamic scenes, and can more accurately identify targets in different postures, action states or with large morphological changes, thereby improving the reliability and intelligent recognition capabilities of the system in actual monitoring and security protection applications.
[0084] 3. The improved YOLOv9 target detection network of this method can improve the detection accuracy of targets under occlusion, so that the network can still effectively identify targets under occlusion conditions. This improvement improves the robustness of target detection, especially in application scenarios such as traffic monitoring and public safety. It can still accurately identify and locate human targets in complex scene backgrounds or crowded scenes, thereby improving the responsiveness of the entire monitoring system.
[0085] Through different experimental methods and their result data:
[0086] (1) Comparative experiment on low-resolution small target detection of people:
[0087] In long-distance video scenes, human targets are small in size and low in resolution. Fig.16 The summary of the impact of different metrics on the test results of the YOLOv9 model on this video scene data is shown.
[0088] Experimental results show that the traditional IoU metric performs relatively poorly in this test, especially in small object detection, where the accuracy is significantly lower than other more advanced metric methods. The index achieved a relatively impressive result (0.776), but its limitations began to emerge when facing more complex scenes. In contrast, GIoU, as an extension of IoU, has some improvements, but the improvement is not significant. After further introducing CIoU and DIoU, the performance of the model has been significantly improved, especially in small objects and complex environments. For example, CIoU The improvement on CIoU is significant, reaching 0.601, which is better than DIoU and GIoU. This shows that CIoU has better optimization in dealing with target position relationship and distance difference, which can effectively enhance the accuracy and robustness of the model. However, the most superior performance comes from the NWD metric method, which shows the best performance in multiple evaluation indicators, especially in In terms of key indicators such as , NWD reached 0.758, 0.785, 0.833 and 0.643 respectively, which is significantly higher than other measurement methods. This shows that NWD can provide more accurate target recognition and positioning in complex multi-scale scenes and different resolutions.
[0089] (2) Comparative experiments on detection of people’s posture changes and irregularly shaped targets:
[0090] The significance of changes in a person's posture is reflected through various dynamic behaviors of the person (such as standing, sitting, walking, turning, etc.). The detection accuracy of YOLOv9 in such scenarios is tested when the person's posture changes and the body shape is irregular. Fig.17 , Fig.18 The summary of the impact of introducing the receptive field enhancement module (RFEM) or deformable convolution (DCN) on the test results of the YOLOv9 model on this dataset is shown.
[0091] Fig.17 The experimental results show that the model accuracy has been improved after the introduction of RFEM. From 0.594 to 0.611, and The number of parameters (Param) has also increased slightly to 39.1M, and FLOPS has dropped to 98.3. After the introduction of the P_RFEM module, the model has further improved in all accuracy indicators. reached 0.784, reached 0.808, The performance is greatly improved, reaching 0.882. The number of parameters increases to 47.4M, and FLOPS decreases to 91.7. Although the amount of calculation increases slightly, the significant improvement in accuracy shows that the model has achieved a good balance between accuracy and efficiency.
[0092] Fig.18 Experimental results show that with the introduction of deformable convolution DCN, the model has been significantly improved in terms of accuracy. When only DCN is added to the Backbone of YOLOv9, From 0.594 to 0.633, and They also reached 0.713 and 0.811 respectively, showing stronger feature extraction capabilities. When DCN is added to the Neck of YOLOv9, Improved to 0.617, and The results are 0.702 and 0.794 respectively. Although the accuracy has been improved, the effect is limited compared to the introduction of DCN in Backbone. The number of parameters is 41.2M, the FLOPS is 93.6, and the amount of calculation is slightly higher than before. When DCN is added to Backbone and Neck of YOLOv9 at the same time, the performance of the model is greatly improved. reached 0.674, is 0.743, The value is 0.830, which shows the combined effect of DCN in two key positions. Although the number of parameters increases to 48.9M and FLOPS drops to 88.7, the overall performance improvement is more significant compared to adding DCN to Backbone or Neck alone.
[0093] (3) Comparative experiment on target detection with partial occlusion by people:
[0094] By artificially adding occlusion elements to the video scene (such as overlapping occlusion between people, occlusion by chairs, TVs, doors, etc.), the feature information of some body parts of the target person is lost, thus simulating the occlusion problem that may be encountered in actual applications. The focus is on evaluating the target positioning and feature extraction capabilities of YOLOv9 under occlusion conditions. Fig.19 It shows a summary of the impact of introducing the spatial enhanced attention module (SEAM) on the test results of the YOLOv9 model on this dataset.
[0095] Fig.19 Experimental results show that adding the Spatial Enhanced Attention Module (SEAM) has different effects at different locations. After adding SEAM at the Backbone_top position of YOLOv9 and the bottleneck position of the backbone network, From 0.594 to 0.558, and The parameters increased to 39.2M, FLOPS slightly increased to 107.4, and the computational overhead increased. When SEAM was added to the Neck position of YOLOv9, Improved to 0.617, and The parameters are reduced to 36.4M, and FLOPS is increased to 111.3. Surface SEAM can effectively improve the detection performance of occluded targets when optimizing the Neck position.
[0096] The visual comparison between YOLOv9 and the proposed method in multi-view surveillance video data is also tested for the effect of detecting people in the conference room. Fig. 20 , 21 The left side shows the original YOLOv9 method and reasoning results, and the right side shows the reasoning results of the improved method. It can be seen that the native YOLOv9 network is prone to false detection in the conference room scene, and there will be some missed detections for some low-resolution personnel or partially occluded personnel. The improved method significantly reduces these errors. The effectiveness of the present invention has also been verified in actual scenarios.
[0097] What is disclosed above is only a preferred embodiment of the personnel target detection network design method for open scene video surveillance of the present invention. Of course, it cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiments and equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A method for designing a human target detection network for open scene video surveillance, characterized in that: The following steps are involved: Use multiple cameras to simulate multi-viewpoint human target detection scenarios and construct training and test sets. Improve the detection model; The detection model is trained and tested by using the training set and the test set to obtain a target detection model; Input videos from different perspectives into the target detection model to generate result videos with detection box information.
2. The method for designing a human target detection network for open scene video surveillance according to claim 1, characterized in that: In "Simulating multi-view human target detection scenarios through multiple cameras and building training and test sets", the following steps are included: Set up multiple cameras for video acquisition; The collected videos are divided into training sets and test sets.
3. The method for designing a human target detection network for open scene video surveillance according to claim 2, characterized in that: In "Setting up multiple cameras for video acquisition", the total video recording time of each camera is 35 minutes, and scene pictures are extracted every 20 frames.
4. The method for designing a human target detection network for open scene video surveillance according to claim 1, characterized in that: In "Improving the detection model", the following steps are included: Use the NWD metric method to replace the IoU of the detection model; Introduce the initial receptive field enhancement module into the detection model; Introducing deformable convolution into detection models; Add the spatial enhanced attention module to the detection model.
5. The method for designing a human target detection network for open scene video surveillance according to claim 4, characterized in that: The NWD metric method is used to model the bounding box as a two-dimensional Gaussian distribution, and the normalized Wasserstein distance (NWD) is used to replace the traditional IoU to reduce the position sensitivity of low-resolution small targets; the initial receptive field enhancement module is based on the multi-branch structure of dilated convolution, captures multi-scale information through branches with different expansion rates, and fuses weighted features; the deformable convolution is used to dynamically adjust the convolution kernel sampling position to improve the detection ability of deformed targets; the spatial enhanced attention module restores occlusion features through channel and spatial mixing modules to enhance the target area response.