Video monitoring dynamic deployment and control method based on multi-target identification

By performing multi-object recognition and dynamic control of the video surveillance system, using HSV color space and area growth algorithm to calculate personnel density, and dynamically adjust camera parameters, the problem of difficult real-time perception of the gathering status of people in key areas in large outdoor venues is solved, and security response speed and resource utilization efficiency are improved.

CN120499347AActive Publication Date: 2025-08-15GUANGZHOU JINHUI COMM TECH CO LTD

Patent Information

Application Number
CN202510817279.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-08-15
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

When existing video surveillance systems deal with multi-objective identification and dynamic deployment, it is difficult to achieve efficient monitoring and optimal resource allocation in complex scenarios. Especially in large outdoor venues, it is difficult to perceive the gathering status of personnel in key areas in real time, resulting in untimely responses to security personnel and inefficient human resource allocation.

Method used

By collecting video data streams of multiple cameras, target detection and cutout processing, HSV color space feature decomposition and area growth algorithms calculate personnel density distribution, combine GIoU loss to optimize target positioning, build a camera priority queue, and dynamically adjust the camera shooting angle and focal length to focus on key areas.

Benefits of technology

Real-time perception of the gathering status of personnel in key areas in large outdoor venues is achieved, the response speed of security personnel and the efficiency of human resource allocation is improved, and the efficient utilization of camera resources is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499347A_ABST
    Figure CN120499347A_ABST
Patent Text Reader

Abstract

The invention provides a video monitoring dynamic deployment and control method based on multi-target recognition, relates to the technical field of computer vision, and aims to detect the personnel density condition of a region through a data stream of a camera and dynamically adjust the shooting angle and the focal length of the camera according to a detection result. The method provided by the invention can solve the technical problems of untimely response of security personnel and low human resource allocation efficiency caused by difficulty in real-time sensing of the gathering state of people in a key area in an outdoor large-scale venue to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a video surveillance dynamic control method based on multi-target recognition. Background Art

[0002] With the rapid development of video surveillance technology, video surveillance systems based on intelligent analysis have become a crucial component in public safety, traffic management, and security for key venues. Traditional video surveillance systems primarily rely on manual real-time monitoring, failing to fully utilize the vast amounts of data collected by cameras and prone to blind spots and information omissions. In recent years, breakthroughs in computer vision and deep learning have led to the widespread application of technologies such as object detection, object tracking, and behavior analysis in the surveillance field. These technologies, through intelligent analysis of video data, can identify and track specific targets, significantly improving surveillance efficiency. However, most current intelligent surveillance technologies still have limitations. For example, they rely too heavily on single-target recognition algorithms, struggle to simultaneously process the dynamic characteristics of multiple targets, and suffer from reduced recognition accuracy and tracking effectiveness in complex scenes with densely populated targets. Furthermore, global camera scheduling strategies remain relatively crude, failing to fine-tune camera angles, focal lengths, and priorities based on real-time scene changes, resulting in inefficient resource utilization.

[0003] Existing multi-target detection algorithms are already capable of real-time detection and labeling of targets in video surveillance scenarios, but they still face numerous deficiencies in practical applications for dynamic surveillance. First, existing algorithms typically extract target features at the simple color or shape level, lacking in-depth analysis of target regions. This makes it difficult to accurately extract key target features in scenes with dense or overlapping targets. Second, dynamic regional analysis based on target features has yet to be fully implemented. Existing technologies often fail to effectively combine the distribution characteristics of key areas within a venue to accurately calculate and dynamically map target density, resulting in insufficient attention to high-risk areas. Furthermore, the dynamic adjustment mechanism for cameras is still imperfect. Existing automatic zoom and angle adjustment algorithms typically rely on fixed logic and are unable to dynamically optimize camera scheduling strategies based on target distribution characteristics and the venue's real-time needs. Therefore, in complex scenarios, existing technologies struggle to achieve efficient multi-target monitoring and optimal resource allocation. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides a video surveillance dynamic control method based on multi-target recognition, which can to a certain extent solve the technical problems that the gathering status of people in key areas of large outdoor venues is difficult to perceive in real time, resulting in untimely response of security personnel and inefficient human resource allocation.

[0005] According to one aspect of the present invention, a method for dynamic video surveillance based on multi-target recognition is provided, which includes:

[0006] Collecting video data streams from multiple cameras in the venue and obtaining monitoring scene images in the video data streams;

[0007] Performing target detection on the monitoring scene image to obtain a target detection frame, performing cutout processing on the target based on the target detection frame to obtain a target image, and performing feature decomposition on the target image according to the HSV color space to obtain a target feature map;

[0008] Based on the target feature map, a region growing algorithm is used to mark the connected domains of the target feature map to obtain a target focus area, and a personnel density distribution map within the target focus area is calculated. At the same time, the personnel density distribution map is spatially mapped with the preset venue key areas to obtain a regional focus coefficient;

[0009] A camera priority queue is constructed according to the area attention coefficient, the shooting angle and focal length of the camera are dynamically adjusted according to the priority queue, and the camera with a higher priority is turned to the target attention area for automatic zooming.

[0010] Furthermore, the target detection uses a target detection network to identify and locate human targets in the monitoring scene image;

[0011] The object detection network optimizes the feature map resolution transition through smooth up- and down-sampling, and uses GIoU loss to improve the positioning accuracy of low-overlap objects.

[0012] Furthermore, the up and down sampling determines the weighting coefficient by calculating the relative distance between the target pixel and the source pixel, and interpolates in steps in the horizontal and vertical directions, combined with boundary filling to ensure sampling continuity, thereby generating a smooth and high-quality target size feature map.

[0013] Furthermore, the GIoU loss includes:

[0014] Calculate the intersection area and union area of the predicted box and the real box to get the IoU value;

[0015] Determine the minimum enclosed region that can simultaneously surround the predicted box and the true box, and calculate the area of the minimum enclosed region;

[0016] Subtract the union area of the predicted box and the true box from the area of the minimum closure area, and divide the result by the area of the minimum closure area to obtain a penalty term;

[0017] The penalty term is added to the value obtained by subtracting the IoU value from 1 to obtain the GIoU loss value.

[0018] Furthermore, the cutout process includes the following steps:

[0019] Get the target frame coordinates output by the target detection network, analyze the target posture and calculate the main direction angle;

[0020] Rotate and align the target frame according to the main direction angle to obtain the minimum bounding rectangle parallel to the image coordinate axis;

[0021] Expand the minimum bounding rectangle boundary, adjust the part beyond the image boundary and fill it with boundary pixels to retain the context information around the target;

[0022] The projection transformation matrix is constructed based on the vertex coordinates of the minimum circumscribed rectangle, and the color value of each pixel of the target image is calculated through bilinear interpolation to complete the cutout processing.

[0023] Furthermore, the region growing algorithm includes:

[0024] On the target feature map, the point with the strongest HSV feature response is selected as the seed point through adaptive grid division;

[0025] Calculate the HSV feature vector of the seed point and use eight-neighborhood search in the region growing process to determine whether the adjacent pixels should be added to the current region based on the weighted combination of HSV spatial distance and position distance;

[0026] When multiple regions are adjacent, the feature similarity matrix is used to determine whether to merge them. When the threshold condition is met, the regions are merged and the statistical features are updated.

[0027] A two-pass scanning algorithm is used to complete the connected domain labeling to ensure that pixels in the same region have the same label;

[0028] The marked area is post-processed to remove areas with too small an area or irregular shape through area filtering and shape optimization to obtain the final marking result.

[0029] Furthermore, a density estimation network is constructed to calculate a personnel density distribution map, wherein the density estimation network adopts a multi-branch structure, including a global density estimation branch and a local detail enhancement branch.

[0030] Furthermore, spatial mapping is performed on the personnel density distribution map and the preset key areas of the venue to obtain a spatial mapping relationship;

[0031] Calculate the cumulative density value in each key area based on the spatial mapping relationship;

[0032] The regional attention coefficient is obtained based on the cumulative density value combined with the regional importance weight.

[0033] Furthermore, the calculation of the cumulative density value is based on the spatial position weight and the temporal smoothing factor, and the cumulative density value of the key area is obtained by layered weighted accumulation.

[0034] Furthermore, the dynamic adjustment adjusts the cameras in sequence according to the priority queue, calculates the horizontal rotation angle, pitch angle and focal length through the spherical coordinate system, ensures that the angle and focal length are within the mechanical limit and optical range, and gives priority to adjusting the camera with the best viewing angle when the target areas overlap.

[0035] Compared to existing technologies, the dynamic video surveillance deployment method based on multi-target recognition provided by this invention uses camera data streams to detect the density of people in an area and dynamically adjusts the camera's shooting angle and focal length based on the detection results. This can, to a certain extent, address the technical issues of difficulty in real-time detection of crowd concentrations in key areas of large outdoor venues, leading to delayed security responses and inefficient human resource deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0037] Figure 1 Flowchart of a method for dynamic video surveillance deployment based on multi-target recognition according to an embodiment of the present invention.

[0038] Figure 2 The figure is a flowchart of bilinear interpolation in a video surveillance dynamic control method based on multi-target recognition according to an embodiment of the present invention. DETAILED DESCRIPTION

[0039] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein.

[0040] Figure 1 FIG. 1 is a flow chart of a method for dynamic video surveillance based on multi-target recognition according to an embodiment of the present invention. Figure 1 As shown, the video surveillance dynamic control method based on multi-target recognition includes:

[0041] S1: Collect video data streams from multiple cameras in the venue and obtain monitoring scene images in the video data streams;

[0042] The video data streams of multiple cameras in the venue are collected, including fixed cameras at the venue entrances and exits, the stand area, the square area and the passage area, as well as rotatable pan-tilt cameras. The video data streams are sampled in real time at a frame rate of 25fps to obtain the monitoring scene images in the video data streams. The resolution of the monitoring scene images is 1920×1080 pixels to ensure the clarity of image details. The monitoring scene images are stored and transmitted in YUV420 format, where the Y component represents brightness information and the UV component represents chrominance information. The monitoring scene images are subjected to image enhancement processing, including brightness equalization and color correction, to eliminate image quality differences caused by changes in lighting at different time periods.

[0043] S2: Target detection is performed on the monitoring scene image to obtain a target detection frame, and the target is cut out based on the target detection frame to obtain a target image, and the target image is feature decomposed according to the HSV color space to obtain a target feature map;

[0044] Target detection is performed on the surveillance scene image, and an improved YOLOv5 target detection network is used to identify and locate human targets in the surveillance scene image. The backbone feature extraction network of the target detection network adopts the CSPDarknet53 structure, and the receptive field range is improved through the SPP spatial pyramid pooling module to output the target detection frame coordinate information with a confidence score. Based on the coordinate information of the target detection frame, the target area is cut out using the minimum enclosing rectangle method to obtain a target image, and the target image is size-normalized and uniformly scaled to 224×224 pixels. The target image is converted from the RGB color space to the HSV color space, where the H component represents hue information, the S component represents saturation information, and the V component represents brightness information. By independently decomposing the three HSV components, the color, texture and brightness features of the target are extracted respectively to obtain a target feature map. The target feature map adopts a multi-scale feature fusion method to weightedly combine HSV features at different scales to form a feature representation with rich details.

[0045] The improved YOLOv5 target detection network is optimized based on the standard YOLOv5 network. When the monitoring scene image is input, the feature extraction is first performed through the improved CSPDarknet53 backbone network, in which the CSPDarknet53 backbone network replaces the standard convolution operation with a depthwise separable convolution and introduces a channel attention mechanism in each CSP module to highlight the weights of important feature channels; the feature map extracted by the backbone network is subjected to multi-scale feature fusion through the improved feature pyramid network FPN. When the scale of the feature map changes, bilinear interpolation is used for upsampling and downsampling operations to ensure that the feature map is separated. An improved spatial attention module is added to the neck layer. When the target changes in scale or is densely distributed, the weight coefficients of different areas of the feature map are adaptively adjusted by calculating the response strength of the spatial position. A dense prediction mechanism is introduced in the detection head layer. When there is a large deviation between the predicted box and the true box, the box position accuracy is improved by refining the prediction offset. At the same time, the original cross entropy loss function is improved to the GIoU loss function based on IOU. When the overlap between the predicted box and the true box is low, more effective gradient information can be provided to guide network optimization. If the target is occluded, the occluded area is secondary detected and optimized through the set cascade detection mechanism.

[0046] The improved feature pyramid network FPN adopts an adaptive feature fusion strategy. When the backbone network outputs multiple feature maps of different scales, a feature pyramid structure is first constructed from deep to shallow. The feature pyramid contains five scale levels from P2 to P6. For the deepest P6 feature map, the number of channels is adjusted through a 1×1 convolution layer and then directly used as the top-level feature of the pyramid. When it is necessary to construct the P5 feature map, the P6 feature map is upsampled by 2 times and adaptively fused with the feature map of the corresponding level of the backbone network. The adaptive fusion includes calculating the channel attention weights and spatial attention weights of the two feature maps. When the importance of different regions and channels of the feature map is different, their contribution is dynamically adjusted through the attention mechanism. P4, P3 and P2 feature maps are constructed in the same way, and the feature fusion of each level adopts an adaptive weighting method.

[0047] After completing the top-down feature fusion, the feature maps of each level are optimized through 3×3 convolution. The optimization includes eliminating the aliasing effect caused by upsampling and enhancing the discriminability of features. If small-scale targets are detected, the high-resolution feature maps of the P2 and P3 levels are mainly used. When detecting large-scale targets, the large receptive field feature maps of the P5 and P6 levels are mainly relied upon. A dense connection mechanism is introduced between feature maps of different scales. When the features of a certain level are insufficient, they can be supplemented by the features of the adjacent levels. At the same time, a spatial attention module is embedded in the feature maps of each level. When the target is occluded or deformed, the feature expression of important areas is dynamically highlighted by calculating the spatial response map.

[0048] It should be noted that when it is necessary to perform an upsampling operation on the feature map, the feature map of the target size is first determined, and for each pixel position in the target feature map, its corresponding floating-point coordinates are found in the source feature map; the floating-point coordinates are usually located between four adjacent pixels in the source feature map, and the relative distance between the target position and the four source pixels is calculated at this time, and the contribution weight of each source pixel to the target position is determined based on the relative distance; when calculating the pixel value of the target position, the values of the four source pixels are weighted averaged according to their weights.

[0049] When performing the downsampling operation, the pixels in adjacent areas of the source feature map are also fused using a distance-weighted method. If the size of the source feature map is an integer multiple of the target feature map, the pixel values in the corresponding area are weighted according to their relative positions. When the ratio of the source feature map to the target feature map is not an integer, the source image area corresponding to each target pixel is determined by reverse mapping, and the pixels in the area are weighted averaged. If sampling occurs in the boundary area of the source feature map, the boundary pixels are filled to ensure the continuity of the sampling operation.

[0050] In the bilinear interpolation process, the interpolation operations in the horizontal and vertical directions can be performed step by step. When performing horizontal interpolation, linear interpolation calculation is first performed on each row; after completing the horizontal interpolation, the obtained intermediate results are then subjected to vertical linear interpolation to finally obtain a feature map of the target size; if the feature map contains multiple channels, the same interpolation operation is performed independently on each channel, as shown in the following example: Figure 2 shown.

[0051] The IOU-based GIoU loss function is an improved loss function designed for the bounding box regression problem in target detection. When the predicted box completely overlaps with the true box, the GIoU value is 1; when the predicted box is completely separated from the true box, the GIoU value is less than or equal to 0; the GIoU loss function first calculates the intersection area and union area of the predicted box and the true box to obtain the IoU value, and at the same time calculates the minimum closure area C that can simultaneously surround the predicted box and the true box; the area of the minimum closure area C is subtracted from the area of the union area of the predicted box and the true box, and then divided by the area of the minimum closure area C to obtain the penalty term; when the IoU value of the predicted box and the true box is the same, by introducing the penalty term, the predicted box is more inclined to approach the true box; if the predicted box is far away from the true box, the penalty term is larger, providing stronger gradient information to guide network optimization; the final GIoU loss value is equal to 1 minus the IoU value plus the penalty term. When the predicted box gradually approaches the true box, the loss value decreases monotonically. In more detail, the loss function L GIoU As shown in the following formula:

[0052]

[0053] Among them, B p is the area of the predicted bounding box, B g is the area of the true bounding box, C is the minimum closure area that surrounds both the predicted box and the true box, and |·| represents the area of the calculation region.

[0054] Compared to the traditional IoU loss function, the GIoU loss function introduces the concept of closure regions, effectively solving the problem of vanishing gradients when the predicted box and the ground-truth box do not overlap. It also considers the overlap and relative position of the predicted and ground-truth boxes, more accurately guiding the network to learn the position information of the bounding box. It exhibits improved convergence and stability, especially when dealing with objects with large scale variations.

[0055] For example, when there are multiple people densely distributed in the monitoring scene, the target detection network will generate a prediction frame for each human target. If a human target is detected in the stand area, the predicted frame and the real frame will partially overlap. At this time, the intersection area and the union area of the two frames are calculated to obtain the IoU value; at the same time, the area of the minimum closure area C that can completely surround the predicted frame and the real frame is calculated; when the distance between the predicted frame and the real frame is close, the area of the minimum closure area C is relatively small, and the calculated penalty value is small. The final GIoU loss value is mainly determined by the IoU value;

[0056] If the predicted box deviates significantly from the true box, resulting in a decrease in the overlap between the two boxes, the IoU value decreases and the area of the minimum closure region C increases, which increases the calculated penalty term and significantly increases the GIoU loss value. Through the change in the loss value, the network obtains clearer gradient information during back propagation, driving the predicted box to adjust to the position of the true box. When the predicted box and the true box do not overlap at all, the traditional IoU loss will experience gradient vanishing, while the GIoU loss function can still calculate a valid loss value through the closure region C, continuing to provide directional guidance for network optimization.

[0057] On the other hand, the specific steps of the minimum bounding rectangle method for cutting out the target area are as follows: after obtaining the target frame coordinates output by the target detection network, the target frame is first subjected to posture analysis. If the target is tilted, the main direction angle of the target is calculated; the main direction angle is based on the pixel distribution characteristics of the target area and is obtained by calculating the second-order matrix eigenvalues and eigenvectors of the regional pixels; after determining the main direction, the target frame is rotated and aligned so that it is parallel to the image coordinate axis, thereby obtaining the minimum bounding rectangle of the target; after obtaining the minimum bounding rectangle, the rectangular boundary is expanded. When the target is close to the edge of the image, the expansion range is adjusted according to the actual situation. If it exceeds the image boundary, the boundary pixels are used for filling; the expansion operation is intended to retain the contextual information around the target. When the target is partially occluded, the contextual information is used to assist subsequent feature extraction; if the width-to-height ratio of the target area is abnormal, it is cropped or filled according to the set threshold to ensure the rationality of the cutout result.

[0058] When performing the cutout operation, the projection transformation matrix from the original image to the target image is first constructed based on the coordinates of the four vertices of the minimum bounding rectangle; when performing pixel mapping, the bilinear interpolation method is used to calculate the color value of each pixel position in the target image.

[0059] The process of constructing the projection transformation matrix from the original image to the target image is as follows: first, obtain the coordinates of the four vertices of the minimum enclosing rectangle and arrange them in the order of upper left, upper right, lower right, and lower left to form a source point coordinate set; at the same time, according to the preset size of the target image (224×224 pixels), construct a corresponding target point coordinate set so that the four corner points of the target image correspond to the four vertices of the minimum enclosing rectangle respectively; after the source point coordinate set and the target point coordinate set are determined, a homography matrix solving method based on the least squares method is used to establish a 3×3 projection transformation matrix H from the source space to the target space; the projection transformation matrix H contains transformation information such as rotation, scaling, and translation. When the source point and target point are both four pairs, 8 linear equations can be listed, and the 9 elements of the matrix H are solved by the singular value decomposition (SVD) method.

[0060] The constructed projection transformation matrix H must satisfy the following properties: when any point in the source image is transformed, the product of its homogeneous coordinates and the matrix H is the position of the corresponding point in the target image; if there is an error in the point position before and after the transformation, the element values of the matrix H are optimized by minimizing the reprojection error; when the matrix H is determined, each pixel in the original image is projected and transformed to obtain the mapping position in the target image coordinate system, thereby achieving geometric correction and size normalization of the image. In more detail, the projection transformation matrix is shown as follows:

[0061]

[0062] Among them, H * is the 3×3 projection transformation matrix, p i is the homogeneous coordinates (x_i, y_i, 1)^T of the i-th vertex in the source image, p′ i is the homogeneous coordinates (x'_i,y'_i,1)^T of the i-th vertex in the target image, α i is the weight coefficient of the i-th vertex pair (calculated based on the reliability of the vertex position), d i is the Euclidean distance between the i-th pair of vertices, β is the distance attenuation factor, μ is the regularization coefficient, ∥*∥ 2 represents the L2 norm, ∥*∥ F represents the Frobenius norm

[0063] S3: Based on the target feature map, a region growing algorithm is used to mark the connected domains of the target feature map to obtain a target focus area, and a personnel density distribution map within the target focus area is calculated. At the same time, the personnel density distribution map is spatially mapped with a preset venue key area to obtain a regional focus coefficient;

[0064] Based on the target feature map, an improved region growing algorithm is used to mark the connected domains of the target feature map. When a seed point in the target feature map is detected, the region is expanded to the surrounding area with the seed point as the center. The seed point is selected based on the similarity measurement of the HSV feature space. When the difference between the HSV feature value of the adjacent pixel point and the seed point is less than the dynamically calculated growth threshold, the pixel point is merged into the current region. If multiple growing regions appear, the region merging criterion is used to determine whether to merge the adjacent regions. The region merging criterion considers the feature similarity and spatial distance relationship between the regions. After completing the region growing, all connected domains are marked to obtain the target focus region.

[0065] The improved region growing algorithm is optimized based on the standard region growing algorithm. First, a set of candidate seed points is established on the target feature map. When a local maximum point exists in the image, it is used as the initial seed point. The seed point selection process uses an adaptive grid partitioning method to divide the feature map into multiple sub-regions. In each sub-region, the point with the strongest HSV feature response is selected as the candidate seed point. After the seed point is determined, the HSV feature vector of the point is calculated, which includes a weighted combination of hue, saturation, and brightness components.

[0066] During the region growing process, an eight-neighborhood search strategy is used. When traversing to an adjacent pixel point, the characteristic distance between the point and the current seed point is calculated. The characteristic distance includes a weighted combination of the HSV spatial distance and the spatial position distance. If the characteristic distance is less than the adaptive growth threshold, the pixel point is added to the current region and the statistical characteristics of the region are updated. When multiple growing regions are adjacent, the region merging mechanism is activated, and the need for merging is determined by calculating the characteristic similarity matrix between regions. The characteristic similarity matrix considers the average HSV characteristics of the region, the regional shape characteristics, and the regional boundary gradient information. If the similarity between two adjacent regions exceeds the merging threshold and the boundary gradient between the regions is weak, the two regions are merged into a new region. After the region merge occurs, the statistical characteristics of the merged region are recalculated, and the region growing parameters are updated. In the connected domain labeling stage, a two-pass scanning algorithm is used to label all growing areas. The first pass scans the image from top to bottom and from left to right, assigning a temporary label to each pixel. The second pass solves the equivalent labeling problem to ensure that pixels in the same area have the same label. After labeling is completed, each connected domain is post-processed, including area filtering and shape optimization, to remove areas with too small an area or irregular shape, and obtain the final target focus area labeling result.

[0067] Based on the target area of interest, a density estimation network is constructed to calculate the personnel density distribution map. The density estimation network adopts a multi-branch structure, including a global density estimation branch and a local detail enhancement branch. When performing density estimation, the global density estimation branch extracts large-scale features through downsampling convolution to obtain the overall density distribution of the scene; the local detail enhancement branch maintains the original resolution and extracts local detail features through void convolution to accurately depict the distribution of personnel in high-density areas. The feature maps of the two branches are adaptively fused to obtain the final personnel density distribution map.

[0068] The global density estimation branch first downsamples the input feature map. When the feature map size is large, it is sequentially downsampled through a maximum pooling operation with a stride of 2 until the feature map size is reduced to 1 / 8 of the original size. The downsampled feature map is processed by multiple residual blocks, each of which contains two 3x3 convolutional layers and a shortcut connection. When the feature map passes through the residual block, the global semantic information of the scene is gradually extracted. A spatial pyramid pooling module is set after the residual block, which captures global contextual information under different receptive fields through adaptive pooling operations at three scales: 1x1, 2x2, and 4x4. If there is uneven density distribution in the scene, the feature response of the high-density area is enhanced through the attention mechanism.

[0069] The local detail enhancement branch maintains the original resolution of the input feature map and uses multi-scale dilated convolution for feature extraction. When performing dilated convolution, a 3x3 convolution kernel with dilation rates of 1, 2, and 4 is used, and an instance normalization layer and a LeakyReLU activation function are added after each convolution layer. The multi-scale features are adaptively fused through a feature selection module. When the features of a certain scale have strong expressive power for the current area, the weight of the features of that scale is enhanced through a soft attention mechanism. After the feature selection module, a detail enhancement module is set to enhance the density details of the local area through residual learning.

[0070] The feature fusion process of the two branches adopts an adaptive weighting strategy. First, the feature map of the global branch is upsampled to the original resolution through bilinear interpolation. When performing feature fusion, the channel attention weights and spatial attention weights of the feature maps of the two branches are calculated. The channel attention weights are calculated through the squeeze-and-excitation module to select important feature channels. The spatial attention weights are calculated based on the spatial response strength of the feature map. When a position has a strong response in both branches, the position obtains a higher fusion weight.

[0071] The feature map after feature fusion is passed through the density regression head for final density estimation. The density regression head contains multiple 1x1 convolutional layers, which is used to map the feature map into a single-channel density map. When generating the density map, the Poisson fusion algorithm is used to ensure the local consistency of the density value. If there are areas with abnormal density values, they are corrected through density value smoothing.

[0072] This dual-branch architecture achieves precise estimation of scene density distribution through the complementary effects of global and local branches. This is particularly true when dealing with scenes with large variations in crowd density, maintaining both global density accuracy and clarity of local details. Furthermore, through an adaptive feature fusion mechanism, it effectively integrates density information at different scales, improving the robustness and accuracy of density estimation.

[0073] The personnel density distribution map is spatially mapped with the preset key areas of the venue. When the key areas are irregular in shape, the density distribution map is aligned to the venue plan coordinate system using perspective transformation. Based on the spatial mapping relationship, the cumulative density value in each key area is calculated, and combined with the regional importance weight to obtain the regional attention coefficient.

[0074] The calculation of the cumulative density value is performed in a hierarchical weighted accumulation manner. When the density distribution map after spatial mapping is obtained, the key area is first divided into multiple grid cells of equal area. For each grid cell, the density value of the corresponding pixel point of the density distribution map is calculated. When a grid cell covers multiple pixels, the average density value of the grid cell is calculated using a bilinear interpolation method. The density value of the grid cell is related to its spatial position in the key area. When the grid cell is located at the center of the area, a higher position weight coefficient is assigned. When the grid cell is located at the edge of the area, a lower position weight coefficient is assigned. If the density value of the grid cell changes continuously over time, a temporal smoothing factor is introduced to avoid drastic fluctuations in the density value. The weighted density values of all grid cells are accumulated to obtain the cumulative density value of the key area. More specifically, the cumulative density value is shown in the following formula:

[0075]

[0076] in,

[0077]

[0078] Among them, D total is the cumulative density value of the key area, N is the total number of grid cells, D i is the average density value of the i-th grid cell, W i is the spatial position weight coefficient, d i is the distance from the grid unit to the center of the region, σ is the spatial weight attenuation parameter, T i is the time series smoothing factor, D i (t) is the density value at the current moment, D i (t-1) is the density value at the previous moment, and α is the time series smoothing adjustment parameter.

[0079] S4: Construct a camera priority queue based on the area attention coefficient, dynamically adjust the shooting angle and focal length of the camera according to the priority queue, and turn the camera with higher priority to the target attention area for automatic zooming, and at the same time push the location information and personnel density distribution map of the target attention area to the security personnel terminal.

[0080] First, each camera is processed sequentially according to the priority queue. When a camera is identified for adjustment, the required horizontal rotation and pitch angles are calculated based on its installation location and the spatial coordinates of the target area of interest. This angle calculation uses a spherical coordinate system conversion method. If the calculated angle exceeds the mechanical limit of the camera, the closest feasible angle is selected. If the target area is within the common coverage area of multiple cameras, the camera with the best viewing angle is selected for adjustment.

[0081] The focal length adjustment is calculated based on the area of the target area and the observation distance. When the target area is large, the field of view is expanded by reducing the focal length value. If the target area is small and the distance is far, the image is magnified by increasing the focal length value. The focal length adjustment must ensure that it is within the optical zoom range of the camera. When the calculated focal length value is out of range, the closest available focal length value is used.

[0082] For example, suppose a stadium is equipped with five spherical cameras (labeled A, B, C, D, and E). According to the current priority queue, the northwest corner of the stands has the highest attention coefficient (covered by cameras A and B), followed by the southeast gate entrance area (covered by camera C), and then the central area (covered by cameras D and E). For the northwest corner area with the highest priority, the system first analyzes the position conditions of cameras A and B: camera A is 12 meters away from the target area and has a viewing angle of 15 degrees; camera B is 18 meters away from the target area and has a viewing angle of 8 degrees. Therefore, camera A with better viewing conditions is selected for adjustment. The specific adjustment process is as follows:

[0083] Horizontal angle adjustment: If the current horizontal angle of camera A is 225 degrees and the target area requires 245 degrees, the horizontal rotation should be completed in three steps (about 7 degrees each time).

[0084] Pitch angle adjustment: The current pitch angle is 35 degrees, and the target area requires 42 degrees. The pitch adjustment is completed in two steps (about 3.5 degrees each time);

[0085] Focal length adjustment: The target area is approximately 80 square meters and the distance is 12 meters. The calculated required focal length is 18mm. The current focal length is 25mm. Adjust to the target focal length in four steps.

[0086] Once Camera A is properly adjusted, the system switches to the second-priority southeast gate area and controls Camera C to make corresponding adjustments. This step-by-step, gradual adjustment ensures the surveillance system maintains optimal focus on key areas. When the priority of an area changes, the system recalculates and executes the new adjustment instructions.

[0087] In summary, the multi-target recognition-based dynamic video surveillance control method of the present invention has been demonstrated. This method uses the camera's data stream to detect the density of people in an area and dynamically adjusts the camera's shooting angle and focal length based on the detection results. This can, to a certain extent, address the technical issues of difficulty in real-time detection of crowd concentrations in key areas of large outdoor venues, leading to delayed security responses and inefficient human resource allocation.

[0088] Here, those skilled in the art will appreciate that the specific operations of each step in the above-mentioned video surveillance dynamic control method based on multi-target recognition have been described in detail in the above reference. Figure 1 and Figure 2 The description of the video surveillance dynamic control method based on multi-target recognition has been introduced in detail, and therefore, its repeated description will be omitted.

Claims

1. A video surveillance dynamic control method based on multi-target recognition, characterized in that: include: Collecting video data streams from multiple cameras in the venue and obtaining monitoring scene images in the video data streams; Performing target detection on the monitoring scene image to obtain a target detection frame, performing cutout processing on the target based on the target detection frame to obtain a target image, and performing feature decomposition on the target image according to the HSV color space to obtain a target feature map; Based on the target feature map, a region growing algorithm is used to mark the connected domains of the target feature map to obtain a target focus area, and a personnel density distribution map within the target focus area is calculated. At the same time, the personnel density distribution map is spatially mapped with the preset venue key areas to obtain a regional focus coefficient; A camera priority queue is constructed according to the area attention coefficient, the shooting angle and focal length of the camera are dynamically adjusted according to the priority queue, and the camera with a higher priority is turned to the target attention area for automatic zooming.

2. The video surveillance dynamic control method based on multi-target recognition according to claim 1 is characterized in that: The target detection uses a target detection network to identify and locate human targets in the monitoring scene image; The object detection network optimizes the feature map resolution transition through smooth up- and down-sampling, and uses GIoU loss to improve the positioning accuracy of low-overlap objects.

3. The video surveillance dynamic control method based on multi-target recognition according to claim 2 is characterized in that: The up and down sampling determines the weighting coefficient by calculating the relative distance between the target pixel and the source pixel, and interpolates in steps in the horizontal and vertical directions, combined with boundary filling to ensure sampling continuity, thereby generating a smooth and high-quality target size feature map.

4. The video surveillance dynamic control method based on multi-target recognition according to claim 3 is characterized in that: The GIoU loss includes: Calculate the intersection area and union area of the predicted box and the real box to get the IoU value; Determine the minimum enclosed region that can simultaneously surround the predicted box and the true box, and calculate the area of the minimum enclosed region; Subtract the union area of the predicted box and the true box from the area of the minimum closure area, and divide the result by the area of the minimum closure area to obtain a penalty term; The value obtained by subtracting the IoU value from 1 is added to the penalty term to obtain the GIoU loss value.

5. The video surveillance dynamic control method based on multi-target recognition according to claim 1 is characterized in that: The cutout process comprises the following steps: Get the target frame coordinates output by the target detection network, analyze the target posture and calculate the main direction angle; Rotate and align the target frame according to the main direction angle to obtain the minimum bounding rectangle parallel to the image coordinate axis; Expand the minimum bounding rectangle boundary, adjust the part beyond the image boundary and fill it with boundary pixels to retain the context information around the target; The projection transformation matrix is constructed based on the vertex coordinates of the minimum circumscribed rectangle, and the color value of each pixel of the target image is calculated through bilinear interpolation to complete the cutout processing.

6. The video surveillance dynamic control method based on multi-target recognition according to claim 1 is characterized in that: The region growing algorithm includes: On the target feature map, the point with the strongest HSV feature response is selected as the seed point through adaptive grid division; Calculate the HSV feature vector of the seed point and use eight-neighborhood search in the region growing process to determine whether the adjacent pixels should be added to the current region based on the weighted combination of HSV spatial distance and position distance; When multiple regions are adjacent, the feature similarity matrix is used to determine whether to merge them. When the threshold condition is met, the regions are merged and the statistical features are updated. A two-pass scanning algorithm is used to complete the connected domain labeling to ensure that pixels in the same region have the same label; The marked area is post-processed to remove areas with too small an area or irregular shape through area filtering and shape optimization to obtain the final marking result.

7. The video surveillance dynamic control method based on multi-target recognition according to claim 1 is characterized in that: A density estimation network is constructed to calculate a personnel density distribution map. The density estimation network adopts a multi-branch structure, including a global density estimation branch and a local detail enhancement branch.

8. The video surveillance dynamic control method based on multi-target recognition according to claim 1 is characterized in that: Performing spatial mapping between the personnel density distribution map and the preset key areas of the venue to obtain a spatial mapping relationship; Calculate the cumulative density value in each key area based on the spatial mapping relationship; The regional attention coefficient is obtained based on the cumulative density value combined with the regional importance weight.

9. The video surveillance dynamic control method based on multi-target recognition according to claim 8 is characterized in that: The calculation of the cumulative density value is based on the spatial position weight and the temporal smoothing factor, and the cumulative density value of the key area is obtained by layered weighted accumulation.

10. The video surveillance dynamic control method based on multi-target recognition according to claim 1, characterized in that: The dynamic adjustment adjusts the cameras in sequence according to the priority queue, calculates the horizontal rotation angle, pitch angle and focal length through the spherical coordinate system, ensures that the angle and focal length are within the mechanical limit and optical range, and gives priority to adjusting the camera with the best viewing angle when the target areas overlap.

Citation Information

Patent Citations

  • Crowd aggregation detection method

    CN109063578A

  • Crowd density estimation system and method under multi-camera condition

    CN110543867A

  • Urban area deploying and controlling system based on body recognition and deploying and controlling method thereof

    CN110852287A

  • Real-time crowd density fusion sensing method and model based on camera cluster

    CN115223102A

  • Self-supervised image depth estimation method based on channel self-attention mechanism

    US20250078299A1

Cited By

  • Camera point complementing method for global tracking

    CN120751279A

  • A camera patching method for global tracking

    CN120751279B