High-definition monitoring camera tracking method and system for multi-angle face detection

Through multi-angle face detection methods and systems, the problems of low accuracy and poor stability of non-front angle face detection in complex monitoring environments are solved, and high-precision face detection and tracking are realized, which is suitable for lighting changes and face occlusion scenes in multi-camera environments.

CN120279064AInactive Publication Date: 2025-07-08SHENZHEN JIKEYUAN ELECTRONIC TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510764784.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has low accuracy in non-frontal angle face detection in complex monitoring environments, making it difficult to capture facial micro-expressions, and poor detection stability under light changes and face occlusion, resulting in a decrease in system reliability.

Method used

The high-definition surveillance camera tracking method using multi-angle face detection is used to achieve high-precision face correlation and trajectory smoothing across cameras through preprocessing, feature extraction and fusion, position-sensitive convolution network, facial area segmentation and angle characteristic analysis, combined with Retinex algorithm, adaptive color temperature estimation and depth feature matching.

Benefits of technology

It significantly improves the detection ability of non-facial angle faces, enhances the ability to capture facial micro-expression, improves the stability and detection accuracy of the system under light changes and face occlusion, and realizes high-precision face tracking in multi-camera environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279064A_ABST
    Figure CN120279064A_ABST
Patent Text Reader

Abstract

The invention relates to a high-definition monitoring camera tracking method and system for multi-angle face detection. The method comprises the following steps: preprocessing an original image collected by a high-definition monitoring camera to obtain a first preprocessed image and a second preprocessed image; performing feature extraction and feature fusion on the first preprocessed image and the second preprocessed image to obtain a multi-scale fusion feature map; face positioning and posture analysis are carried out through a position sensitive convolutional network, and a face detection result is obtained; performing face region subdivision and angle characteristic analysis on the face detection result to obtain enhanced face representation; and calculating a cross-camera matching cost matrix based on the enhanced face representation, and performing trajectory prediction and updating processing to obtain a face tracking result in the multi-camera environment. According to the method, the capability of capturing facial micro-expressions and fine features is improved, high-precision face association and smooth trajectory handover among different cameras are realized, and the problem of trajectory fragmentation in view crossing areas of the cameras is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of surveillance cameras, and particularly to a high-definition surveillance camera tracking method and system for multi-angle face detection. Background Art

[0002] With the rapid development of security surveillance technology, face detection and tracking have become the core functions of intelligent security systems, playing a crucial role in fields such as public security, identity authentication, and behavior analysis. Although methods based on deep learning have made significant progress in face detection tasks, they still face many challenges in actual complex surveillance environments. Existing technical solutions often perform poorly when dealing with non-frontal angle faces, and the accuracy of face detection for side, upward, or downward angle faces drops significantly, making it difficult to meet the requirements of omnidirectional surveillance; at the same time, these methods have limited ability to capture detailed features such as facial micro-expressions, resulting in deficiencies in advanced tasks such as identity recognition and emotion analysis.

[0003] The problem of light change in complex surveillance environments further exacerbates the difficulty of face detection. The dynamic changes in indoor and outdoor light intensity and angle lead to uneven face brightness, shadow interference, and overexposure or underexposure phenomena, greatly reducing the stability of traditional detection algorithms. At the same time, the problem of partial occlusion of faces in crowded scenes is also a major challenge faced by existing technologies. When a face is occluded by objects such as masks, sunglasses, hats, or partially occluded by other body parts, key facial feature points are lost, resulting in detection failure or tracking interruption, seriously affecting the reliability of the system. Summary of the Invention

[0004] The main objective of the present invention is to provide a high-definition surveillance camera tracking method and system for multi-angle face detection. The present invention improves the ability to capture facial micro-expressions and subtle features, realizes high-precision face association and smooth trajectory handover between different cameras, and solves the problem of trajectory fragmentation in the overlapping area of camera fields of view.

[0005] To achieve the above objective, the present invention provides a high-definition surveillance camera tracking method for multi-angle face detection, including the following steps: Preprocess the original images collected by the high-definition surveillance cameras to obtain a first preprocessed image and a second preprocessed image; Extract features and perform feature fusion on the first preprocessed image and the second preprocessed image respectively to obtain a multi-scale fusion feature map; Input the multi-scale fusion feature map into a position-sensitive convolutional network for face localization and pose analysis to obtain a face detection result; Perform facial region subdivision and angle characteristic analysis on the face detection result to obtain an enhanced face representation; Calculate the cross - camera matching cost matrix based on the enhanced face representation, and perform trajectory prediction and update processing to obtain the face tracking result in a multi - camera environment.

[0006] The present invention also provides a high - definition monitoring camera tracking system for multi - angle face detection, including: A pre - processing module for pre - processing the original images collected by the high - definition monitoring camera to obtain a first pre - processed image and a second pre - processed image; A feature extraction module for respectively performing feature extraction and feature fusion on the first pre - processed image and the second pre - processed image to obtain a multi - scale fusion feature map; A face detection module for inputting the multi - scale fusion feature map into a position - sensitive convolutional network for face localization and pose analysis to obtain a face detection result; A face enhancement module for performing facial region subdivision and angle characteristic analysis on the face detection result to obtain an enhanced face representation; A face tracking module for calculating a cross - camera matching cost matrix based on the enhanced face representation and performing trajectory prediction and update processing to obtain the face tracking result in a multi - camera environment.

[0007] In summary, the technical solution provided by the present invention processes images of different scales through a two - stream network structure, combines deformable convolution and multi - scale convolution modules, significantly enhances the face detection ability for non - frontal angle faces, and effectively solves the problem of difficult face detection for side, upward, and downward angle faces. The system uses the Retinex algorithm for illumination equalization processing, combines adaptive color temperature estimation and bilateral filter noise suppression mechanisms, so that the system maintains stable detection performance under different illumination conditions. The high - resolution image stream focuses on capturing detailed features, and cooperates with the facial region subdivision module to divide the face into multiple sub - regions for independent feature extraction, significantly improving the ability to capture facial micro - expressions and subtle features. The self - attention dilated convolution block effectively captures global features with long - distance dependencies, enhancing the robustness of the system to partial face occlusion. The combination of deep feature matching and the adaptive simulated annealing extended Kalman filter algorithm realizes high - precision face association and smooth trajectory handover between different cameras, solving the problem of trajectory fragmentation in the overlapping area of camera fields of view. The simulated annealing strategy avoids the local minimum problem, and cooperates with the adaptive weighted filter correction mechanism, so that the system can still maintain stable tracking in the case of fast face movement or partial occlusion. The feature fusion method of the multi - scale feature pyramid structure and the attention mechanism optimizes the calculation efficiency and is suitable for real - time monitoring system applications. The facial angle information extractor adopts a special feature enhancement strategy for faces of different angles, combined with the temporal consistency constraint module, making the system have stronger adaptability to environmental changes and face pose changes, and is applicable to a wider range of monitoring scenarios. Description of the Drawings

[0008] Figure 1 It is a schematic diagram of the steps of a high-definition surveillance camera tracking method for multi-angle face detection in an embodiment of the present invention; Figure 2 It is a block diagram of the structure of a high-definition surveillance camera tracking system for multi-angle face detection in an embodiment of the present invention.

[0009] The realization, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments

[0010] In order to make the object, technical solution, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0011] Referring to Figure 1 , this embodiment provides a high-definition surveillance camera tracking method for multi-angle face detection, including the following steps: S1, preprocess the original image collected by the high-definition surveillance camera to obtain a first preprocessed image and a second preprocessed image; Among them, multi-view image acquisition is performed by a high-definition surveillance camera array deployed in the target surveillance area. The resolution of the cameras is not less than 1920×1080 pixels and they are distributed at different spatial positions to achieve full coverage of the surveillance area. Each camera is equipped with an independent image acquisition module, and its acquisition frequency is maintained at 25 to 30 frames per second, so as to ensure that continuous and frame-complete raw image data can be captured in dynamic scenarios. All the acquired images are transmitted to the central processing unit in real time through a high-speed channel and uniformly enter the image preprocessing module for subsequent processing. An automatic white balance correction operation is performed on the raw images. An adaptive color temperature estimation algorithm based on the "gray world assumption" is adopted. Through the statistical analysis of the gray distribution in the images, the color temperature parameter matrix T of the current image is deduced, and based on this, the three RGB channels are adjusted and compensated respectively to generate a color-balanced corrected image. This process effectively corrects the color cast phenomenon generated by the images under different color temperature environments (such as sunlight, incandescent lamps or shadow areas), and improves the color stability and discriminability in subsequent recognition processes. To enhance the usability of the images under non-uniform illumination conditions, an illumination equalization processing mechanism based on the Retinex theory is introduced after color correction. This processing method maps the images from the original RGB space to a luminance space sensitive to illumination. By extracting the illumination component L and the reflection component R of the images, the dynamic range of L is compressed through logarithmic domain transformation, and then it is reconstructed and synthesized with R to generate a new image, resulting in a more balanced image under illumination conditions. After the illumination equalization process, a noise suppression operation is performed on the images to weaken the interference factors introduced by low illumination, camera thermal noise or transmission errors in the surveillance environment. This noise reduction process uses a bilateral filtering algorithm, which has the advantage of removing random noise in flat areas while retaining the edge structure information of the images. The parameters of the bilateral filter are set as the spatial distance weight σd is 3.0 and the pixel value difference weight σr is 0.1. This setting can balance the filtering intensity and the edge protection effect, so that the images can obtain a high signal-to-noise ratio while maintaining clarity. The image after noise suppression is the denoised image. The image resolution is standardized for the denoised images. Considering that the multi-scale feature extraction network structure needs to obtain global and local features simultaneously at different scales, the same image is adjusted to two fixed resolutions respectively. The image is scaled to 512×512 pixels and used as the input for the subsequent global semantic feature extraction path. This scale is suitable for quickly overviewing the entire face structure and pose information; the image is scaled to 1024×1024 pixels and used as the input for the detailed feature extraction path. This size is convenient for capturing fine-grained information such as local key points, textures and contours of the face. The images of the two scales respectively form the first preprocessed image and the second preprocessed image, and are jointly input into the dual-stream feature extraction network to ensure that the system can achieve accurate and stable face detection and tracking performance under different perspectives, different illuminations and different scale conditions.

[0012] S2. Feature extraction and feature fusion are respectively performed on the first preprocessed image and the second preprocessed image to obtain a multi-scale fusion feature map; Specifically, the first preprocessed image is input into the first feature extraction path, which uses an improved ResNeXt-101 network as the backbone structure. Its first stage includes several standard convolutional layers and residual modules. After the input image passes through this stage, a first-stage feature map is output, which contains low-level edge, color, and local structure information and has a high spatial resolution. On this basis, the first-stage feature map is input into the second stage of the network and processed through a series of residual blocks to obtain a second-stage feature map, which then enters the third, fourth, and fifth stages of processing in sequence, and the third-stage feature map, fourth-stage feature map, and fifth-stage feature map are output respectively. In the third and fourth stages, in order to enhance the network's adaptability to non-frontal face features, a deformable convolutional structure is introduced in the convolutional operation, and its offset field is adaptively learned through a sub-convolutional module to improve the network's feature response ability under complex pose changes; while in the fourth and fifth stages, multi-scale convolutional branches are integrated in parallel, and the combination of 1×1, 3×3, and 5×5 convolutional kernels is used to expand the receptive field, so as to balance the expression ability of large-scale facial structures and small-scale key points. The feature maps output by the above five stages respectively are uniformly input into the subsequent normalization and activation processing module after being calculated by the backbone network. Batch normalization operations are applied to the feature maps of each stage to alleviate the problem of internal covariate shift and accelerate training convergence by normalizing the activation distribution on each channel, and then the exponential linear unit (ELU) activation function is applied to the normalized results to make the network maintain non-linear responses in both positive and negative intervals and enhance the model's expression ability for complex boundary distributions. The feature maps obtained after this series of processing constitute the first feature set, which contains the full-level semantic information from low-level to high-level extracted from the first preprocessed image. At the same time, the second preprocessed image is input into the second feature extraction path, whose structure is also designed based on the ResNeXt-101 backbone, but the resolution of the input image is 1024×1024, and the higher input scale makes it more suitable for capturing detailed information. The image passes through the first stage of the network to generate a sixth-stage feature map, and then through the second stage of processing to obtain a seventh-stage feature map, and sequentially passes through the third, fourth, and fifth stages to output the eighth, ninth, and tenth-stage feature maps. Similarly, a deformable convolutional structure is introduced in the third and fourth stage convolutions to adapt to face angle and local geometric changes, and the design of the multi-scale receptive field is continued to improve the perception ability of complex multi-angle facial information. Consistent with the first path processing, the feature maps of the five stages are input into the batch normalization and ELU activation module for processing to ensure stability and non-linear response ability in deep semantic expression, and finally a second feature set is formed, representing the multi-scale feature information extracted from the second preprocessed image. The feature fusion of the attention mechanism is performed on the first feature set and the second feature set to obtain a multi-scale fusion feature map.

[0013] Apply one-dimensional convolution dimensionality reduction processing to the fourth-stage feature map and the fifth-stage feature map in the first feature set. Use one-dimensional convolution for linear compression to uniformly reduce their number of channels to 256 dimensions, generating the first dimensionality-reduced feature map. At the same time, apply one-dimensional convolution dimensionality reduction processing to the ninth-stage feature map and the tenth-stage feature map in the second feature set. After dimensionality reduction processing, they are also unified to 256 channels, forming the second dimensionality-reduced feature map. Input the first dimensionality-reduced feature map and the second dimensionality-reduced feature map into three parallel convolutional branches for channel concatenation. Each branch respectively uses a standard 3×3 convolution, a 3×3 dilated convolution with a dilation rate of 2, and a 3×3 dilated convolution with a dilation rate of 4 structure. By adjusting the sampling interval of the convolution kernel, different-sized feature responses can be perceived at the same spatial position. The three-way outputs are then concatenated in the channel dimension to form a complete first multi-scale receptive field feature map and a second multi-scale receptive field feature map. Apply a self-attention module to the first multi-scale receptive field feature map and the second multi-scale receptive field feature map. This module constructs three groups of representation vectors, namely query (Q), key (K), and value (V), through 1×1 convolution. Subsequently, calculate the attention weight matrix, the core of which is the dot-product attention structure. By scaling the product of Q and K and applying the softmax normalization operation, the attention matrix A = softmax(Q·Kᵀ / √d) is obtained, where d represents the vector dimension. Then, use A to perform weighted summation on the value matrix V to obtain an attention-enhanced representation containing global dependency relationships, and add it to the original input features through a residual connection to generate the first enhanced feature map and the second enhanced feature map. Introduce an adaptive weight fusion mechanism to merge the first enhanced feature map and the second enhanced feature map into a group of channel attention-enhanced features. This fusion process is based on feature correlation. Concatenate the two enhanced feature maps in the channel dimension and input them into two parallel small perceptron networks, which respectively correspond to two fusion coefficients α and β. Through linear transformation, ReLU activation, and sigmoid normalization processes, two weight coefficients in the channel dimension are dynamically generated, and these coefficients represent the contribution degrees of features from different channels. The fusion operation is completed in the way of weighted summation, that is, F = α·the first enhanced feature map + β·the second enhanced feature map, so as to achieve an adaptive fusion output driven by data. Input the fused features into the feature pyramid network structure for layer-by-layer integration. This structure starts with the fused features as the top-level P5 features, performs channel alignment through continuous upsampling operations and 1×1 convolutions, and then adds them to the next-layer fused input features to form P4, and repeats until P3, P2, and P1, thus forming a multi-scale feature pyramid set {P1, P2, P3, P4, P5} that covers different spatial resolutions and has continuous semantic associations. The top-down fusion strategy ensures the gradual propagation of high-semantic features and the complete retention of low-level detail features, enabling the system to obtain stronger context support and local perception capabilities in subsequent modules such as face localization and pose estimation.Generate a multi-scale fusion feature map.

[0014] S3. Input the multi-scale fusion feature map into a position-sensitive convolutional network for face localization and pose analysis to obtain face detection results. It should be noted that the multi-scale fusion feature map is input into the position-sensitive convolutional network to guide the context path of the decoder network for context information extraction. In the decoder network, the context path starts from the highest-level fusion feature map P5 and sequentially processes the input through multiple context blocks to extract higher-dimensional global semantic and spatial structure information in the image. Each context block contains a standard 3×3 convolutional operation, followed by batch normalization and ReLU activation functions, and then connects a global context module. This module obtains the global representation of the entire image through global average pooling, compresses the dimension through a 1×1 convolutional layer and applies non-linear activation to generate the context features after semantic enhancement. Subsequently, it adds them to the input feature map pixel by pixel to complete context enhancement. After the multi-layer context blocks are processed in series, a feature map with global context awareness ability is output, denoted as C5, which serves as the main channel of high-level semantics and will be used as the starting point of the decoding path. The context feature C5 is input into the decoding path for layer-by-layer restoration. C5 is adjusted to the same spatial size as the fusion feature map P4 through upsampling operation, and added to P4 after channel adjustment through 1×1 convolution to form the fusion feature map D4. Then, it is processed by a decoding block, which consists of two consecutive 3×3 convolutional layers, a batch normalization layer and a ReLU activation function. In this way, the decoding process of the feature map from the high level to the middle level is completed. This decoding method continues to be passed down, and D3, D2 and D1 are respectively constructed and processed by decoding blocks, and finally a decoding feature pyramid {D1', D2', D3', D4', D5'} consisting of five levels is formed, covering different spatial levels from low-level texture to high-level semantics, constituting a multi-scale response expression of the face region. The features at each level in the decoding feature pyramid are respectively input into the position-sensitive convolutional network for position-sensitive score calculation. This network performs k×k convolutional transformation on each level of decoding feature through the position-sensitive score map generation module to generate a feature map with (k×k)×c channels, where k is the position grid size (set to 7) and c is the number of categories (background and face). These feature maps have explicit position encoding ability, enabling the system to independently model the category response at each spatial position. Based on this score map, foreground probability prediction is performed on all candidate regions in the image to form a candidate region set, and its position information, scale information and category score are recorded. To obtain a more discriminative feature representation, a position-sensitive region pooling operation is applied to each candidate region. This operation divides the candidate region into k×k sub-regions, corresponding to each position-sensitive channel in the score map, and performs pooling operations within each sub-region respectively to extract the maximum response value or average value of each spatial segment, obtaining a two-dimensional feature vector with a fixed dimension. The two-dimensional feature vectors are respectively sent to the classification head and the regression head for final decision-making.The classification head consists of several fully connected layers and outputs a two-dimensional classification score vector using the softmax function. Each component represents the probability that the candidate region belongs to the background or a face, and the system uses this as the confidence index for the candidate region. The regression head, on the other hand, uses another set of fully connected structures to output two vectors. The first is the position regression vector, in the format (tx, ty, tw, th), representing the offsets of the predicted bounding box relative to the initial candidate box in terms of the center position and size. The second is the pose regression vector, used to characterize the face orientation information in the three-dimensional angular direction, including the yaw angle, pitch angle, and roll angle. The decoded outputs from all scales are jointly processed with the corresponding prediction results, and non-maximum suppression is performed to eliminate overlapping candidate boxes, with an overlap threshold set at 0.5, thus retaining the most confident set of face detection results. Each detection result contains complete position information (x, y), size (w, h), and pose angles, constituting the high-precision three-dimensional face detection output.

[0015] S4. Subdivide the face detection results in the facial region and analyze the angular characteristics to obtain an enhanced face representation; Specifically, based on the detection results output by the face detection module, which include the two-dimensional position coordinates, scale information, and three-dimensional pose angles of each face. Each face region is taken as the object to be processed, and a spatial transformation operation is performed. A three-dimensional rotation matrix is constructed using the previously regressed pose angles (including yaw, pitch, and roll angles), and based on this, the original face region is geometrically corrected to transform it into a standard frontal pose. This process achieves spatial normalization through affine transformation or perspective transformation, generating a structurally standardized face image region to eliminate the structural distortion problems caused by angle changes for subsequent feature extraction. After completing the spatial transformation, the standardized face regions are uniformly scaled to a fixed size, such as 128×128 pixels, and a uniform grid division operation is performed on them. The entire image is divided into 25 facial sub-regions in a 5×5 grid. Each sub-region is located in different anatomical regions of the face, such as the corners of the eyes, the wings of the nose, the cheekbones, the mandible, etc. These 25 sub-regions are respectively input into the local feature extraction network, which consists of three consecutive 3×3 convolutions. After each convolution, batch normalization and the LeakyReLU activation function are connected. The negative slope of the activation function is set to 0.1 to enhance the model's ability to express small negative responses. Through the layer-by-layer processing of this local convolutional network, each sub-region is mapped to a 128-dimensional local feature vector, thereby expressing the texture, boundary, and angle change information of this region in the high-dimensional feature space, generating a set of local feature vectors {f1, f2, ..., f25}, covering the local semantic structure of the entire face. A self-attention aggregation module is introduced to fuse the above set of feature vectors. This module projects each local feature vector into three groups of vectors: query (Q), key (K), and value (V) through a linear transformation, and then calculates the attention matrix A = softmax(Q·Kᵀ / √d) to measure the relative importance between sub-regions. Subsequently, the value vector V is weighted and summed using this attention matrix to obtain the fused global facial feature representation F. The face is classified into five types based on the predicted pose angles: frontal (|yaw| < 15° and |pitch| < 15°), left side view (yaw ≤ -15°), right side view (yaw ≥ 15°), looking up (pitch ≥ 15°), and looking down (pitch ≤ -15°). According to the divided angle categories, the system activates the corresponding angle feature enhancement sub-networks. These sub-networks adopt a unified residual structure, including two 1×1 convolutional layers, a batch normalization layer, and a ReLU activation function, with a specially designed angle-specific attention module inserted in the middle. This attention module dynamically adjusts the weights of the facial regions according to the angle type. For example, in the side view, it focuses on enhancing the ear and face edge regions; in the looking-up case, it enhances the tip of the nose and the jawline regions; and in the looking-down angle, it highlights the forehead and eyebrow bone regions' responses.Each sub-network outputs an angle-enhanced feature vector F'. While maintaining the original semantic integrity, it enhances the expression of key regions from a specific perspective, enabling the system to maintain a high recognition accuracy even in the face of large-angle changes. The angle-enhanced feature vector and the global feature vector are added together through a residual connection to form a multi-angle facial feature representation F''. To improve the system's perception of the consistency of human faces in temporally consecutive frames, a temporal consistency constraint mechanism is introduced. The cosine similarity between the feature F''(t) in the current frame t and the feature F''(t−1) in the previous frame t−1 is calculated, defined as s = cos(F''(t), F''(t−1)). This value is used to measure the stability of the features of the same identity over continuous time, generating a temporal consistency score s. This score can be used as a credibility index reference in subsequent multi-camera tracking and identity matching. The enhanced face representation in each frame is output, which includes the position coordinates (x, y), size (w, h), and pose angles (yaw, pitch, roll), as well as the fused multi-angle facial feature vector F'' and the temporal consistency score s, constituting a well-structured, angle-adaptive, and temporally stable face identity description result.

[0016] S5. Calculate the cross-camera matching cost matrix based on the enhanced face representation and perform trajectory prediction and update processing to obtain the face tracking results in a multi-camera environment.

[0017] Among them, cross-camera face similarity calculation is performed on the multi-angle facial feature vectors and temporal consistency scores in the enhanced face representations captured by different cameras to evaluate whether the faces detected by different cameras are of the same target. Calculate the similarity between each pair of cross-camera face features, and these similarities include the matching degree based on the facial feature vectors and the comparison of the temporal consistency scores. The temporal consistency score reflects the stability of the target in consecutive frames and improves the accuracy of matching. Through calculation, a cost matrix containing the matching costs of all face pairs is obtained, reflecting the similarity between each pair of candidate faces. A lower cost indicates that these two faces are more likely to belong to the same target. By calculating the minimum cost matching combination of this cost matrix and using the Hungarian algorithm or other optimization algorithms, the cross-camera face identity association result is obtained. This result assigns a unique identity label to each pair of faces detected in the cameras, thus completing the cross-camera face matching. After the information of the successfully matched face pairs is determined, these information are merged and processed to generate a global face representation. State prediction processing is performed on the global face representation to predict the face position at the current moment. State prediction is based on methods such as Kalman filtering or extended Kalman filtering. Through the Kalman filtering algorithm, based on the global face representation of the previous frame, the position of the face in the current frame is predicted. This process is based on the state transition model, taking the face state (including information such as position, velocity, acceleration, etc.) of the previous frame as input and calculating the preliminary prediction result of the face in the current frame. The preliminary prediction result is compared with the actual measurement value. The actual measurement value comes from the target detection module in the camera, which is the face position obtained through measurement. To improve the prediction accuracy, an adaptive noise estimation mechanism is introduced. This mechanism dynamically adjusts the measurement noise covariance matrix according to the prediction error of the previous frame. This adjustment process is determined according to the motion characteristics of the target and the current detection error, enabling the system to adaptively cope with changes in different environments, such as light changes, occlusion, and other interference factors. Through dynamic noise estimation, more accurate prediction and update are performed, thereby reducing error accumulation. After the optimized trajectory prediction, weighted filtering correction is performed. Weighted filtering correction adjusts the prediction result by weighted averaging and uses an adaptive weighting method based on perspective quality for correction. Calculate the quality weights of each camera perspective and adjust the prediction result according to these weights. The calculation of the weights takes into account factors such as the quality of the camera perspective (such as clarity, occlusion situation, etc.) and the temporal consistency score. Through weighted correction, the deviation caused by poor perspective quality or occasional mis-matching is effectively eliminated, ensuring that the final face trajectory result is more accurate. The face tracking results in a multi-camera environment are obtained, reflecting the position of the target and real-time tracking of the target's motion trajectory. These results include the identifier, three-dimensional position, motion state, pose angle, and historical trajectory of each face, providing target trajectory information.

[0018] Using estimation algorithms such as Kalman filtering, estimate the state of the human face at the current moment (including position, velocity, acceleration, etc.) based on the face state at the previous moment. To improve the prediction accuracy, compare these prediction results with the actual measurement values from the camera, and adjust and optimize the prediction accordingly. The actual measurement values reflect the position of the human face detected by the camera at the current time point, while the preliminary prediction results are speculations based on historical trajectories and the current state. The error between the two is used as a correction signal to guide the optimization of the prediction. After the preliminary comparison, to handle the non-linear characteristics in the prediction results, a non-linear measurement mapping function is used to map the preliminary prediction results into the observation space to calculate the expected observation values. The non-linear measurement mapping function is linearized through the Taylor series. The Taylor series approximation transforms the complex non-linear relationship into a manageable linear form, thus simplifying the subsequent calculation process to obtain the measurement Jacobian matrix, which contains the local linearization approximation of the non-linear measurement function with respect to the state variables. Based on the measurement Jacobian matrix, the predicted state covariance matrix, and the measurement noise covariance matrix, calculate the Kalman gain. The Kalman gain is a weight coefficient that determines the weight distribution between the predicted value and the actual measurement value. In the Kalman filtering framework, a higher Kalman gain indicates a higher degree of trust in the actual measurement value, while a lower Kalman gain indicates a stronger trust in the prediction result. The calculation of the Kalman gain is based on the covariance matrix, which reflects the magnitude of the prediction error, and the measurement noise covariance matrix, which reflects the uncertainty in the measurement process. By accurately calculating the Kalman gain, effectively balance the differences between the prediction and the actual measurement, ensuring that the final state update has a high accuracy. Calculate the state update based on the Kalman gain, the actual measurement value, and the expected observation value. Combine the actual measurement value and the predicted state value through weighted averaging to obtain the updated state estimate. At the same time, update the predicted state covariance matrix to provide new covariance information for the prediction at the next time step. To achieve adaptive noise adjustment, calculate the covariance of the prediction error sequence for the past N frames. By analyzing the change trend of the past prediction errors, evaluate the current level of measurement noise, and dynamically adjust the measurement noise covariance matrix based on this evaluation. Use the exponentially weighted moving average method to update the noise covariance matrix, and newer error data will be given a higher weight, thus responding more quickly to environmental changes. Through the adaptive noise estimation mechanism, adjust the noise model in real time to adapt to different measurement accuracies and dynamic environments, thereby improving the adaptability and accuracy of human face target tracking in complex environments. After completing the state update and noise estimation, use the corrected state estimate and the updated measurement noise covariance matrix for the prediction and update iteration at the next moment. Through continuous iteration, continuously optimize the human face trajectory prediction results in a multi-camera environment, ensuring that human face tracking can still maintain high accuracy under different environmental changes such as illumination, viewing angle, and occlusion.Each round of prediction and update is adjusted according to the latest error and noise estimates, so as to achieve accurate tracking of information such as the position and pose of the face, and finally obtain an optimized face trajectory prediction result.

[0019] In one example, preprocessing the original image collected by a high-definition surveillance camera to obtain a first preprocessed image and a second preprocessed image, including: Collecting multi-view images of the target surveillance area through a high-definition surveillance camera array to obtain the original image; Using an adaptive color temperature estimation algorithm to perform white balance correction on the original image to obtain a color-balanced corrected image; Using the Retinex algorithm to perform illumination equalization processing on the color-balanced corrected image to obtain an illumination-equalized image, and performing noise suppression processing on the illumination-equalized image to obtain a denoised image; Performing image resolution standardization processing on the denoised image, and uniformly adjusting the processed image to 512×512 pixels to obtain the first preprocessed image; Performing image resolution standardization processing on the denoised image, and uniformly adjusting the processed image to 1024×1024 pixels to obtain the second preprocessed image.

[0020] In this example, a set of high-definition surveillance camera arrays are deployed, and these cameras are distributed at different positions in the surveillance area. Through the cooperation of multiple cameras, image data of the same target area are collected from different perspectives. The resolution of each camera is not less than 1920×1080 pixels to ensure the clarity and richness of details of the images, and to ensure that sufficient detailed information can be captured from different surveillance angles. These original images are transmitted in real time to the central processing unit through a high-speed transmission channel and enter the subsequent image preprocessing stage. In the image preprocessing stage, white balance correction using an adaptive color temperature estimation algorithm is performed on the original images. White balance correction is used to adjust the colors in the images so as to maintain color consistency under different lighting conditions. In the natural environment, different light sources (such as sunlight, incandescent lamps, fluorescent lamps, etc.) will cause color deviations in the images, and the role of white balance is to eliminate this deviation so that the colors in the images are as close as possible to the colors of the real scene. The adaptive color temperature estimation algorithm calculates the color temperature parameter matrix T by analyzing the gray world assumption in the images, and then adjusts the color distribution of the RGB three channels according to this matrix, making the colors in the images more balanced and avoiding color distortion caused by color temperature changes. Through this step, the colors in the original images are corrected to natural and real colors, generating color-balanced corrected images. The color-balanced corrected images are subjected to illumination equalization processing to overcome the influence of illumination changes in different environments. In a complex surveillance environment, illumination changes will cause overexposure or underexposure phenomena in the images, thus affecting the extraction of image details and the effect of subsequent processing. The Retinex algorithm is used for illumination equalization. The Retinex algorithm is an image enhancement method that can remove the influence of illumination changes without distortion and retain the reflection component of the images. The Retinex algorithm decomposes the images into a reflection component and an illumination component. After processing, the true reflection characteristics of the images are restored. The logarithmic transformation in the algorithm can compress the illumination component and reduce the influence of illumination changes on the images, while during reconstruction, the reflection component is used to restore the true colors and details of the images, generating an illumination-equalized image. Noise suppression processing is performed on the illumination-equalized images to remove the noise information in the images. Through the bilateral filtering algorithm, filtering is achieved by considering both the spatial distance and pixel value differences of the pixels, thereby retaining the edge information in the images while removing the noise. During the bilateral filtering process, the parameters of the filter kernel are set. The spatial distance parameter σd is 3.0, and the pixel value difference parameter σr is 0.1. These two parameters determine the intensity of the filtering. By adjusting these two parameters, the effects of noise removal and edge retention are balanced, obtaining a denoised image. Image resolution normalization processing is performed on the denoised images. In order to be compatible with different image processing algorithms and deep learning models, the denoised images are adjusted to two resolutions respectively: 512×512 pixels and 1024×1024 pixels.These two resolutions respectively meet the input requirements of different scale models. Images of 512×512 pixels are suitable for fast calculation and global feature extraction, while images of 1024×1024 pixels can provide more detailed information and are suitable for finer feature extraction and detail analysis. Two preprocessed images of different scales are generated, namely images of 512×512 pixels and 1024×1024 pixels.

[0021] In one example, feature extraction and feature fusion are respectively performed on the first preprocessed image and the second preprocessed image to obtain a multi-scale fusion feature map, including: The first preprocessed image is input into the first stage of the first ResNeXt-101 network for processing to obtain a first-stage feature map; The first-stage feature map is input into the second stage, third stage, fourth stage to fifth stage of the first ResNeXt-101 network for processing to obtain a second-stage feature map, a third-stage feature map, a fourth-stage feature map, and a fifth-stage feature map; Batch normalization and an exponential linear unit activation function are applied to the first-stage feature map, the second-stage feature map, the third-stage feature map, the fourth-stage feature map, and the fifth-stage feature map to obtain a first feature set; The second preprocessed image is input into the first stage of the second ResNeXt-101 network for processing to obtain a sixth-stage feature map; The sixth-stage feature map is input into the second stage, third stage, fourth stage to fifth stage of the second ResNeXt-101 network for processing to obtain a seventh-stage feature map, an eighth-stage feature map, a ninth-stage feature map, and a tenth-stage feature map; Batch normalization and an exponential linear unit activation function are applied to the sixth-stage feature map, the seventh-stage feature map, the eighth-stage feature map, the ninth-stage feature map, and the tenth-stage feature map to obtain a second feature set; Feature fusion of the attention mechanism is performed on the first feature set and the second feature set to obtain a multi-scale fusion feature map.

[0022] In this example, the first preprocessed image is input into the first stage of the first ResNeXt-101 network for processing. Through the convolutional layer, spatial information is abstracted and feature learning is carried out. The selection of the convolutional kernel size, stride, and number of channels directly affects the quality and spatial resolution of the feature map. At this stage, the network extracts low-level features, such as basic morphological information like edges and textures, to obtain the feature map of the first stage, which contains the preliminary feature representations at various positions in the image. The feature map obtained in the first stage is input into the second stage of the first ResNeXt-101 network for processing. In this stage, the network continues to perform convolutional operations on the feature map to extract more detailed information. As the network deepens, the convolutional layers in each stage gradually transition from low-level features to high-level features. The feature map of the second stage contains more semantic information, such as the relationship between the contours of objects and regions. After the feature map of the first stage is processed by the second stage, the feature map of the second stage is obtained. Similarly, the network continues to input it into the third stage, fourth stage, and fifth stage for processing. In these stages, the network gradually extracts higher-level abstract features and obtains more complex detailed information in the image, respectively obtaining the feature maps of the third stage, fourth stage, and fifth stage. The feature maps obtained from the first stage to the fifth stage are processed with batch normalization and the exponential linear unit (ELU) activation function. Batch normalization is a technique used to accelerate training and improve the stability of the network. It can normalize the input of each layer so that the mean of the output of each layer is zero and the variance is one, thereby avoiding the problems of gradient disappearance or gradient explosion. Batch normalization can also promote the network's adaptation to changes in data distribution. The exponential linear unit (ELU) activation function is applied. By providing a negative exponential response for the negative value part, the nonlinear characteristics of the network are enhanced, which helps to capture richer features. Through this step, the feature maps of the first stage, second stage, third stage, fourth stage, and fifth stage are fused into the first feature set. The second preprocessed image is input into the first stage of the second ResNeXt-101 network for processing. In the second ResNeXt-101 network, the first stage performs preliminary feature extraction on the input image, similar to the operation of the first network, extracting the basic edges and texture information of the image. After completion, the feature map of the sixth stage is obtained. The feature map of the sixth stage is continued to be input into the second stage, third stage, fourth stage, and fifth stage of the second ResNeXt-101 network for processing. The detailed features in the image are deeply extracted to obtain the feature maps of the seventh stage, eighth stage, ninth stage, and tenth stage. After all stages of feature extraction are completed, similarly, batch normalization and the exponential linear unit (ELU) activation function are applied to process the feature maps obtained in the second ResNeXt-101 network. The application of batch normalization and the ELU activation function ensures the effectiveness of the feature maps, while enhancing the network's nonlinear expression ability and training stability.Through this series of operations, the feature maps of the sixth stage, seventh stage, eighth stage, ninth stage, and tenth stage finally constitute the second feature set. The feature fusion of the attention mechanism is performed on the first feature set and the second feature set to obtain a multi-scale fusion feature map. The attention mechanism dynamically adjusts the weights of each feature map by calculating the importance of each part in the feature map, so as to highlight important regions or information and suppress irrelevant parts. During the feature fusion process, the contribution of each feature map to the final output is determined by calculating the attention weights of each feature map.

[0023] In one example, the feature fusion of the attention mechanism is performed on the first feature set and the second feature set to obtain a multi-scale fusion feature map, including: Apply one-dimensional convolution dimensionality reduction processing to the feature maps of the fourth stage and fifth stage in the first feature set to obtain the first dimensionality-reduced feature map, and apply one-dimensional convolution dimensionality reduction processing to the feature maps of the ninth stage and tenth stage in the second feature set to obtain the second dimensionality-reduced feature map; Input the first dimensionality-reduced feature map and the second dimensionality-reduced feature map into three parallel convolution branches for channel splicing respectively to obtain the first multi-scale receptive field feature map and the second multi-scale receptive field feature map; Apply the self-attention module to the first multi-scale receptive field feature map and the second multi-scale receptive field feature map to obtain the corresponding first enhanced feature map and second enhanced feature map; Apply the adaptive weight fusion mechanism to the first enhanced feature map and the second enhanced feature map to obtain the channel attention enhanced feature, and integrate the channel attention enhanced feature through the feature pyramid network structure to form a multi-scale fusion feature map.

[0024] In this example, one-dimensional convolution dimensionality reduction processing is performed on the fourth-stage feature map and the fifth-stage feature map in the first feature set to obtain the first dimensionality-reduced feature map. The dimensionality reduction of the channel dimension of the feature map is carried out through one-dimensional convolution operations. The main role of one-dimensional convolution is to compress information along the channel dimension, which helps to reduce redundant features and maintain important spatial information. Since in a multi-layer convolutional network, as the depth of the layer increases, the number of channels of the feature map will increase, and excessive channels will increase the computational complexity and bring the risk of overfitting. By performing one-dimensional convolution dimensionality reduction operations, the number of channels of the feature map is reduced while maintaining its key features. After the dimensionality reduction operation, the first dimensionality-reduced feature map is obtained. One-dimensional convolution dimensionality reduction processing is applied to the ninth-stage feature map and the tenth-stage feature map in the second feature set to obtain the second dimensionality-reduced feature map. The first dimensionality-reduced feature map and the second dimensionality-reduced feature map are respectively input into three parallel convolutional branches for channel concatenation to obtain the first multi-scale receptive field feature map and the second multi-scale receptive field feature map. The dimensionality-reduced feature maps are respectively input into three parallel convolutional branches for further feature processing. Each convolutional branch extracts features of different scales through convolutional kernels of different sizes. These features capture local information of different sizes in the input image, thereby enhancing the multi-scale perception ability of the model. Each convolutional branch uses different convolutional kernel sizes (such as 1×1, 3×3, and 5×5) to capture image features at different scales. The output feature maps after processing will respectively represent the features of different receptive fields. The output feature maps from each branch are concatenated along the channel dimension to fuse the information of multiple receptive fields, forming the first multi-scale receptive field feature map and the second multi-scale receptive field feature map. The self-attention module is applied to the first multi-scale receptive field feature map and the second multi-scale receptive field feature map to obtain the corresponding first enhanced feature map and second enhanced feature map. The self-attention module calculates the importance of each position in the feature map and dynamically assigns weights to each region in the feature map, enabling the network to adaptively focus on important regions and suppress the interference of irrelevant regions. In this module, the attention weights of the feature map are calculated, and the attention weights represent the contribution size of each position to the final task. By performing weighted processing on these weights, the network adaptively adjusts the local region information in the feature map, thereby enhancing the feature expression ability of the key regions. After being processed by the self-attention module, the first multi-scale receptive field feature map and the second multi-scale receptive field feature map will be enhanced to generate the first enhanced feature map and the second enhanced feature map. The adaptive weight fusion mechanism is applied to the first enhanced feature map and the second enhanced feature map to obtain the channel attention enhanced feature. The adaptive weight fusion mechanism fuses multiple feature maps by assigning different weights to different feature maps, and this process can effectively combine the information in each feature map.By calculating the similarity between feature maps, the network adaptively assigns appropriate weights to each feature map, enabling the network to focus on features helpful for the task while ignoring irrelevant parts. This mechanism dynamically adjusts the way of feature fusion through the learned parameters obtained during training, so that the finally output features can better support subsequent tasks. The channel attention enhanced features are integrated through the feature pyramid network structure to form a multi-scale fusion feature map. The feature pyramid network is a network structure for multi-scale feature integration. By successively upsampling and weighted fusing feature maps of different scales, a feature map integrating multi-level information is obtained.

[0025] In one example, the multi-scale fusion feature map is input into a position-sensitive convolutional network for face localization and pose analysis to obtain face detection results, including: The multi-scale fusion feature map is input into the context path of the guiding decoder network in the position-sensitive convolutional network to extract context information, obtaining context features; Perform decoding path processing on the context features to obtain a decoded feature pyramid, and input the features at each level in the decoded feature pyramid into the position-sensitive convolutional network respectively for position-sensitive score calculation to obtain candidate regions including two categories: background and face; Perform position-sensitive region pooling operation on the candidate regions to obtain a two-dimensional feature vector; Input the two-dimensional feature vector into the classification head and regression head for processing. The classification head outputs a classification score vector representing the probabilities that the region belongs to the background and face, and the regression head outputs a position regression vector and a pose regression vector to obtain face detection results including position coordinates, size, and three-dimensional pose angles.

[0026] In this example, the multi-scale fusion feature map is input into the context path part of the guiding decoder network. In the context path, the network maps the semantic features from the high level to the semantic space features with stronger expression step by step through a series of context extraction modules. Each context extraction module includes a standard 3×3 convolutional layer, a batch normalization operation, and a ReLU activation function, and nests a global context modeling structure. The global context module extracts the global feature distribution of the entire image through global average pooling operation, then compresses the channels through 1×1 convolution, and then reconstructs the global context response through ReLU activation and another group of 1×1 convolution, and adds it element by element to the original features, so as to embed the global structural information into the local semantic representation to obtain the context features. Enter the decoding path stage. The decoding path upsamples the above context features step by step and fuses them with the shallow features at different scales, so as to gradually restore the spatial detail information and enhance the perception ability of the target boundary while maintaining the deep semantics. The decoder starts from the deepest context output feature C5. First, it restores its spatial resolution to the previous layer feature map P4 through the upsampling operation, then adjusts the channel dimension to be consistent with P4 through 1×1 convolution, and then performs the element-wise addition operation to complete the high-low level feature fusion and generate D4. The fusion result then enters a standard decoding module, which consists of two consecutive 3×3 convolutional layers, followed by a batch normalization layer and a ReLU activation function respectively, to further refine and semantically reconstruct the feature map in the spatial dimension. Proceed in this way successively to form a decoded feature pyramid {D1', D2', D3', D4', D5'}. Each level of the feature map in the decoded feature pyramid is input into the position-sensitive convolutional network to calculate the position-sensitive score map. The goal of this operation is to more accurately distinguish the probability that each position in the image belongs to the face or the background by means of a discriminative mechanism driven by position information. The position-sensitive score module processes each decoded feature map through a convolution operation set to k×k, and outputs a feature map with (k×k)×c channels, where c is the number of categories, which is 2 here, representing two categories: background and face. The network generates k×k position-sensitive response maps for each position, and each response map encodes the local response of whether this area is a face. Through this step, the entire image is divided into multiple candidate regions, and each region is attached with its classification score in the spatial position, forming a preliminary set of detection candidate regions. For the above generated set of candidate regions, perform the position-sensitive region pooling operation. This operation divides each candidate region into k×k sub-regions according to its coordinates, and performs pooling operations on these sub-regions respectively on the corresponding k×k position-sensitive score maps. Each sub-region extracts the local response value on the corresponding channel, ensuring that the spatial position information is strictly aligned and encoded into the features. Each candidate region is represented as a k×k×c-dimensional feature vector, which is flattened to obtain a two-dimensional feature vector. The two-dimensional feature vectors are respectively input into two branches: the classification head and the regression head.The classification head consists of several fully connected layers, and a softmax function is set at the output end to output a probability vector indicating whether the candidate region belongs to the background or a face, forming the final classification score. At the same time, the regression head also consists of several fully connected layers, and its output includes not only the bounding box regression vector of the region, representing the precise position coordinates (x, y) and size (w, h) of the face in the image, but also the regression vector of the three-dimensional pose angle. This set of pose parameters provides the rotation information of the face in space, supporting subsequent angle-aware tracking and feature enhancement. The finally achieved face detection results include the spatial position information (center coordinates and size) of the face of each detected target and the corresponding three-dimensional pose angle information.

[0027] In one example, the face detection results are subjected to facial region subdivision and angle characteristic analysis to obtain an enhanced face representation, including: Perform spatial transformation processing on the face region in the face detection results to obtain a normalized face region, and perform grid division processing on the normalized face region to form multiple facial sub-regions; Input the multiple facial sub-regions into a local feature extraction network respectively for local feature extraction to obtain the local feature vector corresponding to each sub-region; Perform self-attention aggregation processing on the set of local feature vectors to obtain a global feature vector, and classify the face based on the predicted pose angle to obtain an angle-enhanced feature vector; Add the angle-enhanced feature vector and the global feature vector by residual connection, and apply a temporal consistency constraint by calculating the cosine similarity between the current frame feature and the previous frame feature to obtain an enhanced face representation containing multi-angle facial feature vectors and temporal consistency scores.

[0028] In this example, for each face region included in the face detection result, a spatial transformation matrix is constructed using its corresponding three-dimensional pose information, namely, the three angular components of yaw, pitch, and roll output in the detection stage. This matrix is used to remap the detected face region back to a standard frontal pose reference frame, thereby eliminating the geometric distortion effects caused by pose rotation. The spatial transformation is achieved through affine transformation or three-dimensional face alignment methods. By rotating the current face pose back to the standard pose and resampling in the image domain, a standardized face region with a unified scale and pose is obtained. This region is adjusted to a fixed size, such as 128×128 pixels, to ensure consistency and alignment in size for subsequent network processing at all levels. A grid division process is performed on the standardized face region, dividing the standardized face region into 5×5 grid cells at equal intervals, forming 25 equally sized facial sub-regions. Each sub-region represents a local region at a fixed position in the face image, such as segments of regions like the forehead, eyebrows, eyes, nose, mouth, and jaw. In this way, the original face is deconstructed into a series of local blocks with fixed spatial meanings. The above 25 facial sub-regions are respectively input into a local feature extraction network for processing to extract the local representation information of each region. The local feature extraction network is implemented using a shallow convolutional neural network structure, specifically including three consecutive stacked 3×3 convolutional layers, followed by batch normalization and the LeakyReLU activation function in sequence after each convolutional layer. The negative slope is generally set to 0.1 to maintain a weak response ability to negative-valued features. Each local sub-region is encoded as a local feature vector of a fixed dimension, such as a 128-dimensional real-valued vector, representing the manifestation of this region in the feature space. After 25 independent forward propagation operations, a set of local feature vectors is obtained, containing 25 sub-vectors, respectively corresponding to the feature representations of each structural region of the face. The self-attention aggregation module is applied to the above set of local feature vectors to synthesize multiple region vectors into a global feature vector. This module projects all input vectors into three groups of matrices, namely query (Q), key (K), and value (V), through linear mapping, and then calculates the attention weight matrix A = softmax(Q·Kᵀ / √d), where d is the vector dimension, representing the scaling factor. Then, a weighted sum operation is applied to the value matrix V to generate the aggregated global feature vector. This self-attention mechanism significantly enhances the attention to key regions while suppressing the adverse interference of redundant regions on the overall expression. The generated global feature vector will represent a unified description of the entire face and possess region-selective feature weights. Based on the global feature vector, an angle-aware feature enhancement mechanism is introduced, that is, the face is classified according to the pose angle, and the corresponding angle-specific sub-network is called to enhance the global feature.The angle classification rule is usually based on the yaw and pitch angles, and the human face is divided into five categories: front face (|yaw| < 15°, |pitch| < 15°), left side (yaw ≤ -15°), right side (yaw ≥ 15°), looking up (pitch ≥ 15°), and looking down (pitch ≤ -15°). A specialized angle enhancement sub-network is configured for each type of angle. This network adopts a residual structure and consists of two cascaded 1×1 convolutional layers, a batch normalization layer, and a ReLU activation function. At the same time, an angle-specific attention module is inserted. This module emphasizes different facial regions according to different perspectives. For example, when looking at the side face, it highlights the ears and contour lines; when looking up, it highlights the chin and tip of the nose area; when looking down, it strengthens the forehead and eyebrow bone areas, and finally generates an angle-enhanced feature vector. The global feature vector and the angle-enhanced feature vector are fused through a residual connection, and a temporal consistency constraint mechanism is introduced. Denote the fused feature representation generated in the current frame as F(t), and the corresponding feature of the previous frame as F(t - 1). Calculate the cosine similarity s = cos(F(t), F(t - 1)) between the two. This similarity score is the temporal consistency score, representing the degree of stability of the representation of the same target between consecutive frames. This score is used as a dynamic matching reference basis in multi-frame fusion and tracking, for strengthening the matching judgment or guiding the target continuation in the case of low confidence. Through the above process, an enhanced face representation containing multi-angle facial feature vectors and temporal consistency scores is obtained.

[0029] In one example, based on the enhanced face representation, calculate the cross-camera matching cost matrix and perform trajectory prediction and update processing to obtain the face tracking results in a multi-camera environment, including: Perform cross-camera face similarity calculation on the multi-angle facial feature vectors and temporal consistency scores in the enhanced face representations captured by different cameras to obtain a cost matrix containing the matching cost of each pair of faces; Calculate the minimum cost matching combination based on the cost matrix to obtain the cross-camera face identity association result, and merge the information of the successfully matched face pairs according to the face identity association result to obtain the global face representation; Perform state prediction processing on the global face representation to obtain a preliminary prediction result of the face position; Compare the preliminary prediction result with the actual measurement value, and at the same time introduce an adaptive noise estimation mechanism to dynamically adjust the measurement noise covariance matrix to obtain an optimized face trajectory prediction result; Perform weighted filtering correction on the optimized face trajectory prediction result to obtain the face tracking result in a multi-camera environment.

[0030] In this example, two key features are obtained from the enhanced face representations extracted from the video frames captured by each camera: one is the multi-angle face feature vector, which represents the high-dimensional semantic representation of the face encoded by the self-attention and angle enhancement mechanism from the perspective of the current camera; the other is the temporal consistency score, which characterizes the degree of feature stability between the current frame and the previous frame and reflects the continuity of the target in the time domain. For this information, starting from the face sets in any two different cameras (e.g., c1 and c2), denoted as Ec1 and Ec2 respectively, assuming they contain n and m enhanced face representations, then a matching cost matrix C of size n×m is constructed, where each element Cᵢⱼ represents the matching cost between the i-th face in camera c1 and the j-th face in camera c2. The calculation of the cost comprehensively considers multiple dimensions, including feature similarity, spatial consistency, and temporal stability. Among them, the feature similarity uses the cosine distance to calculate the angular difference between two multi-angle face feature vectors, the spatial consistency estimates the distance error of the 3D face position based on the projection geometric relationship between the two cameras, and the temporal consistency uses the difference |sᵢ−sⱼ| as a quantization index of the dynamic difference. The calculation formula for each cost element Cᵢⱼ is: Cᵢⱼ = λ1·(1−cos(Fᵢ,Fⱼ)) + λ2·d(Pᵢ,Pⱼ) + λ3·|sᵢ−sⱼ|, where F is the face feature vector, P is the 3D position, s is the temporal consistency score, and λ1, λ2, λ3 are weighting coefficients satisfying λ1 + λ2 + λ3 = 1, which are used to balance the contributions of multi-source information. The cross-camera face identity association result is obtained by solving the minimum-cost matching combination of the cost matrix C. Optimization methods such as the Hungarian algorithm are used to find the optimal cross-view face matching pairs on the premise of ensuring the uniqueness of the matching and the minimum overall cost. The successfully matched face pairs are considered to be the same real person individuals observed by different cameras. The information of these matching pairs is merged, that is, a unified global identity identifier id is established for each pair of successfully matched faces, and the feature information, position information, and temporal information from the two perspectives are integrated to form a global face representation g. This global representation includes a unique identity number id, a global 3D position P, a fused face feature vector F, a comprehensive confidence score s (which can be generated by weighting multi-source scores), and a historical trajectory set H. To achieve dynamic continuous tracking of the face, state prediction processing is performed on the above global face representation. This process is carried out based on the structure of the extended Kalman filter. Its core is to define a state vector x to describe the motion state and pose changes of the target in space-time. This vector includes spatial coordinates (Px,Py,Pz), velocity (vx,vy,vz), acceleration (ax,ay,az), and 3D pose angles (yaw,pitch,roll).State prediction follows the non - linear dynamics equation \(x^{(t|t - 1)}=f(x^{(t - 1|t - 1)},u(t)) + w(t)\), where \(f\) is the state transition function, \(u(t)\) is the external control input, and \(w(t)\) is the process noise, which is assumed to follow a Gaussian distribution with covariance matrix \(Q\). The preliminary prediction results of the face position and pose in the current frame \(t\) are obtained through this prediction model. To improve the trajectory accuracy and robustness, the above - mentioned prediction results are compared with the measurement values actually extracted from the current frame in the image, and the measurement noise covariance matrix \(R\) is dynamically adjusted according to the error between the two, thereby introducing an adaptive noise estimation mechanism. This mechanism is based on historical prediction residuals, and the mean square or standard deviation of the error is statistically calculated through a sliding window to update the value of \(R\) in real - time, enabling the filter to automatically correct the confidence in the measurement data according to the current data stability. This method avoids the under - fitting or over - fitting problems caused by fixed parameter settings and enhances the model's adaptability to scene changes and data perturbations. The updated value of \(R\) is used to recalculate the Kalman gain, and the state estimate is updated accordingly, so that the finally output state result integrates the optimal estimates of prediction and observation. After obtaining the optimized face trajectory prediction results, weighted filtering correction processing is performed on them to adapt to the fusion error caused by the perspective differences of multiple cameras. Especially in the areas where there is cross - field of view or target switching between different cameras, problems such as trajectory jumps or redundancies are likely to occur. Therefore, a weighted strategy based on perspective quality is used to integrate multi - source observation data. A perspective weight function \(w_i=\exp(-|yaw_i| / \sigma_1-|pitch_i| / \sigma_2)\cdot s_i\) is constructed, where \(yaw_i\) and \(pitch_i\) are the pose angles corresponding to this observation, \(s_i\) is its confidence level, and \(\sigma_1\) and \(\sigma_2\) are smoothing factors that control attenuation. Through this weight function, the influence of non - frontal and low - confidence observation data in the fusion is effectively suppressed, and the stability and accuracy of the filtering estimation are improved, obtaining the face tracking results in a multi - camera environment.

[0031] In one example, the preliminary prediction results are compared with the actual measurement values, and at the same time, an adaptive noise estimation mechanism is introduced to dynamically adjust the measurement noise covariance matrix to obtain the optimized face trajectory prediction results, including: The preliminary prediction results are mapped to the observation space through a non - linear measurement mapping function, the expected observation value is calculated, and the non - linear measurement mapping function is linearized through the Taylor series to calculate the measurement Jacobian matrix; The Kalman gain is calculated according to the measurement Jacobian matrix, the predicted state covariance matrix, and the measurement noise covariance matrix; The state update is calculated based on the Kalman gain, the actual measurement value, and the expected observation value, and at the same time, the predicted state covariance matrix is updated to obtain the corrected state estimate; Perform covariance calculation on the prediction error sequence of the historical N frames, update the measurement noise covariance matrix through the exponentially weighted moving average method, and obtain an adaptively adjusted measurement noise covariance matrix; Use the corrected state estimate and the adaptively adjusted measurement noise covariance matrix for the prediction and update iteration at the next moment to obtain an optimized face trajectory prediction result.

[0032] In this example, the preliminary prediction result is mapped to the observation space through a non-linear measurement mapping function, and the current predicted state is converted into an expected observation value with the same dimension and meaning as the actual observed data. The non-linear function includes operations such as spatial projection, coordinate transformation, and angle transformation, and its specific form depends on the adopted sensor model and camera calibration parameters. The theoretically expected observation value is obtained at the current predicted state point through this non-linear function, and it is compared with the measurement value actually extracted from the image in the current frame to obtain the residual term for state correction. Since the non-linear function cannot be directly used in the linear matrix calculation framework of Kalman update, the Taylor series is used to perform local linearization on this non-linear function, that is, the first-order partial derivative is calculated near the predicted state to form a linear approximation expression of this non-linear mapping. This derivative matrix is the measurement Jacobian matrix, which reflects the degree of change caused by a small perturbation in each dimension of the predicted state in the observation space. The Kalman gain is calculated based on the measurement Jacobian matrix, the predicted state covariance matrix, and the measurement noise covariance matrix. The Kalman gain is an important coefficient that weighs the prediction result and the observation result, and determines the degree of trust in the observed data when the system corrects the state. The state covariance matrix generated by the previous state prediction and the set or estimated measurement noise covariance matrix are used to calculate this gain. The state covariance matrix describes the system's uncertainty estimation of the current predicted state, while the measurement noise covariance matrix reflects the reliability of the sensor observation. When the system prediction uncertainty is large and the observation reliability is strong, the Kalman gain will tend to assign a greater weight to the observed value. Through algebraic operations on the measurement Jacobian matrix, the state covariance matrix, and the measurement noise covariance matrix, a gain matrix in the multi-dimensional feature space is obtained, and this matrix will be used as an adjustment factor for subsequent state updates. The predicted state is corrected using this Kalman gain. The difference between the actual measurement value and the expected observation value is calculated to form an observation residual, which is multiplied by the Kalman gain and added to the predicted state to obtain the corrected state estimate. The predicted value and the observed value are fused proportionally, where the fusion ratio is adaptively determined by the Kalman gain to ensure that the updated state has the minimum mean square error. After the state vector is updated, the state covariance matrix is updated to reflect the system's re-estimation of the uncertainty of the current state after this round of observation correction. This update process will cause the state covariance to shrink, thereby enhancing the confidence in the prediction of the next moment and forming a complete "prediction-update" cycle. To improve the system's adaptability in a dynamic environment, especially when the quality of the observed data is unstable, the measurement noise covariance matrix is dynamically adjusted. This process relies on the statistical analysis of the prediction residuals in several historical frames. The observation residuals of each of the previous N frames are extracted, that is, the difference between the actual measurement value and the corresponding predicted observation value, and the covariance matrix of these N error vectors is calculated. This covariance represents the distribution of the observation errors of the system during the current period.To avoid sudden noise from causing severe disturbances to the system stability, the exponential weighted moving average method is adopted to smoothly update the covariance sequence. That is, a larger weight is assigned to the covariance of the new frame, while the weight assigned to the covariance of historical frames gradually decreases, forming an adaptive measurement noise covariance matrix update mechanism that is continuous in time and smooth in value. In this way, it is ensured that the system gradually reduces its dependence on measurement values when the observation quality deteriorates, and can automatically enhance its reference value after the observation data returns to stability, reflecting the fast adaptability of the filter to the environmental dynamics. The corrected state estimate and the latest adaptive measurement noise covariance matrix are used in the prediction and update iteration of the next moment, that is, re-entering the state prediction stage, relying on the state transition model to predict the face position and pose at the next moment, and continuing to perform a new round of residual calculation and state update after the next round of observation arrives. This iterative cycle continues, enabling the entire system to continuously fuse historical information and new observations in the time series, gradually approaching the true face trajectory, achieving continuous and high-precision tracking of the target, and obtaining an optimized face trajectory prediction result.

[0033] Referring to Figure 2 , this embodiment provides a high-definition surveillance camera tracking system for multi-angle face detection, including: A preprocessing module 1, configured to preprocess the original image collected by the high-definition surveillance camera to obtain a first preprocessed image and a second preprocessed image; A feature extraction module 2, configured to perform feature extraction and feature fusion on the first preprocessed image and the second preprocessed image respectively to obtain a multi-scale fusion feature map; A face detection module 3, configured to input the multi-scale fusion feature map into a position-sensitive convolutional network for face localization and pose analysis to obtain a face detection result; A face enhancement module 4, configured to perform facial region subdivision and angle characteristic analysis on the face detection result to obtain an enhanced face representation; A face tracking module 5, configured to calculate a cross-camera matching cost matrix based on the enhanced face representation and perform trajectory prediction and update processing to obtain a face tracking result in a multi-camera environment.

[0034] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to that described in the above method embodiment, and details will not be elaborated here.

[0035] It should be noted that in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, system, article or method comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, system, article or method. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, system, article or method comprising such element.

[0036] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall equally be included in the patent protection scope of the present invention.

Claims

1. A high-definition monitoring camera tracking method for multi-angle face detection, characterized in that, Including: Preprocessing the original images collected by high-definition surveillance cameras to obtain a first preprocessed image and a second preprocessed image; Performing feature extraction and feature fusion on the first preprocessed image and the second preprocessed image respectively to obtain a multi-scale fusion feature map; Inputting the multi-scale fusion feature map into a position-sensitive convolutional network for face localization and pose analysis to obtain a face detection result; Performing facial region subdivision and angular characteristic analysis on the face detection result to obtain an enhanced face representation; Calculating a cross-camera matching cost matrix based on the enhanced face representation and performing trajectory prediction and update processing to obtain a face tracking result in a multi-camera environment.

2. The high-definition monitoring camera tracking method for multi-angle face detection according to claim 1, wherein The preprocessing the original images collected by high-definition surveillance cameras to obtain a first preprocessed image and a second preprocessed image includes: Collecting multi-view images of a target surveillance area through a high-definition surveillance camera array to obtain original images; Performing white balance correction on the original images using an adaptive color temperature estimation algorithm to obtain color-balanced corrected images; Performing illumination equalization processing on the color-balanced corrected images using the Retinex algorithm to obtain illumination-equalized images, and performing noise suppression processing on the illumination-equalized images to obtain denoised images; Performing image resolution standardization processing on the denoised images, and uniformly adjusting the processed images to 512×512 pixels to obtain a first preprocessed image; Performing image resolution standardization processing on the denoised images, and uniformly adjusting the processed images to 1024×1024 pixels to obtain a second preprocessed image.

3. The high-definition monitoring camera tracking method for multi-angle face detection according to claim 1, characterized in that, The performing feature extraction and feature fusion on the first preprocessed image and the second preprocessed image respectively to obtain a multi-scale fusion feature map includes: Inputting the first preprocessed image into the first stage of a first ResNeXt-101 network for processing to obtain a first-stage feature map; Inputting the first-stage feature map into the second stage, third stage, fourth stage to fifth stage of the first ResNeXt-101 network for processing to obtain a second-stage feature map, a third-stage feature map, a fourth-stage feature map, and a fifth-stage feature map; Applying batch normalization and exponential linear unit activation function processing to the first-stage feature map, the second-stage feature map, the third-stage feature map, the fourth-stage feature map, and the fifth-stage feature map to obtain a first feature set; Inputting the second preprocessed image into the first stage of a second ResNeXt-101 network for processing to obtain a sixth-stage feature map; Inputting the sixth-stage feature map into the second stage, third stage, fourth stage to fifth stage of the second ResNeXt-101 network for processing to obtain a seventh-stage feature map, an eighth-stage feature map, a ninth-stage feature map, and a tenth-stage feature map; Applying batch normalization and exponential linear unit activation function processing to the sixth-stage feature map, the seventh-stage feature map, the eighth-stage feature map, the ninth-stage feature map, and the tenth-stage feature map to obtain a second feature set; Perform feature fusion of the attention mechanism on the first feature set and the second feature set to obtain a multi-scale fusion feature map.

4. The high-definition surveillance camera tracking method for multi-angle face detection according to claim 3, characterized in that, The performing feature fusion of the attention mechanism on the first feature set and the second feature set to obtain a multi-scale fusion feature map includes: Apply one-dimensional convolution dimensionality reduction processing to the fourth-stage feature map and the fifth-stage feature map in the first feature set to obtain a first dimensionality-reduced feature map, and apply one-dimensional convolution dimensionality reduction processing to the ninth-stage feature map and the tenth-stage feature map in the second feature set to obtain a second dimensionality-reduced feature map; Input the first dimensionality-reduced feature map and the second dimensionality-reduced feature map into three parallel convolution branches respectively for channel splicing to obtain a first multi-scale receptive field feature map and a second multi-scale receptive field feature map; Apply a self-attention module to the first multi-scale receptive field feature map and the second multi-scale receptive field feature map to obtain corresponding first enhanced feature maps and second enhanced feature maps; Apply an adaptive weight fusion mechanism to the first enhanced feature map and the second enhanced feature map to obtain a channel attention enhanced feature, and integrate the channel attention enhanced feature through a feature pyramid network structure to form a multi-scale fusion feature map.

5. The high-definition monitoring camera tracking method for multi-angle face detection according to claim 1, characterized in that, Inputting the multi-scale fusion feature map into a position-sensitive convolutional network for face localization and pose analysis to obtain a face detection result includes: Input the multi-scale fusion feature map into the context path of the guiding decoder network in the position-sensitive convolutional network for context information extraction to obtain context features; Perform a decoding path process on the context features to obtain a decoded feature pyramid, and input the features at all levels in the decoded feature pyramid into the position-sensitive convolutional network for position-sensitive score calculation to obtain candidate regions including two categories: background and face; Perform a position-sensitive region pooling operation on the candidate regions to obtain a two-dimensional feature vector; Input the two-dimensional feature vector into a classification head and a regression head for processing. The classification head outputs a classification score vector representing the probabilities that the region belongs to the background and the face, and the regression head outputs a position regression vector and a pose regression vector to obtain a face detection result including position coordinates, size, and three-dimensional pose angles.

6. The high-definition monitoring camera tracking method for multi-angle face detection according to claim 1, characterized in that, The performing facial region subdivision and angle characteristic analysis on the face detection result to obtain an enhanced face representation includes: Perform a spatial transformation process on the face region in the face detection result to obtain a normalized face region, and perform a grid division process on the normalized face region to form multiple facial sub-regions; Input the multiple facial sub-regions into a local feature extraction network for local feature extraction respectively to obtain local feature vectors corresponding to each sub-region; Perform self-attention aggregation processing on the set of local feature vectors to obtain a global feature vector, and perform type division on the face based on the predicted pose angle to obtain an angle-enhanced feature vector; Add the angle-enhanced feature vector and the global feature vector by residual connection, and apply temporal consistency constraints by calculating the cosine similarity between the current frame feature and the previous frame feature to obtain an enhanced face representation containing multi-angle face feature vectors and temporal consistency scores.

7. The high-definition monitoring camera tracking method for multi-angle face detection according to claim 1, wherein Calculate the cross-camera matching cost matrix based on the enhanced face representation and perform trajectory prediction and update processing to obtain the face tracking result in a multi-camera environment, including: Perform cross-camera face similarity calculation on the multi-angle face feature vectors and temporal consistency scores in the enhanced face representations captured by different cameras to obtain a cost matrix containing the matching cost of each pair of faces; Calculate the minimum-cost matching combination based on the cost matrix to obtain the cross-camera face identity association result, and merge the information of successfully matched face pairs according to the face identity association result to obtain the global face representation; Perform state prediction processing on the global face representation to obtain a preliminary prediction result of the face position; Compare the preliminary prediction result with the actual measurement value, and introduce an adaptive noise estimation mechanism to dynamically adjust the measurement noise covariance matrix to obtain an optimized face trajectory prediction result; Perform weighted filtering correction on the optimized face trajectory prediction result to obtain the face tracking result in a multi-camera environment.

8. The high-definition monitoring camera tracking method for multi-angle face detection according to claim 7, characterized in that, The step of comparing the preliminary prediction result with the actual measurement value, and introducing an adaptive noise estimation mechanism to dynamically adjust the measurement noise covariance matrix to obtain an optimized face trajectory prediction result includes: Map the preliminary prediction result to the observation space through a non-linear measurement mapping function, calculate the expected observation value, and linearize the non-linear measurement mapping function through Taylor series to calculate the measurement Jacobian matrix; Calculate the Kalman gain based on the measurement Jacobian matrix, the predicted state covariance matrix, and the measurement noise covariance matrix; Calculate the state update based on the Kalman gain, the actual measurement value, and the expected observation value, and update the predicted state covariance matrix at the same time to obtain the corrected state estimate; Perform covariance calculation on the prediction error sequence of the historical N frames, and update the measurement noise covariance matrix by the exponentially weighted moving average method to obtain an adaptively adjusted measurement noise covariance matrix; Use the corrected state estimate and the adaptively adjusted measurement noise covariance matrix for the prediction and update iteration at the next moment to obtain an optimized face trajectory prediction result.

9. A high-definition monitoring camera tracking system for multi-angle face detection, characterized in that, Steps for implementing the multi-angle face detection high-definition surveillance camera tracking method according to any one of claims 1 to 8, the multi-angle face detection high-definition surveillance camera tracking system includes: A preprocessing module for preprocessing the original image collected by the high-definition surveillance camera to obtain a first preprocessed image and a second preprocessed image; A feature extraction module for respectively performing feature extraction and feature fusion on the first preprocessed image and the second preprocessed image to obtain a multi-scale fusion feature map; A face detection module for inputting the multi-scale fusion feature map into a position-sensitive convolutional network for face localization and pose analysis to obtain a face detection result; A face enhancement module, which is used to perform facial region subdivision and angular feature analysis on the face detection result to obtain an enhanced face representation; A face tracking module, which is used to calculate a cross-camera matching cost matrix based on the enhanced face representation and perform trajectory prediction and update processing to obtain a face tracking result in a multi-camera environment.

Citation Information

Cited By

  • Foreign matter visual detection system and detection equipment thereof

    CN120890997A

  • Face recognition method and system suitable for outdoor electronic instrument

    CN120976991A

  • A method, device, equipment and medium for measuring three-dimensional pose of human face at a large angle

    CN122574923A