Pedestrian target tracking method, device and equipment in complex monitoring scene, medium and product

By introducing attention mechanism module and optimization of loss function into the target detection module, the problem of low pedestrian target tracking accuracy in complex traffic monitoring scenarios is solved, and higher detection accuracy and robustness are achieved.

CN119991737AInactive Publication Date: 2025-05-13BEIJING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510146960.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex traffic monitoring scenarios, it is difficult for the existing technology to ensure the accuracy and real-timeness of pedestrian target tracking, especially in specific scenarios such as dense crowds, complex environments, and small target sizes.

Method used

By introducing an attention mechanism module into the object detection module, including the channel attention submodule and the spatial attention submodule, the loss function is optimized to comprehensively consider the loss function of the intersection ratio, center point distance and aspect ratio, and track and predict based on the optimized object detection box.

Benefits of technology

The accuracy and robustness of target detection have been significantly improved, especially in the detection capabilities of small targets and occlusion targets, thus achieving more stable and accurate pedestrian target tracking in complex monitoring scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991737A_ABST
    Figure CN119991737A_ABST
Patent Text Reader

Abstract

The invention discloses a pedestrian target tracking method and device in a complex monitoring scene, equipment, a medium and a product, and relates to the field of traffic surveillance, the method comprises the following steps: preprocessing a real-time video stream, and determining a preprocessed video frame; inputting the preprocessed video frame into a target detection model into which an attention mechanism module is introduced, and determining a preliminary target detection frame; the attention mechanism module comprises a channel attention sub-module and a space attention sub-module; optimizing the initial detection frame by using the transformed loss function to obtain an optimized target detection frame; based on the optimized target detection frame, tracking prediction is carried out on a target in the real-time video stream, and a tracking result is determined; and the tracking result is the processed video frame picture. According to the method and the device, the pedestrian target tracking precision in a complex monitoring scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of traffic monitoring, and in particular to a method, device, equipment, medium and product for tracking pedestrian targets in complex monitoring scenarios. Background Art

[0002] With the rapid development of information technology, video surveillance systems have become an indispensable technical support in the field of traffic management. One of the core functions of video surveillance systems is video target tracking, which is to detect and track targets, especially pedestrians, in real time in video sequences. This technology plays a vital role in ensuring traffic safety, optimizing traffic flow, and improving traffic monitoring efficiency. However, video target tracking technology faces many challenges in actual traffic monitoring applications, especially in specific scenarios such as dense crowds, complex environments, and small target sizes. The accuracy and real-time performance of tracking are difficult to guarantee.

[0003] In actual traffic monitoring scenarios, the problems that target tracking technology needs to deal with are far more complex than those under laboratory conditions. Especially in urban traffic monitoring scenarios, the rapid movement of pedestrians and their frequent crossing and merging behaviors, as well as the mutual occlusion between pedestrians, all pose huge challenges to the continuous tracking of targets. In addition, since pedestrian targets are smaller in size than cameras and are easily obstructed, this brings additional difficulties to the accurate identification and tracking of targets. In urban traffic monitoring scenarios, the rapid movement of pedestrians and the complex and changeable traffic conditions require that the tracking algorithm must have high real-time and high accuracy. In addition, in severe weather or lighting conditions, such as rain and fog, low light at night, etc., the visibility of the target is reduced, further increasing the difficulty of tracking. Summary of the invention

[0004] The purpose of this application is to provide a method, device, equipment, medium and product for tracking pedestrian targets in complex monitoring scenarios to solve the problem of low accuracy of pedestrian target tracking in existing traffic monitoring scenarios.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides a method for tracking pedestrian targets in complex monitoring scenarios, including:

[0007] Preprocessing the real-time video stream and determining the preprocessed video frame;

[0008] Inputting the preprocessed video frame into a target detection model that introduces an attention mechanism module to determine a preliminary target detection frame; the attention mechanism module includes a channel attention submodule and a spatial attention submodule;

[0009] The modified loss function is used to optimize the preliminary detection frame to obtain the optimized target detection frame; the modified loss function is Among them, CIOU′ is the complete intersection-over-union ratio, which is a loss function that comprehensively considers the intersection-over-union ratio, center point distance and aspect ratio; IOU is the intersection-over-union ratio; l is the weight of the aspect ratio consistency loss; c is the Euclidean distance between the center point of the predicted box and the real box; d is the possible maximum distance between the center point of the predicted box and the real box; μ is the weight adjustment factor; a is the adjustment coefficient; v is the coefficient of the aspect ratio consistency loss; av is the value of the aspect ratio consistency loss;

[0010] Based on the optimized target detection frame, the target in the real-time video stream is tracked and predicted to determine the tracking result; the tracking result is a processed video frame image; the tracking result includes the target, the optimized target detection frame corresponding to the target, the target ID, the position coordinates, the width and the height.

[0011] In the second aspect, the present application provides a pedestrian target tracking device in a complex monitoring scenario, including:

[0012] A preprocessing module, used to preprocess the real-time video stream and determine the preprocessed video frame;

[0013] A preliminary target detection frame determination module, used to input the preprocessed video frame into the target detection model introduced into the attention mechanism module to determine the preliminary target detection frame; the attention mechanism module includes a channel attention submodule and a spatial attention submodule;

[0014] The optimization module is used to optimize the preliminary detection frame using the modified loss function to obtain the optimized target detection frame; the modified loss function is Among them, CIOU′ is the complete intersection-over-union ratio, which is a loss function that comprehensively considers the intersection-over-union ratio, center point distance and aspect ratio; IOU is the intersection-over-union ratio; l is the weight of the aspect ratio consistency loss; c is the Euclidean distance between the center point of the predicted box and the real box; d is the possible maximum distance between the center point of the predicted box and the real box; μ is the weight adjustment factor; a is the adjustment coefficient; v is the coefficient of the aspect ratio consistency loss; av is the value of the aspect ratio consistency loss;

[0015] A tracking result determination module is used to track and predict the target in the real-time video stream based on the optimized target detection frame to determine the tracking result; the tracking result is a processed video frame image; the tracking result includes the target, the optimized target detection frame corresponding to the target, the target ID, the position coordinates, the width and the height.

[0016] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the pedestrian target tracking method in any of the complex monitoring scenarios described above.

[0017] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the pedestrian target tracking method in any of the complex monitoring scenarios described above.

[0018] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the pedestrian target tracking method in any of the complex monitoring scenarios described above.

[0019] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0020] This application introduces an attention mechanism module into the target detection module, and the attention mechanism module includes a channel attention submodule and a spatial attention submodule, which weight the channels and spatial positions in the feature map respectively, highlight important features and suppress unimportant information. This dual attention mechanism enables Attn-YOLO to adaptively adjust the feature map to better adapt to different detection tasks, especially in video surveillance. The detection ability of small targets and occluded targets has been significantly improved, thereby improving the accuracy and robustness of target detection.

[0021] In addition, this application also transforms the loss function and uses the transformed loss function to optimize the preliminary detection frame and the optimized target detection frame. The transformed loss function comprehensively considers the overlapping area, center point distance and aspect ratio difference, and is not limited by the size of the pedestrian target relative to the camera, further enhancing the positioning accuracy of small targets.

[0022] Finally, based on the optimized target detection frame, the target state vector is considered to track and predict the target in the real-time video stream, and the tracking result is determined, so that the final tracking result is closer to the actual situation and adapts to the acceleration and deceleration changes of the target. The target state vector includes a position vector, a velocity vector and an acceleration vector. Even in severe weather or lighting conditions, more stable and accurate target tracking can be achieved in a dynamic and complex traffic environment, thereby improving the pedestrian target tracking accuracy in complex monitoring scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 A flow chart of the pedestrian target tracking method in complex monitoring scenarios provided by this application;

[0025] Figure 2 A schematic diagram of the target detection model structure provided in this application;

[0026] Figure 3 This is a graph of experimental results of the weight adjustment factor μ provided in this application;

[0027] Figure 4 Schematic diagram of the target tracking experiment effect provided in this application. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0029] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0030] The present application embodiment provides a method for tracking pedestrian targets in a complex monitoring scenario. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or can be executed by a terminal and a server together. In the present application embodiment, Figure 1 As shown, the method includes the following steps.

[0031] S1: Preprocess the real-time video stream and determine the preprocessed video frame.

[0032] S2: Input the preprocessed video frame into the target detection model that introduces the attention mechanism module to determine the preliminary target detection frame; the attention mechanism module includes a channel attention submodule and a spatial attention submodule.

[0033] S3: Optimize the preliminary detection frame using the modified loss function to obtain the optimized target detection frame; the modified loss function is: Among them, CIOU′ is the complete intersection-over-union ratio, which is a loss function that comprehensively considers the intersection-over-union ratio, center point distance and aspect ratio; IOU is the intersection-over-union ratio; l is the weight of the aspect ratio consistency loss; c is the Euclidean distance between the center point of the predicted box and the true box; d is the possible maximum distance between the center point of the predicted box and the true box; μ is the weight adjustment factor; a is the adjustment coefficient; v is the coefficient of the aspect ratio consistency loss; av is the value of the aspect ratio consistency loss.

[0034] S4: Based on the optimized target detection frame, track and predict the target in the real-time video stream to determine the tracking result; the tracking result is a processed video frame image; the tracking result includes the target, the optimized target detection frame corresponding to the target, the target unique code (Identity document, ID), location coordinates, width and height.

[0035] In an exemplary embodiment, before S1, the process further includes: inputting a video stream.

[0036] Input: Get real-time video stream from traffic monitoring camera. The video stream format is H.264 or H.265 encoding, supporting common video container formats such as MP4, AVI, etc. The frame rate and resolution of the video stream are determined according to the requirements of the specific monitoring system, usually 30fps and 1080p (1920x1080).

[0037] Source: The video stream comes from fixed cameras, mobile cameras, drones and other devices, ensuring the stability and continuity of the video stream.

[0038] In an exemplary embodiment, S1 may be replaced by the following steps.

[0039] S11: performing frequency adjustment, resolution adjustment, normalization processing, denoising processing and data enhancement on the real-time video stream to determine a pre-processed video frame.

[0040] In practical applications, the data preprocessing process is as follows:

[0041] Input: The original video stream obtained.

[0042] deal with:

[0043] Frame rate adjustment: Adjust the frame rate of the video stream from 30fps to 15fps to reduce the amount of calculation and increase the processing speed.

[0044] Resolution adjustment: The resolution of the video stream is adjusted from 1080p (1920x1080) to 720p (1280x720) to adapt to the input requirements of the model.

[0045] Normalization: Normalize the video frames and adjust the pixel values ​​from the range of [0, 255] to the range of [0, 1] to improve the training effect of the model.

[0046] Denoising: Denoise the video frames to reduce the noise caused by environmental factors (such as rain, fog, and low light) and improve the video quality.

[0047] Data augmentation: Perform data augmentation operations on video frames, including random cropping, rotation, flipping, etc., to improve the generalization ability of the model.

[0048] Output: preprocessed video frames for model training and prediction.

[0049] In an exemplary embodiment, S2 may be replaced by the following steps.

[0050] S21: Construct the target detection model; the target detection model includes a feature extraction module, an attention mechanism module, a feature enhancement module and a detection box output module connected in sequence.

[0051] S22: Utilize the feature extraction module to extract a feature map of the preprocessed video frame.

[0052] S23: Use the attention mechanism module to perform channel attention enhancement and spatial attention enhancement on the feature map in sequence to determine a feature map enhanced with channel attention and spatial attention.

[0053] S24: Use the feature enhancement module to perform feature enhancement processing on the feature map enhanced by the channel attention and spatial attention, and determine the enhanced detection box and category confidence.

[0054] S25: Utilize the detection frame output module to format the enhanced detection frame and category confidence, and output a preliminary target detection frame.

[0055] In practical applications, target detection models such as Figure 2 As shown, for Figure 2 For any cuboid in the feature extraction module, each value corresponds to length, width and height respectively; where length is the width of the feature map, indicating the resolution of the feature map in the horizontal direction; width is the height of the feature map, indicating the resolution of the feature map in the vertical direction; height is the depth of the feature map, indicating the number of channels of the feature map. These values ​​are only example values ​​of this application, and other values ​​can also be selected according to actual conditions, all of which are within the protection scope of this application.

[0056] The processing process is as follows:

[0057] Input: preprocessed video frames.

[0058] Processing and Optimization:

[0059] 1. YOLO detector

[0060] Description: YOLO (You Only Look Once) is a single-stage target detection algorithm known for its fast detection speed and small training cost. The YOLO detector regards the target detection task as a regression problem, directly mapping from image pixels to bounding box coordinates and category probabilities.

[0061] Input: preprocessed video frames.

[0062] deal with:

[0063] Feature extraction module: Use convolutional neural networks (such as Darknet-53) to extract feature maps of video frames. Darknet-53 is a deep convolutional neural network that extracts rich feature information through multiple convolutional layers and residual blocks.

[0064] Prediction: Perform convolution operations on the feature map to generate bounding boxes and category confidence. The YOLO detector maps the feature map to the coordinates (x, y, w, h) of the detection box and the category confidence through a fully connected layer, where (x, y) is the horizontal and vertical coordinates of the target detection box, and (w, h) is the width and height of the target detection box.

[0065] Output: preliminary detection box and category confidence.

[0066] 2.CBAM module

[0067] Description: CBAM (Convolutional Block Attention Module) is an attention mechanism module, which includes two sub-modules: Channel Attention and Spatial Attention. The CBAM module learns the importance weights of different parts of the input data and dynamically adjusts the processing intensity of these parts, thereby improving the model's ability to capture key features.

[0068] Input: Feature map output by the YOLO detector.

[0069] deal with:

[0070] Channel attention submodule: Determine the importance of each channel by weighting its features. The specific formula is: ChannelAttention = σ(MLP(MaxPool(F))+MLP(AvgPool(F))), where F is the input feature map, σ is the sigmoid function, MLP is the multi-layer perceptron, MaxPool and AvgPool are the maximum pooling and average pooling operations respectively.

[0071] Spatial attention submodule: Determine the importance of features at each position in the image by weighting them. The specific formula is: Spatial Attention = σ(Conv(MaxPool(Fc))+Conv(AvgPool(Fc))), where Fc is the feature map after channel attention processing and Conv is the convolution operation.

[0072] The CBAM module first uses the channel attention mechanism to weight different channels in the feature map and strengthen important features; then, it uses the spatial attention mechanism to emphasize the key spatial positions in the feature map, so that the model can focus more on the key information in the image. This combination of dual attention mechanisms enables CBAM to adaptively adjust feature maps to better adapt to different detection tasks.

[0073] Output: Feature map enhanced by channel attention and spatial attention.

[0074] 3. Feature Enhancement Module

[0075] Description: This module uses the feature map enhanced by the CBAM module for target detection. Through the dual attention mechanism, the model can focus more on the key information in the image, improving the quality of feature representation and model performance.

[0076] Input: Enhanced feature map output by the CBAM module.

[0077] Processing: The enhanced feature map is fed into subsequent layers of the detector for more accurate bounding box predictions. These subsequent layers may include more convolutional layers and upsampling layers to further refine the feature map.

[0078] Output: Enhanced detection box and category confidence.

[0079] 4. Detection box output module

[0080] Description: This module outputs preliminary detection boxes and category confidences, which will be used for subsequent loss function calculation and optimization.

[0081] Input: Enhanced detection boxes and category confidences.

[0082] Processing: Format the detection boxes and category confidences into the format required for subsequent processing.

[0083] Output: formatted detection box and category confidence.

[0084] 5. Loss function calculation

[0085] Description: This node calculates the loss value of the detection box, using the modified loss function (Complete Intersection over Union, CIOU) loss function, which comprehensively considers the differences in IoU, center point distance and aspect ratio, to more accurately evaluate the degree of match between the predicted box and the real box.

[0086] Input: formatted detection box and category confidence, as well as the groundtruth box.

[0087] deal with:

[0088] In the YOLO series of algorithms, the loss function is mainly composed of three parts: confidence loss, classification loss, and positioning loss. In the most important positioning loss part, CIOU is used to measure the loss value of the detection box. The formula is as follows:

[0089]

[0090] in,

[0091] CIOU: Complete intersection over union (IoU), is a loss function that takes into account IoU, center point distance, and aspect ratio, and is used to more accurately evaluate the degree of match between the predicted box and the true box.

[0092] IOU: Intersection over Union, which is the ratio of the intersection area of ​​the predicted box and the true box to the union area, is used to measure the degree of overlap between the two boxes.

[0093] l: represents the weight of the aspect ratio consistency loss, which is used to balance the losses of different parts.

[0094] av: represents the value of aspect ratio consistency loss, which is used to measure the difference in aspect ratio between the predicted box and the true box.

[0095] c: represents the Euclidean distance between the center points of the predicted box and the true box.

[0096] d: represents the maximum possible distance between the center points of the predicted box and the true box, that is, the diagonal length of the minimum circumscribed rectangle containing the two boxes.

[0097] a: represents the adjustment coefficient.

[0098] v: Coefficient representing the aspect ratio consistency loss.

[0099] α: represents the adjustment coefficient of the aspect ratio consistency loss, which is used to dynamically adjust the weight of the aspect ratio consistency loss according to the IoU value.

[0100] Specifically, the CIOU loss function can be expressed as IOU minus the ratio of the center point distance to the diagonal length of the minimum circumscribed rectangle, minus a correction term related to IOU, which is used to adjust the effect of the aspect ratio (mainly to eliminate the difference in the shape and direction of the target box). This enables the model to show better detection performance in practical applications, especially when the target is occluded or the shape changes greatly.

[0101] However, for small targets, the size of the detection box itself and the offset from the actual box will be smaller than those of regular targets, so the difference in the center distance parameter will be magnified, and the width and height similarity adjustment parameters can provide more information for the positioning difference of small targets. Therefore, in the positioning loss part, the two adjustment parameters are weighted to increase the applicability of the algorithm to small targets. The revised formula is as follows:

[0102]

[0103] Where μ is the weight adjustment factor, and its value range is [0,1]. The relationship between μ and AP50 is as follows: Figure 3 As shown, based on the experimental results, the final value of the weight adjustment factor μ in this application is 0.45.

[0104] 6. Optimize and adjust the module

[0105] Description: According to the calculated loss value, the model parameters are optimized and adjusted to improve the detection accuracy of the model.

[0106] Input: The calculated loss value.

[0107] Processing: Use the back propagation algorithm and optimizer (SGD, Adam) to update the model parameters to minimize the loss value. The specific steps include:

[0108] Calculate gradient: Calculate the gradient of the loss function with respect to the model parameters through the back-propagation algorithm.

[0109] Update parameters: Use an optimizer (SGD, Adam) to update the model parameters according to the gradient to minimize the loss value.

[0110] Output: optimized model parameters.

[0111] 7. Output prediction results

[0112] Description: Output the final prediction result, including detection box, category confidence, target ID and other information.

[0113] Input: Optimized detection boxes and category confidences.

[0114] Processing: Format the detection results into the format required by the system, including the target ID, location coordinates (x, y) and width.

[0115] Output: target detection result, that is, the optimized target detection frame, including the target ID, position coordinates (x, y) and width.

[0116] In an exemplary embodiment, S23 may be replaced by the following steps.

[0117] S231: Based on the channel attention submodule, use Channel Attention = σ(MLP(MaxPool(F))+MLP(AvgPool(F))) to perform channel attention enhancement on the feature map, and determine the feature map after channel attention enhancement; wherein σ is a sigmoid function; F is a feature map; MLP is a multi-layer perceptron, MaxPool() is a maximum pooling operation; AvgPool() is an average pooling operation;

[0118] S232: Based on the spatial attention submodule, it is used to perform spatial attention enhancement on the feature map after the channel attention enhancement by using Spatial Attention=σ(Conv(MaxPool(Fc))+Conv(AvgPool(Fc))), and determine the feature map enhanced by channel attention and spatial attention; wherein Fc is the feature map after the channel attention enhancement; Conv() is a convolution operation.

[0119] In an exemplary embodiment, S3 further includes:

[0120] S5: Based on the loss value of the optimized target detection frame, the target detection model is optimized and adjusted using a back propagation algorithm and an optimizer.

[0121] In an exemplary embodiment, S4 may be replaced by the following steps.

[0122] S41: Based on the optimized target detection frame, determine a state equation according to a target state vector in the real-time video stream; the target state vector includes a position vector, a velocity vector and an acceleration vector.

[0123] S42: Determine a state transfer matrix and an observation matrix according to the state equation.

[0124] S43: Determine a state vector expression and an observation expression according to the state transfer matrix and the observation matrix, construct an acceleration and deceleration Kalman filter model, and output a video frame image processed by the acceleration and deceleration Kalman filter model.

[0125] S44: reconstructing the video frame images processed by the acceleration / deceleration Kalman filter model into a frame sequence in chronological order to synthesize a video sequence.

[0126] S45: Using a graphic interface tool to perform target labeling on each video frame in the video sequence, and determining a labeled target detection frame.

[0127] S46: Encode the video sequence containing the marked target detection frame to determine a video file; the video file includes the tracking result.

[0128] S47: Store and display the video file and the real-time video stream.

[0129] In practical applications, tracking prediction is achieved through the acceleration and deceleration Kalman green wave model, that is, the uniform acceleration Kalman filter.

[0130] Input: Optimized object detection box.

[0131] Processing and Optimization:

[0132] The Kalman filter in DeepSort uses a linear filter, and the state vector has only eight dimensions, that is, it only has the position and direction of the movement speed, but no acceleration information. This will reduce the prediction effect of the target suddenly stopping or accelerating or decelerating in actual situations, resulting in ID loss. Therefore, this application introduces the acceleration state vector to predict the movement state of the object.

[0133] First, in order to establish the acceleration and deceleration Kalman filter model, the state vector is split first.

[0134] p=(x,y,r,h) T ,

[0135] Among them, p is the position vector, v is the velocity vector, γ is the acceleration vector, and r is the width of the object; is the velocity of the object in the x direction; is the velocity of the object in the y direction; is the rate of change of the object width; is the rate of change of the object's height; is the acceleration of the object in the x direction; is the acceleration of the object in the y direction; is the acceleration rate of change of the object width; is the rate of change of acceleration of the object's height.

[0136] The interval between two frames is T, and the state equation can be derived from the physical speed and distance formula as follows:

[0137] p t =p t-1 +v t-1 T+0.5γT 2

[0138] v t =v t-1 +γT

[0139] γ t =γ t-1

[0140] Thus, we can get the new state transfer matrix A,

[0141] The observation matrix H is [1, 0, 0], so the observation value Z at time t is t It can be expressed as:

[0142] Among them, p t is the position vector at time t, v t is the velocity vector at time t, γ t is the acceleration vector at time t.

[0143] Based on the above derivation, we can get the state vector expression and observation expression of the acceleration and deceleration Kalman filter:

[0144]

[0145] Among them, ω t-1 is the process noise vector, which represents the uncertainty or error of the system model. It is usually assumed to be zero-mean Gaussian white noise, which is used to describe the deviation caused by external interference or inaccurate model during the state transfer process. t is the observation noise vector, which represents the error in the observation process and is usually assumed to be zero-mean Gaussian white noise, used to describe the deviation between the observed value and the true state.

[0146] After updating Kalman's acceleration, deceleration and constant speed models, the tracking system will adapt to the acceleration, deceleration and constant speed tracking targets, thereby achieving higher accuracy.

[0147] Output: Tracking trajectory of the target (feature file + final framed video frame).

[0148] In practical applications, outputting tracking results specifically includes the following steps.

[0149] Input: Video frame image obtained by target detection algorithm (Attn-YOLO) with attention mechanism + uniformly accelerated Kalman filter processing.

[0150] Processing process:

[0151] Frame sequence reconstruction: Reassemble the processed video frames into a video sequence in time order. Ensure that the order and timestamp of each frame are correct to maintain the smoothness and continuity of the video.

[0152] Target annotation: In each frame of video, a graphical interface tool (OpenCV) is used to draw a detection box around the detected target, and the target ID and confidence information are displayed near the detection box. The color and style of the detection box can be distinguished according to the type or state of the target.

[0153] Video encoding: Encode the reconstructed frame sequence into a video file. Select the appropriate video encoding format (H.264) and container format (MP4) to ensure the compatibility and playback quality of the video file.

[0154] Result storage: Store the generated video files in the specified storage medium (hard disk or network storage device) for subsequent playback and analysis.

[0155] Real-time display: The processed video is displayed in real time on the monitoring system interface, supporting the simultaneous display of the original video stream and the processed video stream, making it convenient for monitoring personnel to compare and analyze.

[0156] Output: processed video file and real-time video stream. The video contains the detected target and its corresponding detection box and confidence information. The target ID, position coordinates (x, y), width and height (w, h) and other information are clearly marked in the video frame, such as Figure 4 As shown, Figure 4 (a)-(f) in the figure are all intercepted target tracking experiment effect pictures. Figure 4 The numerical values ​​in (a)-(f) are the target IDs.

[0157] 1. More suitable for complex scenarios.

[0158] Aiming at the complex scene of traffic monitoring, this scheme successfully built a YOLO detector (Attn-YOLO) that integrates the attention mechanism, and enhanced the model's ability to capture key features through the CBAM module. The CBAM module contains two sub-modules: channel attention and spatial attention, which weight the channels and spatial positions in the feature map respectively, highlighting important features and suppressing unimportant information. This dual attention mechanism enables Attn-YOLO to adaptively adjust the feature map to better adapt to different detection tasks, especially in video surveillance. The detection ability of small targets and occluded targets has been significantly improved, thereby improving the accuracy and robustness of target detection.

[0159] 2. Small target positioning accuracy is higher.

[0160] In the YOLO series of algorithms, this solution proposes a parameter optimization strategy to address the challenges of small target detection. By adjusting the CIOU loss function and introducing the weight adjustment factor μ, the localization loss calculation of small targets is optimized. The CIOU loss function takes into account the overlapping area, center point distance, and aspect ratio difference, and the adjustment in this study further enhances the localization accuracy of small targets. The μ value determined experimentally enables the model to more accurately predict the position of small targets, reducing the detection error caused by the small size of the target, thereby significantly improving the detection accuracy of small targets in scenarios such as traffic monitoring.

[0161] 3. Stronger prediction capabilities, closer to reality.

[0162] The improved Kalman filter can predict the acceleration changes of the target, making the tracking model closer to the actual situation and adapting to the acceleration and deceleration changes of the target, thereby achieving more stable and accurate target tracking in a dynamic and complex traffic environment.

[0163] This application optimizes the video target tracking algorithm to improve its performance in the field of traffic monitoring, especially for pedestrian tracking. The core of this application is to introduce an attention mechanism and optimize the loss function to enhance the algorithm's ability to detect small-target pedestrians. The attention mechanism can identify and emphasize the most important parts of the feature map, thereby enhancing the expressive power of the feature map and obtaining more feature information. The modification of the loss function makes the model more suitable for small-target scenarios. These optimizations are intended to address the shortcomings of existing algorithms in dealing with small targets, complex backgrounds, and real-time requirements, in order to achieve more efficient and accurate video target tracking in actual traffic monitoring applications.

[0164] Through these optimization measures, this application is expected to significantly improve the performance and reliability of the multi-target tracking system in the field of traffic monitoring, especially in pedestrian tracking. This will not only improve the intelligence level of the traffic monitoring system, but also provide more accurate and real-time data support for traffic flow analysis, accident prevention and emergency response, thereby contributing to building a safer and more efficient urban traffic environment.

[0165] Based on the same inventive concept, the embodiment of the present application also provides a device for tracking pedestrian targets in complex monitoring scenarios for implementing the method for tracking pedestrian targets in complex monitoring scenarios involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of one or more devices for tracking pedestrian targets in complex monitoring scenarios provided below can refer to the limitations of the method for tracking pedestrian targets in complex monitoring scenarios above, and will not be repeated here.

[0166] In an exemplary embodiment, a pedestrian target tracking device in a complex monitoring scenario is provided, comprising:

[0167] The preprocessing module is used to preprocess the real-time video stream and determine the preprocessed video frames.

[0168] The module for determining a preliminary target detection frame is used to input the preprocessed video frame into a target detection model that introduces an attention mechanism module to determine a preliminary target detection frame; the attention mechanism module includes a channel attention submodule and a spatial attention submodule.

[0169] The optimization module is used to optimize the preliminary detection frame using the modified loss function to obtain the optimized target detection frame; the modified loss function is Among them, CIOU′ is the complete intersection-over-union ratio, which is a loss function that comprehensively considers the intersection-over-union ratio, center point distance and aspect ratio; IOU is the intersection-over-union ratio; l is the weight of the aspect ratio consistency loss; c is the Euclidean distance between the center point of the predicted box and the true box; d is the possible maximum distance between the center point of the predicted box and the true box; μ is the weight adjustment factor; a is the adjustment coefficient; v is the coefficient of the aspect ratio consistency loss; av is the value of the aspect ratio consistency loss.

[0170] A tracking result determination module is used to track and predict the target in the real-time video stream based on the optimized target detection frame to determine the tracking result; the tracking result is a processed video frame image; the tracking result includes the target, the optimized target detection frame corresponding to the target, the target ID, the position coordinates, the width and the height.

[0171] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store pedestrian target tracking data in complex monitoring scenarios. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a pedestrian target tracking method in a complex monitoring scenario is implemented.

[0172] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the above method is implemented when the processor executes the computer program.

[0173] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the above method when executed by a processor.

[0174] In an exemplary embodiment, a computer program product is provided, including a computer program, which implements the above method when executed by a processor.

[0175] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ReadOnlyMemory, ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (Magnetoresistive RandomAccess Memory, MRAM), ferroelectric random access memory (Ferroelectric RandomAccess Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (RandomAccess Memory, RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0176] In this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0177] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0178] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0179] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A pedestrian target tracking method in a complex monitoring scene, characterized in that: The pedestrian target tracking method in the complex monitoring scene includes: Preprocessing the real-time video stream and determining the preprocessed video frame; Inputting the preprocessed video frame into a target detection model that introduces an attention mechanism module to determine a preliminary target detection frame; the attention mechanism module includes a channel attention submodule and a spatial attention submodule; The modified loss function is used to optimize the preliminary detection frame to obtain the optimized target detection frame; the modified loss function is Among them, CIOU′ is the complete intersection-over-union ratio, which is a loss function that comprehensively considers the intersection-over-union ratio, center point distance and aspect ratio; IOU is the intersection-over-union ratio; l is the weight of the aspect ratio consistency loss; c is the Euclidean distance between the center point of the predicted box and the real box; d is the possible maximum distance between the center point of the predicted box and the real box; μ is the weight adjustment factor; a is the adjustment coefficient; v is the coefficient of the aspect ratio consistency loss; av is the value of the aspect ratio consistency loss; Based on the optimized target detection frame, the target in the real-time video stream is tracked and predicted to determine the tracking result; the tracking result is a processed video frame image; the tracking result includes the target, the optimized target detection frame corresponding to the target, the target ID, the position coordinates, the width and the height.

2. The pedestrian target tracking method in a complex monitoring scene according to claim 1 is characterized in that: Preprocess the real-time video stream and determine the preprocessed video frame, specifically including: The real-time video stream is subjected to frequency adjustment, resolution adjustment, normalization processing, denoising processing and data enhancement to determine a pre-processed video frame.

3. The pedestrian target tracking method in a complex monitoring scene according to claim 1 is characterized in that: The preprocessed video frame is input into the target detection model that introduces the attention mechanism module to determine the preliminary target detection frame, specifically including: Constructing the target detection model; the target detection model includes a feature extraction module, an attention mechanism module, a feature enhancement module and a detection frame output module connected in sequence; Extracting a feature map of the preprocessed video frame using the feature extraction module; Using the attention mechanism module to sequentially perform channel attention enhancement and spatial attention enhancement on the feature map, and determine a feature map enhanced with channel attention and spatial attention; Using the feature enhancement module to perform feature enhancement processing on the channel attention and spatial attention enhanced feature maps, and determine the enhanced detection box and category confidence; The enhanced detection frame and category confidence are formatted using the detection frame output module to output a preliminary target detection frame.

4. The pedestrian target tracking method in a complex monitoring scene according to claim 3 is characterized in that: The attention mechanism module is used to sequentially perform channel attention enhancement and spatial attention enhancement on the feature map to determine the feature map enhanced with channel attention and spatial attention, specifically including: Based on the channel attention submodule, it is used to perform channel attention enhancement on the feature map by using Channel Attention=σ(MLP(MaxPool(F))+MLP(AvgPool(F))) to determine the feature map after channel attention enhancement; wherein σ is a sigmoid function; F is a feature map; MLP is a multi-layer perceptron, MaxPool() is a maximum pooling operation; AvgPool() is an average pooling operation; Based on the spatial attention submodule, it is used to perform spatial attention enhancement on the feature map after the channel attention enhancement by using Spatial Attention=σ(Conv(MaxPool(Fc))+Conv(AvgPool(Fc))) to determine the feature map of channel attention and spatial attention enhancement; wherein Fc is the feature map after channel attention enhancement; Conv() is a convolution operation.

5. The pedestrian target tracking method in a complex monitoring scene according to claim 1 is characterized in that: The preliminary detection frame is optimized using the modified loss function, and the optimized target detection frame further includes: Based on the loss value of the optimized target detection frame, the target detection model is optimized and adjusted using a back propagation algorithm and an optimizer.

6. The pedestrian target tracking method in a complex monitoring scene according to claim 1 is characterized in that: Based on the optimized target detection frame, tracking and predicting the target in the real-time video stream to determine the tracking result specifically includes: Based on the optimized target detection frame, a state equation is determined according to a target state vector in the real-time video stream; the target state vector includes a position vector, a velocity vector and an acceleration vector; Determine a state transfer matrix and an observation matrix according to the state equation; Determine a state vector expression and an observation expression according to the state transfer matrix and the observation matrix, construct an acceleration and deceleration Kalman filter model, and output a video frame image processed by the acceleration and deceleration Kalman filter model; Reconstructing the video frame images processed by the acceleration / deceleration Kalman filter model into a frame sequence in chronological order to synthesize a video sequence; Using a graphical interface tool to mark each video frame in the video sequence, and determining a marked target detection frame; Encoding a video sequence containing the marked target detection frame to determine a video file; the video file includes a tracking result; The video file and the real-time video stream are stored and displayed.

7. A pedestrian target tracking device in a complex monitoring scene, characterized in that: The pedestrian target tracking device in the complex monitoring scene includes: A preprocessing module, used to preprocess the real-time video stream and determine the preprocessed video frame; A preliminary target detection frame determination module, used to input the preprocessed video frame into the target detection model introduced into the attention mechanism module to determine the preliminary target detection frame; the attention mechanism module includes a channel attention submodule and a spatial attention submodule; The optimization module is used to optimize the preliminary detection frame using the modified loss function to obtain the optimized target detection frame; the modified loss function is Among them, CIOU′ is the complete intersection-over-union ratio, which is a loss function that comprehensively considers the intersection-over-union ratio, center point distance and aspect ratio; IOU is the intersection-over-union ratio; l is the weight of the aspect ratio consistency loss; c is the Euclidean distance between the center point of the predicted box and the real box; d is the possible maximum distance between the center point of the predicted box and the real box; μ is the weight adjustment factor; a is the adjustment coefficient; v is the coefficient of the aspect ratio consistency loss; av is the value of the aspect ratio consistency loss; A tracking result determination module is used to track and predict the target in the real-time video stream based on the optimized target detection frame to determine the tracking result; the tracking result is a processed video frame image; the tracking result includes the target, the optimized target detection frame corresponding to the target, the target ID, the position coordinates, the width and the height.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the pedestrian target tracking method in a complex monitoring scenario described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the pedestrian target tracking method in a complex monitoring scenario described in any one of claims 1-6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the pedestrian target tracking method in a complex monitoring scenario described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Unmanned ship target detection tracking method and system

    CN114596335A

  • Multi-target detection and tracking method based on improved YOLO-V5s

    CN114882351A

  • Target detection and tracking system based on improved YOLOv7 and DeepSORT

    CN117423031A

  • Substation peripheral potential safety hazard identification method and system based on target detection

    CN118397545A