Tracking method for passengers in car type elevator based on improved StrongSORT algorithm

Through the improved StrongSORT algorithm, combined with motion compensation and efficient object detection model, the problems of occlusion and fast movement in occupant tracking in van elevators are solved, and stable and accurate occupant tracking is achieved, suitable for real-time tracking tasks with limited resources.

CN120375013APending Publication Date: 2025-07-25SPECIAL EQUIP SAFETY SUPERVISION INSPECTION INST OF JIANGSU PROVINCE
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510456129.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing occupant tracking method in van elevators is difficult to achieve stable and accurate target tracking when target occlusion, rapid movement or similar targets appear side by side, resulting in a decrease in tracking accuracy.

Method used

The improved StrongSORT algorithm is used, combined with the enhanced correlation coefficient maximization algorithm for motion compensation, the YOLOv8-MPDIoU+C2fGhostV2 model is used for object detection, and appearance features are extracted through the Kalman filter and the ResNet50 model, and tracking trajectory management is combined with the exponential moving average strategy to optimize the robustness and accuracy of the tracking algorithm.

Benefits of technology

Real-time and high-precision occupant tracking is realized in complex elevator environments, improving detection accuracy and speed, enhancing the applicability and robustness of the algorithm, and suitable for real-time tracking tasks with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375013A_ABST
    Figure CN120375013A_ABST
Patent Text Reader

Abstract

The invention discloses a method for tracking passengers in a car elevator based on an improved StrongSORT algorithm, and relates to the technical field of elevator safety management, and the method comprises the steps: obtaining a monitoring video stream in the car elevator, and obtaining a current frame image after motion compensation based on an enhanced correlation coefficient maximization algorithm; processing the compensated current frame image by using a target detection model to obtain bounding box coordinates, categories and confidence of each passenger target, and taking the bounding box coordinates, categories and confidence as detection results; based on a detection result, tracking a passenger target of each frame of image in the monitoring video stream by using a StrongSORT algorithm; generating a tracking trajectory, and performing tracking trajectory state management; and outputting the final tracking trajectory and trajectory ID of the passenger target, thereby realizing stable and accurate target tracking in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of elevator safety management, and particularly to an in-car passenger tracking method for a van elevator based on an improved StrongSORT algorithm. Background Art

[0002] Although van elevators bring great convenience to people's lives and work, there are also some potential safety hazards, such as abnormal situations like overloading of passengers, violent behavior, and sudden illnesses. Traditional elevator safety monitoring mainly relies on manual on-duty, which has defects such as high work intensity and low monitoring efficiency. Therefore, there is an urgent need for an intelligent multi-object tracking method for elevator safety monitoring that can track and analyze the status and behavior of passengers in the elevator car in real time and accurately.

[0003] The object tracking task generally consists of two major parts: a detector and a tracker. The detector is responsible for locating the bounding boxes of objects in the video and passing this information to the tracker. Based on these detection results, the tracker analyzes the motion and appearance features of the objects, identifies the same object in consecutive video frames, and assigns a unique trajectory identifier to it, thereby achieving the locking of the object's motion. Currently, multi-object tracking algorithms have been widely applied in the field of video surveillance, and representative algorithms include the SORT algorithm and the DeepSORT algorithm. The object tracking task generally consists of two major parts: a detector and a tracker. The detector is responsible for locating the bounding boxes of objects in the video and passing this information to the tracker. Based on these detection results, the tracker analyzes the motion and appearance features of the objects, identifies the same object in consecutive video frames, and assigns a unique trajectory identifier to it, thereby achieving the locking of the object's motion. The key steps of this tracking mode include: First, perform object detection, send the video frame sequence into the detection model to obtain the bounding boxes of the objects; Second, generate trajectories, use the detection results of the first frame to initialize the tracking trajectories; Then, perform association matching, extract the features of the objects in the current frame, and match them with the tracking trajectories of the previous frame; Finally, update the trajectories, assign a trajectory ID to the successfully matched detection results, and update the status of the trajectories. However, due to factors such as the small internal space of the van elevator, large changes in lighting, dense personnel, and frequent occlusions, directly applying these general multi-object tracking algorithms to the elevator scenario often fails to achieve ideal results.

[0004] Existing in-car passenger tracking methods for van elevators will encounter problems of lost or incorrect object tracking when dealing with object occlusion, fast object movement, or the appearance of similar objects side by side, which affects the accuracy of tracking. Summary of the Invention

[0005] This application aims to solve at least one of the technical problems in the related art to some extent. For this purpose, an object of this application is to propose an occupant tracking method in a van elevator based on an improved StrongSORT algorithm, which can effectively handle problems such as target occlusion, fast movement, and high similarity between targets, so as to achieve stable and accurate target tracking in complex scenarios.

[0006] One aspect of this application provides an occupant tracking method in a van elevator based on an improved StrongSORT algorithm, including:

[0007] Step S100: Obtain the monitoring video stream in the van elevator, and obtain the current frame image after motion compensation based on the enhanced correlation coefficient maximization algorithm;

[0008] Step S200: Process the compensated current frame image using an object detection model to obtain the bounding box coordinates, category, and confidence of each occupant target as the detection result;

[0009] Step S300: Based on the detection result, use the StrongSORT algorithm to track the occupant targets in each frame image of the monitoring video stream;

[0010] Step S400: Generate tracking trajectories and perform tracking trajectory status management;

[0011] Step S500: Output the tracking trajectories and trajectory IDs of the final occupant targets.

[0012] The specific method for obtaining the monitoring video stream in the van elevator and obtaining the current frame image after motion compensation based on the enhanced correlation coefficient maximization algorithm is as follows:

[0013] Step S110: Obtain the monitoring video stream in the van elevator, determine the reference frame image, use the current frame image as the distorted frame image, perform zero-mean processing on the reference frame image and the distorted frame image to obtain the zero-mean reference frame and the zero-mean distorted frame, and calculate the maximized enhanced correlation coefficient;

[0014] Optionally, the reference frame image is the previous frame image or the first frame image of the current frame image;

[0015] Step S120: Minimize the maximized enhanced correlation coefficient ECC(p), optimize the distortion parameter p, and perform motion compensation on the distorted frame image based on the distortion parameter;

[0016] Step S130: Output the current frame image after motion compensation and the distortion parameter p;

[0017] The specific method for processing the compensated current frame image using an object detection model to obtain the bounding box coordinates, category, and confidence of each occupant target as the detection result is:

[0018] Step S210: Construct a target detection model;

[0019] Step S220: Define the input of the target detection model as the frame image after motion compensation, and the output as the bounding box coordinates, bounding box categories, and confidence levels of the occupant targets in each frame image after compensation;

[0020] Step S230: Use the output bounding box coordinates, bounding box categories, and confidence levels of the occupant targets as the detection results of the current frame image;

[0021] The target detection model selects the YOLOv8-MPDIoU+C2fGhostV2 model, takes the image after motion compensation as the input data, and the output is the bounding box coordinates, class probabilities, and confidence levels of each occupant target in the image after motion compensation, where the class corresponding to the maximum class probability is the bounding box category;

[0022] The YOLOv8-MPDIoU+C2fGhostV2 model is improved based on the YOLOv8 architecture. Its core modules include a backbone network, a feature fusion layer, and a detection head. The MPDIoU loss function is used to optimize the bounding box regression accuracy. The C2f module is replaced with the C2fGhostV2 module in the backbone network and the feature fusion layer. The MOT17 pedestrian detection dataset with annotated pedestrian bounding boxes is used as the training set to train the target detection model, and minimizing the MPDIoU loss function is used as the training objective.

[0023] The specific method for tracking the occupant targets using the StrongSORT algorithm based on the detection results is as follows:

[0024] Step S310: For the detection results of the first frame image, initialize each detected occupant target as a tracking trajectory in an unconfirmed state and assign it a trajectory ID;

[0025] Step S320: For each subsequent frame image, use the detection results of the current frame and the tracking trajectories of the previous frame as the input, and use the Kalman filter to predict the next tracking box corresponding to the tracking trajectory of the previous frame;

[0026] Step S330: For each predicted tracking box, calculate the motion feature similarity between it and the bounding box in the detection results using the Mahalanobis distance;

[0027] Step S340: Extract the appearance features of the predicted tracking box and the bounding box in the detection results, and calculate the minimum cosine distance between the appearance features of the bounding box and the predicted tracking box as the appearance feature similarity between the predicted tracking box and the bounding box;

[0028] Step S350: Based on the motion feature similarity and appearance feature similarity, construct a cost matrix, solve the cost matrix through the Hungarian algorithm, find the bounding box in the detection results that best matches the tracking box, and update the tracking trajectory. For the successfully matched tracking trajectories, use the EMA method to update the appearance features;

[0029] Step S360: For the detection results that are not successfully matched, initialize the occupant targets therein as new tracking trajectories;

[0030] The Kalman filter adopts the NSA Kalman filter algorithm in GIAOTracker, and calculates the detection noise covariance based on the confidence of the detection results and a preset fixed value of the detection noise covariance;

[0031] The specific method for using the Kalman filter to predict the next tracking box corresponding to the tracking trajectory of the previous frame with the detection results of the current frame and the tracking trajectory of the previous frame as inputs for each subsequent frame of image is as follows:

[0032] Step S321: Characterize the tracking box with motion features;

[0033] Step S322: Predict the predicted value of the tracking box at time k based on the tracking box at time k - 1;

[0034] Step S323: Update the Kalman filter with the bounding box in the detection results at time k, associate the predicted value of the tracking box at time k with the bounding box in the detection results at time k, and obtain the tracking box at time k;

[0035] Specifically, the appearance features of the predicted tracking box and the bounding box in the detection results are extracted using the feature extraction unit BoT. The feature extraction unit BoT uses the ResNet50 model as the backbone network and is pre-trained on the DukeMTMC-reID dataset; the ResNet50 model consists of convolutional layers, pooling layers, and residual blocks, and deep features are extracted between the residual blocks through the skip connection mechanism;

[0036] The specific method for generating the tracking trajectory and managing the tracking trajectory status is as follows:

[0037] Step S410: Preset a first matching threshold. When the number of consecutive successful matches of the newly generated tracking trajectory reaches the first matching threshold, it is considered an effective trajectory;

[0038] Step S420: Preset a second matching threshold. When the number of consecutive unsuccessful matches of the effective trajectory reaches the second matching threshold, the effective trajectory is deleted;

[0039] The method for tracking passengers in a van elevator based on the improved StrongSORT algorithm proposed in this application has the following advantages compared with the prior art:

[0040] The YOLOv8-MPDIoU+C2fGhostV2 model of this application, by introducing the MPDIoU loss function and the C2fGhostV2 module, significantly reduces the number of model parameters and computational overhead while improving the detection accuracy, and is very suitable for real-time tracking tasks with limited resources. Applying it to passenger detection significantly improves the accuracy and speed of passenger detection in van elevators, providing high-quality input for tracking and recognition.

[0041] The StrongSORT algorithm of this application can achieve real-time, long-term, and high-precision passenger tracking in the complex environment of van elevators. An adaptive NSA strategy is introduced to dynamically adjust the observation noise, and a lighter-weight OSNet is used as the ReID network, which significantly reduces the computational overhead while maintaining the feature extraction performance, improving the robustness and efficiency of the algorithm. The optimization for the elevator scenario further enhances the applicability of the algorithm, enabling it to meet the requirements of practical applications.

[0042] This application introduces ResNet50 into the ReID network. By using its deeper network structure and residual connection design, it can learn more refined and discriminative appearance features. Using the pre-trained ResNet50 as the base network can better adapt to different tracking environments, enabling the ResNet50 ReID network to provide better appearance features for the StrongSORT algorithm and comprehensively improving the tracking performance.

[0043] The StrongSORT algorithm of this application abandons the traditional feature library method for storing the appearance features of targets and instead uses an exponential moving average strategy to dynamically update the appearance state of passenger targets. Although the feature library can accumulate long-term feature information of multiple frames, it is more sensitive to noise in the detection process. In contrast, the exponential moving average strategy reduces the interference of noise by utilizing the changes in inter-frame features, thereby improving the accuracy of matching.

[0044] This application designs a systematic solution for the specific requirements of passenger tracking in van elevators. The organic integration and targeted optimization of each module improve the accuracy, real-time performance, and robustness of tracking. At the same time, the standardized algorithm process also makes this solution have good engineering application value and is convenient for deployment and promotion in actual systems. Description of the Drawings

[0045] Figure 1 It is the method flow chart of the method for tracking passengers in a van elevator based on the improved StrongSORT algorithm provided by this application;

[0046] Figure 2 Schematic diagram of the full-scale feature learning architecture of the OSNet_x1_0 model provided by this application;

[0047] Figure 3 Schematic diagram of the building block design of the OSNet architecture provided by this application;

[0048] Figure 4 Frame and performance comparison diagram between DeepSORT and StrongSORT provided by this application;

[0049] Figure 5 Scene diagram of the 125th frame of the test video sequence;

[0050] Figure 6 Scene diagram of the 182nd frame of the test video sequence;

[0051] Figure 7 Scene diagram of the 196th frame of the test video sequence;

[0052] Figure 8 Test effect diagram of the 2nd frame of the monitoring video stream of the van elevator;

[0053] Figure 9 Test effect diagram of the 36th frame of the monitoring video stream of the van elevator;

[0054] Figure 10 Test effect diagram of the 63rd frame of the monitoring video stream of the van elevator. Detailed implementation manners

[0055] To better understand this application, various aspects of this application will be described in more detail with reference to the accompanying drawings. It should be understood that these detailed descriptions are only descriptions of the exemplary embodiments of this application and do not limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.

[0056] In the accompanying drawings, for ease of illustration, the sizes, dimensions, and shapes of the elements have been slightly adjusted. The drawings are only examples and are not drawn to strict scale. As used herein, terms such as "substantially", "approximately", and similar terms are used as terms indicating approximation and not as terms indicating degree, and are intended to illustrate the inherent deviations in measured or calculated values that would be recognized by those of ordinary skill in the art. Additionally, in this application, the order of description of the various step processes does not necessarily represent the order in which these processes occur in actual operation, unless otherwise clearly defined or derivable from the context.

[0057] It should also be understood that expressions such as "including", "comprising", "having", "containing" and / or "comprising of" are open-ended rather than closed-ended expressions in this specification, which means the presence of the stated features, elements and / or components is indicated, but the presence of one or more other features, elements, components and / or combinations thereof is not excluded. In addition, when an expression such as "at least one of..." appears after a list of listed features, it modifies the entire list of features rather than just individual elements in the list. In addition, when describing embodiments of the present application, the use of "may" means "one or more embodiments of the present application". And the term "exemplary" is intended to refer to an example or illustration.

[0058] Unless otherwise defined, all terms used herein (including engineering terms and scientific and technical terms) have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs. It should also be understood that unless clearly stated in this application, words defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense.

[0059] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0060] Embodiment 1

[0061] As Figure 1 shown, the method for tracking occupants in a van elevator based on an improved StrongSORT algorithm provided by this application includes:

[0062] Step S100: Obtain the monitoring video stream in the van elevator, and obtain the current frame image after motion compensation based on the enhanced correlation coefficient maximization algorithm;

[0063] The monitoring video stream in the van elevator is composed of multiple consecutive frame images;

[0064] The enhanced correlation coefficient maximization (ECC) algorithm is used to compensate for the global camera motion between adjacent frames and reduce the tracking error caused by camera jitter or perspective change. As a parameterized image alignment technique, the ECC algorithm can accurately quantify the warping distortion in the image and estimate the global rotation and translation between adjacent frames.

[0065] The specific method for obtaining the monitoring video stream in the van elevator and obtaining the current frame image after motion compensation based on the enhanced correlation coefficient maximization algorithm is:

[0066] Step S110: Obtain the monitoring video stream inside the van elevator, determine the reference frame image, use the current frame image as the distorted frame image, perform zero-mean normalization on the reference frame image and the distorted frame image to obtain the zero-mean reference frame and the zero-mean distorted frame, and calculate the maximized enhanced correlation coefficient;

[0067] The calculation formula for the maximized enhanced correlation coefficient is: where ECC(p) is the maximized enhanced correlation coefficient, ||.|| is the Euclidean norm, p is the distortion parameter, is the zero-mean reference frame, is the zero-mean distorted frame, i r is the reference frame image, i w (p) is the distorted frame image;

[0068] Optionally, the reference frame image is the previous frame image or the first frame image of the current frame image;

[0069] Step S120: Minimize the maximized enhanced correlation coefficient ECC(p), optimize the distortion parameter p, and perform motion compensation on the distorted frame image based on the distortion parameter;

[0070] By minimizing the maximized enhanced correlation coefficient, the distortion parameter can be effectively optimized, thereby solving the image alignment problem between adjacent frames. This method can efficiently reduce the influence of motion noise caused by factors such as camera movement on video frames.

[0071] Step S130: Output the current frame image after motion compensation and the distortion parameter p;

[0072] Step S200: Process the current frame image after compensation using the object detection model to obtain the bounding box coordinates, category, and confidence of each occupant target as the detection result;

[0073] The specific method of processing the current frame image after compensation using the object detection model to obtain the bounding box coordinates, category, and confidence of each occupant target as the detection result is:

[0074] Step S210: Construct an object detection model;

[0075] Optionally, the object detection model selects the YOLOv8-MPDIoU+C2fGhostV2 model. Using the image after motion compensation as the input data, the output is the bounding box coordinates, category probability, and confidence of each occupant target in the image after motion compensation, where the category corresponding to the maximum value of the category probability is the bounding box category;

[0076] The category probability refers to the probability that each bounding box belongs to each category, and the category with the maximum probability value is used as the category corresponding to the bounding box;

[0077] The confidence level refers to the probability of the presence of the member object in the bounding box;

[0078] Specifically, the YOLOv8-MPDIoU+C2fGhostV2 model is improved based on the YOLOv8 architecture. Its core modules include the backbone network, the feature fusion layer, and the detection head. The MPDIoU loss function is used to optimize the bounding box regression accuracy. The C2f module is replaced by the C2fGhostV2 module in the backbone network and the feature fusion layer. The MOT17 pedestrian detection dataset with annotated pedestrian bounding boxes is used as the training set to train the object detection model, with minimizing the MPDIoU loss function as the training objective.

[0079] Step S220: Define the input of the object detection model as the motion-compensated frame image, and the output as the bounding box coordinates, bounding box categories, and confidence levels of the occupant objects in each compensated frame image;

[0080] Step S230: Take the output bounding box coordinates, bounding box categories, and confidence levels of the occupant objects as the detection results of the current frame image;

[0081] Step S300: Based on the detection results, use the StrongSORT algorithm to track the occupant objects in each frame image of the surveillance video stream;

[0082] The specific method of using the StrongSORT algorithm to track the occupant objects based on the detection results is as follows:

[0083] Step S310: For the detection results of the first frame image, initialize each detected occupant object as a tracking trajectory in an unconfirmed state and assign it a trajectory ID;

[0084] Step S320: For each subsequent frame image, take the detection results of the current frame and the tracking trajectories of the previous frame as inputs, and use the Kalman filter to predict the next tracking box corresponding to the tracking trajectory of the previous frame;

[0085] The Kalman filter adopts the NSA Kalman filter algorithm in GIAOTracker, and calculates the detection noise covariance based on the confidence level of the detection results and a preset fixed value of the detection noise covariance;

[0086] The calculation formula for the detection noise covariance is: R k =(1 - c k )R k ; where, R k represents the detection noise covariance, c k is the confidence level of the detection results, and R k is a preset fixed value of the detection noise covariance;

[0087] The Kalman filter enhances performance by adaptively calculating the detection noise covariance, enabling StrongSORT to more accurately adapt to changes in detection noise, thereby improving the stability and accuracy of tracking. According to the calculation formula of the detection noise covariance, a lower detection confidence corresponds to a higher detection noise covariance. Combining with the motion state update mechanism in DeepSORT, it can be learned that as the detection noise covariance increases, the weight of the detection result in the Kalman filter state update process will decrease accordingly. Such a mechanism effectively weakens the influence of low-quality detection results on state updates, ensuring an improvement in the accuracy of motion state updates.

[0088] For each subsequent frame of the image, the specific method of using the Kalman filter to predict the next tracking box corresponding to the tracking trajectory of the previous frame with the detection result of the current frame and the tracking trajectory of the previous frame as inputs is as follows:

[0089] Step S321: Characterize the tracking box with motion features;

[0090] The motion features are expressed as where x k is the tracking box at time k, (u, v) represents the center point coordinates of the tracking box, s and r are the aspect ratio and height of the tracking box respectively, and are the derivatives of the center point abscissa, center point ordinate, aspect ratio and height of the tracking box;

[0091] The derivatives are used to characterize the velocity parameters of the tracking box;

[0092] Step S322: Predict the predicted value of the tracking box at time k with the tracking box at time k - 1;

[0093] The expression for predicting the tracking box at time k with the tracking box at time k - 1 is: where x′ k is the predicted value of the tracking box at time k, F k is the state transition equation using a constant velocity linear motion model, x k-1 is the tracking box at time k - 1, B k and u k are quantities related to external disturbances, P′ k is the covariance matrix of x′ k , and Q k is the covariance matrix of the noise;

[0094] Step S323: Update the Kalman filter with the bounding box in the detection result at time k, associate the predicted value of the tracking box at time k with the bounding box in the detection result at time k, and obtain the tracking box at time k;

[0095] The expression for updating the Kalman filter with the bounding box in the detection result at time k is as follows: where K is the Kalman gain, z k is the bounding box in the detection result at time k, and H k is the observation matrix, and R k is the noise covariance matrix of z k , and P k is the covariance matrix of x k ;

[0096] Step S330: For each predicted tracking box, calculate the similarity of the motion features between it and the bounding box in the detection result using the Mahalanobis distance;

[0097] The calculation formula for the similarity of the motion features is as follows: where d j is the position information of the j-th bounding box in the detection result, and y i is the position information of the tracking box predicted by the i-th tracking trajectory, and S i is the covariance matrix obtained through the Kalman filter;

[0098] Step S340: Extract the appearance features of the predicted tracking box and the bounding box in the detection result, and calculate the minimum cosine distance between the appearance features of the bounding box and the predicted tracking box as the similarity of the appearance features between the predicted tracking box and the bounding box;

[0099] The tracking trajectory consists of the tracking boxes of the occupant target;

[0100] The calculation formula for the similarity of the appearance features is as follows: where is the appearance feature of the tracking box at time t predicted by the i-th tracking trajectory, and r j is the appearance feature of the j-th bounding box in the detection result;

[0101] Specifically, the feature extraction unit BoT is used to extract the appearance features of the predicted tracking box and the bounding box in the detection result. The feature extraction unit BoT uses the ResNet50 model as the backbone network and is pre-trained on the DukeMTMC-reID dataset;

[0102] The ResNet50 model consists of convolutional layers, pooling layers, and residual blocks. Deep features are extracted between the residual blocks through the skip connection mechanism;

[0103] StrongSORT replaces the relatively basic CNN architecture used in DeepSORT with a more advanced feature extraction unit, BoT (Bag of Tricks), thus enhancing the feature extraction ability. The BoT model uses ResNeSt50 as the backbone and is pre-trained on the DukeMTMC-reID dataset, enabling it to extract more discriminative features. In the field of van elevator occupant recognition (Re-ID), this improvement is particularly important as it directly relates to the performance and accuracy of the algorithm.

[0104] The introduction of the ResNet50 model brings significant advantages to the StrongSORT algorithm. First, through its unique skip connection mechanism, ResNet50 effectively addresses the common problem of vanishing gradients in deep learning networks. This structure allows the network to maintain stable performance during training, avoiding the degradation of the output feature quality even when the network depth increases. Second, the deep structure of the ResNet50 model enables StrongSORT to capture richer and more refined features. These deep features are crucial for differentiating individuals as they can provide more information about the subtle differences in pedestrian appearance. This not only improves the distinctiveness of the features but also enhances the robustness of the algorithm in complex scenarios.

[0105] Through the extraction of such deep features, the StrongSORT algorithm can achieve more precise target matching and more stable tracking performance when dealing with the van elevator occupant recognition task. This deep learning-based improvement brings new possibilities to the field of multi-object tracking, especially in real-world applications that require handling large crowds and complex backgrounds.

[0106] However, the ResNet50 model has a heavy computational burden during training. In addition, due to the large number of parameters in the ResNet50 neural network, it faces a high risk of overfitting when trained on small datasets.

[0107] Based on the specific requirements of the target and experimental environment, this application chooses to introduce the OSNet_x1_0 model in the appearance feature extraction section of the StrongSORT algorithm to replace the original ResNet50 model. OSNet_x1_0 achieves an ideal compromise between the parameter scale and high accuracy. It neither has high resource requirements like ResNet50 nor can it maintain high recognition accuracy. As one of the well-known architectures in the field of van elevator occupant recognition, OSNet has been widely applied in diverse application scenarios due to its lightweight and parameter-lean characteristics.

[0108] A notable advantage of the OSNet architecture lies in the design of its convolutional layers with different kernel sizes, and by concatenating the features extracted from each layer, the richness of features is achieved. This design allows the network to capture information at different scales, thereby enhancing the feature representation ability. In addition, due to the differences in the number of input channels of each convolutional block, the OSNet model also varies in the number of parameters, which provides flexibility for experimental environments with different computing resources. As Figure 2 shown, it is a schematic diagram of the full-scale feature learning architecture of the OSNet_x1_0 model. OSNet consists of 5 convolutional blocks, Conv1, Conv2, Conv3, Conv4, and Conv5, which are specifically used for feature extraction. This modular design not only improves the efficiency of the network but also helps to maintain the diversity and richness of features during training. By introducing the OSNet_x1_0 model, the StrongSORT algorithm of this application can more effectively handle the task of identifying occupants in a van elevator. Especially in complex scenarios, it can provide more accurate target matching and more stable tracking performance.

[0109] The architecture design of OSNet is known for its simplicity and efficiency. It stacks carefully designed lightweight bottleneck structures layer by layer, eliminating the complexity of customizing blocks for different depths in traditional networks. This innovative construction method not only simplifies the network structure but also significantly improves performance. As Figure 3 shown, it is a schematic diagram of the building block design of the OSNet architecture. The detailed architecture of OSNet shows how it achieves efficient feature extraction through streamlined stacking. Compared with traditional standard convolutional networks, OSNet has achieved significant optimization in terms of the number of parameters and computational volume. The standard convolutional network architecture usually contains approximately 6.9 million parameters and up to 3.3849 billion multiply-accumulate operations, while OSNet adopts the Lite3x3 convolutional layer design, and its model scale is only one-third of the former. This design not only reduces the complexity of the model but also significantly reduces the computational cost while maintaining the performance of the network.

[0110] In the field of the design of multi-stream architectures, OSNet shows obvious differences from the Inception and ResNeXt architectures. The design concept of OSNet is based on a strict scale-increasing principle, which enables each stream to have a unique receptive field, and all streams are constructed using a unified Lite3x3 layer. This design method shows extremely high efficiency in capturing multi-scale features. In contrast, the Inception architecture reduces the computational cost by sharing computing resources, and the combination of its convolutional and pooling operations is carefully hand-designed. While ResNeXt uses multiple streams of the same scale to learn the feature representation of the same scale.

[0111] Furthermore, OSNet adopts an innovative Aggregation Gate (AG) in feature aggregation. This design not only promotes the effective learning of multi-scale features but also makes the feature fusion process more flexible and dynamic, enabling it to adapt to various unique input images, thereby enhancing the model's adaptability and generalization ability to different scenarios. This dynamic adaptability is a key advantage of OSNet when dealing with diverse inputs. In contrast, Inception and ResNeXt usually aggregate features in a cascaded or additive manner. Additionally, the lightweight nature of OSNet benefits from its decomposed convolution technique, which keeps the building blocks and the entire network lightweight. This design significantly reduces the model's complexity and computational requirements while maintaining network performance. Compared with SENet, OSNet shows fundamental differences in design concepts. SENet mainly strengthens feature channels by adjusting activation values in a single stream, while OSNet focuses on comprehensively learning multi-scale features by consciously integrating multiple feature streams with different receptive fields. This strategy enables OSNet to capture richer feature information at both the global and local levels. This method of full-scale feature learning allows OSNet to more effectively capture and utilize the multi-scale information of the target in tasks such as identifying occupants in a van elevator, enhancing the accuracy and stability of the model in identification tasks.

[0112] In summary, by reconstructing and optimizing the occupant identification network of the StrongSORT model, changing the original ResNet50 model to the OSNet_x1_0 model, the improvement of the target tracking model of this application is successfully achieved. This improvement not only enhances the performance of the model but also provides new ideas for future research in related fields.

[0113] Step S350: Based on the similarity of motion features and appearance features, construct a cost matrix, solve the cost matrix through the Hungarian algorithm, find the bounding box in the detection results that best matches the tracking box, and update the tracking trajectory. For the successfully matched tracking trajectories, use the EMA method to update the appearance features;

[0114] Each element in the cost matrix represents the cost of matching the j-th bounding box in the detection results with the i-th tracking trajectory;

[0115] The calculation formula for the cost matrix is: c = (1 - λ) × d (1) (i, j) + λ × d (2) (i, j), where d (1) (i, j) is the similarity of motion features, d (2) (i, j) is the similarity of appearance features, and λ is the weight;

[0116] The weights are used to balance the proportions of motion similarity and appearance similarity, and are set by those skilled in the art according to experience;

[0117] Based on the Hungarian algorithm, the cost matrix is processed to achieve the optimal matching between the detection results and the tracking trajectories. The Hungarian algorithm ensures that the matching is completed with the minimum total matching cost, thereby improving the accuracy of tracking.

[0118] The calculation formula for updating the appearance features of the successfully matched tracking trajectories using the EMA method is as follows: Where, represents the appearance state of the i-th tracking trajectory at the t-th frame, represents the appearance features of the bounding box matched by the i-th tracking trajectory at the t-th frame, and α represents the momentum parameter;

[0119] The momentum parameter is set by those skilled in the art according to experience.

[0120] StrongSORT abandons the traditional feature library method for storing the appearance features of the target and instead adopts the exponential moving average strategy EMA to dynamically update the appearance state of the target. Although the feature library can accumulate long-term feature information of multiple frames, it is more sensitive to the noise in the detection process. In contrast, the EMA method reduces the interference of noise by utilizing the changes in the inter-frame features, thereby improving the accuracy of matching.

[0121] Step S360: For the detection results that are not successfully matched, initialize the occupant targets therein as new tracking trajectories;

[0122] Step S400: Generate tracking trajectories and perform tracking trajectory state management;

[0123] The specific method for generating tracking trajectories and performing tracking trajectory state management is as follows:

[0124] Step S410: Preset a first matching threshold. When the number of consecutive successful matches of the newly generated tracking trajectories reaches the first matching threshold, it is considered an effective trajectory;

[0125] Step S420: Preset a second matching threshold. When the number of consecutive unsuccessful matches of the effective trajectory reaches the second matching threshold, the effective trajectory is deleted;

[0126] The values of the first matching threshold and the second matching threshold are set by those skilled in the art according to experience. Preferably, the first matching threshold is 3 times and the second matching threshold is 30 times;

[0127] Step S500: Output the tracking trajectories and trajectory IDs of the final occupant targets.

[0128] Embodiment 2

[0129] As a classic multi-object tracking algorithm, DeepSORT has received extensive attention and applications since its proposal due to its excellent performance. The StrongSORT algorithm developed on its basis has shown better performance through further optimization and improvement. Therefore, this application selects StrongSORT as the core of the tracking algorithm in order to achieve a higher-precision multi-object tracking effect. By comprehensively using the Mahalanobis distance and the minimum cosine distance in this way, the StrongSORT algorithm can effectively handle problems such as target occlusion, fast movement, and high similarity between targets, thus achieving stable and accurate target tracking in complex scenarios.

[0130] StrongSORT adopts the YOLOx-x model in object detection and selects the ResNet50 architecture to replace the traditional CNN model in order to more effectively extract and represent the original image features. In addition, StrongSORT optimizes the Kalman filter and executes the camera motion compensation algorithm (ECC) prior to processing each frame of the image. Compared with DeepSORT, StrongSORT improves the accuracy of the tracking results by integrating motion metrics. Further, StrongSORT integrates two lightweight enhancement modules: the appearance-free linking model AFLink and the Gaussian smoothing interpolation GSI. AFLink connects short trajectory segments into a complete trajectory sequence by fusing spatio-temporal information. And GSI optimizes the effect of linear interpolation by providing a more accurate and motion-aware trajectory estimate, effectively filling the gaps in the trajectory. These innovations have significantly improved the performance of StrongSORT. Based on the above improvements, StrongSORT has achieved excellent performance in terms of HOTA and IDF metrics on multiple MOTChallenge benchmark datasets and has reached real-time processing speed when using a fast object detector. Figure 4This is a framework and performance comparison chart between DeepSORT and StrongSORT provided by this application. The StrongSORT algorithm has been deeply optimized and improved based on DeepSORT, mainly reflected in the following three aspects: First, in terms of appearance feature extraction, StrongSORT replaces the relatively basic CNN architecture used in DeepSORT by introducing a more advanced feature extraction unit BoT (Bag of Tricks), thereby enhancing the feature extraction ability. StrongSORT abandons the traditional feature library method to store the appearance features of targets and instead adopts the exponential moving average (EMA) strategy to dynamically update the appearance state of targets; Second, in terms of motion feature extraction, the enhanced correlation coefficient maximization ECC technology is introduced to achieve camera motion compensation, and the Kalman filter is improved. Given that the standard Kalman filter fails to fully consider the impact of detection noise and may thus be interfered by low-quality detection results, StrongSORT adopts the NSA (Non-Stationary State Augmentation) Kalman filter algorithm in GIAOTracker. This algorithm enhances performance by adaptively calculating the detection noise covariance; StrongSORT also adjusts the construction of the cost matrix during the matching process. Different from the method in DeepSORT that only relies on appearance information to generate the cost matrix, StrongSORT adopts a fusion strategy that combines appearance and motion information. In StrongSORT, the motion feature similarity between a trajectory and a detected target is no longer just a threshold for screening matching results but is incorporated into the construction of the cost matrix together with the appearance feature similarity. This fusion strategy enables the cost matrix to more comprehensively reflect the accuracy of matching, thereby improving the overall performance of the matching process; Third, in terms of removing cascade matching, StrongSORT abandons the cascade matching strategy adopted by DeepSORT and instead selects a more general global linear matching method. This change aims to improve the flexibility and adaptability of tracking, thereby achieving a more accurate tracking effect in variable scenarios.

[0131] Example 3

[0132] An implementation manner of this application also selects the multi-object tracking accuracy MOTA, the higher-order tracking accuracy HOTA, and the proportion IDF1 of correctly identified detections to the average of true detections and calculated detections as the key indicators for evaluating the performance of the target tracking model. These indicators comprehensively reflect the performance of the model in terms of tracking accuracy, recognition effect, and overall performance.

[0133] MOTA comprehensively measures the ability of a model to maintain accurate tracking in object tracking tasks. It encompasses three main errors that may occur during tracking: missed detection of objects, false detection, and incorrect switching of object identity labels.

[0134] The specific calculation method for multi-object tracking accuracy is as follows: Among them, FN t represents the number of objects not detected in frame t, FP t represents the number of falsely detected objects, IDs t represents the number of objects with incorrect identity label switches, GT t refers to the total number of all ground truth objects.

[0135] HOTA is a comprehensive evaluation metric in the field of multi-object tracking. It measures the overall performance of a tracker by integrating the performance of three sub-tasks: detection, association, and localization. The calculation of HOTA involves statistics of the number of correct associations TPA, the number of false associations FPA, and the number of non-associated objects FNA, as well as the calculation of IoU (Intersection over Union), and then obtains the detection score DetA, the association score AssA, and the localization score LocA. These scores are finally combined into a single HOTA score. The characteristic of HOTA is that it can integrate over different IoU thresholds IoU TPA and incorporate the localization accuracy into the final score, thus balancing the accuracy of detection and association. Compared with MOTA and IDF1, HOTA can more comprehensively reflect the performance of a tracker in terms of detection, association, and localization. Its value ranges from 0 to 1, and the higher the value, the higher the tracking accuracy. The application of HOTA is to evaluate the performance of multi-object tracking algorithms, providing a more comprehensive evaluation method.

[0136] The calculation formula for the detection score is as follows:

[0137] The calculation formula for the association score is as follows:

[0138] The calculation formula for the localization score is as follows:

[0139] The calculation formula for the higher-order tracking accuracy is as follows:

[0140] IDF1 is an important performance evaluation metric in the field of multi-object tracking. It measures the ability of a tracking algorithm to identify and maintain the consistency of object identities.

[0141] Specifically, the calculation formula for IDF1 is as follows: Among them, IDTP represents the number of correctly matched identities, that is, the cases where the tracker correctly identifies and tracks the targets. IDFP represents the number of mis-matched identities, that is, the cases where the tracker mis-identifies different targets as the same target. IDFN represents the number of missed identities, that is, the number of targets that the tracker fails to identify and track.

[0142] The value range of IDF1 is [0, 1]. The closer the value is to 1, the higher the target recognition accuracy rate. This metric particularly focuses on the consistency of target identities during the tracking process and is especially important for evaluating the performance of multi-target tracking algorithms in complex scenarios.

[0143] Furthermore, this application explores the motivation and rationality for replacing the original ReID network resNet50 in StrongSORT with OSNet_x1_0. To compare the effects of the improvements in target tracking technology, it is necessary to train OSNet_x1_0 and resnet50 on the same dataset to generate pedestrian re-identification network models. Therefore, this application selects the Market1501 dataset to perform these comparative experiments. This dataset was proposed by ZhengL et al. in 2015 and has now become one of the standard datasets in the field of pedestrian re-identification. Market1501 contains pedestrian images captured by 6 cameras on the campus of Tsinghua University, with a total of 1501 pedestrians labeled. Specifically, 751 pedestrians are used for the training set, while 750 pedestrians are used for the test set, ensuring that there are no overlapping pedestrian IDs between the training set and the test set.

[0144] Finally, ReID models of resnet50_market1501.pt and osnet_x1_0_market1501.pt are trained. Table 3.5 shows the comparison of the number of parameters and GFLOPs of the two models. OSNet_x1_0 has a significant reduction in both of these metrics.

[0145] Table 3.5 Comparison of ReID parameter quantity and GFLOPs

[0146]

[0147] In one implementation of this application, all tracking experiments use the YOLOv8-MPDIoU+C2fGhostV2 model as the target detector. On the MOT17 dataset, the embodiments verify the performance of the proposed multi-target tracking algorithm and conduct a comparative analysis of its performance with the original algorithm and the performance of other current mainstream tracking algorithms under the same conditions. Table 3.6 shows the results of the pedestrian tracking comparative experiment.

[0148] Table 3.6 Results of pedestrian tracking comparative experiment

[0149]

[0150] From the results of the pedestrian tracking comparison experiment, it can be seen that the resnet50_market1501.pt model, with its 23.5M parameters and 2.7GFLOPs of computational volume, demonstrates powerful performance in the field of multi-object tracking. The osnet_x1_0_market1501.pt model, on the other hand, with its 2.2M parameters and 0.98GFLOPs of computational volume, has become a model of lightweight design and is particularly suitable for environments with limited computing resources. In terms of the MOTA metric, which measures the overall accuracy of the tracking algorithm, both StrongSORT-Resnet50 and StrongSORT-OsNet_x_10 outperform the DeepSORT and ByteTrack algorithms. Among them, StrongSORT-Resnet50 slightly leads StrongSORT-OsNet_x_10 with 33.581% of MOTA compared to 33.596% of StrongSORT-OsNet_x_10. In terms of the HOTA metric, which comprehensively considers the accuracy of detection, association, and localization, StrongSORT-Resnet50 leads other methods with 39.857%, while StrongSORT-OsNet_x_10 follows closely with 39.533%. In terms of the IDF1 metric for identity recognition accuracy, StrongSORT-Resnet50 leads with 46.654%, and StrongSORT-OsNet_x_10 also performs well with 46.077%. Considering these metrics comprehensively, StrongSORT-Resnet50 demonstrates excellent accuracy and identity recognition ability in multi-object tracking tasks. Although the osnet_x1_0_market1501.pt model has fewer parameters and less computational volume, its performance does not decline significantly, especially in terms of HOTA and IDF1, where it is not much different from StrongSORT-Resnet50. Therefore, for application scenarios with high requirements for computing resources, the osnet_x1_0_market1501.pt model is an ideal choice; while for applications that pursue higher tracking accuracy and identity recognition ability, the resnet50_market1501.pt model is undoubtedly a more suitable choice. For the application scenario of this article - tracking the entry and exit of occupants in a van elevator, considering the particularity of the environment and the limitation of computing resources, the StrongSORT-OsNet_x_10 algorithm, with its lightweight advantage and performance close to StrongSORT-Resnet50, is a more suitable choice.

[0151] This embodiment also demonstrates the test effect of using the StrongSORT-OsNet_x_10 algorithm for pedestrian tracking in a test video. Figure 5 It is the scene graph of the 125th frame of the test video sequence.Figure 6 It is the scene diagram of the 182nd frame of the test video sequence. Figure 7 It is the scene diagram of the 196th frame of the test video sequence. It can be seen that after the pedestrians with id15 and id18 overlap and occlude, the tracking of the pedestrian with id15 has a short interruption, but at the 196th frame, re-identification can be performed and tracking can continue. This result fully demonstrates the excellent performance of the OSNet_x1_0 re-identification network in optimizing the number of parameters and computational efficiency. Although the number of parameters and computational load are reduced during the model simplification process, OSNet_x1_0 can still maintain excellent pedestrian re-identification performance. This ability is extremely important in practical applications, especially in dealing with complex scenarios and resource-constrained environments such as van elevators. OSNet_x1_0 demonstrates its reliability and effectiveness in pedestrian tracking tasks.

[0152] In addition, an embodiment of the present application uses the improved object tracking model StrongSORT-OsNet_x_10 and the object detection model YOLOv8-MPDIoU+C2fGhostV2 to conduct a series of detection and tracking tests on the surveillance video stream of the van elevator. The test results are respectively reflected in the 2nd frame, 36th frame and 63rd frame of the video. Figure 8 It is the test effect diagram of the 2nd frame of the surveillance video stream of the van elevator. Figure 9 It is the test effect diagram of the 36th frame of the surveillance video stream of the van elevator. Figure 10 It is the test effect diagram of the 63rd frame of the surveillance video stream of the van elevator. In these frames, the occupant targets are not only accurately identified, but also continuously and stably tracked throughout the video sequence. Even when encountering detection challenges at certain moments, the system can quickly resume tracking with the help of the re-identification network to ensure the consistency of the target trajectory ID. This series of tests fully demonstrates the high performance and robustness of the tracking algorithm, and also verifies the effectiveness of the improved object tracking model in practical applications. Through these tests, the practicality of the model in complex environments is demonstrated, providing an effective solution for the van elevator monitoring system.

[0153] In addition, the parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.

[0154] As described above in the specific embodiments, the purpose, technical solutions and beneficial effects of the present invention are further described in detail. It should be understood that the above are only the specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for tracking passengers in a van elevator based on an improved StrongSORT algorithm, characterized in that, Including: Obtain the monitoring video stream inside the van elevator, and obtain the current frame image after motion compensation based on the enhanced correlation coefficient maximization algorithm; Use the object detection model to process the compensated current frame image, and obtain the bounding box coordinates, category, and confidence of each occupant object as the detection result; Based on the detection result, use the StrongSORT algorithm to track the occupant objects in each frame image of the monitoring video stream; Generate tracking trajectories and perform tracking trajectory status management; Output the tracking trajectories and trajectory IDs of the final occupant objects.

2. The method for tracking passengers in a van elevator based on the improved StrongSORT algorithm according to claim 1, wherein The specific method for obtaining the monitoring video stream inside the van elevator and obtaining the current frame image after motion compensation based on the enhanced correlation coefficient maximization algorithm is as follows: Obtain the monitoring video stream inside the van elevator, determine the reference frame image, and use the current frame image as the distorted frame image. Zero-mean the reference frame image and the distorted frame image to obtain the zero-mean reference frame and the zero-mean distorted frame, and calculate the maximized enhanced correlation coefficient; Minimize the maximized enhanced correlation coefficient ECC(p), optimize the distortion parameter p, and perform motion compensation on the distorted frame image based on the distortion parameter; Output the current frame image after motion compensation and the distortion parameter p.

3. The method for tracking occupants in a van elevator based on the improved StrongSORT algorithm according to claim 2, characterized in that, The specific method for using the object detection model to process the compensated current frame image and obtaining the bounding box coordinates, category, and confidence of each occupant object as the detection result is as follows: Construct an object detection model; Define the input of the object detection model as the frame image after motion compensation, and the output as the bounding box coordinates, bounding box category, and confidence of the occupant objects in each frame image after compensation; Take the output bounding box coordinates, bounding box category, and confidence of the occupant objects as the detection result of the current frame image.

4. The method for tracking occupants in a van elevator based on the improved StrongSORT algorithm according to claim 3, wherein, The object detection model selects the YOLOv8-MPDIoU+C2fGhostV2 model, uses the image after motion compensation as the input data, and the output is the bounding box coordinates, category probability, and confidence of each occupant object in the image after motion compensation. The category corresponding to the maximum category probability is the bounding box category; The YOLOv8-MPDIoU+C2fGhostV2 model is improved based on the YOLOv8 architecture. The core modules include the backbone network, feature fusion layer, and detection head. Use the MPDIoU loss function to optimize the bounding box regression accuracy. Replace the C2f module with the C2fGhostV2 module in the backbone network and the feature fusion layer. Use the MOT17 pedestrian detection dataset with labeled pedestrian bounding boxes as the training set to train the object detection model, and use minimizing the MPDIoU loss function as the training objective.

5. The method for tracking passengers in a van elevator based on the improved StrongSORT algorithm according to claim 4, characterized in that The specific method for tracking the occupant objects using the StrongSORT algorithm based on the detection result is as follows: For the detection result of the first frame image, initialize each detected occupant object as a tracking trajectory in an unconfirmed state and assign a trajectory ID to it; For each subsequent frame image, use the detection result of the current frame and the tracking trajectory of the previous frame as the input, and use the Kalman filter to predict the next tracking box corresponding to the tracking trajectory of the previous frame; For each predicted tracking box, the Mahalanobis distance is used to calculate the similarity of motion features between it and the bounding boxes in the detection results; Extract the appearance features of the predicted tracking box and the bounding boxes in the detection results, and calculate the minimum cosine distance between the appearance features of the bounding boxes and the predicted tracking box as the appearance feature similarity between the predicted tracking box and the bounding boxes; Based on the motion feature similarity and the appearance feature similarity, construct a cost matrix, solve the cost matrix through the Hungarian algorithm, find the bounding box in the detection results that best matches the tracking box, and update the tracking trajectory. For the successfully matched tracking trajectories, use the EMA method to update the appearance features; For the detection results that are not successfully matched, initialize the occupant targets therein as new tracking trajectories.

6. The method for tracking passengers in a van elevator based on the improved StrongSORT algorithm according to claim 5, characterized in that, The specific method of using the Kalman filter to predict the next tracking box corresponding to the tracking trajectory of the previous frame with the detection results of the current frame and the tracking trajectory of the previous frame as inputs for each subsequent frame of image is as follows: Characterize the tracking box with motion features; Predict the predicted value of the tracking box at time k with the tracking box at time k-1; Update the Kalman filter with the bounding box in the detection results at time k, and associate the predicted value of the tracking box at time k with the bounding box in the detection results at time k to obtain the tracking box at time k.

7. The method for tracking occupants in a van elevator based on the improved StrongSORT algorithm according to claim 6, characterized in that, The Kalman filter adopts the NSA Kalman filter algorithm in GIAOTracker, and calculates the detection noise covariance based on the confidence of the detection results and a preset fixed value of the detection noise covariance.

8. The method for tracking occupants in a van elevator based on the improved StrongSORT algorithm according to claim 7, characterized in that, The feature extraction unit BoT is used to extract the appearance features of the predicted tracking box and the bounding boxes in the detection results. The feature extraction unit BoT uses the ResNet50 model as the backbone network and is pre-trained on the DukeMTMC-reID dataset; the ResNet50 model consists of convolutional layers, pooling layers and residual blocks, and deep features are extracted through the skip connection mechanism between the residual blocks.

9. The method for tracking occupants in a van elevator based on the improved StrongSORT algorithm according to claim 8, characterized in that, The specific method of generating tracking trajectories and managing the states of the tracking trajectories is as follows: Preset a first matching threshold. When the number of consecutive successful matches of the newly generated tracking trajectory reaches the first matching threshold, it is considered an effective trajectory; Preset a second matching threshold. When the number of consecutive unsuccessful matches of the effective trajectory reaches the second matching threshold, the effective trajectory is deleted.

Citation Information

Cited By

  • Coal mining machine roller tracking system and method

    CN121010931A

  • Operation site safety supervision method based on multi-target tracking

    CN121121661A

  • A job site safety supervision method based on multi-target tracking

    CN121121661B

  • Natural gas pipeline camera video personnel abnormal key frame extraction method

    CN121214309A