Campus scene pedestrian detection and tracking method and system

Through the improved YOLOv8 and OcSORT algorithm and self-built campus dataset, the pedestrian detection and tracking algorithm is optimized, and the problems of frame rate decline and poor applicability of campus environment in the existing technology are solved, and efficient and real-time pedestrian monitoring is achieved.

CN120014535APending Publication Date: 2025-05-16INST OF ADVANCED TECH UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411970619.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The significant decline in frame rates in the prior art leads to the difficulty in meeting real-time requirements and the poor applicability of the system in campus environments.

Method used

The improved YOLOv8 and OcSORT algorithm are adopted, combined with self-built campus scene data sets, and the pedestrian detection and tracking algorithm is optimized. Improve detection accuracy through the SCConv module and EMA module, and track it using Kalman filtering and pedestrian re-identification algorithms.

Benefits of technology

It improves the accuracy of pedestrian detection and tracking accuracy, solves the problems of insufficient detection accuracy and low tracking accuracy, and performs excellently in cross-camera tracking and ID matching, meeting the needs of efficient and real-time pedestrian monitoring in campus scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014535A_ABST
    Figure CN120014535A_ABST
Patent Text Reader

Abstract

The invention provides a campus scene pedestrian detection and tracking method and system. The method comprises the following steps: constructing a campus scene data set; according to the campus scene data set, a pedestrian detection model is set and optimized to detect pedestrian targets, and the pedestrian detection module comprises an SCConv module, an EMA module and a YOLOv8 module; the pedestrian tracking module is based on a staged tracking strategy, utilizes an OcSORT algorithm, and uses Kalman filtering to track a pedestrian target of which the IoU confidence is greater than or equal to a preset threshold value; and carrying out confirmation operation on the pedestrian target of which the IoU confidence coefficient is smaller than a preset threshold value by utilizing a pedestrian re-recognition algorithm. Wherein in the motion data association stage, association matching is performed in combination with inter-frame shape similarity and IoU confidence; in the track updating stage, an ID matching method based on pedestrian re-identification is utilized to initialize pedestrian IDs. The technical problems that the real-time requirement is difficult to meet due to the fact that the frame rate is obviously reduced, and the applicability of the system in the campus environment is poor are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and artificial intelligence technology, and in particular to a method and system for detecting and tracking pedestrians in campus scenes. Background Art

[0002] With the rapid development of deep learning technology, pedestrian detection and multi-target tracking technologies have been widely used in traffic monitoring, campus safety warning, public safety protection, and autonomous driving. These technologies not only play a key role in ensuring driving safety and achieving real-time detection and accurate tracking of pedestrians, but also provide tracking support for specific pedestrians for public safety maintenance. However, how to efficiently and accurately detect and track pedestrians in complex environments is still a technical challenge in practical applications.

[0003] In recent years, target detection technology has made significant progress, especially the YOLO series of algorithms have attracted much attention due to their high efficiency and accuracy. For example, the existing invention patent application document "A road pedestrian and vehicle detection algorithm based on improved YOLOv5s" with publication number CN116740679A includes: selecting the KITTI dataset, reclassifying the original dataset and converting the format, and dividing the obtained dataset into a training set and a validation set in a ratio of 8:2; step 2) transmitting the training set image into the input end of YOLO v5s, and the input end uses Mosaic data enhancement; step 3) the feature map after the input end enters the Backbone structure of YOLO v5s for processing, step 4) the feature map will enter the Neck of YOLO v5s from different places of the Backbone to generate a feature map of multi-scale information; step 5) the aggregated features are input into the Head, and loss optimization and evaluation are performed in the Head. And the existing invention patent application document "A lightweight pedestrian detection method based on YOLOv5 algorithm" with publication number CN116740532A, in which YOLO uses multi-layer convolution for image detection, in which 3×3 convolution occupies the main part of the calculation amount. Usually, the detector based on convolutional neural network consists of three parts, backbone, neck, and head. The backbone is used to extract the features of the input image, which is used to better allocate and merge the features to the head and neck. The neck is generally responsible for strengthening the features, and then the head is responsible for prediction. However, there is still room for improvement in the accuracy of pedestrian detection in monitoring scenarios by existing algorithms. In addition, after combining the multi-target tracking technology with the pedestrian re-identification method, although it shows excellent performance under lighting changes, occlusion and complex background environments, due to the complex reasoning process of the deep learning network, it often leads to a significant decrease in frame rate, which is difficult to meet real-time requirements.

[0004] In summary, the existing technology has technical problems such as a significant decrease in frame rate, which makes it difficult to meet real-time requirements, and poor applicability of the system in a campus environment. Summary of the invention

[0005] The technical problem to be solved by the present invention is: how to solve the technical problems in the prior art that the frame rate is significantly reduced, making it difficult to meet the real-time requirements, and the system has poor applicability in a campus environment.

[0006] The present invention adopts the following technical solutions to solve the above technical problems: A method for detecting and tracking pedestrians in a campus scene includes:

[0007] S1. Build a campus scene dataset;

[0008] S2. According to the campus scene dataset, set up and tune the pedestrian detection model to detect pedestrian targets. The pedestrian detection module includes: SCConv module, EMA module and YOLOv8 module;

[0009] S3. Based on the phased tracking strategy, the OcSORT algorithm is used to track pedestrian targets with an IoU confidence greater than or equal to the preset threshold using Kalman filtering; for pedestrian targets with an IoU confidence less than the preset threshold, the pedestrian re-identification algorithm is used for confirmation. Specifically, in the motion data association stage, the inter-frame shape similarity and IoU confidence are combined for association matching; in the trajectory update stage, the pedestrian ID is initialized using the ID matching method based on pedestrian re-identification.

[0010] The present invention performs pedestrian detection in campus scenes based on improved YOLOv8 and OcSORT algorithms. Different from traditional pedestrian detection and multi-target tracking algorithms, the present invention optimizes the algorithm according to the specific needs of the campus environment, and improves the detection and tracking algorithms by introducing a self-built campus scene data set, thereby improving the accuracy of pedestrian detection and tracking. The present invention effectively solves the problems of insufficient detection accuracy and low tracking accuracy of existing algorithms in complex campus environments, and solves the problems of ID mismatch and cross-camera tracking when pedestrians reappear after leaving for a short time, thereby providing a reliable technical solution for efficient and real-time pedestrian monitoring in campus scenes.

[0011] In a more specific technical solution, S1 includes:

[0012] S11, setting acquisition equipment, acquisition environment, and data format to acquire diversified image data;

[0013] S12. Perform a multi-target tracking data set labeling operation on the diversified image data frame by frame to obtain a labeled data set, and divide the labeled data set into a training set and a test set.

[0014] Aiming at the actual needs of pedestrian detection and tracking in campus environments, the present invention independently constructs a pedestrian detection and multi-target tracking dataset for campus scenes. This dataset contains scene features and crowd distribution patterns that are unique to campus environments, providing reliable data support for model optimization, enabling the algorithm to better adapt to the application needs of campus scenes, thereby improving detection and tracking performance. This customized dataset greatly improves the application effect of the algorithm in campus scenes, and provides a solid data foundation for security and crowd behavior analysis in smart campuses, thereby improving its practicality and value.

[0015] In a more specific technical solution, in S2, an SCConv module is set and used to perform self-calibration convolution, divide the convolution filter to perform adaptive expansion of the receptive field, exchange information between channels, and obtain context information and local features;

[0016] The EMA module is set up and used to perform channel grouping operations and multi-scale parallel convolution to enhance feature expression with high computational efficiency. The EMA module also includes: a local channel relationship preservation branch and a spatial dependency relationship capture branch to model multi-scale feature information.

[0017] Set up and use the YOLOv8 module to build a pedestrian detection network, and use multi-level feature layers P3, P4, and P5 for multi-scale target detection.

[0018] The present invention optimizes the YOLOv8 target detection algorithm, significantly improves the detection accuracy by introducing an attention mechanism and replacing the original C2f module with a plug-and-play SCConv module.

[0019] In a more specific technical solution, in S3, a matching operation is performed based on the shape similarity and IoU confidence of the pedestrian detection frame, wherein the shape similarity includes: the aspect ratio of the detection frame and the shape change information of adjacent frames.

[0020] In a more specific technical solution, the shape similarity is calculated using the following logic:

[0021]

[0022] In the formula, (W D ,H D ) and (W T ,H T ) represent the width and height of the detection frame of the current frame and the detection frame of the tracked target in the previous frame, respectively.

[0023] In a more specific technical solution, the following logic is used to obtain the associated matching score and perform the matching operation accordingly:

[0024] S=(1-λ)SIoU +λS shape

[0025] In the formula, S IoU represents the spatial overlap score based on IoU confidence, S shape represents the shape similarity score, and λ is the adjustment coefficient.

[0026] In a more specific technical solution, in S3, when a pedestrian target enters the camera field of view for the first time, a feature vector of the pedestrian target is calculated by a pedestrian re-identification algorithm, and cosine similarity calculation is performed between the feature vector and the feature vector of a preset feature library.

[0027] In terms of multi-target tracking, this paper proposes a new tracking strategy, combined with a pedestrian re-identification method, to improve tracking accuracy while maintaining a high frame rate. The tracking accuracy is further improved by introducing the matching criteria of inter-frame detection box shape similarity and IoU confidence.

[0028] In a more specific technical solution, when the cosine similarity is greater than or equal to a preset similarity threshold, the pedestrian target is determined to be an existing pedestrian and an original ID is assigned;

[0029] When the cosine similarity is less than a preset similarity threshold, the pedestrian target is determined to be a new pedestrian, a new ID is assigned, and the new ID and feature vector are stored in a preset feature library.

[0030] In a more specific technical solution, in S3, in the cross-camera situation, an ID initialization operation is performed on the pedestrian feature library, each camera independently executes the Kalman filter-based tracking algorithm frame by frame, and all cameras share the pedestrian feature library.

[0031] This paper proposes a new solution to the ID mismatch problem caused by pedestrians leaving the camera for a short time in a single-camera scenario: when pedestrians first enter the camera's field of view, pedestrian re-identification technology is used to match pedestrian IDs, thereby ensuring that even if pedestrians temporarily leave the field of view, their IDs can be correctly identified and matched. This solution is also applicable to multi-target tracking across cameras, significantly enhancing the adaptability and application value of the system in complex environments.

[0032] In a more specific technical solution, a campus scene pedestrian detection and tracking system includes:

[0033] Campus dataset construction module, used to construct campus scene dataset;

[0034] The pedestrian detection module is used to set and tune the pedestrian detection model according to the campus scene data set to detect pedestrian targets. The pedestrian detection module includes: SCConv module, EMA module and YOLOv8 module. The pedestrian detection module is connected to the campus data set construction module;

[0035] The phased tracking module is used to track pedestrian targets with IoU confidence greater than or equal to the preset threshold using the OcSORT algorithm based on the phased tracking strategy, using Kalman filtering; for pedestrian targets with IoU confidence less than the preset threshold, the pedestrian re-identification algorithm is used for confirmation. The phased tracking module is connected to the pedestrian detection module. Specifically, in the motion data association stage, the inter-frame shape similarity and IoU confidence are combined for association matching; in the trajectory update stage, the pedestrian ID is initialized using the ID matching method based on pedestrian re-identification.

[0036] Compared with the prior art, the present invention has the following advantages:

[0037] The present invention performs pedestrian detection in campus scenes based on improved YOLOv8 and OcSORT algorithms. Different from traditional pedestrian detection and multi-target tracking algorithms, the present invention optimizes the algorithm according to the specific needs of the campus environment, and improves the detection and tracking algorithms by introducing a self-built campus scene data set, thereby improving the accuracy of pedestrian detection and tracking. The present invention effectively solves the problems of insufficient detection accuracy and low tracking accuracy of existing algorithms in complex campus environments, and solves the problems of ID mismatch and cross-camera tracking when pedestrians reappear after leaving for a short time, thereby providing a reliable technical solution for efficient and real-time pedestrian monitoring in campus scenes.

[0038] Aiming at the actual needs of pedestrian detection and tracking in campus environments, the present invention independently constructs a pedestrian detection and multi-target tracking dataset for campus scenes. This dataset contains scene features and crowd distribution patterns that are unique to campus environments, providing reliable data support for model optimization, enabling the algorithm to better adapt to the application needs of campus scenes, thereby improving detection and tracking performance. This customized dataset greatly improves the application effect of the algorithm in campus scenes, and provides a solid data foundation for security and crowd behavior analysis in smart campuses, thereby improving its practicality and value.

[0039] The present invention optimizes the YOLOv8 target detection algorithm, significantly improves the detection accuracy by introducing an attention mechanism and replacing the original C2f module with a plug-and-play SCConv module.

[0040] In terms of multi-target tracking, this paper proposes a new tracking strategy, combined with the pedestrian re-identification method, to improve the tracking accuracy while maintaining a high frame rate. The tracking accuracy is further improved by introducing the inter-frame detection box shape similarity and IoU matching criteria.

[0041] This paper proposes a new solution to the ID mismatch problem caused by pedestrians leaving the camera for a short time in a single-camera scenario: when pedestrians first enter the camera's field of view, pedestrian re-identification technology is used to match pedestrian IDs, thereby ensuring that even if pedestrians temporarily leave the field of view, their IDs can be correctly identified and matched. This solution is also applicable to multi-target tracking across cameras, significantly enhancing the adaptability and application value of the system in complex environments.

[0042] The present invention solves the technical problems existing in the prior art that the frame rate is significantly reduced, making it difficult to meet the real-time requirements, and the system has poor applicability in a campus environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a schematic diagram of the basic steps of a method for detecting and tracking pedestrians in a campus scene according to Embodiment 1 of the present invention;

[0044] Figure 2 This is a schematic diagram of the MOT17 data set of Example 1 of the present invention;

[0045] Figure 3 This is a schematic diagram of the MOT20 data set of Example 1 of the present invention;

[0046] Figure 4 This is a schematic diagram of a DanceTrack data set according to Embodiment 1 of the present invention;

[0047] Figure 5 This is a schematic diagram of a self-built campus data set according to Embodiment 1 of the present invention;

[0048] Figure 6 This is a schematic diagram of the SCConv module diagram of Example 1 of the present invention;

[0049] Figure 7 This is a schematic diagram of an EMA module diagram of Example 1 of the present invention;

[0050] Figure 8 Schematic diagram of an improved pedestrian detection network architecture based on YOLOv8 according to Embodiment 1 of the present invention;

[0051] Fig. 9 This is a schematic diagram of data flow processing based on the Kalman filter tracking algorithm according to Embodiment 1 of the present invention;

[0052] Fig.10 This is a schematic diagram of data stream processing based on Kalman filtering + ReID tracking algorithm according to Embodiment 1 of the present invention;

[0053] Fig.11 This is a schematic diagram of specific process steps of the improved tracking algorithm of Example 1 of the present invention;

[0054] Fig.12This is a schematic diagram of data flow processing of the ID matching solution of Example 1 of the present invention;

[0055] Fig.13 This is a schematic diagram of data flow processing of a tracking solution improved based on the OcSORT algorithm according to Example 1 of the present invention;

[0056] Fig.14 This is a schematic diagram of data stream processing for a cross-camera multi-target tracking solution according to Embodiment 1 of the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] Example 1

[0059] like Figure 1 As shown, the present invention provides a campus scene pedestrian detection and tracking method, including the following basic steps:

[0060] S1. Build a campus scene dataset;

[0061] In the multi-target tracking task of this embodiment, the core goal is to achieve real-time, accurate positioning and continuous identification of multiple targets by analyzing video sequences, so as to obtain the spatiotemporal trajectory of each target. Multi-target tracking has a wide range of applications in intelligent video surveillance, behavior analysis, autonomous driving, smart campuses, etc. It not only supports traffic statistics and abnormal behavior detection, but also can realize interactive analysis of multiple subjects in complex scenarios, thus meeting the diverse needs in practical applications.

[0062] like Figure 2 , Figure 3 , Figure 4 As shown in the figure, in this embodiment, the number, density, appearance difference and other conditions of the targets in each application scenario are different, which makes the selection and construction of the data set particularly critical in the multi-target tracking (MOT) research. The multi-target tracking data sets commonly used in academia currently include MOT17, MOT20 and DanceTrack, etc. These data sets have their own characteristics in terms of scale, scene setting, number of sequences, number of frames and application fields, and are widely used in different experimental verifications and algorithm comparisons. The specific data set comparison is shown in Table 1, and the data set examples can be found in Figure 2 , Figure 3 , Figure 4 .

[0063] Table 1 Comparison of public multi-target tracking datasets

[0064]

[0065] Although datasets such as MOT17, MOT20, and DanceTrack have played a key role in multi-target tracking research, their application scenarios are relatively limited and difficult to meet the specific needs of campus environments. Campus scenes have significant uniqueness, such as diverse target types, obvious age differences, large density changes, and frequent partial occlusions. In addition, scene switching is frequent in campus environments and is easily affected by external conditions such as lighting and weather. Existing datasets cannot fully cover these characteristics, resulting in insufficient data for multi-target tracking research in smart campuses. Therefore, building a multi-target tracking dataset specifically for campus scenes has important research value and practical significance for improving functions such as accurate security monitoring, traffic statistics, and anomaly detection in smart campuses.

[0066] In terms of data collection, in order to ensure the comprehensiveness and diversity of the data, the constructed dataset has been designed in detail in terms of collection equipment, collection environment and data format. The collection equipment uses a high-definition camera with a resolution of 1920×1080 and a frame rate of 30 frames per second to meet the technical requirements of the multi-target tracking task. Data collection covers different time periods (such as daytime and evening), and selects multiple typical scenes on campus (such as the cafeteria entrance, outside the dormitory building, playground, dormitory door, outside the teaching building, etc.) to capture the diverse characteristics of the campus environment under different conditions. The data format adopts a unified standard to keep the video consistent in encoding format, resolution, frame rate and duration, which is convenient for subsequent processing and annotation, and ensures the standardization and high availability of data.

[0067] In the data annotation stage, the present invention uses DarkLabel, an open source tool designed specifically for target detection and tracking tasks. The tool provides an intuitive user interface and an efficient annotation process, which is very suitable for the annotation needs of multi-target tracking data sets. The annotation content is performed frame by frame through a graphical interface, recording the category, location and unique ID of each target to ensure the integrity and accuracy of the annotation information. DarkLabel supports shortcut key operations, which greatly improves the annotation speed and reduces the manual workload, thereby ensuring the annotation quality and providing reliable data support for the training and testing of multi-target tracking algorithms.

[0068] In this embodiment, the detailed information of the self-built data set is shown in Table 2. Figure 5 As shown in Figure 2, the annotated dataset is divided into a training set and a test set to support comprehensive experimental verification of the multi-target tracking task.

[0069] Table 2 Self-built campus dataset information

[0070]

[0071] S2. According to the campus scene dataset, set up and tune the pedestrian detection model to detect pedestrian targets. The pedestrian detection module includes: SCConv module, EMA module and YOLOv8 module;

[0072] Campus scenes have unique environmental characteristics, such as large changes in pedestrian density, frequent target occlusion, and complex and changeable lighting and weather conditions. These factors pose severe challenges to existing pedestrian detection algorithms. Specifically, traditional algorithms often suffer from insufficient detection accuracy, high false detection rate and missed detection rate in campus scenes. Therefore, it is necessary to make special improvements to pedestrian detection algorithms to adapt to the special needs of campus scenes, thereby improving the accuracy and robustness of detection.

[0073] In this embodiment, the application performance of the YOLOv8 algorithm is relatively stable, especially YOLOv8s achieves a good balance between detection accuracy, parameter quantity and processing speed, and is suitable for scenarios with high real-time requirements. Based on the balance between accuracy and speed, this embodiment chooses to further improve and optimize YOLOv8s as the core to better meet the actual needs of pedestrian detection in a campus environment.

[0074] In this embodiment, the pedestrian detection model includes: an SCConv module, an EMA module, and an improved pedestrian detection network based on YOLOv8;

[0075] like Figure 6 As shown, the SCConv module (Self-Calibrated Convolution) of this embodiment is a convolution operation that enhances feature conversion capability through a self-calibration mechanism. This module significantly improves the modeling capability of convolutional neural networks for spatial features without increasing computational complexity.

[0076] In this embodiment, the SCConv module divides the convolution filter into multiple parts, so that each part can adaptively expand the receptive field and achieve sufficient information exchange between channels, thereby obtaining richer contextual information. This design not only enhances the channel association during the convolution process, but also captures the detail information in the input image more comprehensively. Compared with the original C2f module of YOLOv8, the SCConv module has significant advantages in feature resolution and detail retention, and can more keenly capture subtle changes and local features in the image.

[0077] like Figure 7As shown, the EMA module of this embodiment is an efficient multi-scale attention mechanism that aims to maintain the integrity of channel information while reducing computational overhead. The EMA module enhances feature expression through channel grouping and multi-scale parallel convolution, and has high computational efficiency. Specifically, the EMA module consists of two parallel sub-network branches, using 1x1 and 3x3 convolution kernels respectively: the 1x1 convolution branch is used to preserve local channel relationships, while the 3x3 convolution branch is used to capture a wider range of spatial dependencies. This design combines short-range and long-range dependencies and can effectively model multi-scale feature information.

[0078] In the pedestrian detection task of this embodiment, the cross-space interaction mechanism of the EMA module can capture the spatial relationship at the pixel level, which is particularly important for pedestrian detection in complex backgrounds. Pedestrian detection is often interfered by similar objects in the background, and the EMA module can better distinguish pedestrian features from background noise, thereby improving detection accuracy. In addition, since pedestrians may appear at different scales and positions in the image, the multi-scale convolution design of the EMA module helps capture these multi-scale features and further improve the detection effect.

[0079] like Figure 8 As shown, in the improved pedestrian detection network based on YOLOv8 in this embodiment, the architecture replaces the original C2f module with the SCConv module and introduces the EMA module at the end of the backbone network.

[0080] At the detection level, YOLOv8 uses multi-level feature layers (P3, P4, P5) for multi-scale target detection. After the introduction of SCConv and EMA modules, the perception ability of each level of feature layer for pedestrians of different scales has been significantly improved, which is especially suitable for multi-scale pedestrian detection tasks. The SCConv module enhances feature resolution and retains details, making the model perform better in detecting small-scale targets; at the same time, the EMA module optimizes the capture of spatial features through multi-scale convolution and channel grouping mechanisms, further improving detection accuracy. The combination of SCConv and EMA modules enables the improved YOLOv8 architecture to have stronger multi-scale feature representation capabilities and more accurate spatial relationship modeling capabilities, thereby achieving higher accuracy and robustness in pedestrian detection tasks.

[0081] S3, based on the phased tracking strategy, using the OcSORT algorithm, for pedestrian targets whose IoU confidence is greater than or equal to the preset threshold, Kalman filtering is used for tracking; for pedestrian targets whose IoU confidence is less than the preset threshold, the pedestrian re-identification algorithm is used for confirmation operation;

[0082] like Fig. 9As shown, in this embodiment, the OcSORT algorithm performs well in tracking accuracy and real-time performance, and is an ideal choice as a benchmark for tracking algorithms. The traditional tracking method based on Kalman filtering predicts the position of the target in the current frame through a motion model, and has a high tracking speed, but the tracking accuracy is relatively low.

[0083] like Fig.10 As shown, in this embodiment, although the scheme combining Kalman filtering with pedestrian re-identification can improve the tracking accuracy, since the pedestrian feature vector needs to be calculated for each frame, the tracking speed is significantly reduced, which is difficult to meet the actual application requirements.

[0084] Based on the above problems, this embodiment proposes an improved method to improve the tracking accuracy and maintain a high tracking speed while introducing a pedestrian re-identification algorithm. For specific solutions, see Fig.11 .

[0085] like Fig.11 As shown, in this embodiment, the improved tracking algorithm also includes the following specific steps:

[0086] S31, target detection;

[0087] S32, position prediction;

[0088] S33, determining whether the IoU confidence is greater than or equal to the confidence threshold;

[0089] S34, if yes, then perform data association;

[0090] S35, if not, perform feature extraction;

[0091] S36: Update the trajectory.

[0092] In this embodiment, the above method adopts a staged tracking strategy: for pedestrians with higher IoU confidence, Kalman filtering is directly used for tracking to ensure the tracking speed; for pedestrians with lower IoU confidence (the confidence may decrease due to occlusion or walking side by side), they are further confirmed in combination with the pedestrian re-identification algorithm, thereby significantly enhancing the tracking speed while improving the tracking accuracy.

[0093] In order to improve the matching accuracy of pedestrians with the same ID in the tracking and association process, especially in the case of occlusion, parallel pedestrians or other situations that may lead to reduced IoU confidence, this scheme introduces a matching strategy that combines the shape similarity of the pedestrian detection box with the IoU confidence. The detection boxes of the same pedestrian in adjacent frames are usually consistent in shape, so the matching accuracy can be further enhanced by shape features. Traditional IoU matching only focuses on the spatial overlap of targets between frames, but in complex scenes, such as occlusion or perspective changes, IoU alone cannot accurately associate the same pedestrian. To solve this problem, the shape similarity matching method combines multi-dimensional features such as the aspect ratio of the detection box and the shape changes of adjacent frames.

[0094] The calculation formula for shape similarity is as follows:

[0095]

[0096] Among them, (W D ,H D ) and (W T ,H T ) represent the width and height of the detection frame of the current frame and the detection frame of the tracked target in the previous frame, respectively. This formula limits the similarity value to the range of 0 to 1 in exponential form, and the closer the similarity value is to 1, the more similar the shapes of the two detection frames are.

[0097] The final correlation matching score calculation formula is as follows:

[0098] S=(1-λ)S IoU +λS shepe

[0099] Among them, S IoU represents the spatial overlap score based on IoU, S shape represents the shape similarity score, and λ is the adjustment coefficient, which is used to balance the weights of IoU and shape similarity in matching. This formula can still effectively associate the same target when the detection boxes have the same shape but small spatial overlap, thereby improving the matching accuracy in low IoU situations.

[0100] The current multi-target tracking method is mainly applicable to single-camera scenarios. The basic process is: if the detection frame in the current frame cannot be associated with the tracking frame in the previous frame, the detection frame is regarded as a new pedestrian and a new ID is assigned (usually the current Kalman filter maximum value plus 1. In addition, the tracking algorithm usually sets a parameter max_age, which indicates the maximum number of frames that the tracked object can remain active without a new detection frame match. For example, when a tracked object is not updated within 30 frames, the system will determine that the pedestrian has left the monitoring screen and delete its track from the memory. However, this method is prone to ID inconsistency when the pedestrian reappears after a short absence because the previous track has been deleted.

[0101] One solution is to set max_age to a larger value, but this method will cause the trajectories of pedestrians who have actually left the monitoring area and will not return to remain, reducing the efficiency of the algorithm and may cause memory overflow when processing a large number of trajectories.

[0102] To solve the problem of inconsistent IDs when pedestrians reappear after leaving for a short time, this patent proposes a new solution, such as Fig.12 As shown in the figure. When a pedestrian enters the camera's field of view for the first time, the pedestrian re-identification algorithm is used to calculate the pedestrian's feature vector, and the cosine similarity is calculated with the feature vector in the feature library. If the similarity exceeds the set threshold, the system determines that the pedestrian is a previously appeared pedestrian and assigns the original ID; otherwise, it is regarded as a new pedestrian, assigned a new ID, and the ID and its feature vector are stored in the feature library. This method can effectively improve the consistency and accuracy of the pedestrian ID in the system in multi-target tracking.

[0103] In summary, the improved tracking solution based on the OcSORT algorithm is implemented as follows Fig.13 shown.

[0104] The specific process is as follows: the system first obtains all pedestrian detection frames from the current frame and extracts the position information and confidence. Subsequently, the detection frames are preprocessed to obtain high-confidence detection frames, and the predicted positions of all current trackers are obtained. Next, the system performs motion data association on the high-confidence detection frames, using a comprehensive matching method of IoU confidence and shape similarity to improve the accuracy of association. For the detection frames that fail to match after motion association, the system further associates them through pedestrian re-identification technology to cope with matching requirements in complex situations such as occlusion. Based on the association results, the system updates the trajectories of the successfully matched detection frames and trackers; for the unmatched detection frames, they are regarded as new pedestrians, and the ID matching method is used to find out whether there are similar pedestrian features in the feature library. If the corresponding pedestrian is found, its original ID is initialized; if there is no match, a new ID is assigned to the pedestrian and its trajectory is initialized. Finally, the system outputs the ID data of all tracked objects in the current frame, including position, trajectory and status information.

[0105] like Fig.14 As shown, in this embodiment, the method of initializing the matching ID in a single-camera scenario is also applicable to a cross-camera scenario. After initializing the ID through the pedestrian feature library, each camera can independently execute the Kalman filter-based tracking algorithm frame by frame, while all cameras share the same pedestrian feature library. Compared with the traditional method of calculating the pedestrian feature vector frame by frame in a cross-camera scenario, this solution greatly improves the processing speed and ensures the consistency of the pedestrian ID in a multi-camera system.

[0106] Example 2

[0107] In this embodiment, the self-built campus data set is divided into a training set, a validation set, and a test set as shown in Table 3 for subsequent YOLOv8 training and verification. The test set is consistent with the multi-target tracking data set, which facilitates the subsequent direct verification of the multi-target tracking effect of the improved algorithm.

[0108] Table 3 Division of self-built campus dataset

[0109]

[0110] In order to comprehensively evaluate the performance of the proposed pedestrian detection network, ablation experiments and analysis were conducted. In the ablation experiment, the SCConv module, EMA module and their combination were gradually introduced into the original YOLOv8 detection network, and the detection accuracy of different models was compared to verify the specific contribution of each module to the improvement of algorithm performance. During the experiment, all training hyperparameters were kept consistent. The training parameters of each experiment were set to Epoch 50, Batch Size 32, and Adam was selected as the optimizer to ensure the consistency of experimental conditions. The comparison results of the ablation experiment are shown in Table 4.

[0111] Table 4. Comparison of ablation experiments

[0112]

[0113] The experimental results show that the YOLOv8s (SCConv) + EMA combination performs best in all experiments, with mAP50 reaching 0.937 and mAP50-95 reaching 0.554, which is one percentage point higher than the original solution. This shows that the combination of SCConv and EMA significantly improves the detection accuracy of the model. Although additional parameters and calculations are added, the improvement in accuracy has high practical value in practical applications.

[0114] 6.2 Improved tracking algorithm testing

[0115] In the improved tracking algorithm, the detection part uses the optimized detection network and its weight file. The pedestrian re-identification algorithm uses the SBS algorithm and performs 150 rounds of training on the self-built dataset. The division of the dataset is shown in Table 5.

[0116] Table 5 ReID division of self-built campus dataset

[0117]

[0118] After completing 150 rounds of training on the SBS person re-identification algorithm, the verification results on the test set are shown in Table 6.

[0119] Table 6. Training and testing results of the self-built campus dataset on SBS

[0120]

[0121] The overall improved test results based on the OcSORT tracking algorithm are shown in Table 7.

[0122] Table 7 Test results of the improved tracking algorithm on the self-built campus dataset

[0123]

[0124] The results show that the introduction of ID matching, shape similarity and ReID modules significantly improves the algorithm performance.

[0125] Since sequence 11 of the self-built dataset contains a sports scene on a football field, in which the players move with large amplitudes and sometimes leave for a short time and then reappear, the effectiveness of the improved scheme in solving the ID mismatch problem when pedestrians leave for a short time and then reappear can be verified, and compared with the currently recognized tracking algorithm with better results. The results are shown in Table 8.

[0126] Table 8 Test results of different tracking algorithms on sequence 11

[0127]

[0128] The test results show that the improved scheme is also effective in scenarios with large movements and pedestrians who leave for a short time and then reappear. The scheme can significantly improve the HOTA and IDF1 indicators, and its performance is better than all the compared algorithms. In addition, in terms of the MOTA indicator, the performance of the improved scheme is very close to the current optimal indicator, further verifying its superiority and stability in multi-target tracking tasks.

[0129] The improved tracking algorithm was compared with the currently recognized tracking algorithms with better results on the test set of the self-built dataset. The results are shown in Table 9.

[0130] Table 9 Test results of different tracking algorithms on self-built data sets

[0131]

[0132] Compared with the currently recognized tracking algorithms with better effects, the improved OcSORT algorithm shows better performance in all indicators, especially after the introduction of ReID, compared with DeepOcSORT, BoTSORT and StrongSORT, the frame rate of the algorithm has not been significantly reduced. The test results on the self-built data set show that the optimized tracking algorithm has significantly improved the accuracy of the original OcSORT algorithm, with only a slight decrease in frame rate, achieving a good balance between accuracy and speed.

[0133] Example 3

[0134] The implementation of the present invention is based on a single card 24GB NVIDIARTX3090 GPU, using PyTorch 2.2.2 version and CUDA12.1 version to implement the improved pedestrian detection and tracking algorithm. In order to verify the effectiveness of the improved algorithm, the present invention uses a self-built campus dataset to train and test the improved detection and tracking algorithm.

[0135] In this embodiment, for the improved pedestrian detection algorithm, a self-built campus data set is used for training to obtain a detection model weight file adapted to the campus environment. The training process makes full use of the data characteristics unique to the campus scene to improve the performance of the detection model in practical applications.

[0136] In the evaluation phase of the tracking algorithm of this embodiment, the above-mentioned trained detection weight file is used to ensure that the tracking system has accurate target recognition capabilities. At the same time, in order to optimize the tracking performance, the present invention sets several new hyperparameters in the tracking algorithm. Among them, the appearance threshold in the ID matching phase is set to 0.8, the IoU threshold in the data association phase is set to 0.3, and the weight λ of the appearance similarity is set to 0.3.

[0137] In this embodiment, after completing the hyperparameter setting, the improved detection and tracking algorithm can be experimentally verified. The method of the present invention demonstrates high detection accuracy and tracking accuracy in a complex campus environment, achieves stable real-time performance, and provides reliable technical support for the pedestrian monitoring system of the smart campus.

[0138] In summary, the present invention performs pedestrian detection in campus scenes based on improved YOLOv8 and OcSORT algorithms. Different from traditional pedestrian detection and multi-target tracking algorithms, the present invention optimizes the algorithm according to the specific needs of the campus environment, and improves the detection and tracking algorithms by introducing a self-built campus scene data set, thereby improving the accuracy of pedestrian detection and tracking. The present invention effectively solves the problems of insufficient detection accuracy and low tracking accuracy of existing algorithms in complex campus environments, and solves the problems of ID mismatch and cross-camera tracking when pedestrians reappear after leaving for a short time, thereby providing a reliable technical solution for efficient and real-time pedestrian monitoring in campus scenes.

[0139] Aiming at the actual needs of pedestrian detection and tracking in campus environments, the present invention independently constructs a pedestrian detection and multi-target tracking dataset for campus scenes. This dataset contains scene features and crowd distribution patterns that are unique to campus environments, providing reliable data support for model optimization, enabling the algorithm to better adapt to the application needs of campus scenes, thereby improving detection and tracking performance. This customized dataset greatly improves the application effect of the algorithm in campus scenes, and provides a solid data foundation for security and crowd behavior analysis in smart campuses, thereby improving its practicality and value.

[0140] The present invention optimizes the YOLOv8 target detection algorithm, significantly improves the detection accuracy by introducing an attention mechanism and replacing the original C2f module with a plug-and-play SCConv module.

[0141] In terms of multi-target tracking, this paper proposes a new tracking strategy, combined with a pedestrian re-identification method, to improve tracking accuracy while maintaining a high frame rate. The tracking accuracy is further improved by introducing the matching criteria of inter-frame detection box shape similarity and IoU confidence.

[0142] This paper proposes a new solution to the ID mismatch problem caused by pedestrians leaving the camera for a short time in a single-camera scenario: when pedestrians first enter the camera's field of view, pedestrian re-identification technology is used to match pedestrian IDs, thereby ensuring that even if pedestrians temporarily leave the field of view, their IDs can be correctly identified and matched. This solution is also applicable to multi-target tracking across cameras, significantly enhancing the adaptability and application value of the system in complex environments.

[0143] The present invention solves the technical problems existing in the prior art that the frame rate is significantly reduced, making it difficult to meet the real-time requirements, and the system has poor applicability in a campus environment.

[0144] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting and tracking pedestrians in a campus scene, characterized in that: The method comprises: S1. Build a campus scene dataset; S2. According to the campus scene data set, a pedestrian detection model is set and tuned to detect pedestrian targets, wherein the pedestrian detection module includes: an SCConv module, an EMA module, and a YOLOv8 module; S3. Based on the staged tracking strategy, the OcSORT algorithm is used to track the pedestrian target whose IoU confidence is greater than or equal to the preset threshold using Kalman filtering; for the pedestrian target whose IoU confidence is less than the preset threshold, a pedestrian re-identification algorithm is used to perform a confirmation operation; wherein, in the motion data association stage, association matching is performed by combining the inter-frame shape similarity and the IoU confidence; in the trajectory update stage, the pedestrian ID is initialized using the ID matching method based on pedestrian re-identification.

2. According to the method for detecting and tracking pedestrians in a campus scene according to claim 1, it is characterized in that: The S1 includes: S11, setting acquisition equipment, acquisition environment, and data format to acquire diversified image data; S12. Perform a multi-target tracking data set labeling operation on the diversified image data frame by frame to obtain a labeled data set, and divide the labeled data set into a training set and a test set.

3. The method for detecting and tracking pedestrians in a campus scene according to claim 1, characterized in that: In S2, the SCConv module is set and used to perform self-calibration convolution, divide the convolution filter to perform adaptive expansion of the receptive field, exchange information between channels, and obtain context information and local features; The EMA module is set and used to perform channel grouping operations and multi-scale parallel convolutions to enhance feature expression, which has high computational efficiency. The EMA module also includes: a local channel relationship preservation branch and a spatial dependency relationship capture branch to model multi-scale feature information; The YOLOv8 module is set and used to construct the pedestrian detection network, and multi-level feature layers P3, P4, and P5 are used for multi-scale target detection.

4. The method for detecting and tracking pedestrians in a campus scene according to claim 1, characterized in that: In S3, a matching operation is performed according to the shape similarity of the pedestrian detection frame and the IoU confidence, wherein the shape similarity includes: the aspect ratio of the detection frame and the shape change information of adjacent frames.

5. The method for detecting and tracking pedestrians in a campus scene according to claim 4, characterized in that: The shape similarity is calculated using the following logic: In the formula, (W D ,H D ) and (W T ,H T ) represent the width and height of the detection frame of the current frame and the detection frame of the tracked target in the previous frame, respectively.

6. The method for detecting and tracking pedestrians in a campus scene according to claim 4, characterized in that: The following logic is used to obtain the associated matching score, and the matching operation is performed accordingly: S=(1-λ)S IoU +λS shape In the formula, S IoU represents the spatial overlap score based on the IoU confidence, S shape represents the shape similarity score, and λ is the adjustment coefficient.

7. The method for detecting and tracking pedestrians in a campus scene according to claim 1, characterized in that: In S3, when the pedestrian target enters the camera field of view for the first time, the feature vector of the pedestrian target is calculated by the pedestrian re-identification algorithm, and the cosine similarity calculation is performed between the feature vector and the feature vector of the preset feature library.

8. The method for detecting and tracking pedestrians in a campus scene according to claim 7, characterized in that: When the cosine similarity is greater than or equal to a preset similarity threshold, the pedestrian target is determined to be an existing pedestrian, and an original ID is assigned; When the cosine similarity is less than the preset similarity threshold, the pedestrian target is determined to be a new pedestrian, a new ID is assigned, and the new ID and feature vector are stored in the preset feature library.

9. The method for detecting and tracking pedestrians in a campus scene according to claim 1, characterized in that: In S3, in the cross-camera situation, an ID initialization operation is performed on the pedestrian feature library, each camera independently executes a tracking algorithm based on Kalman filtering frame by frame, and all the cameras share the pedestrian feature library.

10. A campus scene pedestrian detection and tracking system, characterized in that: The system comprises: Campus dataset construction module, used to construct campus scene dataset; A pedestrian detection module, used to set and tune a pedestrian detection model according to the campus scene dataset, so as to detect pedestrian targets. The pedestrian detection module includes: an SCConv module, an EMA module and a YOLOv8 module. The pedestrian detection module is connected to the campus dataset construction module; The phased tracking module is used to track the pedestrian target whose IoU confidence is greater than or equal to the preset threshold by using the OcSORT algorithm based on the phased tracking strategy; the pedestrian target whose IoU confidence is less than the preset threshold is confirmed by using the pedestrian re-identification algorithm, and in the motion data association stage, the inter-frame shape similarity and the IoU confidence are combined for association matching; in the trajectory update stage, the pedestrian ID is initialized by using the ID matching method based on pedestrian re-identification. The phased tracking module is connected to the pedestrian detection module.

Citation Information

Patent Citations

  • Lightweight pedestrian detection method based on yov5 algorithm

    CN116740532A

  • Road pedestrian and vehicle detection algorithm based on improved YOLOv5s

    CN116740679A