Multi-target detection tracking method, device and equipment and readable storage medium
By introducing a multi-objective tracking and detection model of multi-objective classification detection network, target re-identification network and data-related network, the problems of high computing complexity and reduced detection accuracy in the prior art are solved, efficient and real-time multi-objective detection and tracking are achieved, and the generalization ability of the model is enhanced.
Patent Information
- Application Number
- CN202411477856.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2025-06-06
AI Technical Summary
The existing multi-object detection and tracking technologies for pedestrian and vehicle are highly complex in processing high-resolution images or complex scenes, which is difficult to meet real-time requirements. The detection and tracking accuracy decreases under target occlusion, lighting changes, and viewing angle changes, and lacks generalization capabilities.
A multi-objective tracking and detection model consisting of a multi-objective classification detection network, a target re-identification network and a data correlation network is adopted. Through the two-stage target detection and tracking process, candidate target areas are quickly screened and fine classification and position adjustment are carried out, robust features are extracted and feature matching is performed, and combined with space-time correlation processing is used to determine the target action trajectory.
It improves the real-time and accuracy of detection and tracking, enhances the generalization ability of the model, can adapt to the detection and tracking needs of different scenarios and targets, and solves the tracking failure caused by target occlusion and perspective changes.
Smart Images

Figure CN120107307A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of visual processing, and more specifically, to a multi-target detection and tracking method, device, equipment and readable storage medium. Background Art
[0002] As a key technology in the field of computer vision and pattern recognition, the core of pedestrian and vehicle multi-target detection and tracking technology is to extract the features of multiple targets of multiple categories (such as pedestrians, vehicles, etc.) from the input original image frame, and accurately obtain the location and unique identifier (ID) information of these targets in the image frame. This technology provides a reliable data foundation for upper-level business applications such as pedestrian counting, area intrusion detection, and line crossing counting. However, the following problems usually exist in the existing processing methods:
[0003] First, when processing high-resolution images or complex scenes, the computational complexity of traditional image processing techniques and feature-based methods increases significantly, resulting in a decrease in processing speed and difficulty in meeting real-time requirements. This limits the application of these technologies in situations that require fast response.
[0004] Secondly, the detection and tracking accuracy of these methods is often seriously affected in complex situations such as target occlusion, lighting changes, and perspective changes. For example, when the target is occluded by other objects, traditional algorithms may not be able to accurately track the target; and in the case of lighting changes or perspective changes, the stability of the algorithm will also be greatly reduced.
[0005] Finally, for the detection and tracking of different scenes and targets, these methods usually require parameter adjustment or model training for specific problems, and lack sufficient generalization ability. This means that every time a new scene or target is encountered, the algorithm parameters need to be readjusted or the model needs to be retrained, which undoubtedly increases the complexity and cost of the application.
[0006] In summary, although the existing pedestrian and vehicle multi-target detection and tracking technologies provide a reliable data foundation for upper-level business applications, they still have significant deficiencies in computational complexity, detection and tracking accuracy, and generalization ability. Summary of the invention
[0007] In view of this, the present application provides a multi-target detection and tracking method, apparatus, device and readable storage medium. By introducing a multi-target tracking detection model composed of a multi-target classification detection network, a target re-identification network and a data association network, the real-time and accuracy of detection and tracking are improved, and the generalization ability of the model is enhanced.
[0008] A multi-target detection and tracking method, comprising:
[0009] Acquire video frame data captured by a camera device, and input the video frame data into a multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network, and a data association network;
[0010] Performing target detection and positioning and information normalization processing on the video frame data through the multi-target classification detection network to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target;
[0011] Performing re-identification feature extraction and pooling alignment processing on the video frame processing data through the target re-identification network, determining the same target in each of the video frames, and generating a target re-identification data set containing the same target association information;
[0012] The target information file and the target re-identification data set are subjected to spatiotemporal association and matching processing by combining the preset ID data with the data association network to determine the target action trajectory corresponding to each of the IDs.
[0013] Optionally, the target re-identification network includes a feature extraction network, a pooling alignment network and a similarity calculation network;
[0014] The target re-identification network performs re-identification feature extraction and pooling alignment processing on the video frame processing data, determines the same target in each of the video frames, and generates a target re-identification data set containing the same target association information, including:
[0015] The feature extraction network performs re-identification feature extraction on the video frame processing data to obtain a video frame feature map;
[0016] The pooling alignment network performs fine-grained component alignment and overlapping feature fusion on the video frame feature map to generate fine-grained segmentation feature data;
[0017] The similarity calculation network performs cosine similarity calculation on the fine-grained division feature data, determines the same target in each of the video frames based on the similarity calculation result, and generates a target re-identification data set containing the same target association information.
[0018] Optionally, the pooling alignment network includes a feature vertical division module and a feature alignment fusion module;
[0019] The pooling alignment network performs fine-grained component alignment and overlapping feature fusion on the video frame feature map to generate fine-grained segmentation feature data, including:
[0020] The feature vertical division module performs feature division on the video frame feature map by using a feature division method based on the vertical direction to form a plurality of vertical feature subsets;
[0021] The feature alignment and fusion module uses a preset alignment algorithm to align the vertical feature subsets and fuse overlapping features to generate fine-grained segmentation feature data.
[0022] Optionally, the training process of the target re-identification network includes:
[0023] Acquire sample video frame processing data and a sample target re-identification data set containing the same target association information in the sample video frame processing data;
[0024] Inputting the sample video frame processed data into a preset initial target re-identification network to obtain the same target in each sample video frame output by the initial target re-identification network;
[0025] The initial target re-identification network is trained with the goal that the loss function of the same target in each sample video frame output by the model and the same target association information recorded in the sample target re-identification dataset meets the preset deviation;
[0026] The trained initial target re-identification network is used as the target re-identification network.
[0027] Optionally, the multi-target classification detection network includes a first feature extraction network, a second feature extraction network and a feature channel splicing network;
[0028] The multi-target classification detection network performs target detection positioning and information normalization processing on the video frame data to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target, including:
[0029] The first feature extraction network uses the small feature map to detect the large target and obtains a first predicted positioning result;
[0030] The first feature extraction network uses the large feature map to detect the small target and obtains a second predicted positioning result;
[0031] The feature channel stitching network performs feature fusion based on the first predicted positioning result and the second predicted positioning result, and performs information normalization processing to generate video frame processing data marked with target bounding boxes and target categories, and a target information file that records the position and category information of each target.
[0032] Optionally, the training process of the multi-target classification detection network includes:
[0033] Acquire a multi-scale sample image, wherein the multi-scale sample image is annotated with a target bounding box and a target category of a sample target present therein;
[0034] Inputting the multi-scale sample image into a preset initial multi-target classification detection network to obtain sample image processing data output by the initial multi-target classification detection network and annotated with a target bounding box and a target category;
[0035] Training the initial multi-target classification detection network with the goal of making the annotations in the sample image processing data consistent with the annotations in the multi-scale sample image;
[0036] When the initial multi-target classification detection network meets the preset training conditions, the trained initial multi-target classification detection network is used as the multi-target classification detection network.
[0037] Optionally, the data association network includes a feature matching module, a spatiotemporal association module and a trajectory construction module;
[0038] The data association network combines the preset ID data to perform spatiotemporal association and matching processing on the target information file and the target re-identification data set to determine the target action trajectory corresponding to each ID, including:
[0039] The feature matching module compares the features in the target information file and the target re-identification data set, and determines a feature set corresponding to each ID match in the preset ID data;
[0040] The spatiotemporal association module performs spatiotemporal association verification on the matching results to generate a verification matching result that satisfies the spatiotemporal constraints;
[0041] The trajectory construction module connects the verification matching results in series to construct a target action trajectory corresponding to each of the IDs.
[0042] A multi-target detection and tracking device, comprising:
[0043] A video frame acquisition unit, used to acquire video frame data captured by a camera device, and input the video frame data into a multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network, and a data association network;
[0044] A detection and positioning unit, configured to perform target detection and positioning and information normalization processing on the video frame data through the multi-target classification detection network, and generate video frame processing data annotated with target bounding boxes and target categories, and a target information file recording information on the positions and target categories of each target;
[0045] A re-identification extraction unit, configured to perform re-identification feature extraction and pooling alignment processing on the video frame processing data through the target re-identification network, determine the same target in each of the video frames, and generate a target re-identification data set containing the same target association information;
[0046] The association matching unit is used to perform spatiotemporal association and matching processing on the target information file and the target re-identification data set through the data association network in combination with preset ID data, so as to determine the target action trajectory corresponding to each ID.
[0047] A multi-target detection and tracking device, comprising a memory and a processor;
[0048] The memory is used to store programs;
[0049] The processor is used to execute the program to implement each step of the multi-target detection and tracking method as described in any one of the above items.
[0050] A readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, each step of the multi-target detection and tracking method as described in any one of the above items is implemented.
[0051] It can be seen from the above technical solutions that the embodiment of the present application provides a multi-target detection and tracking method, device, equipment and readable storage medium. First, the video frame data captured by the camera device is obtained, and the video frame data is input into the multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network and a data association network. After that, the video frame data is subjected to target detection positioning and information normalization processing by the multi-target classification detection network to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target. The video frame processing data is subjected to re-identification feature extraction and pooling alignment processing by the target re-identification network to determine the same target in each of the video frames, and to generate a target re-identification data set containing the same target association information. The target information file and the target re-identification data set are subjected to spatiotemporal association and matching processing by the data association network in combination with the preset ID data to determine the target action trajectory corresponding to each of the IDs.
[0052] This application introduces a multi-target tracking detection model composed of a multi-target classification detection network, a target re-identification network, and a data association network. Based on a two-stage target detection and tracking process, the candidate target areas are quickly screened out in the first stage to reduce the amount of calculation; the candidate areas are finely classified and position adjusted in the second stage to improve the detection accuracy. The target re-identification network can extract the robust features of the target and match the features between different frames, thereby solving the problem of tracking failure caused by target occlusion and perspective changes. At the same time, it can accurately identify the same target between different frames to achieve long-term, cross-perspective tracking. At the same time, the algorithm and model proposed in this application have strong generalization capabilities and can adapt to the detection and tracking needs of different scenes and targets. By adjusting the model parameters and training data, it can be easily applied to new scenes and tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0054] Figure 1 A flowchart of a multi-target detection and tracking method disclosed in an embodiment of the present application;
[0055] Figure 2 A schematic diagram of a two-stage target detection and tracking process disclosed in an embodiment of the present application;
[0056] Figure 3 A schematic diagram of the YOLO target detection algorithm and the corresponding data set disclosed in the embodiment of the present application;
[0057] Figure 4 A schematic diagram of a target re-identification network processing flow disclosed in an embodiment of the present application;
[0058] Figure 5 A schematic diagram of a pooling alignment method disclosed in an embodiment of the present application;
[0059] Figure 6 A schematic diagram of a multi-target detection and tracking device disclosed in an embodiment of the present application;
[0060] Figure 7 This is a hardware structure block diagram of a multi-target detection and tracking device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0062] The present application can be used in many general or special computing device environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multi-processor devices, distributed computing environments including any of the above devices or equipment, etc.
[0063] An embodiment of the present application provides a multi-target detection and tracking method, which can be applied to various systems or platforms involving multi-target detection and tracking, and can also be applied to various computer terminals or smart terminals. The execution subject can be a processor or server of the computer terminal or smart terminal.
[0064] Next, the present application scheme is introduced. The present application proposes the following technical scheme, please see below for details.
[0065] Figure 1 A flowchart of a multi-target detection and tracking method disclosed in an embodiment of the present application.
[0066] Figure 2 A schematic diagram of a two-stage target detection and tracking process disclosed in an embodiment of the present application.
[0067] like Figure 1 and Figure 2 As shown, the method may include:
[0068] Step S1, obtaining video frame data captured by a camera device, and inputting the video frame data into a multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network, and a data association network.
[0069] Specifically, the video frame data captured by the camera device is obtained. The captured data may include multiple categories and multiple targets, such as outdoor road monitoring data sets and indoor and outdoor data sets, which may include multiple target categories such as pedestrians, bicycles, electric vehicles (motorcycles), cars, trucks, buses, tricycles, etc. These data are input into the multi-target tracking detection model, which consists of three parts: a multi-target classification detection network, a target re-identification network, and a data association network. Before inputting the model, the video frame data can be preprocessed, such as denoising, contrast enhancement, etc., to improve the detection accuracy of the model.
[0070] Step S2: performing target detection positioning and information normalization processing on the video frame data through the multi-target classification detection network to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and category information of each target.
[0071] Specifically, Figure 2 As shown in FIG. 1 , the video frame data is processed by a multi-target classification detection network to generate video frame processing data annotated with target bounding boxes and target categories, and a target information file recording the position and category information of each target is generated. The process of target detection and positioning is to use a multi-target classification detection network to identify and locate the target in the video frame. Advanced detection algorithms such as YOLO and SSD can be used to improve the detection speed and accuracy. Information normalization processing is to normalize the detected target information, including target position, size, category, etc. Normalization is helpful for subsequent target re-identification and data association processing. Appropriate normalization methods and parameters can be selected according to specific application scenarios and requirements. The processed target information (including target position, category, etc.) is recorded in the target information file to ensure that the format and content of the target information file match the subsequent processing flow. Standardized file formats (such as JSON, XML, etc.) can be used to facilitate data storage and sharing.
[0072] like Figure 3 As shown in the figure, the YOLO target detection algorithm and the corresponding data set are used as the final implementation solution, in which all targets in a picture correspond to a text TXT file, which normalizes the position and category information of each target in the original image and records them accordingly.
[0073] The training process of the multi-target classification detection network includes:
[0074] ① Obtain a multi-scale sample image, wherein the multi-scale sample image is annotated with a target bounding box and a target category of a sample target present therein;
[0075] ② Inputting the multi-scale sample image into a preset initial multi-target classification detection network to obtain sample image processing data output by the initial multi-target classification detection network and annotated with target bounding boxes and target categories;
[0076] ③ With the goal of making the annotations in the sample image processing data consistent with the annotations in the multi-scale sample image, training the initial multi-target classification detection network;
[0077] ④ When the initial multi-target classification detection network meets the preset training conditions, the trained initial multi-target classification detection network is used as the multi-target classification detection network.
[0078] The training process of the multi-target classification detection network includes collecting and labeling multi-scale sample images, inputting these images into the preset initial network, and adjusting the network parameters through iterative optimization to minimize the difference between the predicted annotations and the actual annotations. When the preset training conditions are met, the training is completed and a finalized multi-target classification detection network is obtained. The network can accurately predict the location and category of the target in the image, providing high-precision and stable target detection support for practical applications.
[0079] Furthermore, the multi-target classification detection network includes a first feature extraction network, a second feature extraction network and a feature channel splicing network.
[0080] The multi-target classification detection network performs target detection positioning and information normalization processing on the video frame data to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target, including:
[0081] ① The first feature extraction network uses a small feature map to detect a large target and obtains a first prediction positioning result;
[0082] ② The first feature extraction network uses the large feature map to detect small targets and obtains a second prediction positioning result;
[0083] ③ The feature channel stitching network performs feature fusion based on the first predicted positioning result and the second predicted positioning result, and performs information normalization processing to generate video frame processing data marked with target bounding boxes and target categories, as well as a target information file that records the position and category information of each target.
[0084] Specifically, the multi-target classification detection network integrates the first feature extraction network, the second feature extraction network and the feature channel splicing network. This network can efficiently perform deep processing on video frame data to achieve accurate detection and classification of targets.
[0085] When processing video frame data, the network first performs target detection and positioning and information normalization operations. This process aims to accurately capture the target object from the complex video frame and make a preliminary judgment and normalization of its position, size and category.
[0086] The first feature extraction network uses a smaller feature map, through fine feature extraction and analysis, to quickly lock on to large targets in the video frame and generate the first prediction and positioning result. When facing small targets, the second feature extraction network uses the large feature map mode and richer detail information to capture the features of small targets, thereby obtaining the second prediction and positioning result.
[0087] The feature channel stitching network plays a crucial role. It is responsible for fusing the predicted positioning results output by the first feature extraction network and the second feature extraction network. This process not only combines the advantages of both in detecting targets of different sizes, but also further improves the accuracy and stability of the detection results through information normalization. Finally, the feature channel stitching network generates video frame processing data annotated with target bounding boxes and target categories, as well as target information files that record the location and category information of each target in detail.
[0088] Step S3: performing re-identification feature extraction and pooling alignment processing on the video frame processing data through the target re-identification network, determining the same target in each of the video frames, and generating a target re-identification data set containing the same target association information.
[0089] Specifically, the target re-identification network is used to perform deep processing on the video frame processing data, including re-identification feature extraction and pooling alignment processing, aiming to accurately identify and track each target from the video frame, especially the same target across frames. Through this processing, it is possible to determine which targets are the same in different video frames, and generate a special target re-identification dataset (ReID dataset) based on this.
[0090] The structure of the ReID dataset is significantly different from that of conventional object detection datasets. It is usually divided into three main parts: training set, query set, and gallery set. This division helps us evaluate the performance of the re-identification model more comprehensively.
[0091] Training set: The training set contains a certain number of target IDs, each of which represents all the data of different individuals (such as pedestrians) under a single camera.
[0092] Query set: The query set contains all target IDs except those in the training set. For each target ID, at least one image taken by a single camera is selected as query data. These query data will be used in the model evaluation phase to test whether the model can accurately identify targets from a large number of candidate images.
[0093] Gallery: The target ID of the gallery is the same as the query set, but the data under each target ID is all the data except the query data. These data constitute the candidate image library for model evaluation, which is used to test the model's re-identification ability in complex scenes.
[0094] It should be noted that the target ID of the training set is mutually exclusive with the target ID of the query set and the gallery set, which ensures the fairness and accuracy of the model evaluation.
[0095] The training process of the target re-identification network includes:
[0096] ① Obtaining sample video frame processing data and a sample target re-identification data set containing the same target association information in the sample video frame processing data;
[0097] ② Inputting the sample video frame processed data into a preset initial target re-identification network to obtain the same target in each sample video frame output by the initial target re-identification network;
[0098] ③ The initial target re-identification network is trained with the goal that the loss function of the same target in each sample video frame output by the model and the same target association information recorded in the sample target re-identification dataset meets the preset deviation;
[0099] ④ Use the trained initial target re-identification network as the target re-identification network.
[0100] The training process of the target re-identification network involves obtaining sample video frame processing data and the corresponding sample target re-identification dataset, inputting the sample data into a preset initial network, and optimizing the network accordingly by calculating the loss function of the association information between the same target output by the network and the same target recorded in the dataset until the preset training conditions are met. Finally, the trained network is finalized as a target re-identification network.
[0101] In the design of the target re-identification network, in order to improve performance and ensure the stability of feature extraction to meet the needs of accurate feature extraction of each target, the network's image input size was optimized, and the common 256*256 was changed to 32*96 (width*height). This adjustment significantly improved the model's reasoning performance while maintaining the target feature scale.
[0102] In the training phase, the cross entropy loss function is used to treat the cropped and uniformly scaled target image as a classification problem. By increasing the difference between classes, the feature extraction network is able to learn feature information that effectively distinguishes different categories. Specifically, the target image is divided into multiple parts, each part is pooled and feature vectors are extracted, and these feature vectors are used for subsequent feature matching. The corresponding ID loss function measures the model's ability to distinguish target categories by calculating the probability value between the feature vector of each part and the weight of the fully connected layer.
[0103] Step S4: performing spatiotemporal association and matching processing on the target information file and the target re-identification data set through the data association network in combination with preset ID data, and determining the target action trajectory corresponding to each ID.
[0104] Specifically, the data association network combines the preset ID data, processes the target information file and the target re-identification data set, and accurately determines the target action trajectory corresponding to each ID through spatiotemporal association and matching. The data association network includes a feature matching module, a spatiotemporal association module, and a trajectory construction module.
[0105] The data association network combines the preset ID data to perform spatiotemporal association and matching processing on the target information file and the target re-identification data set to determine the target action trajectory corresponding to each ID, including:
[0106] ① The feature matching module compares the features in the target information file and the target re-identification data set, and determines a feature set corresponding to each ID in the preset ID data;
[0107] ② The spatiotemporal correlation module performs spatiotemporal correlation verification on the matching results to generate a verification matching result that satisfies the spatiotemporal constraints;
[0108] ③ The trajectory construction module connects the verification matching results in series to construct the target action trajectory corresponding to each ID.
[0109] The feature matching module is used to compare the feature information in the target information file and the target re-identification data set. It uses advanced feature matching algorithms to accurately find the feature set that matches each ID in the preset ID data. Based on feature matching, the spatiotemporal association module further verifies the matching results. It uses spatiotemporal constraints to screen and filter the matching results to ensure that only matching results that meet spatiotemporal consistency can be retained. This step effectively reduces the impact of mismatches and noise data and improves the accuracy of the target's action trajectory. The trajectory construction module connects the matching results that have been verified by spatiotemporal association in series to construct the target action trajectory corresponding to each ID. These trajectories record the movement of the target in time and space.
[0110] It can be seen from the above technical solutions that the embodiment of the present application provides a multi-target detection and tracking method, device, equipment and readable storage medium. First, the video frame data captured by the camera device is obtained, and the video frame data is input into the multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network and a data association network. After that, the video frame data is subjected to target detection positioning and information normalization processing by the multi-target classification detection network to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target. The video frame processing data is subjected to re-identification feature extraction and pooling alignment processing by the target re-identification network to determine the same target in each of the video frames, and to generate a target re-identification data set containing the same target association information. The target information file and the target re-identification data set are subjected to spatiotemporal association and matching processing by the data association network in combination with the preset ID data to determine the target action trajectory corresponding to each of the IDs.
[0111] This application introduces a multi-target tracking detection model composed of a multi-target classification detection network, a target re-identification network, and a data association network. Based on a two-stage target detection and tracking process, the candidate target areas are quickly screened out in the first stage to reduce the amount of calculation; the candidate areas are finely classified and position adjusted in the second stage to improve the detection accuracy. The target re-identification network can extract the robust features of the target and match the features between different frames, thereby solving the problem of tracking failure caused by target occlusion and perspective changes. At the same time, it can accurately identify the same target between different frames to achieve long-term, cross-perspective tracking. At the same time, the algorithm and model proposed in this application have strong generalization capabilities and can adapt to the detection and tracking needs of different scenes and targets. By adjusting the model parameters and training data, it can be easily applied to new scenes and tasks.
[0112] In some embodiments of the present application, the target re-identification network includes a feature extraction network, a pooling alignment network and a similarity calculation network.
[0113] On the basis of the above, step S3, performing re-identification feature extraction and pooling alignment processing on the video frame processing data through the target re-identification network, determining the same target in each of the video frames, and generating a target re-identification data set containing the same target association information, may specifically include:
[0114] ① The feature extraction network performs re-identification feature extraction on the video frame processing data to obtain a video frame feature map;
[0115] ② The pooling alignment network performs fine-grained component alignment and overlapping feature fusion on the video frame feature map to generate fine-grained segmentation feature data;
[0116] ③ The similarity calculation network performs cosine similarity calculation on the fine-grained division feature data, and determines the same target in each of the video frames based on the similarity calculation results, and generates a target re-identification data set containing the same target association information.
[0117] Specifically, Figure 4 As shown in the figure, the video frame processing data first enters the feature extraction network, which uses its powerful feature extraction capability to extract key re-identification features from the video frame and then generates a video frame feature map.
[0118] Subsequently, these video frame feature maps are fed into the pooling alignment network. In this network, the feature maps undergo a process of fine-grained component alignment and overlapping feature fusion. Specifically, the pooling alignment network draws on the feature alignment method of the PCB network and subdivides the feature map into multiple parts by dividing the features based on the vertical direction. Figure 5 As shown in the figure, taking the human body as an example, the pedestrian target is visually divided into six key parts: head, shoulders, waist, hips, legs, and feet. This fine-grained division not only helps to capture the local features of pedestrian targets, but also effectively addresses the problem of feature alignment in ReID.
[0119] Finally, the fine-grained feature data is fed into a similarity calculation network. The network uses the cosine similarity algorithm to calculate the similarity of the feature data. Based on the calculation results, the same target in each video frame can be accurately identified, and a target re-identification dataset containing the same target association information can be generated accordingly.
[0120] Wherein, the pooling alignment network includes a feature vertical division module and a feature alignment fusion module;
[0121] The pooling alignment network performs fine-grained component alignment and overlapping feature fusion on the video frame feature map to generate fine-grained segmentation feature data, including:
[0122] The feature vertical division module performs feature division on the video frame feature map by using a feature division method based on the vertical direction to form a plurality of vertical feature subsets;
[0123] The feature alignment and fusion module uses a preset alignment algorithm to align the vertical feature subsets and fuse overlapping features to generate fine-grained segmentation feature data.
[0124] Specifically, the feature vertical partitioning module is responsible for partitioning the input video frame feature map based on the vertical direction. It uses a sophisticated algorithm to partition the feature map into multiple vertical feature subsets, each of which represents a vertical region or component in the feature map.
[0125] The feature alignment and fusion module aligns and fuses the vertical feature subsets output by the feature vertical partitioning module. It uses a preset alignment algorithm to ensure that each vertical feature subset remains consistent in spatial position, thereby facilitating subsequent feature fusion. At the same time, the module also has the ability to fuse overlapping features, and can effectively fuse adjacent or similar feature subsets to generate more representative fine-grained partition feature data.
[0126] When the video frame feature map enters the pooling alignment network, the feature vertical division module first performs feature division based on the vertical direction. The purpose of this step is to divide the feature map into smaller, easier-to-process parts or areas. Next, the feature alignment and fusion module uses a preset alignment algorithm to align the divided vertical feature subsets to ensure the consistency of each feature subset in spatial position. Finally, the feature alignment and fusion module also performs overlapping feature fusion on the aligned vertical feature subsets. By fusing adjacent or similar feature subsets, the module can generate more representative fine-grained division feature data. These data not only contain the key information in the original feature map, but also improve the distinguishability and robustness of the features through fine-grained division and feature fusion.
[0127] In this application, features are usually represented in the form of vectors, so when analyzing the similarity between two feature vectors, cosine similarity is often used to represent it. , and its cosine similarity is calculated as:
[0128]
[0129] Before calculating the cosine similarity, we usually do the vector Normalize the norm, so , .
[0130] Therefore, the calculation formula for cosine distance is:
[0131] Since the value range of cosine similarity is , so the value range of cosine distance is When the cosine similarity is 1, that is, the cosine distance is 0, it indicates that the two vectors are exactly the same; when the cosine similarity is 0, that is, the cosine distance is 1, it indicates that the two vectors are orthogonal and independent; when the cosine similarity is -1, that is, the cosine distance is 2, it indicates that the two vectors are completely opposite.
[0132] A multi-target detection and tracking device provided in an embodiment of the present application is described below. The multi-target detection and tracking device described below and the multi-target detection and tracking method described above can refer to each other.
[0133] See also Figure 6 , Figure 6 A schematic diagram of a multi-target detection and tracking device disclosed in an embodiment of the present application.
[0134] like Figure 6 As shown, the multi-target detection and tracking device may include:
[0135] The video frame acquisition unit 110 is used to acquire the video frame data captured by the camera device and input the video frame data into a multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network, and a data association network;
[0136] The detection and positioning unit 120 is used to perform target detection and positioning and information normalization processing on the video frame data through the multi-target classification detection network, and generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target;
[0137] A re-identification extraction unit 130 is used to perform re-identification feature extraction and pooling alignment processing on the video frame processing data through the target re-identification network, determine the same target in each of the video frames, and generate a target re-identification data set containing the same target association information;
[0138] The association matching unit 140 is used to perform spatiotemporal association and matching processing on the target information file and the target re-identification data set through the data association network in combination with preset ID data, so as to determine the target action trajectory corresponding to each ID.
[0139] It can be seen from the above technical solutions that the embodiment of the present application provides a multi-target detection and tracking method, device, equipment and readable storage medium. First, the video frame data captured by the camera device is obtained, and the video frame data is input into the multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network and a data association network. After that, the video frame data is subjected to target detection positioning and information normalization processing by the multi-target classification detection network to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target. The video frame processing data is subjected to re-identification feature extraction and pooling alignment processing by the target re-identification network to determine the same target in each of the video frames, and to generate a target re-identification data set containing the same target association information. The target information file and the target re-identification data set are subjected to spatiotemporal association and matching processing by the data association network in combination with the preset ID data to determine the target action trajectory corresponding to each of the IDs.
[0140] This application introduces a multi-target tracking detection model composed of a multi-target classification detection network, a target re-identification network, and a data association network. Based on a two-stage target detection and tracking process, the candidate target areas are quickly screened out in the first stage to reduce the amount of calculation; the candidate areas are finely classified and position adjusted in the second stage to improve the detection accuracy. The target re-identification network can extract the robust features of the target and match the features between different frames, thereby solving the problem of tracking failure caused by target occlusion and perspective changes. At the same time, it can accurately identify the same target between different frames to achieve long-term, cross-perspective tracking. At the same time, the algorithm and model proposed in this application have strong generalization capabilities and can adapt to the detection and tracking needs of different scenes and targets. By adjusting the model parameters and training data, it can be easily applied to new scenes and tasks.
[0141] Optionally, the target re-identification network includes a feature extraction network, a pooling alignment network and a similarity calculation network;
[0142] The target re-identification network performs re-identification feature extraction and pooling alignment processing on the video frame processing data, determines the same target in each of the video frames, and generates a target re-identification data set containing the same target association information, including:
[0143] The feature extraction network performs re-identification feature extraction on the video frame processing data to obtain a video frame feature map;
[0144] The pooling alignment network performs fine-grained component alignment and overlapping feature fusion on the video frame feature map to generate fine-grained segmentation feature data;
[0145] The similarity calculation network performs cosine similarity calculation on the fine-grained division feature data, determines the same target in each of the video frames based on the similarity calculation result, and generates a target re-identification data set containing the same target association information.
[0146] Optionally, the pooling alignment network includes a feature vertical division module and a feature alignment fusion module;
[0147] The pooling alignment network performs fine-grained component alignment and overlapping feature fusion on the video frame feature map to generate fine-grained segmentation feature data, including:
[0148] The feature vertical division module performs feature division on the video frame feature map by using a feature division method based on the vertical direction to form a plurality of vertical feature subsets;
[0149] The feature alignment and fusion module uses a preset alignment algorithm to align the vertical feature subsets and fuse overlapping features to generate fine-grained segmentation feature data.
[0150] Optionally, the multi-target detection and tracking device may further include a first model training unit, wherein the process of the first model training unit training the target re-identification network includes:
[0151] Acquire sample video frame processing data and a sample target re-identification data set containing the same target association information in the sample video frame processing data;
[0152] Inputting the sample video frame processed data into a preset initial target re-identification network to obtain the same target in each sample video frame output by the initial target re-identification network;
[0153] The initial target re-identification network is trained with the goal that the loss function of the same target in each sample video frame output by the model and the same target association information recorded in the sample target re-identification dataset meets the preset deviation;
[0154] The trained initial target re-identification network is used as the target re-identification network.
[0155] Optionally, the multi-target classification detection network includes a first feature extraction network, a second feature extraction network and a feature channel splicing network;
[0156] The multi-target classification detection network performs target detection positioning and information normalization processing on the video frame data to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target, including:
[0157] The first feature extraction network uses the small feature map to detect the large target and obtains a first predicted positioning result;
[0158] The first feature extraction network uses the large feature map to detect the small target and obtains a second predicted positioning result;
[0159] The feature channel stitching network performs feature fusion based on the first predicted positioning result and the second predicted positioning result, and performs information normalization processing to generate video frame processing data marked with target bounding boxes and target categories, and a target information file that records the position and category information of each target.
[0160] Optionally, the multi-target detection and tracking device may further include a second model training unit, and the process of the second model training unit training the multi-target classification detection network includes:
[0161] Acquire a multi-scale sample image, wherein the multi-scale sample image is annotated with a target bounding box and a target category of a sample target present therein;
[0162] Inputting the multi-scale sample image into a preset initial multi-target classification detection network to obtain sample image processing data output by the initial multi-target classification detection network and annotated with a target bounding box and a target category;
[0163] Training the initial multi-target classification detection network with the goal of making the annotations in the sample image processing data consistent with the annotations in the multi-scale sample image;
[0164] When the initial multi-target classification detection network meets the preset training conditions, the trained initial multi-target classification detection network is used as the multi-target classification detection network.
[0165] Optionally, the data association network includes a feature matching module, a spatiotemporal association module and a trajectory construction module;
[0166] The data association network combines the preset ID data to perform spatiotemporal association and matching processing on the target information file and the target re-identification data set to determine the target action trajectory corresponding to each ID, including:
[0167] The feature matching module compares the features in the target information file and the target re-identification data set, and determines a feature set corresponding to each ID match in the preset ID data;
[0168] The spatiotemporal association module performs spatiotemporal association verification on the matching results to generate a verification matching result that satisfies the spatiotemporal constraints;
[0169] The trajectory construction module connects the verification matching results in series to construct a target action trajectory corresponding to each of the IDs.
[0170] The multi-target detection and tracking apparatus provided in the embodiments of the present application can be applied to multi-target detection and tracking equipment. Figure 7 The hardware structure diagram of the multi-target detection and tracking device is shown in FIG. Figure 7 ,The hardware structure of the multi-target detection and tracking device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0171] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 communicate with each other through the communication bus 4;
[0172] The processor 1 may be a central processing unit CPU, or an application-specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0173] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), etc., such as at least one disk memory;
[0174] The memory stores a program, and the processor can call the program stored in the memory, wherein the program is used to:
[0175] Acquire video frame data captured by a camera device, and input the video frame data into a multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network, and a data association network;
[0176] Performing target detection and positioning and information normalization processing on the video frame data through the multi-target classification detection network to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target;
[0177] Performing re-identification feature extraction and pooling alignment processing on the video frame processing data through the target re-identification network, determining the same target in each of the video frames, and generating a target re-identification data set containing the same target association information;
[0178] The target information file and the target re-identification data set are subjected to spatiotemporal association and matching processing by combining the preset ID data with the data association network to determine the target action trajectory corresponding to each of the IDs.
[0179] Optionally, the detailed functions and extended functions of the program may refer to the above description.
[0180] The embodiment of the present application further provides a readable storage medium, which may store a program suitable for execution by a processor, wherein the program is used to:
[0181] Acquire video frame data captured by a camera device, and input the video frame data into a multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network, and a data association network;
[0182] Performing target detection and positioning and information normalization processing on the video frame data through the multi-target classification detection network to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target;
[0183] Performing re-identification feature extraction and pooling alignment processing on the video frame processing data through the target re-identification network, determining the same target in each of the video frames, and generating a target re-identification data set containing the same target association information;
[0184] The target information file and the target re-identification data set are subjected to spatiotemporal association and matching processing by combining the preset ID data with the data association network to determine the target action trajectory corresponding to each of the IDs.
[0185] Optionally, the detailed functions and extended functions of the program may refer to the above description.
[0186] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0187] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0188] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-target detection and tracking method, characterized in that: include: Acquire video frame data captured by a camera device, and input the video frame data into a multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network, and a data association network; Performing target detection and positioning and information normalization processing on the video frame data through the multi-target classification detection network to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target; Performing re-identification feature extraction and pooling alignment processing on the video frame processing data through the target re-identification network, determining the same target in each of the video frames, and generating a target re-identification data set containing the same target association information; The target information file and the target re-identification data set are subjected to spatiotemporal association and matching processing by combining the preset ID data with the data association network to determine the target action trajectory corresponding to each of the IDs.
2. The method according to claim 1, characterized in that The target re-identification network includes a feature extraction network, a pooling alignment network and a similarity calculation network; The target re-identification network performs re-identification feature extraction and pooling alignment processing on the video frame processing data, determines the same target in each of the video frames, and generates a target re-identification data set containing the same target association information, including: The feature extraction network performs re-identification feature extraction on the video frame processing data to obtain a video frame feature map; The pooling alignment network performs fine-grained component alignment and overlapping feature fusion on the video frame feature map to generate fine-grained segmentation feature data; The similarity calculation network performs cosine similarity calculation on the fine-grained division feature data, determines the same target in each of the video frames based on the similarity calculation result, and generates a target re-identification data set containing the same target association information.
3. The method according to claim 2, characterized in that The pooling alignment network includes a feature vertical division module and a feature alignment fusion module; The pooling alignment network performs fine-grained component alignment and overlapping feature fusion on the video frame feature map to generate fine-grained segmentation feature data, including: The feature vertical division module performs feature division on the video frame feature map by using a feature division method based on the vertical direction to form a plurality of vertical feature subsets; The feature alignment and fusion module uses a preset alignment algorithm to align the vertical feature subsets and fuse overlapping features to generate fine-grained segmentation feature data.
4. The method according to claim 1, characterized in that The training process of the target re-identification network includes: Acquire sample video frame processing data and a sample target re-identification data set containing the same target association information in the sample video frame processing data; Inputting the sample video frame processed data into a preset initial target re-identification network to obtain the same target in each sample video frame output by the initial target re-identification network; The initial target re-identification network is trained with the goal that the loss function of the same target in each sample video frame output by the model and the same target association information recorded in the sample target re-identification dataset meets the preset deviation; The trained initial target re-identification network is used as the target re-identification network.
5. The method according to claim 1, characterized in that The multi-target classification detection network includes a first feature extraction network, a second feature extraction network and a feature channel splicing network; The multi-target classification detection network performs target detection positioning and information normalization processing on the video frame data to generate video frame processing data marked with target bounding boxes and target categories, and a target information file recording the position and target category information of each target, including: The first feature extraction network uses the small feature map to detect the large target and obtains a first predicted positioning result; The first feature extraction network uses the large feature map to detect the small target and obtains a second predicted positioning result; The feature channel stitching network performs feature fusion based on the first predicted positioning result and the second predicted positioning result, and performs information normalization processing to generate video frame processing data marked with target bounding boxes and target categories, and a target information file that records the position and category information of each target.
6. The method according to claim 1, characterized in that The training process of the multi-target classification detection network includes: Acquire a multi-scale sample image, wherein the multi-scale sample image is annotated with a target bounding box and a target category of a sample target present therein; Inputting the multi-scale sample image into a preset initial multi-target classification detection network to obtain sample image processing data output by the initial multi-target classification detection network and annotated with a target bounding box and a target category; Training the initial multi-target classification detection network with the goal of making the annotations in the sample image processing data consistent with the annotations in the multi-scale sample image; When the initial multi-target classification detection network meets the preset training conditions, the trained initial multi-target classification detection network is used as the multi-target classification detection network.
7. The method according to claim 1, characterized in that The data association network includes a feature matching module, a spatiotemporal association module and a trajectory construction module; The data association network combines the preset ID data to perform spatiotemporal association and matching processing on the target information file and the target re-identification data set to determine the target action trajectory corresponding to each ID, including: The feature matching module compares the features in the target information file and the target re-identification data set, and determines a feature set corresponding to each ID match in the preset ID data; The spatiotemporal association module performs spatiotemporal association verification on the matching results to generate a verification matching result that satisfies the spatiotemporal constraints; The trajectory construction module connects the verification matching results in series to construct a target action trajectory corresponding to each of the IDs.
8. A multi-target detection and tracking device, characterized in that: include: A video frame acquisition unit, used to acquire video frame data captured by a camera device, and input the video frame data into a multi-target tracking detection model, wherein the multi-target tracking detection model is composed of a multi-target classification detection network, a target re-identification network, and a data association network; A detection and positioning unit, configured to perform target detection and positioning and information normalization processing on the video frame data through the multi-target classification detection network, and generate video frame processing data annotated with target bounding boxes and target categories, and a target information file recording information on the positions and target categories of each target; A re-identification extraction unit, configured to perform re-identification feature extraction and pooling alignment processing on the video frame processing data through the target re-identification network, determine the same target in each of the video frames, and generate a target re-identification data set containing the same target association information; The association matching unit is used to perform spatiotemporal association and matching processing on the target information file and the target re-identification data set through the data association network in combination with preset ID data, so as to determine the target action trajectory corresponding to each ID.
9. A multi-target detection and tracking device, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the multi-target detection and tracking method as described in any one of claims 1 to 7.
10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the multi-target detection and tracking method as described in any one of claims 1 to 7 is implemented.