Multi-object Tracking Method, System and Electronic Device in Dense Scenes
By redesigning the DLA encoder and adopting the two-time target character matching strategy, the problem of low multi-objective tracking accuracy in dense scenarios is solved, and the tracking accuracy and model convergence speed are significantly improved.
Patent Information
- Application Number
- CN202310036508.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-01-10
AI Technical Summary
The prior art has low multi-object tracking accuracy in dense scenarios, especially on the MOT20 data set, FairMOT's multi-object tracking accuracy is low.
By redesigning the deep aggregation encoder (DLA encoder), multi-objective tracking is performed in dense scenes, and the matching strategy of two target characters is adopted, including IOU matching and quadratic matching, ensuring the effectiveness of target matching in video image frames.
It significantly improves the accuracy of multi-objective tracking in dense scenarios, speeds up the convergence speed of the model, and ensures the effectiveness of target tracking.
Smart Images

Figure CN116030094B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-object tracking, and in particular to a multi-object tracking method, system and electronic device in a dense scene. Background Art
[0002] With the rapid development of the computer vision field, the accuracy of multi-object tracking algorithms has also been improved. In particular, in ordinary pedestrian tracking scenarios, quite a high level of accuracy has been achieved. However, there is still room for improvement in the mainstream mode in dense scenes. Tracking multiple objects in a crowded scene requires more refined feature representation, so it is necessary to improve the existing network, and the improved network can have better performance in multi-object tracking in dense scenes.
[0003] Multi-object tracking (MOT) is a classic task in computer vision. It attempts to find the relationships between frames in a video and label the same object with a bounding box and an ID. There have been many applications in the research of this field, such as autonomous driving, video analysis, human-computer interaction, etc. In recent years, most MOT algorithms have adopted a method called detection and tracking. This means that it first uses some specific object detection algorithms to obtain the object positions in each frame, and then uses methods such as Kalman filtering, Hungarian matching or other advanced methods to match the same objects with the same ID between two frames. For example, JDE, DeepSort, relational tracking and Simple tracking all adopt this method. FairMOT is an important representative of such products. It uses a center network-based object detection network, which is a classic anchor-free object detection algorithm. In addition, it also has a parallel Re-ID branch. The multi-object tracking accuracy (MOTA) of FairMOT reaches 73.7% on the MOT17 dataset, which was state-of-the-art at the time of submission.
[0004] However, FairMOT does not perform well in dense scenes. Only 61.8% MOTA on the MOT20 dataset contains smaller objects. FairMOT uses Deep Layer Aggregation (DLA) as the backbone, which is a classic keypoint detection network. Since CenterNet is a heatmap-based object detection network, DLA can even obtain higher results than ResNet. In FairMOT, the Re-ID branch shares the same model with the detection branch to aggregate multi-layer features. The results show that the intermediate layer features of the backbone may not be conducive to Re-ID. Summary of the Invention
[0005] Therefore, the technical problem to be solved by the present invention is to overcome the problem of low accuracy of multi-object tracking in crowded scenes in the prior art.
[0006] To solve the above technical problems, the present invention provides a multi-object tracking method in a dense scene, including:
[0007] Step S1: Obtain the T-th frame image, and downsample the T-th frame image to obtain a plurality of downsampled feature maps;
[0008] Step S2: Upsample and perform feature fusion on the plurality of downsampled feature maps to obtain a first feature map, and map the first feature map through convolution to obtain a heatmap detection result;
[0009] At the same time, locate the position of the target person in the (T - 1)-th frame image in the T-th frame image to obtain a positioning result, and obtain a target person detection box based on the heatmap detection result;
[0010] Step S3: Perform IOU matching on the target person detection box and the positioning result. If all target persons in the T-th frame image and the (T - 1)-th frame image are successfully IOU-matched, the target tracking is completed; if there are target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching is not successful, execute Step S4;
[0011] Step S4: Upsample the plurality of downsampled feature maps in Step S1 to obtain a second feature map, map the second feature map to obtain a preset-dimension feature map, and reduce the dimension of the preset-dimension feature map to obtain a dimension-reduced feature map;
[0012] Step S5: Perform secondary matching on the dimension-reduced feature map based on IOU matching. Specifically: match the dimension-reduced feature map with the target persons in the pre-stored (T - 1)-th frame image. If all target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching is not successful are successfully secondarily matched, the target tracking is completed; if the target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching is not successful are not successfully secondarily matched, return to Step S2 until all target persons are successfully matched.
[0013] In an embodiment of the present invention, the method of downsampling the T-th frame image in Step S1 to obtain a plurality of downsampled feature maps is specifically: downsample the T-th frame image through a DLA encoder to obtain a plurality of downsampled feature maps, where the DLA encoder includes a plurality of roots and a plurality of types of convolutional blocks, the root is used to add the convolutional blocks of different types to each other, and the convolutional blocks of different types are used to change the number of channels of the image.
[0014] In an embodiment of the present invention, the convolutional block of different types includes a convolutional layer, a first normalization layer, a depthwise separable convolutional layer, a second normalization layer, a first multi-layer perceptron, a GeLU activation function, and a second multi-layer perceptron connected in sequence, and the convolutional layer and the second multi-layer perceptron are summed;
[0015] The convolutional layer is used to change the number of channels of the feature map;
[0016] Both the first normalization layer and the second normalization layer are used to prevent overfitting and increase generalization;
[0017] The depthwise separable convolutional layer is used to reduce the number of parameters and simulate the self-attention operation;
[0018] Both the first multi-layer perceptron and the second multi-layer perceptron are used to make up for the problem that there is no interaction between channels in the depthwise separable convolution.
[0019] In an embodiment of the present invention, in step S2, the plurality of downsampled feature maps are upsampled and feature fused to obtain a first feature map, specifically: the plurality of downsampled feature maps are upsampled to obtain a plurality of upsampled feature maps with different resolutions, and then the plurality of upsampled feature maps with different resolutions are feature fused to obtain a first feature map.
[0020] In an embodiment of the present invention, in step S2, the position of the target person in the (T - 1)-th frame image in the T-th frame image is located, specifically: the position of the target person in the (T - 1)-th frame image in the T-th frame image is located by Kalman filtering.
[0021] In an embodiment of the present invention, in step S4, the preset-dimension feature map is dimension-reduced to obtain a dimension-reduced feature map, specifically: the length and width of the preset-dimension feature map are combined into one dimension to obtain a dimension-reduced feature map.
[0022] In an embodiment of the present invention, in step S5, the dimension-reduced feature map is secondarily matched based on IOU matching, specifically: the dimension-reduced feature map is subjected to Hungarian matching based on IOU matching, and the Hungarian matching is realized by calculating the cosine distance between target persons. If
[0023] the cosine distance between target persons is less than a preset threshold, it indicates that the target persons are successfully matched; if the cosine distance between target 5 persons is greater than the preset threshold, it indicates that the target persons are not successfully matched.
[0024] To solve the above technical problems, the present invention provides a multi-target tracking system in a dense scene, including:
[0025] A downsampling module: used to obtain the T-th frame image and downsample the T-th frame image to obtain a plurality of downsampled feature maps;
[0026] 0 Feature generation and positioning module: used to upsample the plurality of downsampled feature maps and perform feature
[0027] Fuse to obtain a first feature map, perform mapping on the first feature map through convolution to obtain a heatmap detection result, and obtain a target person detection frame based on the heatmap detection result;
[0028] Meanwhile, it is used to determine the position of the target person in the (T - 1)-th frame image in the T-th frame image
[0029] for positioning to obtain a positioning result;
[0030] 5 The first matching module: used to perform IOU matching on the target person detection frame and the positioning result,
[0031] If all target persons in the T-th frame image and the (T - 1)-th frame image are successfully IOU-matched, then target tracking is completed; if there are target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching fails; execute the process of the feature generation module.
[0032] The feature generation module: used to upsample the plurality of downsampled feature maps to obtain a second feature map, map the second feature map to obtain a feature map of a preset dimension, and reduce the dimension of the feature map of the preset dimension
[0033] to obtain a dimension-reduced feature map;
[0034] The second matching module: used to perform secondary matching on the dimension-reduced feature map based on IOU matching, specifically: match the dimension-reduced feature map with the target persons in the pre-stored (T - 1)-th frame image
[0035] If all the target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching fails are successfully secondarily matched, then target tracking is completed; if the target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching fails are not successfully secondarily matched, then return to the execution process of the feature generation and positioning module until all target persons are successfully matched.
[0036] To solve the above technical problems, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the multi-target tracking method in the above dense scene are implemented.
[0037] To solve the above technical problems, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-target tracking method in the above dense scene are implemented.
[0038] The above technical solution of the present invention has the following advantages compared with the prior art:
[0039] The multi-object tracking method in dense scenes described in the present invention redesigned the Deep Layer Aggregation encoder (DLA encoder), which meets the real-time requirements and has a significant improvement in accuracy on dense scene datasets; the present invention can match the targets in the video image frames as much as possible through two matches of target persons, ensuring the effectiveness of target tracking; the present invention further improves the accuracy of target tracking in dense scenes and speeds up the convergence rate of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to make the content of the present invention easier to be clearly understood, the following further describes the present invention in detail according to specific embodiments of the present invention and in conjunction with the accompanying drawings, where
[0041] Figure 1 is a schematic flow chart of the method of the present invention;
[0042] Figure 2 is a schematic diagram of the object detection network of the present invention;
[0043] Figure 3 is a schematic diagram of the DLA encoder framework of the present invention;
[0044] Figure 4 is a schematic diagram of the structure of the convolutional block-like in the DLA encoder of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.
[0046] Embodiment 1
[0047] Please refer to Figure 1 and Figure 2 , the multi-object tracking method in dense scenes of the present invention includes:
[0048] Step S1: Obtain the T-th frame image and downsample the T-th frame image to obtain a plurality of downsampled feature maps;
[0049] Further, the input image is padded to a unified size of 1088×608, and using such an input size helps to detect small objects.
[0050] Further, the method of downsampling the T-th frame image in step S1 to obtain a plurality of downsampled feature maps is specifically: downsample the T-th frame image through a DLA encoder to obtain a plurality of downsampled feature maps, where the DLA encoder includes a plurality of roots and a plurality of convolutional block-like, and the root is used to add the convolutional block-like to each other, and the convolutional block-like is used to change the number of channels of the image.
[0051] Further, the downsampling module downsamples the T-th frame image through a DLA encoder to obtain a number of downsampled feature maps; as Figure 3 shown, the DLA encoder includes a number of roots and a number of types of convolutional blocks. The structure of the DLA encoder is constructed recursively. The root is responsible for connecting each type of convolutional block so that the entire network can be recursively organized in a tree structure. The root is used to add the convolutional blocks to each other, and the convolutional blocks are used to change the number of channels of the image.
[0052] Further, the convolutional block includes a convolutional layer, a first layer of normalization, a depthwise separable convolutional layer, a second layer of normalization, a first multi-layer perceptron, a GeLU activation function, and a second multi-layer perceptron connected in sequence, and the convolutional layer and the second multi-layer perceptron are summed;
[0053] The convolutional layer is used to change the number of channels of the feature map;
[0054] Both the first layer of normalization and the second layer of normalization are used to prevent overfitting and increase generalization;
[0055] The depthwise separable convolutional layer is used to reduce the number of parameters, simulate self-attention operations, and correlate each pixel point;
[0056] Both the first multi-layer perceptron (MLP) and the second multi-layer perceptron (MLP) are used to make up for the problem that there is no interaction between channels in the depthwise separable convolution.
[0057] Specifically, as Figure 4 shown, in this embodiment, the number of channels of the feature map is first changed through a 3×3 convolution. Creatively, the first layer of normalization is then used, and the GeLU activation function is not immediately used. The GeLU activation function is only used once in the entire convolutional block, demonstrating that depth convolution has a similar effect to self-attention. Then a 7×7 depthwise separable convolution is used. The depthwise separable convolution can significantly reduce the number of parameters, but can also reduce the interaction between channels to a certain extent. Moreover, this embodiment also adds two multi-layer perceptrons (MLPs), which have an effect similar to pointwise convolution. The Gelu activation function is only used once after the MLP. The design method of the convolutional block in this embodiment can effectively extract the features of the image.
[0058] Step S2: Upsample the number of downsampled feature maps, then perform feature fusion on the upsampling result to obtain a first feature map, and map the first feature map through convolution to obtain a heatmap detection result;
[0059] At the same time, the position of the target person in the (T - 1)-th frame image in the T-th frame image is located through Kalman filtering to obtain a positioning result.
[0060] Further, in step S2, the plurality of downsampled feature maps are upsampled and feature fused to obtain a first feature map, specifically: the plurality of downsampled feature maps are upsampled to obtain a plurality of upsampled feature maps with different resolutions, and then the plurality of upsampled feature maps with different resolutions are feature fused to obtain a first feature map.
[0061] Further, in step S2, the first feature map is mapped through convolution to obtain a heatmap detection result, and then a target person detection box is generated based on the heatmap detection result (including three data information), specifically: the heatmap is a Gaussian map of the center position of each category (in this embodiment, there is only the category of person), and the approximate position of the center point of each box is obtained through these Gaussian maps, and then the specific coordinates of each target person detection box can be obtained by subtracting the offset and the output of the box information.
[0062] The heatmap detection result in step S2 includes heatmap regression, regression bounding box size, and regression centroid deviation of the first feature map, and a target person detection box is generated based on the heatmap regression, regression bounding box size, and regression centroid deviation.
[0063] Step S3: Perform IOU matching on the target person detection box and the positioning result (i.e., the overlapping result of the target person). If all the target persons in the T-th frame image and the (T - 1)-th frame image are successfully IOU matched, the target tracking is completed; if there are target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching is not successful; execute step S5;
[0064] The above steps S1 to S3 are implemented through Figure 2 the detection branch in the target detection network (i.e., Figure 2 the lower branch of the target detection network in ). During the detection process, some target persons in the image are not matched in the IOU matching because they are blocked or have a relatively fast moving speed. Therefore, in this embodiment, the target persons in the image that are not successfully IOU matched are re-matched, specifically as follows:
[0065] Step S4: The plurality of downsampled feature maps in step S1 are upsampled (implemented through dilated convolution) to obtain a second feature map (the Re-ID branch directly upsamples the 32-fold downsampled feature map of the last layer output by the DLA encoder layer by layer to obtain the second feature map), the second feature map is mapped to obtain a preset dimension feature map (the dimension size can be set artificially), and the preset dimension feature map is dimension-reduced to obtain a dimension-reduced feature map;
[0066] Further, in step S4, the preset dimensional feature map is dimensionally reduced to obtain a dimensionally reduced feature map, specifically: the length and width of the preset dimensional feature map are combined into one dimension to obtain the dimensionally reduced feature map.
[0067] Step S5: Perform secondary matching on the dimensionally reduced feature map based on IOU matching, specifically: match the dimensionally reduced feature map with the target person in the pre-stored (T-1)-th frame image. If the target persons for whom IOU matching fails in both the T-th frame image and the (T-1)-th frame image are successfully secondarily matched, then target tracking is completed; if the target persons for whom IOU matching fails in both the T-th frame image and the (T-1)-th frame image are not successfully secondarily matched, then return to step S2 until all target persons are successfully matched.
[0068] Further, in step S5, performing secondary matching on the dimensionally reduced feature map based on IOU matching specifically means: performing Hungarian matching on the dimensionally reduced feature map based on IOU matching. The Hungarian matching is achieved by calculating the cosine distance between target persons. If the cosine distance between target persons is less than a preset threshold, it indicates that the target persons are successfully matched; if the cosine distance between target persons is greater than the preset threshold, it indicates that the target persons are not successfully matched.
[0069] The above steps S4 to S5 are implemented through Figure 2 the RE-ID branch in the Figure 2 target detection network (i.e., the lower branch of the target detection network in
[0070] It should be noted that when the detection branch and the Re-ID branch share the same branch, conflicts often occur. To solve this problem, in this embodiment, the detection branch and the Re-ID branch are decoupled, thereby achieving a better target tracking effect.
[0071] Embodiment 2
[0072] A multi-target tracking system in a dense scene, comprising:
[0073] A downsampling module: used to obtain the T-th frame image and perform downsampling on the T-th frame image to obtain a plurality of downsampled feature maps;
[0074] A feature generation and localization module: used to perform upsampling and feature fusion on the plurality of downsampled feature maps to obtain a first feature map, and perform mapping on the first feature map through convolution to obtain a heatmap detection result, and obtain a target person detection frame based on the heatmap detection result;
[0075] At the same time, it is used to localize the position of the target person in the (T-1)-th frame image in the T-th frame image to obtain a localization result;
[0076] The first matching module: It is used to perform IOU matching on the target person detection box and the positioning result. If all target persons in the T-th frame image and the (T-1)-th frame image are successfully IOU-matched, the target tracking is completed; if there are target persons in the T-th frame image and the (T-1)-th frame image whose IOU matching fails, the process of the feature generation module is executed;
[0077] The feature generation module: It is used to upsample the several downsampled feature maps to obtain a second feature map, map the second feature map to obtain a feature map with a preset dimension, and reduce the dimension of the feature map with the preset dimension to obtain a dimension-reduced feature map;
[0078] The second matching module: It is used to perform secondary matching on the dimension-reduced feature map based on IOU matching. Specifically, it matches the dimension-reduced feature map with the target persons in the pre-stored (T-1)-th frame image. If all the target persons whose IOU matching fails in the T-th frame image and the (T-1)-th frame image are successfully secondarily matched, the target tracking is completed; if the target persons whose IOU matching fails in the T-th frame image and the (T-1)-th frame image are not successfully secondarily matched, it returns to the execution process of the feature generation and positioning module until all target persons are successfully matched.
[0079] Embodiment III
[0080] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the multi-target tracking method in a dense scene as described in Embodiment I.
[0081] Embodiment IV
[0082] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the multi-target tracking method in a dense scene as described in Embodiment I.
[0083] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0084] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0085] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0087] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0088] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to exhaustively list all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A multi-object tracking method in a dense scene, characterized in that: It includes: Step S1: Obtain the T-th frame image and downsample the T-th frame image to obtain a number of downsampled feature maps. The method is as follows: Downsample the T-th frame image through a DLA encoder to obtain a number of downsampled feature maps. Among them, the DLA encoder includes a number of roots and a number of types of convolutional blocks. The root is used to add the convolutional blocks of different types to each other, and the convolutional blocks of different types are used to change the number of channels of the image; The convolutional block of different types includes a convolutional layer, a first normalization layer, a depthwise separable convolutional layer, a second normalization layer, a first multi-layer perceptron, a GeLU activation function, and a second multi-layer perceptron connected in sequence, and the convolutional layer and the second multi-layer perceptron are summed; The convolutional layer is used to change the number of channels of the feature map; Both the first normalization layer and the second normalization layer are used to prevent overfitting and increase generalization; The depthwise separable convolutional layer is used to reduce the number of parameters and simulate self-attention operations; Both the first multi-layer perceptron and the second multi-layer perceptron are used to make up for the problem that there is no interaction between channels in the depthwise separable convolution; Step S2: Upsample and perform feature fusion on the number of downsampled feature maps to obtain a first feature map, and map the first feature map through convolution to obtain a heatmap detection result, and obtain a target person detection box based on the heatmap detection result; At the same time, locate the position of the target person in the (T - 1)-th frame image in the T-th frame image to obtain a positioning result; Step S3: Perform IOU matching on the target person detection box and the positioning result. If all target persons in the T-th frame image and the (T - 1)-th frame image are successfully IOU matched, the target tracking is completed; if there are target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching is not successful, execute Step S4; Step S4: Upsample the number of downsampled feature maps in Step S1 to obtain a second feature map, map the second feature map to obtain a preset dimension feature map, and reduce the dimension of the preset dimension feature map to obtain a dimension-reduced feature map; Step S5: Perform secondary matching on the dimension-reduced feature map based on IOU matching. Specifically: match the dimension-reduced feature map with the target persons in the pre-stored (T - 1)-th frame image. If all the target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching is not successful are successfully secondarily matched, the target tracking is completed; if the target persons in the T-th frame image and the (T - 1)-th frame image whose IOU matching is not successful are not successfully secondarily matched either, return to Step S2 until all target persons are successfully matched.
2. The multi-object tracking method in a dense scene according to claim 1, characterized in that: In Step S2, when upsampling and performing feature fusion on the number of downsampled feature maps to obtain a first feature map, specifically: Upsample the number of downsampled feature maps to obtain a number of upsampled feature maps with different resolutions, and then perform feature fusion on the number of upsampled feature maps with different resolutions to obtain a first feature map.
3. The multi-object tracking method in a dense scene according to claim 1, It is characterized in that: In step S2, the position of the target person in the (T-1)-th frame image in the T-th frame image is located, specifically: the position of the target person in the (T-1)-th frame image in the T-th frame image is located by Kalman filtering.
4. The multi-object tracking method in a dense scene according to claim 1, It is characterized in that: In step S4, the preset-dimensional feature map is dimension-reduced to obtain a dimension-reduced feature map, specifically: the length and width of the preset-dimensional feature map are combined into one dimension to obtain a dimension-reduced feature map.
5. The multi-object tracking method in a dense scene according to claim 1, It is characterized in that: In step S5, the dimension-reduced feature map is secondarily matched based on IOU matching, specifically: the dimension-reduced feature map is subjected to Hungarian matching based on IOU matching, and the Hungarian matching is realized by calculating the cosine distance between target persons. If the cosine distance between target persons is less than a preset threshold, it indicates that the target persons are successfully matched; if the cosine distance between target persons is greater than the preset threshold, it indicates that the target persons are not successfully matched.
6. A multi-object tracking system in a dense scene for implementing the multi-object tracking method in a dense scene according to any one of claims 1 to 5, It is characterized in that: Comprising: A downsampling module: used to obtain the T-th frame image and downsample the T-th frame image to obtain a plurality of downsampled feature maps; A feature generation and localization module: used to upsample and fuse features of the plurality of downsampled feature maps to obtain a first feature map, map the first feature map through convolution to obtain a heatmap detection result, and obtain a target person detection frame based on the heatmap detection result; At the same time, it is used to locate the position of the target person in the (T-1)-th frame image in the T-th frame image to obtain a localization result; A first matching module: used to perform IOU matching on the target person detection frame and the localization result. If all target persons in the T-th frame image and the (T-1)-th frame image are successfully IOU-matched, the target tracking is completed; if there are target persons in the T-th frame image and the (T-1)-th frame image that are not successfully IOU-matched; the process of the feature generation module is executed; A feature generation module: used to upsample the plurality of downsampled feature maps to obtain a second feature map, map the second feature map to obtain a preset-dimensional feature map, and dimension-reduce the preset-dimensional feature map to obtain a dimension-reduced feature map; A second matching module: used to secondarily match the dimension-reduced feature map based on IOU matching, specifically: match the dimension-reduced feature map with the target persons in the pre-stored (T-1)-th frame image. If all the target persons in the T-th frame image and the (T-1)-th frame image that are not successfully IOU-matched are secondarily matched successfully, the target tracking is completed; if the target persons in the T-th frame image and the (T-1)-th frame image that are not successfully IOU-matched are not secondarily matched successfully, return to the execution process of the feature generation and localization module until all target persons are successfully matched.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the multi-target tracking method in the dense scenario according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps of the multi-target tracking method in the dense scenario according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Target tracking method and device and electronic equipment
CN112967315A
Real-time multi-target tracking method and system in automatic driving
CN114445453A
Multi-target tracking method and device, equipment and storage medium
CN115527143A