Target tracking method and device based on joint model, and program product

By using a joint model that shares the feature extraction network in both the object detection and tracking models, the high hardware requirements and long processing time caused by multi-model deployment are resolved, resulting in more efficient object detection and tracking and improved real-time tracking performance.

CN121190786APending Publication Date: 2025-12-23HEFEI YINGJU INNOVATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511347719.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

In existing technologies, when object detection and object tracking algorithms are deployed on the same platform, multiple AI models need to be run simultaneously, resulting in high hardware computing power requirements, long processing time, poor tracking performance, and the inability to share data, which affects the generalization ability of the backbone network feature extraction.

Method used

By adopting a joint model-based approach, the target detection model and the target tracking model share the feature extraction network, and target detection and tracking are performed through joint inference, which reduces computational resource consumption, reduces time consumption, and improves real-time tracking performance.

Benefits of technology

By using a joint model with a shared feature extraction network, computational resource consumption is reduced, time consumption is decreased, the real-time performance of target detection and tracking is improved, and the generalization ability of the backbone network feature extraction is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190786A_ABST
    Figure CN121190786A_ABST
Patent Text Reader

Abstract

The invention discloses a target tracking method and device based on a joint model and a program product, the joint model comprises a target detection model and a target tracking model, the method comprises the following steps: obtaining a current frame input image of the target detection model, and the target detection model and the target tracking model share a feature extraction network; extracting a current feature map in the current frame input image through the feature extraction network; acquiring initial target feature data of the tracking target; and based on the current feature map and the initial target feature data, using a tracking network in a target tracking model to track a tracking target in a current frame to obtain current tracking data of the tracking target.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a target tracking method based on a joint model, a target tracking device and a computer program product. BACKGROUND

[0002] The detection and tracking function of the target object of interest has been widely used in various fields, such as optoelectronics, pods, etc. In the prior art, when deploying two algorithms on the same platform, multiple AI models need to be run simultaneously, which puts high demands on hardware computing power. In addition, multiple model running increases the time consumption and has a great adverse effect on tracking effect. The detection and single target tracking algorithm are in an isolated state, and are respectively independent modules, and the feature maps extracted by the backbone network cannot be shared. Therefore, when tracking a single target, the full-frame feature map extracted by the backbone network of the detection cannot be reused, and a separate backbone network needs to be run to repeatedly extract features. Running two sets of detection and tracking algorithms simultaneously is time-consuming and requires high hardware demand, which affects the tracking effect. The isolation of algorithms leads to data isolation, and the data of the two cannot be well shared to train the same backbone network at the same time and improve the generalization ability of the feature extraction of the backbone network. SUMMARY

[0003] In order to solve the existing technical problems, the present application provides a target tracking method based on a joint model, a target tracking device and a computer program product, which can reduce the consumption of computing resources, reduce the time consumption, and improve the real-time tracking effect.

[0004] In a first aspect, a target tracking method based on a joint model is provided, the joint model including a target detection model and a target tracking model, and the method includes: obtaining a current frame input image of the target detection model, wherein the target detection model and the target tracking model share a feature extraction network; extracting a current feature map in the current frame input image through the feature extraction network; obtaining initial target feature data of a tracking target; based on the current feature map and the initial target feature data, using a tracking network in the target tracking model to track the tracking target in the current frame to obtain current tracking data of the tracking target.

[0005] In a second aspect, a target tracking device is provided, including a memory and a processor, the memory storing a computer program, and the computer program being executed by the processor to make the processor execute the target tracking method based on the joint model provided in the first aspect of the present application.

[0006] In a third aspect, a computer program product is provided, including a computer program, which, when executed, implements the target tracking method based on the joint model provided in the first aspect of the present application.

[0007] The application obtains a current frame input image, extracts a current feature map in the current frame input image through a feature extraction network shared by a target detection model and a target tracking model, then takes the current feature map as an input of a tracking network and a monitoring network respectively, detects and tracks a tracking target, and obtains initial target feature data of the tracking target, determines current tracking data of the tracking target in the current feature map according to the initial target feature data, so that when target detection and target tracking are performed in an actual scene, the target detection model and the target tracking model share the feature extraction network, the target detection model and the target tracking model are not two isolated models, joint inference is performed, and the calculation resource consumption and time consumption are reduced, so that the real-time tracking effect is improved. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 An application environment diagram of a target tracking method based on a joint model in an embodiment;

[0009] Figure 2 A flowchart of a target tracking method based on a joint model in an embodiment;

[0010] Figure 3 A network structure diagram of a joint model in an embodiment;

[0011] Figure 4 A network diagram for calculating initial target feature data in an embodiment;

[0012] Figure 5 A network diagram of a target tracking method based on a joint model in an embodiment;

[0013] Figure 6 A schematic diagram of a target tracking device based on a joint model in an embodiment. DETAILED DESCRIPTION

[0014] The technical solutions of the present application are further described in detail below in combination with the drawings and specific embodiments of the present application.

[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0016] In the following description, the expression "some embodiments" describes a subset of all possible embodiments, however, it is to be understood that "some embodiments" can be the same subset or different subsets as each other and as other embodiments, and that the expression "some embodiments" can be combined with each other and with other embodiments, without conflict.

[0017] Referring to Figure 1 , an application environment diagram of a joint model based target tracking method in an embodiment. The joint model based target tracking method is applied in a target tracking device 10, which includes an image acquisition device 12, a processor 13 and a memory 14. The image acquisition device 12 is configured to acquire continuous frame images in front of the target tracking device 10, and the processor 13 is configured to detect a gas region in the continuous frame images according to the continuous frame images acquired by the image acquisition device 12. The memory 14 is configured to store programs and data corresponding to the joint model based target tracking method.

[0018] The target tracking device 10 includes but is not limited to a handheld detection device, a non-handheld self-moving detection device, and a non-handheld non-self-moving detection device. The handheld detection device includes but is not limited to a handheld imaging device with infrared thermal imaging function and a handheld imaging device with visible light imaging function. The non-handheld self-moving detection device includes but is not limited to a self-moving detection device with infrared thermal imaging function and a self-moving detection device with visible light imaging function. The non-handheld non-self-moving detection device includes but is not limited to a non-handheld non-self-moving device with infrared thermal imaging function and a non-handheld non-self-moving device with visible light imaging function. Through the handheld detection device and the non-handheld self-moving detection device, a user can perform joint tracking processing on a target in a moving scene or a fixed scene. Through the non-handheld non-self-moving detection device, joint tracking processing can be performed on a target in a fixed scene. Therefore, the joint model based target tracking method provided in the embodiments of the present application can be applied to various complex moving scenes and fixed scenes. For example, the target tracking device 10 can be an infrared imaging device, a monitoring device with image shooting function, and the like.

[0019] The image acquisition device 12 can be one or more sensors. The image acquisition device 12 can be a monocular vision sensor or a multi-view vision sensor. For example, it can be a combination of one or more sensors such as a thermal imaging sensor, a visible light image sensor, a millimeter wave sensor, a laser radar sensor, an infrared thermal imaging sensor, and a depth sensor. For example, the image acquisition device 12 is an infrared image acquisition device with high resolution, such as a handheld target tracking device or a gimbal, to obtain clear image data. For example, the image acquisition device 12 is a dual-optical camera.

[0020] The processor 13 can be one or more, and when there are multiple processors 13, the multiple processors can be integrated on one chip or independently arranged on each chip. The target tracking device 10 can be an electronic device installed on any type of moving body, and the target tracking device 10 can also be an autonomous moving device, for example, the autonomous moving device or the moving body includes but is not limited to: a vehicle, an electric vehicle, a hybrid electric vehicle, a motorcycle, a bicycle, a personal moving device, an airplane, a drone, a ship or a robot, etc. The target tracking device 10 can also be an electronic device fixedly installed on a certain device or a fixed position of a fixed scene.

[0021] The target tracking device 10 can also include other sensor modules, including but not limited to environmental perception sensors and motion posture sensors. The environmental perception sensors include but are not limited to one or more of the following: brightness sensors, temperature sensors, haze sensors, and other environmental sensors. The motion posture sensors include but are not limited to one or more combinations of the following: inertial sensors (IMU), speed sensors, acceleration sensors, gyroscope sensors, geomagnetic sensors, rotation vector sensors, steering wheel angle sensors, level sensors, tilt sensors, vibration sensors, displacement sensors, and gravity sensors.

[0022] The target tracking device 10 can also include a display terminal for displaying images.

[0023] Referring to Figure 2 A flowchart of a target tracking method based on a joint model provided by an embodiment of the present application. The target tracking method based on a joint model is applied to a target tracking device, and the target tracking method based on a joint model includes the following steps:

[0024] S10, obtaining a current frame input image of a target detection model, wherein the target detection model shares a feature extraction network with a target tracking model.

[0025] In this embodiment, the current frame input image can be obtained by processing a current scene image. The current scene image is an image collected from the current scene. The current scene image can be an infrared image, a visible light image, a fusion image of infrared and visible light, etc. For example, Figure 2As shown, the joint model includes a target detection model and a target tracking model, wherein the target detection model and the target tracking model share a feature extraction network. The feature extraction network can include a down-sampling layer and an up-sampling layer. The feature extraction network outputs a current feature map. For example, the current feature map is an 8 times down-sampled feature map. The feature extraction network is connected to a detection network and a tracking network. The detection network is used to perform target detection based on the current feature map, and outputs a current target detection result. The current target detection result includes position data of a detection box of a tracking target, etc. The tracking network is used to perform target tracking based on the current feature map, and outputs current tracking data.

[0026] S11, extracting a current feature map in a current frame input image through a feature extraction network.

[0027] In this embodiment, the feature extraction network can include a down-sampling layer and an up-sampling layer. The feature extraction network outputs a current feature map. For example, the current feature map is an 8 times down-sampled feature map. The current feature map can be used as an input of the detection network and the tracking network, so that the detection network and the tracking network can each use a respective feature extraction network. Sharing the same feature extraction network can improve the generalization ability of the backbone network feature extraction.

[0028] S12, obtaining initial target feature data of a tracking target.

[0029] In this embodiment, the initial target feature data represents the feature data of the tracking target. The initial target feature data includes but is not limited to size data of the tracking target, image features, etc. The initial target feature data can be obtained from an initial image.

[0030] S13, based on the current feature map and the initial target feature data, using a tracking network in the target tracking model to track the tracking target in the current frame, and obtaining current tracking data of the tracking target.

[0031] In this embodiment, the current feature map represents the feature map extracted in the current frame. Therefore, it is necessary to determine the current tracking data of the tracking target in the current frame based on the initial target feature data in the current feature map. The current tracking data represents position data of the tracking target in the current frame. The current tracking data includes but is not limited to position data of a bounding box, coordinates of a center point of the bounding box, motion trajectory data, time sequence association information, scale and shape change information, etc.

[0032] In the above embodiment, the current frame input image is acquired, a current feature map in the current frame input image is extracted through a feature extraction network shared by a target detection model and a target tracking model, and then the current feature map is taken as input of a tracking network and a monitoring network respectively to detect and track the tracking target, and initial target feature data of the tracking target is acquired. The current tracking data of the tracking target is determined in the current feature map according to the initial target feature data. In this way, when target detection and target tracking are performed in an actual scene, the target detection model and the target tracking model share the feature extraction network, and the target detection model and the target tracking model are not two isolated models. Through joint inference, the calculation resource consumption can be reduced, the time consumption can be reduced, and the real-time tracking effect can be improved.

[0033] In some embodiments, the tracking of the tracking target in the current frame by using the tracking network in the target tracking model based on the current feature map and the initial target feature data to obtain the current tracking data of the tracking target comprises:

[0034] calculating a current search region feature map based on the current feature map;

[0035] calculating a current target feature map based on the current search region feature map and the initial target feature data;

[0036] tracking the tracking target by using the tracking network in the target tracking model based on the current target feature map to obtain the current tracking data of the tracking target.

[0037] In this embodiment, the current search region feature map represents a feature map obtained from a current search region indicated by search region position data in the current feature map. Therefore, the current search region feature map is a search region further reduced from the current feature map. In target tracking, the search region position data in the current frame can be determined according to the tracking data of the last frame. In the search region feature map, the current target feature map in the current frame is determined according to the initial target feature data. Since the search region feature map is obtained, the position coordinates of the tracking target in the search region are finally required. The current search region feature map and the initial target feature data are subjected to convolution operation, such as ROIAlign operation, to obtain the current target feature map. The current target feature map is input to a tracking operation head of the tracking network implemented by the convolutional neural network for operation to obtain the current tracking data.

[0038] Optionally, the calculation of the current search region feature map based on the current feature map comprises:

[0039] acquiring search region position data;

[0040] based on the current feature map and the search region position data, a current search region feature map is calculated.

[0041] Optionally, the search region position data comprises:

[0042] When the current frame is not the first frame for tracking the tracking target, based on the tracking data of the tracking target in the last frame of the current frame, a tracking region is obtained, the tracking region is expanded to obtain the search region position data;

[0043] When the current frame is the first frame for tracking the tracking target, initial position data of the tracking target is obtained, based on the initial position data, a tracking region is obtained, the tracking region is expanded to obtain the search region position data.

[0044] In the embodiment, when the current frame is not the first frame for tracking the tracking target, the tracking region is determined according to the position data of the tracking target in the last frame, and then the tracking region is expanded. For example, if the coordinates of the upper left corner and the coordinates of the lower right corner of the tracking target in the last frame are (x1, y1) and (x2, y2). After expansion, the search region position data is obtained by expanding 32 pixels up, down, left and right on the basis of the position coordinates of the tracking target in the last frame, and the search region position data is [x1-32, y1-32, x2+32, y2+32]. However, after expansion, it may exceed the image boundary. Assuming that the width and height of the image are W and H respectively, the actual coordinate values are [max(x1-32, 0), max(y1-32, 0), min(x2+32, W), min(y2+32, H)]. Similarly, when the current frame is the first frame, the initial position data is expanded according to the initial position data. The initial position data is the initial coordinates corresponding to the tracking target in the initialization frame.

[0045] In the above embodiment, the current frame input image is obtained, the current feature map in the current frame input image is extracted through the feature extraction network shared by the target detection model and the target tracking model, then the current feature map is taken as the input of the tracking network and the monitoring network respectively, the tracking target is detected and tracked, in tracking, the search region is further determined in combination with the tracking data of the last frame, the current target feature map is determined according to the search position data and the current feature map, and the current tracking data of the tracking target is determined in the current target feature map. In this way, in the actual scene of target detection and target tracking, the target detection model and the target tracking model share the feature extraction network, the target detection model and the target tracking model are not two isolated models, through joint inference, the calculation resource consumption can be reduced, the time consumption can be reduced, and the real-time tracking effect can be improved.

[0046] In some embodiments, the obtaining the initial target feature data of the tracking target comprises:

[0047] obtaining an initial image;

[0048] extracting an initial feature map in the initial image by the feature extraction network;

[0049] obtaining initial position data of the tracking target;

[0050] determining the initial target feature data according to the initial feature map and the initial position data.

[0051] In the embodiment, the initial image represents an initialization frame, i.e. an image when the tracking process has not yet started to track the tracking target. The initial feature map is a feature map extracted from the initial image. The initial position data represents the position coordinates of the tracking target in the initial image. The initial feature map and the initial position data are taken as inputs of convolution operation, e.g. ROIAlign operation, to perform correlation convolution operation on the initial feature map with the initial position data as the convolution kernel, so as to obtain the initial target feature data.

[0052] Optionally, the obtaining the initial position data of the tracking target comprises at least one of the following:

[0053] outputting the initial position data of the tracking target by a detection network in the target detection model based on the initial feature map; or

[0054] outputting detected initial targets and initial position data of each of the initial targets by the detection network in the target detection model based on the initial feature map, and labeling the initial targets on the initial image displayed on the user interface, obtaining the tracking target selected from the initial targets based on the user interface, and obtaining the initial position data of the tracking target; or

[0055] displaying the initial image on the user interface, obtaining a configured setting target and a target region of the setting target based on the user interface, taking the setting target as the tracking target, and taking the target region of the setting target as the initial position data of the tracking target.

[0056] In the embodiment, the initial position data can be initial position data of an initial target detected by the target detection model from the initial feature map, and the detected initial target is directly taken as the tracking target, and the initial position data of the initial target is directly taken as the initial position data of the tracking target. The detected initial target can also be displayed on the user interface, and the user can select the tracking target to be tracked, and the selected initial target is taken as the tracking target. The tracking target can also be a target in a target region circled by the user in the initial image, and the target is the setting target.

[0057] In the above embodiment, the initialization frame image is detected in the target detection model to obtain initial target feature data of the tracking target, and the subsequent target tracking model can determine the tracking target in a search region of the current frame according to the initial target feature data and track the tracking target. The present application can jointly use the target detection model and the target tracking model in the same process, effectively reduce the consumption of computing resources, and make the algorithm more easily realize real-time deployment.

[0058] In some embodiments, the present application also provides a training method based on a joint model, the joint model comprising a target detection model and a target tracking model, the method comprising:

[0059] obtaining a training data set, each sample in the training data set comprising sample images of consecutive frames and label data of the sample images, wherein the label data comprises target label data of the sample images and tracking label data in the sample images;

[0060] based on the training data set, iteratively training the joint model, outputting sample detection data and sample tracking data in the current iteration, and using a loss function to calculate a current total loss value in the current iteration based on the sample detection data and the sample tracking data in the current iteration, and based on the current total loss value, continuing to obtain sample images from the training data set for training until a pre-trained joint model is obtained, wherein the target detection model and the target tracking model share the same feature extraction network.

[0061] Optionally, the use of the loss function to calculate the current total loss value in the current iteration based on the sample detection data and the sample tracking data in the current iteration comprises:

[0062] calculating a current detection loss based on the sample detection data in the current iteration and the target label data;

[0063] calculating a current tracking loss based on the sample tracking data in the current iteration and the tracking label data;

[0064] calculating the current total loss value based on the current detection loss and the current tracking loss.

[0065] In the present embodiment, the training data set can be consecutive frame sample images collected in different scenes. The tracking label data includes but is not limited to position data of a bounding box of a sample target, coordinates of a center point of the bounding box of the sample target, motion trajectory data of the sample target, time sequence association information of the sample target, scale and morphological change information of the sample target, etc. The joint model and the target detection model are trained in the same process, which can effectively reduce the consumption of computing resources and make the algorithm more easily realize real-time deployment. Figure 3The network structure diagram shown in the middle is the same. In turn, the continuous frame sample image is obtained from the training data set and input into the joint model for training, and the target detection model and the target tracking model are jointly trained. The feature extraction network extracts the sample feature map in the continuous frame sample image, and the sample feature map is input into the detection network and the tracking network during training, respectively. The sample detection data in the current iteration is output through the detection network in the current iteration, and the sample tracking data in the current iteration is output through the tracking network in the current iteration. The sample detection data represents the position data of the detected sample target, including but not limited to: the position coordinates of the bounding box, the coordinates of the center point of the bounding box, etc. The sample tracking data represents the tracking data of the sample target.

[0066] In this embodiment, the loss function includes but is not limited to the mean square error function and the cross-entropy function. In the training process, the target detection model and the target tracking model are jointly trained. The core of the target detection model is to learn the spatial appearance features (such as shape, color, and texture) and positioning ability of the sample target from a single frame image; the target tracking model needs to learn the time sequence correlation features (such as motion trend, position change rule, and shape evolution under occlusion) of the sample target in the video sequence. When jointly training, the two types of models can share the bottom layer feature extractor (such as the shallow features of the convolutional neural network), and the high-level features penetrate each other. The accurate appearance features learned by the target detection model can help the tracking model better cope with the "appearance mutation" of the tracked target (such as the angle change of the vehicle when turning, the posture switching of the pedestrian); the time sequence information learned by the tracking can be fed back to the target detection model, so that it can better understand the "reasonable existence area of the target" (such as in the video, the moving pedestrian will not suddenly appear in the air), and reduce the "false detection" (such as misjudging the background as the target) or "missed detection" (such as the target being temporarily occluded) of the detection.

[0067] By joint training, the total loss value of the two models during the iteration process is calculated, and the network parameters in the joint model are updated according to the total loss value, so that better complementarity between the two models can be formed. The supervision focuses of the detection loss and the tracking loss are different, and after being added, they can form a comprehensive supervision signal covering more dimensions of target perception. The detection loss (such as the bounding box regression loss) can supervise the model to learn the spatial geometric features of the target (such as width, height, and center coordinates), ensuring the accuracy of single-frame positioning. The tracking loss (such as the appearance similarity loss) can supervise the model to learn the time-invariant features of the target (such as the core features that can be stably extracted even if the target rotates or is partially occluded), ensuring the reliability of cross-frame association. The optimization of the neural network relies on gradient backpropagation, and the loss function is the source of the gradient. After adding the detection loss and the tracking loss, the gradient of the total loss will contain the optimization requirements of both tasks, guiding the feature extractor (such as CNN, Transformer, etc.) to learn spatial-temporal fusion features: the gradient of the detection loss will push the feature extraction network to focus on local details (such as the edges and textures of the target), improving the positioning accuracy; the gradient of the tracking loss will push the feature extraction network to focus on global stability (such as the overall contour and motion pattern of the target), improving the cross-frame association ability. This collaborative gradient update enables the shared feature extraction network (such as the backbone network) to have both spatial resolution and temporal coherence. For example, in an autonomous driving scenario, the feature extraction network of the joint model can learn both the detailed features of the vehicle (such as license plates and headlights) through the detection loss (for accurate positioning) and the vehicle's trajectory features (such as the motion patterns of straight driving and turning) through the tracking loss, ultimately outputting features that support both accurate detection and stable tracking.

[0068] In real-world scenarios (such as video surveillance, autonomous driving, and drone tracking), targets often face challenges such as occlusion, rapid movement, sudden changes in lighting, and background interference, making individual detection or tracking models prone to failure. For example, a single detection model is sensitive to "occluded targets", and if a target is completely occluded, single-frame detection may directly miss it. A single tracking model is easily affected by cumulative errors: if there is a deviation in the initial frame detection, or the target's appearance changes too much after a long period of occlusion, the tracking may drift (such as mistaking target A for target B). During joint training, the two models can form a mutual correction mechanism. The temporal association information of tracking (such as target motion trajectory prediction) can provide candidate region hints for the target detection model, reducing the detection range in occlusion scenarios and improving the recall rate of detection; the accurate positioning results of the target detection model can periodically calibrate the tracking trajectory, avoiding drift caused by error accumulation in tracking (for example, updating the target position of tracking every N frames using the detection results).

[0069] The separately trained detection and tracking models are usually executed in series: first run detection on each frame (or interval frame) of the video, and then initialize / update tracking using the detection results. This approach has two aspects of redundancy: feature extraction redundancy: detection and tracking need to extract features respectively, which repeats the calculation; inference step redundancy: the high frequency of detection (such as detection per frame) will consume a large amount of computing resources (especially in high-resolution videos). Joint training can optimize efficiency through shared feature networks and dynamic scheduling mechanisms: share the bottom feature extractor (such as CNN backbone), reduce the total amount of parameters and calculation; the tracking model can predict the target position based on the temporal continuity, and only trigger detection when the tracking confidence is low (such as the target suddenly accelerates, the end of occlusion), which reduces the calling frequency of detection while ensuring accuracy.

[0070] In dynamic scenes (such as crowded streets, high-speed vehicles), the identity consistency of the target (i.e. the same target is continuously identified as the same object in multiple frames) is crucial. When trained separately, the detection model may judge the same target as multiple different objects due to changes in target appearance (such as angle, distance); the tracking model can maintain identity, but relies on the accuracy of the initial detection. Joint training can force the model to learn spatiotemporal consistency features: the detection model will refer to the historical identity information of the tracking model when predicting the target class and position, avoiding the same target being labeled repeatedly; the tracking model will combine the current appearance features of the detection when updating the trajectory, ensuring that even if the target appearance changes greatly (such as a pedestrian wearing a hat), the identity consistency can still be maintained.

[0071] In practical applications (such as security monitoring, robot vision), separately deploying detection and tracking models requires additional design of interaction logic (such as how to pass the detection results to tracking, how to re-trigger detection when tracking fails), increasing the complexity of the system. Joint training can integrate the two models into an end-to-end unified framework (such as a joint model based on Transformer), with the output directly containing target position and trajectory data, without the need for additional interaction logic. Reducing the engineering cost of model deployment (such as not needing to separately optimize the interface adaptation of the two models); reducing the delay in real-time scenarios (such as in autonomous driving, a millisecond-level delay may affect the safety of decision-making).

[0072] The joint model-based target tracking method of the present application can be applied to autonomous driving scenarios and video monitoring scenarios. Autonomous driving scenario: when the vehicle detection is occluded by pedestrians, the motion trajectory prediction of the tracking can assist the detection in positioning the vehicle position; the detection result corrects the trajectory deviation of the tracking, avoiding the risk of collision. Video monitoring scenario: in a crowded scene, the joint model can simultaneously ensure accurate positioning detection and continuous tracking of each person without loss, improving the reliability of behavior analysis.

[0073] In summary, joint training, through three core mechanisms—spatial-temporal feature fusion, mutual error correction, and efficient resource utilization—upgrades target detection and tracking from independent operation to collaborative perception, ultimately achieving comprehensive improvements in robustness, efficiency, and consistency, and becoming more adaptable to complex dynamic scenarios in the real world.

[0074] like Figure 4 As shown, Figure 4 This is a schematic diagram of a network for calculating initial target feature data in one embodiment. First, the initial image is input into the feature extraction network to obtain an initial feature map. The initial feature map is then input into the detection network to output the initial position data of the initial target. Based on the initial position data of the initial target, the initial target can be selected as the tracking target, or the user can set the target as the tracking target in the initial image to obtain the initial position data of the tracked face. The initial feature map and the initial position data of the tracking target are used as inputs to the ROIAlign operation to obtain the initial target feature data.

[0075] like Figure 5 As shown, Figure 5 This is a network diagram of a target tracking method based on a joint model in one embodiment. After initializing the initial target feature data, it can be used in subsequent tracking processes. The current frame input image is used as input to the feature extraction network to obtain the current feature map. The current feature map is then used as input to the detection network to output the current target detection result. The current feature map and the search region position data obtained from the tracking data of the previous frame are used as input to the ROIAlign operation, performing a convolution operation to obtain the current search region feature map. The current search region feature map and the initial target feature data are then convolved again to obtain the current target feature map. Finally, the tracking head performs the operation to obtain the current tracking data, continuing to update the current frame and continue tracking.

[0076] In another aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the joint model-based target tracking method described in any embodiment of this application.

[0077] In the computer program product, the optional implementation form of the program module architecture for implementing the steps of the joint model-based target tracking method can be a joint model-based target tracking method apparatus. Please refer to [link to relevant documentation]. Figure 6An embodiment of the present application provides a target tracking method and device based on a joint model, comprising: a obtaining module 61 configured to obtain a current frame input image of a target detection model, wherein the target detection model shares a feature extraction network with a target tracking model; an extracting module 62 configured to extract a current feature map in the current frame input image through the feature extraction network; the obtaining module 61 is further configured to obtain initial target feature data of a tracking target; a tracking module 63 configured to track the tracking target in a current frame based on the current feature map and the initial target feature data, and obtain current tracking data of the tracking target by using a tracking network in the target tracking model.

[0078] Optionally, the tracking module 63 is further configured to:

[0079] calculate a current search region feature map based on the current feature map;

[0080] calculate a current target feature map based on the current search region feature map and the initial target feature data;

[0081] track the tracking target based on the current target feature map by using the tracking network in the target tracking model, and obtain current tracking data of the tracking target.

[0082] Optionally, the tracking module 63 is further configured to:

[0083] obtain search region position data;

[0084] calculate a current search region feature map based on the current feature map and the search region position data.

[0085] Optionally, the tracking module 63 is further configured to:

[0086] when the current frame is not a first frame for tracking the tracking target, obtain tracking region data of the tracking target in a previous frame of the current frame, expand the tracking region based on the tracking region data, and obtain the search region position data;

[0087] when the current frame is the first frame for tracking the tracking target, obtain initial position data of the tracking target, and obtain the search region position data based on the initial position data.

[0088] Optionally, the obtaining module 61 is further configured to:

[0089] obtain an initial image;

[0090] extract an initial feature map in the initial image through the feature extraction network;

[0091] obtaining initial position data of the tracking target;

[0092] determining the initial target feature data according to the initial feature map and the initial position data.

[0093] Optionally, the obtaining module 61 is further configured to:

[0094] outputting the initial position data of the tracking target by using a detection network in the target detection model based on the initial feature map; or

[0095] outputting detected initial targets and initial position data of each initial target by using a detection network in the target detection model based on the initial feature map, and labeling the initial targets on the initial image displayed on the user interface, obtaining the tracking target selected from the initial targets based on the user interface, and obtaining the initial position data of the tracking target; or

[0096] displaying the initial image on the user interface, obtaining a configured setting target and a target region of the setting target based on the user interface, taking the setting target as the tracking target, and taking the target region of the setting target as the initial position data of the tracking target.

[0097] The above mainly describes the solutions provided by the embodiments of the present application from the perspective of methods. In order to implement the above functions, the target tracking method and device based on a joint model comprises a hardware structure and / or a software module corresponding to each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0098] The embodiments of the present application can divide the function modules of the target tracking method and device based on a joint model according to the above method. For example, the target tracking method and device based on a joint model can comprise function modules corresponding to each function division, or two or more functions can be integrated in one processing module. The above integrated module can be implemented in the form of hardware or software function module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. When actually implemented, another division mode can be used.

[0099] The processor 13 is a control center, which connects various parts of the target tracking device through various interfaces and lines, and performs various functions of the target tracking device and processes data by running or executing software programs and / or modules stored in the memory 14 and calling data stored in the memory 14. Optionally, the processor 13 can include one or more processing cores; preferably, the processor 13 can integrate an application processor and a modem processor, wherein the application processor mainly processes operating systems, user pages and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 13.

[0100] The memory 14 can be used to store software programs and modules, and the processor 13 executes various function applications and data processing by running the software programs and modules stored in the memory 14. The memory 14 can mainly include a program storage area and a data storage area, wherein the program storage area can store operating systems, application programs required by at least one function (such as sound playing function, image playing function, etc.), etc.; and the data storage area can store data created according to the use of a target tracking device, etc. In addition, the memory 14 can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory 14 can also include a memory processor to provide access of the processor 13 to the memory 14.

[0101] In another aspect, the embodiment of the present application also provides a computer readable non-volatile storage medium storing a computer program, and the computer program is executed by a processor to make the processor execute the steps of the joint model-based target tracking method provided by any of the above-mentioned embodiments of the present application.

[0102] In another aspect, the embodiment of the present application provides a computer program product including a computer program, and the computer program is executed by a processor to implement the steps of the joint model-based target tracking method and / or the joint model-based training method according to any of the embodiments of the present application.

[0103] In another aspect, the embodiment of the present application provides a computer program product including a computer program, and the computer program is executed by a processor to implement the steps of the joint model-based target tracking method and / or the joint model-based training method according to any of the embodiments of the present application.

[0104] Those skilled in the art can understand that all or part of the processes in the methods provided by the above embodiments can be completed by instructing the relevant hardware by a computer program. The program can be stored in a non-volatile computer readable storage medium, and when the program is executed, the processes of the above embodiments of the methods can be included. Any reference to memory, storage, database, or other medium used in the embodiments of the present application can include non-volatile and / or volatile memory. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. The volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0105] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered by the protection scope of the present application. The protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A target tracking method based on a joint model, characterized in that, The joint model includes a target detection model and a target tracking model, and the method includes: Obtain the current frame input image of the target detection model, wherein the target detection model and the target tracking model share a feature extraction network; The feature extraction network extracts the current feature map from the current frame input image; Obtain initial target feature data of the tracked target; Based on the current feature map and the initial target feature data, the tracking network in the target tracking model is used to track the target in the current frame to obtain the current tracking data of the target.

2. The target tracking method based on a joint model as described in claim 1, characterized in that, The step of tracking the target in the current frame using the tracking network in the target tracking model based on the current feature map and the initial target feature data to obtain the current tracking data of the target includes: Based on the current feature map, calculate the feature map of the current search region; Calculate the current target feature map based on the current search region feature map and the initial target feature data; Based on the current target feature map, the tracking network in the target tracking model is used to track the target, thereby obtaining the current tracking data of the target.

3. The target tracking method based on a joint model as described in claim 2, characterized in that, The step of calculating the current search region feature map based on the current feature map includes: Obtain the location data of the search area; Based on the current feature map and the search region location data, calculate the current search region feature map.

4. The target tracking method based on a joint model as described in claim 3, characterized in that, The acquisition of search area location data includes: When the current frame is not the first frame for tracking the target, the tracking area is obtained based on the tracking data of the target in the previous frame of the current frame, and the tracking area is expanded to obtain the search area position data. When the current frame is the first frame for tracking the target, the initial position data of the target is obtained, the tracking area is obtained based on the initial position data, and the tracking area is expanded to obtain the search area position data.

5. The target tracking method based on a joint model as described in claim 1, characterized in that, The initial target feature data for acquiring the tracked target includes: Get the initial image; The initial feature map is extracted from the initial image using the feature extraction network. Obtain the initial position data of the tracked target; The initial target feature data is determined based on the initial feature map and the initial position data.

6. The target tracking method based on the joint model as described in claim 5, characterized in that, The acquisition of the initial position data of the tracked target includes at least one of the following: Based on the initial feature map, the initial position data of the tracked target is output using the detection network in the target detection model; or Based on the initial feature map, the detection network in the target detection model outputs the detected initial targets and the initial position data of each initial target, and the initial targets are marked on the initial image displayed on the user interface. Based on the user interface, the tracking target selected from the initial targets is obtained, and the initial position data of the tracking target is obtained. or The initial image is displayed on the user interface. Based on the user interface, the configured target and the target area of ​​the target are obtained. The set target is used as the tracking target, and the target area of ​​the set target is used as the initial position data of the tracking target.

7. A training method based on a joint model, characterized in that, The joint model includes a target detection model and a target tracking model, and the method includes: Obtain a training dataset, wherein each sample in the training dataset includes sample images of consecutive frames and label data of the sample images, wherein the label data includes target label data of the sample images and tracking label data in the sample images; Based on the training dataset, the joint model is trained iteratively, and the sample detection data and sample tracking data of the current iteration are output. Using the loss function, the current total loss value of the current iteration is calculated based on the sample detection data and sample tracking data of the current iteration. Based on the current total loss value, sample images are obtained from the training dataset for training until a pre-trained joint model is obtained, wherein the target detection model and the target tracking model share the same feature extraction network.

8. The training method based on the joint model as described in claim 7, characterized in that, The calculation of the current total loss value in the current iteration using the loss function, based on the sample detection data and sample tracking data in the current iteration, includes: Calculate the current detection loss based on the sample detection data and the target label data in the current iteration; Calculate the current tracking loss based on the sample tracking data and tracking position label data in the current iteration; The current total loss value is calculated based on the current detection loss and the current tracking loss.

9. A target tracking device, characterized in that, The system includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the joint model-based target tracking method as described in any one of claims 1 to 6 and / or the joint model-based training method as described in any one of claims 7 or 8.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the target tracking method based on the joint model as described in any one of claims 1 to 6 and / or the training method based on the joint model as described in any one of claims 7 or 8.