Multi-target tracking method and system, electronic equipment and readable storage medium
By combining the self-attention mechanism of Yolo model and Transformer structure, the problem of low data processing reliability and accuracy in multi-objective tracking is solved, and a more efficient and accurate multi-objective tracking effect is achieved.
Patent Information
- Application Number
- CN202411908197.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the data processing method of multi-objective tracking is poor in reliability and low in accuracy, making it difficult to effectively solve the problem of tracking error and target interference in complex scenarios in multi-objective tracking.
Using a multi-objective tracking method based on the self-attention mechanism of the Yolo model and the Transformer structure, the object detection is performed through the Yolo model, the initial detection box and feature information are obtained, and the degree of correlation between the targets is calculated using the Transformer structure, and the status information is input into the actor-critic model to correct the detection box and trajectory prediction is performed.
It improves the accuracy and reliability of multi-objective tracking, enhances the processing ability of complex scenarios, reduces tracking errors, and is suitable for a variety of task scenarios.
Smart Images

Figure CN120047484A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and computer vision, and in particular, to a multi-object tracking method, system, electronic device and readable storage medium. Background Art
[0002] Object tracking is a technology that synthesizes the context information of a video or image sequence, models the appearance and motion information of an object, and predicts the motion state and calibrates the position of the object. It is an important problem in the field of computer vision and has important research significance and application value. It is applied in video surveillance, human-computer interaction, visual navigation, etc.
[0003] Object tracking can be divided into single-object tracking and multi-object tracking. In the multi-object tracking task, without knowing the number of objects in advance, it is necessary to detect, assign IDs and track multiple objects such as people, cars, animals, etc. in the video. Different objects have different IDs for tasks such as trajectory prediction. In the multi-object tracking task, various complex problems need to be faced. In addition to the problems existing in single-object tracking such as occlusion, deformation, motion blur, crowded scenes, illumination changes, scale changes, etc., it is also necessary to face problems such as interference between similar objects and trajectory initialization and termination in multi-object tracking. Therefore, this research direction is still challenging.
[0004] In the related art, the reliability of the data processing method for multi-object tracking is poor and the accuracy is not high. How to provide a stable, reliable and highly accurate multi-object tracking method is an urgent problem to be solved at present. Summary of the Invention
[0005] In order to solve or improve at least one of the above technical problems, an object of the present invention is to provide a multi-object tracking method.
[0006] Another object of the present invention is to provide a multi-object tracking system.
[0007] Another object of the present invention is to provide an electronic device.
[0008] Another object of the present invention is to provide a readable storage medium.
[0009] To achieve the above object, a first aspect of the present invention provides a multi-object tracking method, and the steps include:
[0010] In the first step, based on the Yolo (i.e., YOLO, You only Look Once) model, perform object detection on each frame of the video or image sequence, and determine the detection result; wherein, the detection result includes the initial detection boxes and feature information of multiple objects.
[0011] In the second step, based on the self-attention mechanism of the Transformer structure (i.e., the transducer structure or the conversion structure), the correlation degree between multiple targets is calculated according to the feature information, so that information interaction is carried out between the multiple targets, and the state information is determined; wherein, the state information includes the distribution information and the correlation degree of multiple targets in each frame of image.
[0012] In the third step, the state information is input into the actor-critic (policy neural network and critic neural network) model. The actor-critic model includes a policy neural network. The policy neural network corrects the position and size of the initial detection box according to the state information and determines the corrected detection box.
[0013] In the fourth step, the action trajectories of multiple targets are determined according to the corrected detection boxes, and trajectory prediction is performed to track the multiple targets.
[0014] The present invention aims to provide a multi-target tracking method. By introducing the self-attention mechanism of the Transformer structure, calculating the correlation degree between the feature information of each target and the feature information of other targets is beneficial to improving the ability to read the global information of the current frame of image, thereby improving the accuracy of the tracking results (action trajectories and trajectory prediction), and the reliability is higher. In addition, the multi-target tracking method of the present invention is based on the Yolo model and the actor-critic model, and adopts the method of deep reinforcement learning, which can track targets with different motion characteristics and is applicable to a variety of task scenarios.
[0015] In addition, the above technical solution provided by the present invention may also have the following additional technical features:
[0016] In the above technical solution, optionally, target detection is performed on each frame of image in a video or image sequence based on the Yolo model, and the detection result is determined; wherein, the detection result includes the initial detection boxes and feature information of multiple targets. The steps include: The Yolo model includes a convolutional neural network. Each frame of image in the video or image sequence is divided into multiple grids through the convolutional neural network, and the bounding box and class probability of the target in each grid are predicted; the detection result is determined according to the bounding box and class probability; wherein, the detection result includes the initial detection boxes and feature information of multiple targets.
[0017] In this technical solution, the convolutional neural network divides each frame of image into multiple grids, and takes the grid as a unit to perform preliminary analysis on the image to determine the positioning and recognition of the target. The convolutional neural network extracts features from the image through convolutional layers, pooling layers and fully connected layers, and predicts the bounding box and class probability of the target in each grid. The convolutional neural network determines the detection result (the initial detection boxes and feature information of multiple targets) according to the bounding box and class probability of the target.
[0018] By introducing the Yolo model for target detection, the processing performance of this method is improved, global information can be considered during detection, local interference is reduced, and the detection has higher accuracy.
[0019] In the above technical solution, optionally, the convolutional neural network includes a convolutional layer, a pooling layer and a fully connected layer. The convolutional neural network extracts features of the image through the convolutional layer, the pooling layer and the fully connected layer, and predicts the bounding box and category probability of the target in each grid.
[0020] In this technical solution, the convolutional layer extracts features from the image through a series of convolution kernels. These convolution kernels slide on the image according to the set step size to perform convolution operations, thereby extracting local features and determining feature maps. The convolutional layer can effectively capture local features of the target in the image, such as edges, textures, etc.
[0021] The pooling layer is used to downsample the feature map. This method can reduce the data dimension while retaining key feature information to reduce the amount of subsequent calculations.
[0022] The feature map obtained after the convolution layer may be large in size. Through the downsampling operation of the pooling layer, the size of the feature map can be reduced. This not only reduces the amount of calculation of the subsequent fully connected layer, but also can extract more representative features to a certain extent.
[0023] The fully connected layer is used to integrate the feature maps extracted by the convolutional layer and the pooling layer, and map the integrated feature maps to the dimensions corresponding to the grid.
[0024] For each grid, the convolutional neural network predicts the relevant parameters of the target's bounding box through a fully connected layer. The relevant parameters of the bounding box include but are not limited to the offset of the center coordinates of the bounding box relative to the grid, and the width and height of the bounding box.
[0025] In addition, the convolutional neural network predicts the category probability of the target through the fully connected layer. The category probability here can be understood as the probability distribution of the category. In this way, the probability value of each target in the grid belonging to each category such as human, vehicle, animal, etc. can be predicted.
[0026] In the above technical solution, optionally, the steps of the multi-target tracking method also include: performing target detection on each frame image in the video or image sequence based on the Yolo model and before determining the detection result, obtaining image data related to the task scene, the image data including the video and / or image sequence; constructing a data set based on the image data, and training the Yolo model through the data set.
[0027] In this technical solution, according to the specific task scenario, relevant videos and / or image sequences are collected to construct a dataset suitable for training and testing the model.
[0028] For the target detection requirements in a specific task scenario, the constructed dataset is used to train the Yolo model. By adjusting the parameters of the model, it can accurately detect the targets in the task scenario and improve the accuracy of bounding box prediction and class classification.
[0029] In the above technical solution, optionally, the steps of the multi-object tracking method further include: before constructing the dataset according to the image data and training the Yolo model through the dataset, performing data cleaning and data augmentation on the image data.
[0030] In this technical solution, data cleaning is used to remove noise, misannotations, and redundant information in the image data.
[0031] The ways of data augmentation include but are not limited to flipping, cropping, and adding noise.
[0032] The purpose of data augmentation is to expand the scale of the dataset, increase data diversity, thereby enhancing the robustness and generalization ability of the model.
[0033] In the above technical solution, optionally, the state information is input into the actor-critic model. The actor-critic model includes a policy neural network. The policy neural network modifies the position and size of the initial detection box according to the state information and determines the corrected detection box. The steps include: inputting the state information into the actor-critic model; the actor-critic model includes a policy neural network, and the policy neural network determines the horizontal coordinate change amount, vertical coordinate change amount, and size change amount of the initial detection box according to the state information; according to the horizontal coordinate change amount, vertical coordinate change amount, and size change amount, modifying the position and size of the initial detection box and determining the corrected detection box.
[0034] In this technical solution, the policy neural network determines the change amount of the initial detection box according to the state information. The change amount of the initial detection box includes the horizontal coordinate change amount Δx of the upper left vertex of the initial detection box. The change amount of the initial detection box also includes the vertical coordinate change amount Δy of the upper left vertex of the initial detection box. The change amount of the initial detection box also includes the size change amount Δs of the initial detection box.
[0035] The policy neural network modifies the position and size of the initial detection box according to the calculation formula and determines the corrected detection box.
[0036] The calculation formula is x’ = x + Δx; y’ = y + Δy; h’ = h + Δs × h; w’ = w + Δs × w.
[0037] Among them, x' represents the abscissa of the corrected detection box; x represents the abscissa of the initial detection box. y' represents the ordinate of the corrected detection box; y represents the ordinate of the initial detection box. h' represents the height of the corrected detection box; h represents the height of the initial detection box. w' represents the width of the initial detection box; w represents the width of the corrected detection box.
[0038] The corrected detection box can accurately locate the position and range of the target in the current frame image, providing a more reliable data basis for subsequent tracking steps, which is conducive to continuously and stably performing precise tracking on multiple targets. Even in the case of variable target motion states and complex scenes, it can still maintain high tracking accuracy and robustness.
[0039] In the above technical solution, optionally, the steps of the multi-target tracking method further include: before determining the action trajectories of multiple targets based on the corrected detection box and performing trajectory prediction to track multiple targets, calculating the intersection over union (IoU) between the corrected detection box of the target in the current frame image and the corrected detection box of the target in the next frame image; determining the reward value according to the intersection over union; and adjusting the model parameters according to the reward value to optimize the actor-critic model.
[0040] In this technical solution, the intersection over union (IoU) is a metric widely used in fields such as object detection and object tracking.
[0041] The larger the intersection over union, the higher the degree of overlap between the predicted target position and the actual target position, which also means higher tracking accuracy.
[0042] When the intersection over union is greater than or equal to 0.7, the reward value is 1. In this case, it will guide the actor-critic model to adjust the model parameters in the direction of making the intersection over union higher, so as to track the target more accurately.
[0043] When the intersection over union is less than 0.7, the reward value is -1. In this case, it will guide the actor-critic model to avoid making tracking actions that result in too low an intersection over union.
[0044] In this way, the intersection over union, as a feedback signal, effectively helps the actor-critic model perform gradient updates of the model parameters, thereby optimizing the actor-critic model.
[0045] The second aspect of the present invention provides a multi-target tracking system, including an object detection module, a state information determination module, a detection box correction module, and an action trajectory determination module.
[0046] The target detection module is used to perform target detection on each frame of an image in a video or image sequence based on the Yolo model and determine the detection results. Among them, the detection results include the initial detection boxes and feature information of multiple targets.
[0047] The state information determination module is used to calculate the degree of association between multiple targets based on the self-attention mechanism of the Transformer structure according to the feature information, so that information interaction occurs between multiple targets and determine the state information. Among them, the state information includes the distribution information and degree of association of multiple targets in each frame of the image.
[0048] The detection box correction module is used to input the state information into the actor-critic model. The actor-critic model includes a policy neural network, and the policy neural network corrects the position and size of the initial detection box according to the state information and determines the corrected detection box.
[0049] The action trajectory determination module is used to determine the action trajectories of multiple targets according to the corrected detection boxes and perform trajectory prediction to track multiple targets.
[0050] The present invention aims to provide a multi-target tracking system. By introducing the self-attention mechanism of the Transformer structure to calculate the degree of association between the feature information of each target and the feature information of other targets, it is beneficial to improve the ability to read the global information of the current frame image, thereby improving the accuracy of the tracking results (action trajectories and trajectory predictions) and having higher reliability. In addition, the multi-target tracking system of the present invention is based on the Yolo model and the actor-critic model and uses the method of deep reinforcement learning, which can track targets with different motion characteristics and is applicable to various task scenarios.
[0051] The third aspect of the present invention provides an electronic device, including a memory and a processor. Among them, a program or instruction that can run on the processor is stored on the memory, and when the processor executes the program or instruction, the steps of the multi-target tracking method in any of the above technical solutions are implemented. Therefore, the electronic device has the beneficial effects of any of the above technical solutions and will not be described in detail here.
[0052] The fourth aspect of the present invention provides a readable storage medium. The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the multi-target tracking method in any of the above technical solutions are implemented. Therefore, the readable storage medium has the beneficial effects of any of the above technical solutions and will not be described in detail here.
[0053] The additional aspects and advantages of the technical solutions of the present invention will become obvious in the following description part or be understood through the practice of the present invention. Description of the Drawings
[0054] Figure 1 The flowchart of a multi - target tracking method according to an embodiment of the present invention is shown;
[0055] Figure 2 The flowchart of a multi - target tracking method according to another embodiment of the present invention is shown;
[0056] Figure 3 The flowchart of a multi - target tracking method according to another embodiment of the present invention is shown;
[0057] Figure 4 The flowchart of a multi - target tracking method according to another embodiment of the present invention is shown;
[0058] Figure 5 The flowchart of a multi - target tracking method according to another embodiment of the present invention is shown;
[0059] Figure 6 The flowchart of a multi - target tracking method according to another embodiment of the present invention is shown;
[0060] Figure 7 The structural block diagram of a multi - target tracking system according to an embodiment of the present invention is shown;
[0061] Figure 8 The structural block diagram of an electronic device according to an embodiment of the present invention is shown;
[0062] Figure 9 The flowchart of a method for determining the comprehensive state of multiple targets according to an embodiment of the present invention is shown;
[0063] Figure 10 The flowchart of a data processing method based on an actor - critic model according to an embodiment of the present invention is shown.
[0064] Wherein, Figure 7 and Figure 8 The corresponding relationship between the reference numerals and the component names in the figure is:
[0065] 200: multi - target tracking system; 210: target detection module; 220: state information determination module; 230: detection box correction module; 240: action trajectory determination module; 300: electronic device; 310: memory; 320: processor. Detailed implementation manners
[0066] In order to more clearly understand the above - mentioned objects, features, and advantages of the embodiments of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0067] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, embodiments of the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present application is not limited to the limitations of the specific embodiments disclosed below.
[0068] Reference is now made to Figures 1 to 10 Describe a multi-object tracking method, a multi-object tracking system, an electronic device, and a readable storage medium provided according to some embodiments of the present invention.
[0069] In one embodiment of the present invention, as Figure 1 shown, the steps of the multi-object tracking method include:
[0070] S102, performing object detection on each frame of an image sequence in a video or an image based on a Yolo model, and determining a detection result; wherein, the detection result includes initial detection frames and feature information of multiple objects.
[0071] It should be noted that the Yolo (i.e., YOLO, You only Look Once) model is an object detection model based on a convolutional neural network (CNN, Convolutional Neural Network). The Yolo model adopts an end-to-end architecture, starting directly from the input image pixels, passing through a neural network, and outputting the class probabilities and bounding box positions of the objects at one time. This approach gives it a great advantage in data processing speed and can quickly complete object detection.
[0072] Optionally, the Yolo model includes a convolutional neural network. Each frame of an image sequence in a video or an image is divided into multiple grids through the convolutional neural network, the bounding boxes and class probabilities of the objects in each grid are predicted, and the detection result is determined.
[0073] The convolutional neural network is a deep learning model specifically used to process data with a grid structure. The Yolo model utilizes the powerful feature extraction ability of the convolutional neural network to process each frame of an image sequence in a video or an image.
[0074] The classes of the objects include but are not limited to people, vehicles, and animals.
[0075] Optionally, the convolutional neural network divides each frame of an image into multiple grids, and performs a preliminary analysis of the image in units of grids to determine the positioning and recognition of the objects.
[0076] By dividing each frame of an image into multiple grids, the convolutional neural network can perform a local-to-global analysis of the image in units of grids.
[0077] Optionally, a convolutional neural network extracts features from an image through convolutional layers, pooling layers, and fully connected layers, and predicts the bounding boxes and class probabilities of objects in each grid.
[0078] Through operations such as multi-layer convolution and pooling of the convolutional neural network, deep feature extraction is performed on the image regions within each grid, and then the bounding boxes and class probabilities of the objects are predicted. This grid-based prediction mechanism can, to a certain extent, take into account the detection of objects of different positions and sizes in the image, and due to the parameter sharing feature of the convolutional neural network, it can reduce the complexity of the model while improving the detection efficiency.
[0079] Specifically, the convolutional layer is used to extract local features of the image; the pooling layer reduces the dimensionality of the features to reduce the computational amount and retain key information; the fully connected layer integrates and maps the extracted features, and finally predicts the bounding boxes and class probabilities of each grid.
[0080] Optionally, the convolutional neural network determines the detection results (initial detection boxes and feature information of multiple objects) based on the bounding boxes and class probabilities of the objects.
[0081] The purpose of this step is to determine the positions (bounding boxes) where objects may exist in the image and the classes to which the objects belong, so as to obtain the initial detection boxes of the objects and provide basic data for subsequent tracking.
[0082] It should be noted that the feature information includes but is not limited to shape features, color features, texture features, and class probabilities.
[0083] Shape features are used to reflect the general shape of the object. These shape features can assist in judging whether the object is still the object being tracked before during the subsequent tracking process when the object is partially occluded or the pose changes.
[0084] Color features are used to reflect the color differences between different objects. During the object tracking process, color features are helpful for distinguishing similar objects or re-identifying objects after a short occlusion. For example, in a multi-person tracking scenario where different people wear different colored clothes, color features can be an important distinguishing factor.
[0085] Texture is an important feature of the object's appearance. By performing texture analysis on the image region within the initial detection box, the texture features of the object are extracted. Texture features can enhance the recognition of the object. When the object is deformed or partially occluded, texture features are helpful for re-identifying the object.
[0086] During the multi-object tracking process, when the appearance of the object is blurred or some features are occluded, the class probability can assist in judging the class attribution of the object, and can be combined with other feature information to improve the accuracy of the tracking result.
[0087] By introducing the Yolo model for object detection, the processing performance of the multi-object tracking method is improved. Global information can be considered during detection, local interference is reduced, and the detection has high accuracy.
[0088] S104. Based on the self-attention mechanism of the Transformer structure, calculate the degree of association between multiple objects according to the feature information, so that information interaction occurs between multiple objects, and determine the state information; wherein, the state information includes the distribution information and the degree of association of multiple objects in each frame of image.
[0089] It should be noted that the Transformer structure is also called the transformer structure or conversion structure.
[0090] The purpose of this step is to extract state information from the detection results through the Transformer structure. Specifically, the self-attention mechanism (Self-Attention Mechanism) of the Transformer structure is used to mine the degree of association between objects (calculate the degree of association between the feature information of each object and the feature information of other objects), and promote information interaction between objects, and then determine the state information.
[0091] The advantage of using the self-attention mechanism is that it calculates the degree of association between the feature information of each object and the feature information of other objects in parallel, which is beneficial to improving the data processing efficiency.
[0092] Since the Transformer structure can capture complex mutual relationships between objects, such as the relative positions and motion trends of objects. During the tracking process, when objects are occluded, cross, or have complex motion changes, the method provided by the present invention (multi-object tracking method) can anticipate the motion changes of objects in advance, so as to adjust the tracking strategy more accurately.
[0093] Optionally, the Transformer structure also has a multi-head attention mechanism (Multi-Head Attention). The multi-head attention mechanism can extract relationship information between objects more comprehensively.
[0094] Based on the multi-head attention mechanism of the Transformer structure, extract the relationship information between objects. The relationship information includes but is not limited to the spatial position relationship between objects and the class similarity of objects.
[0095] S106. Input the state information into the actor-critic model. The actor-critic model includes a policy neural network. The policy neural network corrects the position and size of the initial detection box according to the state information and determines the corrected detection box.
[0096] It should be noted that the actor-critic (policy neural network and critic neural network) model is a deep reinforcement learning architecture. Among them, the actor (policy neural network) generates actions according to the state information extracted from the Transformer structure, and this action is the adjustment of the initial detection box. The critic (critic neural network) can evaluate whether the action of the actor is effective and whether it helps to better track the target based on the state information and the action through function operations.
[0097] The policy neural network determines the change amount of the initial detection box according to the state information. The change amount of the initial detection box includes the horizontal coordinate change amount Δx of the upper left vertex of the initial detection box. The change amount of the initial detection box also includes the vertical coordinate change amount Δy of the upper left vertex of the initial detection box. The change amount of the initial detection box also includes the size change amount Δs of the initial detection box.
[0098] The policy neural network corrects the position and size of the initial detection box according to the calculation formula and determines the corrected detection box.
[0099] The calculation formula is x’ = x + Δx; y’ = y + Δy; h’ = h + Δs × h; w’ = w + Δs × w.
[0100] Among them, x’ represents the horizontal coordinate of the corrected detection box; x represents the horizontal coordinate of the initial detection box. y’ represents the vertical coordinate of the corrected detection box; y represents the vertical coordinate of the initial detection box. h’ represents the height of the corrected detection box; h represents the height of the initial detection box. w’ represents the width of the initial detection box; w represents the width of the corrected detection box.
[0101] The corrected detection box can accurately locate the position and range of the target in the current frame image, providing a more reliable data basis for subsequent tracking steps, which is conducive to continuously and stably tracking multiple targets accurately. Even in the case of variable motion states of the targets and complex scenes, it can still maintain a high tracking accuracy and robustness.
[0102] S108, determine the action trajectories of multiple targets according to the corrected detection box and perform trajectory prediction to track multiple targets.
[0103] The corrected detection box can accurately locate the position and range of the target in the current frame image. By analyzing the corrected detection boxes in multiple consecutive frames of images, determine the action trajectories of multiple targets and perform trajectory prediction to track multiple targets.
[0104] Optionally, by analyzing the corrected detection boxes in multiple consecutive frames of images, determine the change rules of the position and size of the corrected detection box in the time series, and then determine the action trajectories of multiple targets.
[0105] Predict the future trajectory of a target based on the target's action trajectory, motion characteristics, and historical data.
[0106] The present invention aims to provide a multi-target tracking method. By introducing the self-attention mechanism of the Transformer structure, the correlation degree between the feature information of each target and the feature information of other targets is calculated, which is beneficial to improving the ability to read the global information of the current frame image, and further improving the accuracy of the tracking results (action trajectory and trajectory prediction), with higher reliability. In addition, the multi-target tracking method of the present invention is based on the Yolo model and the actor-critic model, and adopts the method of deep reinforcement learning, which can track targets with different motion characteristics and is applicable to a variety of task scenarios.
[0107] Target tracking is a technology that synthesizes the context information of a video or image sequence, models the appearance and motion information of the target, and predicts the motion state and calibrates the position of the target. It is an important problem in the field of computer vision and has important research significance and application value. It is applied in video surveillance, human-computer interaction, visual navigation, etc.
[0108] Target tracking can be divided into single-target tracking and multi-target tracking. In the multi-target tracking task, without knowing the number of targets in advance, it is necessary to detect, assign IDs (Identifiers, identification numbers) to, and track multiple targets such as people, cars, and animals in the video. Different targets have different IDs for tasks such as trajectory prediction. In the multi-target tracking task, various complex problems need to be faced. In addition to the problems existing in single-target tracking such as occlusion, deformation, motion blur, crowded scenes, illumination changes, and scale changes, problems such as interference between similar targets and trajectory initialization and termination in multi-target tracking also need to be faced. Therefore, this research direction is still challenging.
[0109] The feature-based multi-target tracking method extracts the features of the target (such as color, texture, shape, etc.) for target detection and tracking. This method is easy to implement but has poor robustness. In actual application scenarios, when multiple complex situations occur simultaneously, such as the coexistence of illumination changes and occlusion, it is difficult for this method to maintain stable tracking performance, and it is easy to lose the target or mis-track, unable to meet the high reliability requirements of multi-target tracking in complex environments.
[0110] The multi-object tracking method provided by the present invention improves the ability to read global information of the current frame by means of the self-attention mechanism, so that the state and scene information of the target can be grasped more comprehensively and accurately during the tracking process, and then more accurate tracking can be carried out. When facing complex problems such as occlusion, deformation, motion blur, crowded scenes, illumination changes, scale changes, etc., as well as unique problems in multi-object tracking such as interference between similar targets, trajectory initialization and termination, the multi-object tracking method of the present invention can effectively reduce the tracking error, improve the success rate and accuracy of tracking, and has high robustness.
[0111] It should be noted that robustness refers to the ability of a system (such as an algorithm, a model, a program, etc.) to maintain its performance stable and reliable in the face of various interference factors (such as noise, abnormal data, small deviations in model assumptions, etc.), without significant performance degradation or failure.
[0112] The multi-object tracking method based on filtering uses dynamic system estimation methods to estimate the motion trajectories of targets and uses filters to correct the predicted trajectories. This method relies on assumptions about the target motion model, such as linear motion assumptions, Gaussian noise assumptions, etc. However, in actual scenarios, the motion of targets is often non-linear and non-Gaussian. For example, in traffic scenarios, the acceleration, deceleration, turning and other actions of vehicles are complex and variable, and it is difficult to accurately describe them with a simple linear model. When the actual motion of the target deviates greatly from the assumed model, the correction effect of the filter will be greatly reduced, resulting in an increase in tracking error.
[0113] The multi-object tracking method provided by the present invention, based on the Yolo model and the actor-critic model, adopts the method of deep reinforcement learning, can track targets with different motion characteristics, and is applicable to a variety of task scenarios. The actor-critic model includes a policy neural network, and the policy neural network corrects the position and size of the initial detection box according to the state information and determines the corrected detection box. In this way, the tracking strategy can be adaptively adjusted according to the motion patterns and scene characteristics of different targets, and the versatility is strong.
[0114] In some embodiments, optionally, as Figure 2 shown, the steps of S102 (performing object detection on each frame image in a video or image sequence based on the Yolo model and determining the detection result; wherein the detection result includes the initial detection boxes and feature information of multiple targets) include:
[0115] S1022, the Yolo model includes a convolutional neural network, and each frame image in the video or image sequence is divided into multiple grids through the convolutional neural network, and the bounding boxes and class probabilities of the targets in each grid are predicted.
[0116] Optionally, the convolutional neural network divides each frame of the image into multiple grids, and performs a preliminary analysis of the image in units of grids to determine the target for positioning and recognition.
[0117] By dividing each frame of the image into multiple grids, the convolutional neural network can perform a local-to-global analysis of the image in units of grids.
[0118] Optionally, the convolutional neural network extracts features from the image through convolutional layers, pooling layers, and fully connected layers, and predicts the bounding box and class probability of the target in each grid.
[0119] Through operations such as multi-layer convolution and pooling of the convolutional neural network, deep feature extraction is performed on the image region within each grid, and then the bounding box and class probability of the target are predicted. This grid-based prediction mechanism can, to a certain extent, take into account the detection of targets with different positions and sizes in the image, and due to the parameter sharing feature of the convolutional neural network, it can improve the detection efficiency while reducing the complexity of the model.
[0120] Specifically, the convolutional layer is used to extract the local features of the image; the pooling layer performs dimensionality reduction on the features to reduce the amount of calculation and retain key information; the fully connected layer integrates and maps the extracted features, and finally predicts the bounding box and class probability of each grid.
[0121] S1024, determine the detection result according to the bounding box and class probability; wherein, the detection result includes the initial detection boxes and feature information of multiple targets.
[0122] The purpose of this step is to determine the position (bounding box) where the target may exist in the image and the class to which the target belongs, so as to obtain the initial detection box of the target and provide basic data for subsequent tracking.
[0123] By introducing the Yolo model for target detection, the processing performance of this method is improved, global information can be considered during detection, local interference is reduced, and the detection has high accuracy.
[0124] In a specific embodiment, the convolutional neural network includes convolutional layers, pooling layers, and fully connected layers. The convolutional neural network extracts features from the image through convolutional layers, pooling layers, and fully connected layers, and predicts the bounding box and class probability of the target in each grid.
[0125] Among them, the convolutional layer (Convolutional Layer) extracts features from the image through a series of convolutional kernels. These convolutional kernels slide on the image according to a set stride for convolutional operations, thereby extracting local features and determining the feature map. The convolutional layer can effectively capture the local features of the target in the image, such as edges, textures, etc.
[0126] The pooling layer is used to downsample the feature map. This method can reduce the data dimension while retaining key feature information to reduce the amount of subsequent calculations.
[0127] The feature map obtained after the convolution layer may be large in size. Through the downsampling operation of the pooling layer, the size of the feature map can be reduced. This not only reduces the amount of calculation of the subsequent fully connected layer, but also can extract more representative features to a certain extent.
[0128] The fully connected layer is used to integrate the feature maps extracted by the convolutional layer and the pooling layer, and map the integrated feature maps to the dimensions corresponding to the grid.
[0129] For each grid, the convolutional neural network predicts the relevant parameters of the target's bounding box through a fully connected layer. The relevant parameters of the bounding box include but are not limited to the offset of the center coordinates of the bounding box relative to the grid, and the width and height of the bounding box.
[0130] In addition, the convolutional neural network predicts the category probability of the target through the fully connected layer. The category probability here can be understood as the probability distribution of the category. In this way, the probability value of each target in the grid belonging to each category such as human, vehicle, animal, etc. can be predicted.
[0131] In some embodiments, optionally, Figure 3 As shown, before S102 (target detection is performed on each frame image in the video or image sequence based on the Yolo model, and the detection result is determined; wherein the detection result includes initial detection frames and feature information of multiple targets), the steps of the multi-target tracking method also include:
[0132] S1012, acquiring image data related to the task scene, where the image data includes video and / or image sequence.
[0133] The purpose of this step is to collect relevant video and / or image sequences according to the specific task scenario and build a dataset suitable for training and testing the model.
[0134] S1014, constructing a data set according to the image data, and training the Yolo model through the data set.
[0135] To meet the target detection requirements in specific mission scenarios, the Yolo model is trained using the constructed dataset. By adjusting the model parameters, it can accurately detect targets in mission scenarios and improve the accuracy of bounding box prediction and category classification.
[0136] In some embodiments, optionally, Figure 4As shown, before S1014 (constructing a dataset based on image data and training the Yolo model with the dataset), the steps of the multi-object tracking method further include:
[0137] S1013, performing data cleaning and data augmentation on the image data.
[0138] Data cleaning is used to remove noise, misannotations, and redundant information from the image data.
[0139] The ways of data augmentation include but are not limited to flipping, cropping, and adding noise.
[0140] The purpose of data augmentation is to expand the scale of the dataset, increase data diversity, thereby improving the robustness and generalization ability of the model.
[0141] In some embodiments, optionally, as Figure 5 shown, the steps of S106 (inputting the state information into the actor-critic model, where the actor-critic model includes a policy neural network, and the policy neural network modifies the position and size of the initial detection box according to the state information and determines the corrected detection box) include:
[0142] S1062, inputting the state information into the actor-critic model.
[0143] The actor-critic model analyzes the state information extracted from the Transformer structure as input data.
[0144] S1064, the actor-critic model includes a policy neural network, and the policy neural network determines the horizontal coordinate change amount, vertical coordinate change amount, and size change amount of the initial detection box according to the state information.
[0145] The actor (policy neural network) generates an action according to the state information extracted from the Transformer structure, and this action is the adjustment of the initial detection box.
[0146] The policy neural network determines the change amount of the initial detection box according to the state information. The change amount of the initial detection box includes the horizontal coordinate change amount Δx of the upper left vertex of the initial detection box. The change amount of the initial detection box also includes the vertical coordinate change amount Δy of the upper left vertex of the initial detection box. The change amount of the initial detection box also includes the size change amount Δs of the initial detection box.
[0147] S1066, according to the horizontal coordinate change amount, vertical coordinate change amount, and size change amount, modify the position and size of the initial detection box and determine the corrected detection box.
[0148] The policy neural network corrects the position and size of the initial detection box according to the calculation formula and determines the corrected detection box.
[0149] The calculation formula is x’ = x + Δx; y’ = y + Δy; h’ = h + Δs × h; w’ = w + Δs × w.
[0150] Among them, x’ represents the abscissa of the corrected detection box; x represents the abscissa of the initial detection box. y’ represents the ordinate of the corrected detection box; y represents the ordinate of the initial detection box. h’ represents the height of the corrected detection box; h represents the height of the initial detection box. w’ represents the width of the initial detection box; w represents the width of the corrected detection box.
[0151] The corrected detection box can accurately locate the position and range of the target in the current frame image, providing a more reliable data basis for subsequent tracking steps, which is conducive to continuously and stably tracking multiple targets accurately. Even in the case of variable motion states of the targets and complex scenes, it can still maintain high tracking accuracy and robustness.
[0152] In some embodiments, optionally, as Figure 6 shown, before S108 (determining the action trajectories of multiple targets according to the corrected detection box and performing trajectory prediction to track multiple targets), the steps of the multi-target tracking method further include:
[0153] S1072, calculating the intersection over union (IoU) between the corrected detection box of the target in the current frame image and the corrected detection box of the target in the next frame image.
[0154] It should be noted that the intersection over union (IoU) is a metric widely used in fields such as object detection and object tracking.
[0155] The larger the intersection over union, the higher the degree of overlap between the predicted target position and the actual target position, which also means higher tracking accuracy.
[0156] S1074, determining the reward value according to the intersection over union.
[0157] Optionally, the actor-critic model has a reward mechanism, and the reward mechanism is based on the intersection over union.
[0158] Optionally, when the intersection over union is greater than or equal to 0.7, the reward value is 1. When the intersection over union is less than 0.7, the reward value is -1.
[0159] S1076, adjusting the model parameters according to the reward value to optimize the actor-critic model.
[0160] When the intersection over union (IoU) is greater than or equal to 0.7, the reward value is 1. In this case, it will guide the actor-critic model to adjust the model parameters in the direction of making the IoU higher, so as to more accurately track the target.
[0161] When the IoU is less than 0.7, the reward value is -1. In this case, it will guide the actor-critic model to avoid making tracking actions that result in too low an IoU.
[0162] In this way, the IoU, as a feedback signal, effectively helps the actor-critic model to update the gradients of the model parameters, thereby optimizing the actor-critic model.
[0163] In some embodiments, optionally, as Figure 9 shown, the steps of the method for determining the multi-target comprehensive state (state information) include:
[0164] S802, determine multiple targets.
[0165] Determine multiple targets that need to be detected according to requirements.
[0166] It should be noted that Figure 9 the "..." in it represents the targets not listed.
[0167] S804, input mapping (determine the mapping relationship through the Transformer structure).
[0168] S806, multi-head self-attention.
[0169] Based on the multi-head attention mechanism of the Transformer structure, track multiple targets.
[0170] S808, feed-forward neural network.
[0171] Input the relationship information between multiple targets into the feed-forward neural network.
[0172] S810, multi-target comprehensive state.
[0173] Output the multi-target comprehensive state.
[0174] In some embodiments, optionally, as Figure 10 shown, the steps of the data processing method based on the actor-critic model include:
[0175] S902, initialize the detection box state S.
[0176] The purpose of this step is to determine the initial detection boxes of multiple targets.
[0177] S904, Self-attention Feature Extraction.
[0178] Based on the self-attention mechanism of the Transformer structure, feature extraction is performed and the state information is determined.
[0179] S906, Policy Network.
[0180] Input the state information into the policy network.
[0181] It should be noted that the policy network is the policy neural network of the actor-critic model.
[0182] S908, Update the state S’ according to the output offset.
[0183] The policy neural network determines the change amount of the initial detection box according to the state information, and determines the corrected detection box according to the change amount.
[0184] S910, The evaluation network calculates V(S) and V(S’), and updates the parameters.
[0185] It should be noted that the evaluation network is the evaluation neural network of the actor-critic model.
[0186] The evaluation network is used to adjust the model parameters to optimize the actor-critic model.
[0187] In an embodiment according to the present invention, as Figure 7 shown, the multi-object tracking system 200 includes an object detection module 210, a state information determination module 220, a detection box correction module 230, and a movement trajectory determination module 240.
[0188] The object detection module 210 is used to perform object detection on each frame of the video or image sequence based on the Yolo model, and determine the detection result. Among them, the detection result includes the initial detection boxes and feature information of multiple objects.
[0189] Optionally, the Yolo model includes a convolutional neural network. Each frame of the video or image sequence is divided into multiple grids through the convolutional neural network, the bounding boxes and class probabilities of the objects in each grid are predicted, and the detection result is determined.
[0190] The convolutional neural network is a deep learning model specifically used to process data with a grid structure. The Yolo model uses the powerful feature extraction ability of the convolutional neural network to process each frame of the video or image sequence.
[0191] The categories of the objects include but are not limited to people, vehicles, and animals.
[0192] Optionally, the convolutional neural network divides each frame of the image into multiple grids, and performs a preliminary analysis of the image in units of grids to determine the target for localization and recognition.
[0193] By dividing each frame of the image into multiple grids, the convolutional neural network can perform a local-to-global analysis of the image in units of grids.
[0194] Optionally, the convolutional neural network extracts features from the image through convolutional layers, pooling layers, and fully connected layers, and predicts the bounding box and class probability of the target in each grid.
[0195] Through operations such as multi-layer convolution and pooling of the convolutional neural network, deep feature extraction is performed on the image regions within each grid, and then the bounding box and class probability of the target are predicted. This grid-based prediction mechanism can, to a certain extent, take into account the detection of targets with different positions and sizes in the image, and due to the parameter sharing feature of the convolutional neural network, it can improve the detection efficiency while reducing the complexity of the model.
[0196] Specifically, the convolutional layer is used to extract local features of the image; the pooling layer performs dimensionality reduction on the features to reduce the computational amount and retain key information; the fully connected layer integrates and maps the extracted features, and finally predicts the bounding box and class probability of each grid.
[0197] Optionally, the convolutional neural network determines the detection results (initial detection boxes and feature information of multiple targets) based on the bounding box and class probability of the target.
[0198] The purpose of this step is to determine the position (bounding box) where the target may exist in the image and the class to which the target belongs, so as to obtain the initial detection box of the target and provide basic data for subsequent tracking.
[0199] It should be noted that the feature information includes but is not limited to shape features, color features, texture features, and class probabilities.
[0200] Shape features are used to reflect the general shape of the target. These shape features can assist in determining whether the target is still the object being tracked before during the subsequent tracking process when the target is partially occluded or its pose changes.
[0201] Color features are used to reflect the color differences between different targets. During the target tracking process, color features are helpful for distinguishing similar targets or re-identifying targets after a short occlusion. For example, in a multi-person tracking scenario where different people wear different colored clothes, color features can be an important distinguishing factor.
[0202] Texture is an important feature of the target appearance. By performing texture analysis on the image region within the initial detection box, the texture features of the target are extracted. The texture features can enhance the recognition of the target. When the target is deformed or partially occluded, the texture features contribute to the re-identification of the target.
[0203] During the multi-target tracking process, when the appearance of the target is blurred or some features are occluded, the class probability can assist in determining the class attribution of the target and can be combined with other feature information to improve the accuracy of the tracking result.
[0204] By introducing the Yolo model for target detection, the processing performance of the multi-target tracking system 200 is improved. Global information can be considered during detection, local interference is reduced, and the detection has high accuracy.
[0205] The status information determination module 220 is used to calculate the association degree between multiple targets based on the self-attention mechanism of the Transformer structure according to the feature information, so that information interaction occurs between multiple targets and the status information is determined. Among them, the status information includes the distribution information and association degree of multiple targets in each frame of image.
[0206] The status information is extracted from the detection results through the Transformer structure. Specifically, the self-attention mechanism of the Transformer structure is used to mine the association degree between targets (calculate the association degree between the feature information of each target and the feature information of other targets), and promote information interaction between targets, and then determine the status information.
[0207] The advantage of adopting the self-attention mechanism is that it calculates the association degree between the feature information of each target and the feature information of other targets in parallel, which is beneficial to improving the data processing efficiency.
[0208] Since the Transformer structure can capture the complex mutual relationships between targets, such as the relative positions and motion trends of targets. During the tracking process, when the targets are occluded, cross or have complex motion changes, the system provided by the present invention (multi-target tracking system 200) can anticipate the motion changes of the targets in advance, so as to adjust the tracking strategy more accurately.
[0209] Optionally, the Transformer structure also has a multi-head attention mechanism. The multi-head attention mechanism can extract the relationship information between targets more comprehensively.
[0210] The multi - head attention mechanism based on the Transformer structure extracts the relationship information between targets. The relationship information includes, but is not limited to, the spatial position relationship between targets and the class similarity of targets.
[0211] The detection box correction module 230 is used to input the state information into the actor - critic model. The actor - critic model includes a policy neural network. The policy neural network corrects the position and size of the initial detection box according to the state information and determines the corrected detection box.
[0212] The policy neural network determines the change amount of the initial detection box according to the state information. The change amount of the initial detection box includes the abscissa change amount Δx of the upper - left vertex of the initial detection box. The change amount of the initial detection box also includes the ordinate change amount Δy of the upper - left vertex of the initial detection box. The change amount of the initial detection box also includes the size change amount Δs of the initial detection box.
[0213] The policy neural network corrects the position and size of the initial detection box according to the calculation formula and determines the corrected detection box.
[0214] The calculation formula is x’ = x + Δx; y’ = y + Δy; h’ = h + Δs×h; w’ = w + Δs×w.
[0215] Among them, x’ represents the abscissa of the corrected detection box; x represents the abscissa of the initial detection box. y’ represents the ordinate of the corrected detection box; y represents the ordinate of the initial detection box. h’ represents the height of the corrected detection box; h represents the height of the initial detection box. w’ represents the width of the initial detection box; w represents the width of the corrected detection box.
[0216] The corrected detection box can accurately locate the position and range of the target in the current - frame image, providing a more reliable data basis for subsequent tracking steps, which is conducive to continuously and stably tracking multiple targets accurately. Even in the case of variable motion states of targets and complex scenes, it can still maintain a high tracking accuracy and robustness.
[0217] The action trajectory determination module 240 is used to determine the action trajectories of multiple targets according to the corrected detection box and perform trajectory prediction to track multiple targets.
[0218] The corrected detection box can accurately locate the position and range of the target in the current - frame image. By analyzing the corrected detection boxes in consecutive multiple frames of images, the action trajectories of multiple targets are determined and trajectory prediction is performed to track multiple targets.
[0219] Optionally, by analyzing the corrected detection boxes in consecutive multiple frames of images, the change rules of the positions and sizes of the corrected detection boxes in the time series are determined, and then the action trajectories of multiple targets are determined.
[0220] Predict the future trajectory of the target based on the action trajectory of the target, as well as the motion characteristics and historical data of the target.
[0221] The present invention aims to provide a multi-target tracking system 200. By introducing the self-attention mechanism of the Transformer structure, the degree of association between the feature information of each target and the feature information of other targets is calculated, which is beneficial to improving the ability to read the global information of the current frame image, and further improving the accuracy of the tracking results (action trajectory and trajectory prediction), with higher reliability. In addition, the multi-target tracking system 200 of the present invention is based on the Yolo model and the actor-critic model, and adopts the method of deep reinforcement learning, which can track targets with different motion characteristics and is applicable to a variety of task scenarios.
[0222] In an embodiment according to the present invention, as Figure 8 shown, the electronic device 300 includes a memory 310 and a processor 320. Among them, a program or instruction that can run on the processor 320 is stored on the memory 310. When the processor 320 executes the program or instruction, the steps of the multi-target tracking method in any of the above embodiments are implemented. Therefore, the electronic device 300 has the beneficial effects of any of the above embodiments, which will not be elaborated here.
[0223] In an embodiment according to the present invention, the readable storage medium stores a program or instruction. When the program or instruction is executed by the processor, the steps of the multi-target tracking method in any of the above embodiments are implemented. Therefore, the readable storage medium has the beneficial effects of any of the above embodiments, which will not be elaborated here.
[0224] In the present invention, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance; the term "plural" refers to two or more, unless otherwise clearly defined. Terms such as "installation", "connection", "connection", and "fixation" should all be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; "connection" can be a direct connection or an indirect connection through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0225] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "front", "rear", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or unit referred to must have a specific direction, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.
[0226] In the description of this specification, the descriptions of terms such as "one embodiment", "some embodiments", "specific embodiments", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or instance. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0227] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-target tracking method, characterized in that it is used for include: Perform target detection on each frame of a video or image sequence based on the Yolo model and determine the detection result; wherein the detection result includes initial detection frames and feature information of multiple targets; Based on the self-attention mechanism of the Transformer structure, the correlation degree between the multiple targets is calculated according to the feature information, so that the multiple targets can interact with each other and determine the state information; wherein the state information includes the distribution information and the correlation degree of the multiple targets in each frame of the image; Inputting the state information into an actor-critic model, wherein the actor-critic model includes a policy neural network, wherein the policy neural network corrects the position and size of the initial detection frame according to the state information and determines a corrected detection frame; The movement trajectories of the multiple targets are determined according to the corrected detection frame, and trajectory prediction is performed to track the multiple targets.
2. The multi-target tracking method according to claim 1, characterized in that: The target detection is performed on each frame of the video or image sequence based on the Yolo model, and the detection result is determined; wherein the detection result includes the initial detection frame and feature information of multiple targets, including: The Yolo model includes a convolutional neural network, which divides each frame image in the video or the image sequence into a plurality of grids, and predicts a bounding box and a category probability of an object in each of the grids; A detection result is determined according to the bounding box and the category probability; wherein the detection result includes the initial detection boxes and the feature information of the plurality of targets.
3. The multi-target tracking method according to claim 2, characterized in that: The convolutional neural network includes a convolutional layer, a pooling layer and a fully connected layer. The convolutional layer, the pooling layer and the fully connected layer are used to extract features from the image and predict the bounding box and the category probability of the target in each of the grids.
4. The multi-target tracking method according to any one of claims 1 to 3, characterized in that: Also includes: Before performing target detection on each frame of an image in a video or an image sequence based on the Yolo model and determining a detection result, acquiring image data related to a task scene, wherein the image data includes the video and / or the image sequence; A data set is constructed according to the image data, and the Yolo model is trained using the data set.
5. The multi-target tracking method according to claim 4, characterized in that: Also includes: Before constructing a data set according to the image data and training the Yolo model through the data set, the image data is cleaned and enhanced.
6. The multi-target tracking method according to any one of claims 1 to 3, characterized in that: The step of inputting the state information into an actor-critic model, wherein the actor-critic model includes a strategy neural network, wherein the strategy neural network corrects the position and size of the initial detection frame according to the state information and determines a corrected detection frame, including: Inputting the state information into the actor-critic model; The actor-critic model includes the strategy neural network, and the strategy neural network determines the horizontal coordinate change, the vertical coordinate change and the size change of the initial detection frame according to the state information; According to the horizontal coordinate change amount, the vertical coordinate change amount and the size change amount, the position and size of the initial detection frame are corrected, and the corrected detection frame is determined.
7. The multi-target tracking method according to any one of claims 1 to 3, characterized in that: Also includes: Before determining the movement trajectories of the multiple targets according to the modified detection frame and performing trajectory prediction to track the multiple targets, calculating an intersection-over-union ratio between the modified detection frame of the target in the image of the current frame and the modified detection frame of the target in the image of the next frame; Determine a reward value according to the intersection-combination ratio; The model parameters are adjusted according to the reward value to optimize the actor-critic model.
8. A multi-target tracking system, characterized in that: include: A target detection module (210) is used to perform target detection on each frame of a video or image sequence based on a Yolo model and determine a detection result; wherein the detection result includes initial detection frames and feature information of multiple targets; A state information determination module (220) is used to calculate the degree of association between the multiple targets based on the Transformer structure self-attention mechanism according to the feature information, so as to enable information exchange between the multiple targets, and determine the state information; wherein the state information includes the distribution information and the degree of association of the multiple targets in each frame of the image; A detection frame correction module (230), used for inputting the state information into an actor-critic model, wherein the actor-critic model includes a policy neural network, and the policy neural network corrects the position and size of the initial detection frame according to the state information and determines a corrected detection frame; The action trajectory determination module (240) is used to determine the action trajectories of the multiple targets according to the modified detection frame, and perform trajectory prediction to track the multiple targets.
9. An electronic device, characterized in that: include: A memory (310) and a processor (320), wherein the memory (310) stores a program or instruction that can be run on the processor (320), and when the processor (320) executes the program or the instruction, the steps of the multi-target tracking method as described in any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or an instruction, and when the program or the instruction is executed by a processor, the steps of the multi-target tracking method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Multi-target tracking method and device, related equipment and computer program product
CN120913153A