Multi-target tracking method and device, related equipment and computer program product

By combining a pre-trained detector and a network model to optimize the bounding box offset, the problem of target tracking under complex conditions such as occlusion and deformation in multi-target tracking is solved, achieving efficient and accurate multi-target tracking and improving tracking accuracy and robustness.

CN120913153APending Publication Date: 2025-11-07CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511164710.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing multi-target tracking technologies suffer from insufficient tracking accuracy and robustness under challenges such as target occlusion, interference from similar appearances, sudden changes in motion, and complex environmental changes.

Method used

By combining a pre-trained detector and a first network model, the detection results are optimized by obtaining the bounding boxes of candidate objects, using state vector modeling and network prediction of bounding box offsets, and then combining the tracking algorithm to complete cross-frame identity association.

Benefits of technology

It improves the accuracy and robustness of multi-target tracking in complex scenarios, reduces computational overhead, and is suitable for stable tracking in challenging scenarios such as target occlusion and deformation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913153A_ABST
    Figure CN120913153A_ABST
Patent Text Reader

Abstract

The invention provides a multi-target tracking method and device, related equipment and a computer program product, and relates to the technical field of computer vision. The method comprises the steps that a target video is acquired, the target video comprises multiple frames of images, and the multiple frames of images comprise a first frame of image; performing object bounding box detection on the first frame image through a pre-training detector to obtain at least one candidate object bounding box; determining a first frame state vector corresponding to the first frame image according to each candidate object bounding box; processing the first frame state vector through a first network model to obtain bounding box offset corresponding to each candidate object bounding box; determining a predicted object bounding box corresponding to the first frame image according to each candidate object bounding box and the bounding box offset corresponding to each candidate object bounding box; and inputting the at least one predicted object bounding box into a target tracking algorithm, and determining an identity label of each object in the first frame image to perform cross-frame object association.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer vision, and particularly relates to a multi-target tracking method and device, electronic equipment, computer readable storage medium and computer program product. BACKGROUND

[0002] This section is intended to provide background or context to the embodiments of the disclosure recited in the claims. The description herein does not constitute admission that the prior art is prior art nor does it constitute an admission of any description in this section as prior art to an application.

[0003] At present, although the multi-target tracking (MOT) technology has made significant progress, it still faces many challenges in practical applications, including target occlusion, appearance similarity interference, motion mutation and complex environmental changes. These challenges seriously affect the accuracy and robustness of the tracking method.

[0004] Therefore, developing a new multi-target tracking method that can effectively cope with these challenges to further improve the tracking accuracy and stability has become a key problem to be solved in this field. SUMMARY

[0005] The purpose of the present disclosure is to provide a multi-target tracking method, device, electronic equipment, computer readable storage medium and computer program product, which can improve the accuracy and stability of multi-target tracking.

[0006] Other characteristics and advantages of the present disclosure will become apparent from the following detailed description, or will be learned by practice of the present disclosure.

[0007] The multi-target tracking method provided by the embodiments of the present disclosure comprises: acquiring a target video, the target video comprising a plurality of frames of images, wherein the plurality of frames of images comprise a first frame of image; performing object bounding box detection on the first frame of image by a pre-trained detector to obtain at least one candidate object bounding box; determining a first frame state vector corresponding to the first frame of image according to each candidate object bounding box; processing the first frame state vector by a first network model to obtain a bounding box offset corresponding to each candidate object bounding box; determining a predicted object bounding box corresponding to the first frame of image according to each candidate object bounding box and the bounding box offset corresponding to each candidate object bounding box; inputting the at least one predicted object bounding box into a target tracking algorithm to determine the identity of each object in the first frame of image for cross-frame object association.

[0008] In some embodiments, the method further comprises: obtaining a real object bounding box corresponding to each object in the first frame image; determining an intersection over union between the predicted object bounding box and the real object bounding box; determining a reward evaluation value based on the intersection over union, so as to update the network parameters of the first network model based on the reward evaluation value.

[0009] In some embodiments, the method further comprises: obtaining N frame images before the first frame image as first adjacent frame images; N is an integer greater than or equal to 1; determining a proportion of frames that correctly maintain identity consistency in the first adjacent frame images according to the identity of each object in the first frame image and the identity corresponding to each object in each first adjacent frame image; wherein, based on the intersection over union to determine a reward evaluation value, so as to update the network parameters of the first network model based on the reward evaluation value, comprises: based on the intersection over union and the proportion of correctly maintaining identity consistency, determine the reward evaluation value, so as to update the parameters of the first network model according to the reward evaluation value.

[0010] In some embodiments, wherein the bounding box offset corresponding to each candidate object bounding box is the predicted action output by the first network model; wherein the method further comprises: obtaining k frame images before the first frame image as second adjacent frame images; k is an integer greater than or equal to 1; obtaining a second adjacent frame state vector corresponding to each second adjacent frame image; performing prediction processing on the predicted action of the first network model, the first frame state vector and each second adjacent frame state vector through a second network model to obtain a first value evaluation value, the first value evaluation value is used to evaluate the value corresponding to the predicted action and the first frame state vector; update the parameters of the first network model and the second network model based on the first value evaluation value.

[0011] In some embodiments, updating the parameters of the first network model and the second network model based on the first value evaluation value comprises: obtaining a reward evaluation value determined based on the output value of the first network model; determining a TD error corresponding to the second network model based on the reward evaluation value; determining a mean square loss based on the TD error and the first value evaluation value; updating the parameters of the first network model based on the TD error; updating the parameters of the second network model through the mean square loss.

[0012] In some embodiments, determining the first frame state vector corresponding to the first frame image according to each candidate object bounding box comprises: determining the object features corresponding to each candidate bounding box; splicing the object features corresponding to each candidate bounding box in the first frame image by object to obtain the first frame state vector corresponding to the first frame image.

[0013] In some embodiments, the at least one candidate object bounding box includes a target candidate object bounding box; wherein determining the object feature corresponding to each candidate bounding box includes: determining spatial position information, motion speed information, appearance feature information, and inter-target speed difference information corresponding to the target candidate object bounding box; and determining the object feature corresponding to the target candidate object bounding box according to the spatial position information, the motion speed information, the appearance feature information, and the inter-target speed difference information corresponding to the target candidate object bounding box.

[0014] Embodiments of the present disclosure provide a multi-target tracking device, which includes a video acquisition module, a bounding box determination module, a first frame state vector determination module, an offset prediction module, a bounding box prediction module, and an object association module.

[0015] The video acquisition module is configured to acquire a target video, the target video including a plurality of frames of images, wherein the plurality of frames of images include a first frame of image; the bounding box determination module is configured to perform object bounding box detection on the first frame of image by using a pre-trained detector to obtain at least one candidate object bounding box; the first frame state vector determination module is configured to determine a first frame state vector corresponding to the first frame of image according to each candidate object bounding box; the offset prediction module is configured to process the first frame state vector by using a first network model to obtain a bounding box offset corresponding to each candidate object bounding box; the bounding box prediction module is configured to determine a predicted object bounding box corresponding to the first frame of image according to each candidate object bounding box and the bounding box offset corresponding to each candidate object bounding box; and the object association module is configured to input the at least one predicted object bounding box into a target tracking algorithm to determine the identity of each object in the first frame of image to perform cross-frame object association.

[0016] Embodiments of the present disclosure provide an electronic device, which includes a memory and a processor; the memory is configured to store computer program instructions; and the processor is configured to invoke the computer program instructions stored in the memory to implement the multi-target tracking method described in any of the above embodiments.

[0017] Embodiments of the present disclosure provide a computer readable storage medium having computer program instructions stored thereon, which implement the multi-target tracking method described in any of the above embodiments.

[0018] Embodiments of the present disclosure provide a computer program product or a computer program, which includes computer program instructions stored in a computer readable storage medium. The computer program instructions are read from the computer readable storage medium, and a processor executes the computer program instructions to implement the multi-target tracking method described above.

[0019] The multi-target tracking method, device, electronic device, computer readable storage medium and computer program product provided by the embodiments of the present disclosure realize efficient and accurate multi-target tracking by combining a pre-trained detector and a first network model. First, a detector is used to obtain a candidate object bounding box, then state vector modeling and network prediction of bounding box offset are used to optimize the detection result, and finally a tracking algorithm is used to complete cross-frame identity association. This scheme effectively improves the accuracy and robustness of multi-target tracking in complex scenes, and reduces the computational overhead through an end-to-end processing flow, and is particularly suitable for stable tracking in challenging scenes such as target occlusion and deformation in videos.

[0020] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0021] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained from these drawings without creative labor for those skilled in the art.

[0022] Figure 1 A scene schematic diagram of a multi-target tracking method or a multi-target tracking device that can be applied to the embodiments of the present disclosure is shown.

[0023] Figure 2 is a flowchart of a multi-target tracking method according to an exemplary embodiment.

[0024] Figure 3 is a flowchart of a first network model parameter updating method according to an exemplary embodiment.

[0025] Figure 4 is a flowchart of a parameter updating method according to an exemplary embodiment.

[0026] Figure 5 is a flowchart of a model parameter updating method according to an exemplary embodiment.

[0027] Figure 6 is a flowchart of a model parameter updating method according to an exemplary embodiment.

[0028] Figure 7 is a flowchart of a state vector determination method according to an exemplary embodiment.

[0029] Figure 8is a flow chart of an object feature determination method according to an example embodiment.

[0030] Figure 9 is a flow chart of a multi-target tracking method according to an example embodiment.

[0031] Figure 10 is a block diagram of a multi-target tracking apparatus according to an example embodiment.

[0032] Figure 11 A structural schematic of an electronic device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0033] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments can be implemented in any

[0034] Those skilled in the art will appreciate that implementing embodiments of the present disclosure can take the form of a system, apparatus, device, method or computer program product. Accordingly, the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects that can all generally be referred to herein as a "circuit," "module" or "system."

[0035] The features, structures or characteristics described in the present disclosure can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the present disclosure. One skilled in the relevant art will recognize, however, that the techniques described herein can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0036] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other relevant parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.

[0037] The accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the present disclosure and, together with the description, serve to explain principles of the present disclosure. In the drawings:

[0038] The flowcharts shown in the drawings are only illustrative and do not necessarily include all contents and steps, nor are they necessarily executed in the order described. For example, some steps can be further divided, and some steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.

[0039] In the description of the present disclosure, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this document only describes the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, "one or more" means one or more, and "multiple" means two or more. "First", "second", and the like do not limit the quantity and execution order, and "first", "second", and the like do not necessarily mean different; the terms "include", "contain" and "have" mean open inclusion and mean that in addition to the listed elements / components / etc. There can be other elements / components / etc.

[0040] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the present application will be further described in detail below in conjunction with the drawings and specific embodiments. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0041] First, some terms related to the embodiments of the present disclosure will be explained below to facilitate understanding by those skilled in the art.

[0042] Kalman filter (KF) is a recursive optimal estimation algorithm, mainly used for dynamically estimating the state of a system from observation data containing noise.

[0043] MOTA (Multiple Object Tracking Accuracy): A comprehensive performance metric that considers the overall performance of three types of errors: False Negatives, False Positives, and ID Switches. MOTA is one of the most commonly used aggregate performance indicators in the field of multiple object tracking (MOT), used to comprehensively evaluate the performance of tracking algorithms in terms of false negatives, false positives, and ID switches. The core idea is to quantify these errors to reflect the overall accuracy of the tracker.

[0044] Multiple object tracking is a core task in the field of computer vision, aiming to continuously monitor and associate multiple targets (or objects) from video sequences, assign a unique ID to each target (or object), and generate trajectories that change over time. The core challenge is to handle occlusions between targets, appearance changes, similar object interference, and real-time requirements.

[0045] IDF1 (ID F1-Score): Mainly measures the proportion of frames in which the identity is correctly maintained. IDF1 is a core indicator in multiple object tracking (MOT) for measuring ID maintenance capability, focusing on whether the tracker can assign consistent and correct IDs to the same target, avoiding ID switching or confusion.

[0046] HOTA (Higher Order Tracking Accuracy): A multi-object tracking evaluation index, aiming to measure both detection accuracy and association quality in the tracking task.

[0047] Replay Buffer (Experience Replay Pool): A mechanism commonly used in reinforcement learning algorithms, which stores "experience" tuples generated during the interaction between the agent and the environment. Experience is randomly stored and randomly sampled in small batches to train the network, which can effectively break the time correlation and improve convergence and stability. In addition, the experience stored in the buffer can be sampled multiple times for multiple network updates.

[0048] The foregoing introduces some concepts related to the embodiments of the present disclosure, and the following introduces the technical features related to the embodiments of the present disclosure.

[0049] Multi-object tracking (MOT) aims to monitor and continuously associate the trajectories of multiple objects in a video sequence, and is widely used in intelligent transportation, video surveillance, sports analysis, and unmanned driving scenarios. With the rapid development of deep learning, MOT has become a research hotspot in the field of computer vision in academia and industry. The mainstream MOT methods can be roughly divided into two paradigms: Tracking-by-Detection: First, a detector is used to extract objects in each frame, and then a Kalman filter prediction algorithm and a Hungarian algorithm are used to complete cross-frame association; Joint Detection And Association: represented by Transformer or directly considering ID prediction as a task completed in context, the detection box and object ID are output in the same network, simplifying the process and improving the robustness in complex scenarios.

[0050] Currently, the challenges of MOT are: 1. Occlusion and dense crowd; 2. Fine-grained appearance similarity; 3. Non-linear and sudden motion; 4. Multi-modal and different environmental transformations.

[0051] In order to solve the above problems, the present application provides a multi-object tracking method.

[0052] The multi-object tracking method related to the present application will be explained and described below in conjunction with the accompanying drawings.

[0053] Figure 1 A scene schematic diagram of a multi-object tracking method or a multi-object tracking device that can be applied to the embodiments of the present disclosure is shown.

[0054] Please refer to Figure 1 which shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present disclosure.

[0055] As Figure 1 shown, the system architecture 100 can include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a communication link medium between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.

[0056] A user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Among them, the terminal devices 101, 102, 103 can be various electronic devices with display screens and support for web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, wearable devices, virtual reality devices, smart home devices, etc.

[0057] The server 105 can be a server that provides various services, such as a background management server that provides support for operations performed by a user using the terminal device 101, 102, or 103. The background management server can analyze and process received request data and the like, and feed back the processing result to the terminal device.

[0058] The server can be a stand-alone physical server, a server cluster composed of multiple physical servers, or a distributed system, and can also be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms, and the like, and the present disclosure does not limit the same.

[0059] The server 105 can obtain, for example, a target video including a plurality of images, wherein the plurality of images include a first image; the server 105 can perform object bounding box detection on the first image by a pre-trained detector to obtain at least one candidate object bounding box; the server 105 can determine a first frame state vector corresponding to the first image according to each candidate object bounding box; the server 105 can process the first frame state vector by a first network model to obtain a bounding box offset corresponding to each candidate object bounding box; the server 105 can determine a predicted object bounding box corresponding to the first image according to each candidate object bounding box and the bounding box offset corresponding to each candidate object bounding box; and the server 105 can input the at least one predicted object bounding box into a target tracking algorithm to determine the identity of each object in the first image for cross-frame object association.

[0060] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above-mentioned embodiments is only illustrative, and the server 105 can be a server of one entity, or can be composed of multiple servers. According to actual needs, the server 105 can have any number of terminal devices, networks, and servers.

[0061] Figure 2 is a flowchart of a multi-target tracking method according to an example embodiment. The method provided by the embodiments of the present disclosure can be executed by any electronic device with computing processing capability, for example, the method can be executed by the server or the terminal device in the above-mentioned embodiments, or can be executed by the server and the terminal device together. In the following embodiments, the server is taken as an example for illustration, but the present disclosure is not limited thereto. Figure 1

[0062] Referring to Figure 2 ​The multi-target tracking method provided by the embodiments of the present disclosure can include the following steps.

[0063] In step S202, a target video is obtained, and the target video includes multiple frames of images, and the multiple frames of images include a first frame of image.

[0064] The target video can be a pedestrian monitoring video in the field of transportation. The pedestrian monitoring video is obtained in compliance with relevant laws and regulations, for example, the pedestrian monitoring video is collected after the pedestrian has given permission.

[0065] In some embodiments, the target video can include multiple frames of images, and the multiple frames of images can include a first frame of image.

[0066] In the following, the present application will take the first frame of image as an example to explain how to perform multi-target tracking, and those skilled in the art can extend the multi-target tracking process of other frames of image without mental effort according to the embodiments.

[0067] In step S204, an object bounding box detector is used to detect an object bounding box in the first frame of image to obtain at least one candidate object bounding box.

[0068] In some embodiments, the pre-trained detector can refer to a deep learning model that has been trained on a large-scale data set, which is used to detect and locate a target object in an image or a frame of video, and outputs a bounding box corresponding to the target object and category information.

[0069] In some embodiments, the pre-trained detector can be a YOLO (You Only Look Once) series detector.

[0070] In some embodiments, the pre-trained detector can be used to detect the first frame of image to obtain at least one object (or target) and a bounding box corresponding to the object (or target) as a candidate object bounding box.

[0071] In some embodiments, the pre-trained detector can output not only the candidate object bounding box, but also an object or an object ID (or target ID) enclosed in the candidate object bounding box.

[0072] In some embodiments, the result output by the pre-trained detector can have errors, and the errors can be reduced through the following steps.

[0073] In step S206, a first frame state vector corresponding to the first frame of image is determined according to each candidate object bounding box.

[0074] In some embodiments, the object features corresponding to the respective object bounding boxes can be extracted respectively, and then the object features corresponding to the respective candidate object bounding boxes can be fused to obtain the first frame state vector.

[0075] At step S208, the first frame state vector is processed by a first network model to obtain a bounding box offset corresponding to each candidate object bounding box.

[0076] The first network model can be any trainable network model, for example, a convolutional network model (such as CNN), or an Actor network model in an Actor-Critic network model.

[0077] Actor-Critic is a method combining policy gradient (Policy Gradient) and value function (Value Function) in reinforcement learning, which includes two main parts: Actor (Actor): responsible for learning policy (Policy), deciding what action should be taken in a given state; Critic (Critic): responsible for evaluating the value of the current policy, i.e., evaluating the value of the state or state-action pair.

[0078] The core idea of Actor-Critic is that the Actor selects an action according to the current policy, the Critic evaluates the value of the action, and the Actor updates the policy according to the evaluation of the Critic to improve the policy towards a direction of obtaining higher evaluation.

[0079] At step S210, the predicted object bounding box corresponding to the first frame image is determined according to the respective candidate object bounding boxes and the bounding box offset corresponding to each candidate object bounding box.

[0080] In some embodiments, the offset of the candidate object bounding box can be superimposed on the respective candidate object bounding box to determine the predicted object bounding box corresponding to the first frame image.

[0081] In some embodiments, the predicted object bounding box corresponding to the first frame object can include at least one predicted object bounding box.

[0082] At step S212, the at least one predicted object bounding box is input into a target tracking algorithm to determine the identity of each object in the first frame image for cross-frame object association.

[0083] In some embodiments, the at least one predicted object bounding box can be input into a target tracking algorithm to determine the identity of each object in the first frame image for cross-frame object association.

[0084] In some embodiments, a unique identity (ID) can be assigned to each object (target) framed by the target object bounding box in the first frame image through the target tracking algorithm described above.

[0085] In some embodiments, the object association across frames can be achieved by the identity of the object in each frame image in the target video.

[0086] In some embodiments, in the multi-target tracking (MOT) task, the core of the cross-frame object association is to correctly match the same target between different frames and ensure that the identity (ID) remains consistent.

[0087] In some embodiments, the target tracking algorithm can be a DeepSORT (Deep Simple Online and Realtime Tracking) algorithm.

[0088] In some embodiments, the multi-target tracking method described above can be applied to real-world scenarios such as intelligent security and traffic monitoring.

[0089] The multi-target tracking method described above uses a pre-trained detector (such as YOLO) to detect targets in the frame images of the video to obtain candidate bounding boxes, optimizes the bounding box offset using the Actor-Critic network model in reinforcement learning to improve detection accuracy, and then uses a tracking algorithm such as DeepSORT to assign a unique ID to each target and achieve cross-frame association.

[0090] In traffic monitoring, the motion of pedestrians has a very high degree of uncertainty, often with occlusion, interaction, and complex walking patterns. Due to the large number of pedestrians, the distance between targets varies greatly, and there is a strong similarity in appearance, and the pedestrian motion trajectory has a sudden change, which makes the traditional multi-target tracking algorithm (Kalman filter + Hungarian algorithm) often face the problems of pedestrian ID switching, missed detection, and false detection in the pedestrian tracking scene.

[0091] Through the technical solutions provided in this embodiment, the accuracy and robustness of multi-target tracking can be significantly improved in scenarios such as traffic monitoring, and through dynamic correction of detection errors and continuous identity matching, the target loss problem under complex conditions such as occlusion and deformation is effectively solved.

[0092] In summary, the multi-target tracking method provided by the embodiment realizes efficient and accurate multi-target tracking by combining a pre-trained detector and a first network model. First, the detector is used to obtain candidate object bounding boxes, then the state vector modeling and network prediction bounding box offset are used to optimize the detection results, and finally the tracking algorithm is used to complete the cross-frame identity association. This scheme effectively improves the accuracy and robustness of multi-target tracking in complex scenes, while reducing the computational overhead through an end-to-end processing flow, and is particularly suitable for stable tracking in challenging scenarios such as target occlusion and deformation in videos.

[0093] Figure 3 is a flowchart of a first network model parameter updating method according to an exemplary embodiment.

[0094] Reference Figure 3 The first network model parameter updating method can include the following steps.

[0095] Step S302, obtaining real object bounding boxes corresponding to each object in the first frame image.

[0096] In some embodiments, the real object bounding box can refer to the accurate position and range of the target object provided by manual annotation or high-precision annotation tools in the image or video frame.

[0097] Step S304, determining the intersection over union between the predicted object bounding box and the real object bounding box.

[0098] The intersection over union (IoU) is an indicator that measures the degree of overlap between two bounding boxes (predicted box and real box). Its value range is between [0, 1]: IoU = 1: complete overlap (predicted box and real box are completely consistent); IoU = 0: no overlap (two boxes are disjoint).

[0099] Step S306, determining a reward evaluation value based on the intersection over union, so as to update the network parameters of the first network model based on the reward evaluation value.

[0100] In some embodiments, a loss function can be determined according to the reward evaluation value, and then the first network model is updated according to the loss function.

[0101] In some embodiments, the reward evaluation value can also update the parameters of the first network model in combination with the following first value evaluation value.

[0102] In summary, the present application does not limit how to update the network parameters of the first network model through the reward evaluation value.

[0103] The technical scheme provided by the embodiment significantly improves the positioning accuracy of multi-target tracking and the adaptability to complex scenes, effectively reduces the target loss and ID switching rate, and provides a more reliable high-precision tracking solution for the field of intelligent monitoring and the like.

[0104] Figure 4 is a flowchart of a parameter updating method according to an exemplary embodiment.

[0105] Reference Figure 4 The parameter updating method can include the following steps.

[0106] Step S402, obtaining real object bounding boxes corresponding to each object in the first frame image.

[0107] Step S404, determining the intersection over union between the predicted object bounding box and the real object bounding box.

[0108] Step S406, obtaining N frames of images before the first frame of image as first adjacent frame images. N is an integer greater than or equal to 1.

[0109] Step S408, determining the proportion of frames that correctly maintain consistent identities in the first adjacent frame images according to the identities of each object in the first frame image and the identities corresponding to each object in each first adjacent frame image.

[0110] In some embodiments, the identity consistency of the first frame image and its adjacent frames (first adjacent frame images) can be analyzed to measure the short-term stability of the tracking algorithm.

[0111] Specific explanation: In the first frame of the video, all the targets to be tracked (such as pedestrians, vehicles, etc.) will be detected and assigned a unique identity ID (such as ID = 1, 2, 3…).

[0112] In the frames (adjacent frames) before the first frame image, all the targets to be tracked (such as pedestrians, vehicles, etc.) will also be detected and assigned a unique identity ID, and the tracking algorithm needs to maintain the ID of these targets unchanged. For example, the person with ID = 1 in the first frame should still be correctly identified as ID = 1 in the subsequent frames.

[0113] In some embodiments, the proportion of frames in these adjacent frames that can correctly maintain the initial ID (without ID switching or loss) can be counted. For example, IDF1 can be calculated according to the identity ID of the adjacent k frames.

[0114] Step S410, determining a reward evaluation value based on the intersection over union and the proportion of frames that correctly maintain consistent identities, so as to update the parameters of the first network model according to the reward evaluation value.

[0115] In some embodiments, the intersection over union and the proportion of consistent identities can be weighted and summed to serve as the reward evaluation value.

[0116] The technical scheme provided by the embodiment constructs a multi-dimensional reward function to optimize the network model parameters by fusing the dual evaluation mechanism of the intersection over union (IoU) and the adjacent frame identity consistency (IDF1), and significantly improves the performance of the target tracking system: while maintaining high positioning accuracy, effectively reducing the ID switching and target loss problems. The method innovatively combines the spatial detection accuracy (IoU) and the time continuity (ID consistency) indicators, realizes end-to-end optimization through the reinforcement learning framework, and provides a more accurate and stable and reliable solution for real-time multi-target tracking in complex scenes (such as traffic monitoring and crowd analysis).

[0117] Figure 5 is a flowchart of a model parameter updating method according to an example embodiment.

[0118] Reference Figure 5 The model parameter updating method can include the following steps.

[0119] In some embodiments, the boundary box offset corresponding to each candidate object boundary box is a predicted action output by the first network model.

[0120] Step S502, k frames of images before the first frame of image are obtained as second adjacent frame images. k is an integer greater than or equal to 1.

[0121] In some embodiments, the k frames of images before the first frame of image in the target video can be taken as the second adjacent frame images.

[0122] Step S504, a second adjacent frame state vector corresponding to each second adjacent frame image is obtained.

[0123] In some embodiments, the state vector corresponding to each second adjacent frame image can be determined as the second adjacent frame state vector by referring to the state vector determination method of the first frame of image. One second adjacent frame image corresponds to one second adjacent frame state vector.

[0124] Step S506, the predicted action of the first network model, the first frame state vector, and each second adjacent frame state vector are processed by the second network model for prediction to obtain a first value evaluation value, which is used to evaluate the value corresponding to the predicted action and the first frame state vector.

[0125] In some embodiments, the second network model can be any network model capable of training, for example, can be a convolutional network model (such as a CNN), and can also be a Critic model in an Actor-Critic network model.

[0126] In some embodiments, the first value evaluation value can be obtained by performing prediction processing on the predicted action of the first network model, the first frame state vector, and each second adjacent frame state vector through the second network model (such as the Critic model).

[0127] Step S508, updating parameters of the first network model and the second network model based on the first value evaluation value.

[0128] In some embodiments, the parameters of the first network model and the second network model can be updated based on the first value evaluation value, for example, the Actor-Critic composed of the first network model and the second network model can be updated based on the first value evaluation value.

[0129] The model parameter updating method introduces historical frame information (state vectors of the previous k frames) and combines the current frame state, evaluates the value of the predicted action of the first network model (the first value evaluation value) by using the second network model (such as the Critic model), and then jointly optimizes the parameters of the two models. The core advantages are: 1) using the temporal context to improve the continuity of action prediction, adapting to dynamic scenes such as video tracking; 2) realizing the collaborative optimization of policy and evaluation through the Actor-Critic framework, enhancing the robustness of the model; 3) balancing the historical dependence and the calculation efficiency, suitable for complex situations such as occlusion and motion mutation. This method is particularly suitable for tasks that require temporal modeling and dynamic decision-making, such as video target tracking or continuous control in reinforcement learning.

[0130] Figure 6 is a flowchart of a model parameter updating method according to an example embodiment.

[0131] Reference Figure 6 The above model parameter updating method can include the following steps.

[0132] Step S602, obtaining a reward evaluation value determined based on an output value of the first network model.

[0133] In some embodiments, the reward evaluation value can be determined according to the intersection over union between the real box and the predicted box and the proportion of correct identity consistency. For example, the reward evaluation value can be determined by Figure 4

[0134] Step S604, determining a TD error corresponding to the second network model based on the reward evaluation value. ​

[0135] TD error can be used to measure the difference between the value estimate of the current state and a more accurate "target estimate". It drives the model to gradually optimize the value function or policy through Temporal Difference Learning.

[0136] The general form of TD error is: y t = r t + γV(S t+1 ) - V(S t ).

[0137] Where y t is the TD error, r t is the immediate reward, γ is the discount factor, V(S t+1 ) is the value estimate of the next state, and V(S t ) is the value estimate of the current state.

[0138] The discount factor γ, γ ∈ [0, 1]. It determines the weight of future rewards in the current estimate. The larger γ, the more emphasis on long-term returns.

[0139] Where V(S t+1 ) is predicted by the second network model according to the action output by the first network model and the current state S t .

[0140] Step S606, determine the mean square loss based on the TD error and the first value estimate.

[0141] Where the mean square loss can refer to the formula L Q = E[(y t - Q(s t , a t )] 2 . Where Q(s t , a t ) can be the first value estimate, s t is the current state (i.e. the first frame state vector), a t is the action output by the first network model for the first frame state vector. E is the expectation, y t is the TD error.

[0142] Step S608, update the parameters of the first network model based on the TD error.

[0143] In some embodiments, the policy gradient method (such as gradient ascent in the Actor-Critic framework) can be used to update the parameters of the first network model (such as Actor) based on the advantage function (or TD error) provided by the Critic.

[0144] Step S610, the second network model is updated in parameters through the mean square loss of the time difference error.

[0145] In some embodiments, the Critic model can be updated in parameters through the mean square loss (MSE) of the time difference error (TD error).

[0146] The model parameter updating method is based on the Actor-Critic framework, and the collaborative optimization of the policy network (Actor) and the value network (Critic) is realized through the time difference (TD) error: the Critic improves the accuracy of the value estimation by minimizing the mean square loss of the TD error, and the Actor updates the parameters using the TD error provided by the Critic as the policy gradient direction, thereby optimizing the action selection policy. At the same time, the method combines task-specific rewards (such as intersection over union and identity consistency), so that the model can optimize short-term returns and long-term performance simultaneously in tasks such as object detection or tracking, and finally realize efficient and stable policy learning and value convergence.

[0147] Figure 7 is a flowchart of a state vector determination method according to an example embodiment.

[0148] Reference Figure 7 The above state vector determination method can include the following steps.

[0149] Step S702, determine the object feature corresponding to each candidate bounding box.

[0150] Step S704, splice the object features corresponding to each candidate bounding box in the first frame image by object to obtain a first frame state vector corresponding to the first frame image.

[0151] The state vector determination method extracts and splices the object features of the candidate bounding boxes to construct a feature vector that comprehensively represents the target state of the first frame image. The technical effects are as follows: 1) the object-level feature extraction retains the detailed information of each detection target; 2) the feature splicing method fuses multi-target features to form a global state representation containing comprehensive information of all objects in the scene; 3) the generated first frame state vector can be used as the input of the reinforcement learning model, providing rich environmental state information for subsequent action decision (such as target tracking), thereby improving the perception and understanding ability of the model for complex multi-target scenes.

[0152] Figure 8 is a flowchart of an object feature determination method according to an example embodiment.

[0153] In some embodiments, the at least one candidate object bounding box includes a target candidate object bounding box. In the following, the present application will take the target candidate object bounding box as an example to explain how to determine the object features corresponding to the candidate object bounding box.

[0154] Reference Figure 8 The above object feature determination method can include the following steps.

[0155] In step S802, spatial position information, motion speed information, appearance feature information, and target-to-target speed difference information corresponding to the target candidate object bounding box are determined.

[0156] The spatial position information is usually expressed in the coordinates of the bounding box (such as the center point coordinates (x, y), the width w, and the height h) or the normalized relative coordinates (relative to the image size). It can be directly output by a target detection model (such as Faster R-CNN or YOLO) or calculated by feature map regression.

[0157] The motion speed information can be obtained by dividing the coordinate difference of the target center point in two consecutive frames by the time interval to obtain the instantaneous speed.

[0158] The appearance feature information can be extracted from the image region within the bounding box by a CNN (such as ResNet) to obtain a high-dimensional feature vector.

[0159] The target-to-target speed difference information is the difference in the motion speed of different targets in the same scene, which is used to analyze the relative motion relationship between the targets. The motion speed of each target can be calculated (in the same way as above), and then the speed vectors of each pair of targets are subtracted.

[0160] In step S804, the object features corresponding to the target candidate object bounding box are determined according to the spatial position information, the motion speed information, the appearance feature information, and the target-to-target speed difference information corresponding to the target candidate object bounding box.

[0161] The object feature determination method fuses multi-dimensional information such as the spatial position (geometric attribute) of the target, the motion speed (dynamic characteristic), the appearance feature (visual identifier), and the target-to-target speed difference (interaction relationship) to construct a comprehensive and robust object feature representation. This method not only accurately describes the static attributes of the target, but also captures its dynamic behavior and interaction relationship, providing a composite feature support with spatial, temporal, and semantic information for subsequent reinforcement learning decisions (such as target tracking), significantly improving the model perception ability and decision accuracy in complex scenes.

[0162] Figure 9 is a flowchart of a multi-target tracking method according to an example embodiment.

[0163] Reference Figure 9The multi-target tracking method can include the following steps.

[0164] The technical method provided by the embodiment is an optimization algorithm architecture of a multi-target tracking method based on reinforcement learning on traffic monitoring, which can specifically include the following steps.

[0165] Step S901, input the monitoring video stream.

[0166] Step S902, use a Yolox detector to extract all pedestrian targets in the video stream.

[0167] Step S903, extract the spatial position, motion speed, appearance feature, and speed difference of each pedestrian target to splice into a state vector.

[0168] Step S904, the reinforcement learning Actor network predicts the execution action and pedestrian target posture information.

[0169] Step S905, the DeepSORT network performs data association and calculates the composite reward.

[0170] Step S906, the Critic network performs calculation.

[0171] Step S907, the Actor-Critic framework jointly updates and optimizes the association result.

[0172] The specific implementation process can include the following steps.

[0173] Step one: preprocess the video data input in the traffic monitoring video and extract the features and state representation of the pedestrians.

[0174] Step two: build a reinforcement learning strategy network and a reward function.

[0175] Step three: give the tracking module the predicted state obtained after reinforcement learning training.

[0176] Step four: the tracking module processes one step at a time and compares the reward function to reward and punish.

[0177] Step five: training cycle, constantly optimize the overall tracking effect through the reward function.

[0178] Step one: video preprocessing and state feature extraction.

[0179] 1. Video input.

[0180] Get the real-time or offline video stream of the traffic monitoring camera and read it by frame.

[0181] 2. Target detection and feature extraction.

[0182] All pedestrian target bounding boxes are identified using the pre-trained detector YOLOX for each frame of image.

[0183] The spatial position (center coordinates, width and height), motion speed (calculated by adjacent frame position difference), appearance feature (extracted deep feature vector using lightweight CNN), and target interaction encoding (speed difference) of each pedestrian target are extracted.

[0184] 3. State representation.

[0185] The above features are spliced by target to form the state vector s of the current frame t ; At the same time, the state sequence of the last k frames of the current frame is maintained for the time series encoding of the Critic.

[0186] Step two: build Actor-Critic network and reward function.

[0187] 1. Actor network design.

[0188] Input layer: state vector s t .

[0189] Hidden layer: fully connected 256→ReLU, fully connected 128→ReLU.

[0190] Output layer: action vector a t : bounding box offset (Δx, Δy, Δw, Δh). Where Δx is the horizontal coordinate, Δy is the vertical coordinate, Δw is the width of the bounding box, and Δh is the height of the bounding box.

[0191] 2. Critic network design.

[0192] Input layer: concatenate state s t action a t .

[0193] Time series encoding layer: this embodiment designs a single-layer LSTM (hidden 128) to encode the state sequence of the last k frames and capture the motion time series dependence of the pedestrian target.

[0194] Multi-layer perceptron: fully connected 256→ReLU→BatchNorm→Dropout (0.2), fully connected 128→ReLU→BatchNorm→Dropout (0.2).

[0195] Output layer: output Q(s t , a t ).

[0196] 3. Reward function design.

[0197] Short-term reward r IOU: The Intersection over Union (IoU) between the predicted bounding box of all pedestrian targets in the current frame and the real labeled box.

[0198] Long-term reward r ID : Calculate the IDF1 increment every 10 frames.

[0199] Composite real-time reward: r t = αr IOU + (1-α) r ID , which takes into account both real-time positioning accuracy and cross-frame consistency.

[0200] Step three: integration with the tracking module.

[0201] 1. Action execution.

[0202] The actor generates an action a t from the current state s t , and applies a delta offset to the bounding box position of each target. When a pedestrian target encounters occlusion, the actor network prioritizes historical trajectory and appearance similarity, reducing ID switching problems when occluded.

[0203] 2. Tracking module.

[0204] The predicted pedestrian target pose generated by the actor, i.e., the updated position and scale, is input into the standard tracking algorithm (DeepSORT) to continue data association.

[0205] 3. Reward evaluation.

[0206] After the tracking module outputs the association result, the reward r t is calculated and fed back to the Critic network for subsequent updates.

[0207] Step four: joint update of Critic and Actor.

[0208] 1. Experience replay: store the interaction quadruple (s t , r t , a t , s t+1 ) in the experience replay buffer (ReplayBuffer). Store historical experience through Replay Buffer, perform batch sampling, optimize tracking strategy, especially when pedestrians are dense or interactions occur, update the strategy to reduce false matches. Here, s t is the current state vector, r t is the reward under the current state, a t is the action obtained by processing the current state vector with the first network model, and s t+1 is the state vector corresponding to the next state of the current state.

[0209] 2. Batch sampling: Randomly sample mini-batches from the experience replay buffer, shuffle the time order to reduce correlation.

[0210] 3. Critic update.

[0211] Compute TD error: y t = r t + γV(S t+1 ) - V(S t ), where the discount factor γ, γ ∈ [0, 1]. Determines the weight of future rewards in the current estimate, the greater the γ, the more emphasis on long-term returns. V(S t+1 ) is the value estimate of the next state, V(S t ) is the value estimate of the current state.

[0212] Minimize mean squared loss: L Q = E[(y t - Q(s t , a t )] 2 . Where Q(s t , a t ) can be the first value estimate, s t is the current state (i.e. the first frame state vector), a t is the action output by the first network model for the first frame state vector. E is the expectation, y t is the TD error.

[0213] Update the Actor through the TD error and update the Critic network by minimizing the mean squared loss.

[0214] Step five: training cycle and convergence.

[0215] Initialization: Randomly initialize the parameters of the Actor, Critic and their target networks, and set an empty Replay Buffer.

[0216] Iterative interaction: Repeat "Step 1→ Step 2→ Step 3→ Step 4" until the global round number or the performance indicator (such as MOTA or IDF1) converges.

[0217] Online optimization: Combine domain randomization, adversarial pre-training and other means to further strengthen the generalization and robustness of the model.

[0218] The following advantageous means are adopted in this embodiment.

[0219] 1. Multi-object tracking optimization architecture based on Actor-Critic.

[0220] Technical means: Introduce deep Actor-Critic algorithm to optimize traditional multi-target tracking algorithm, jointly train policy network (Actor) and value network (Critic), and use the TD error calculated by Critic to directly guide the policy update of Actor.

[0221] Beneficial effects: Compared with the traditional Hungarian algorithm and Kalman filter (KF) which need to manually design association rules and adjust parameters, the embodiment can automatically learn the optimal association strategy through sample interaction, greatly improve the training efficiency and pedestrian target tracking stability, and reduce the cost of human intervention. At the same time, an optimization algorithm that can adapt to occlusion, interaction and appearance similarity is designed, which improves the stability and precision of pedestrian tracking in real environment.

[0222] 2. Deep Critic network with LSTM time series encoding.

[0223] Technical means: Add a single-layer LSTM to the input end of the Critic network to extract time series features for continuous k-frame state sequences; then connect a multilayer perceptron (256→128 neurons), supplemented by BatchNorm and Dropout regularization.

[0224] Beneficial effects: LSTM can accurately capture the motion trend of pedestrians between multiple frames of monitoring video, and combined with deep network structure, it can more accurately estimate future value in mutation or occlusion scenarios, dynamically adjust the association strategy when occlusion occurs, reduce target loss problems caused by occlusion, and improve tracking robustness and precision.

[0225] 3. Compound long-term and short-term reward function design.

[0226] Technical means: Take the IoU precision of each frame as the short-term reward, and take the IDF1 increment every several frames as the long-term reward, and mix r t =αr IOU +(1-α)r ID .

[0227] Beneficial effects: It balances the immediate association accuracy and cross-frame identity consistency, can reduce the ID switching rate, improve the comprehensive evaluation indexes such as MOTA, IDF1 and HOTA, and can balance the accuracy and stability better than the single reward function.

[0228] The embodiment adopts a multi-target tracking framework based on reinforcement learning, and especially optimizes the tracking strategy of pedestrians in a traffic monitoring scene. In a scene with dense pedestrians and serious occlusion, the reinforcement learning agent can adjust the association strategy in real time through interaction with the environment, and preferentially selects pedestrian targets with similar historical trajectories, instead of simply relying on the IoU (Intersection over Union) of the current detection box. Through the reinforcement learning (RL) optimization algorithm, the tracking robustness and efficiency of pedestrians in a complex dynamic environment can be significantly improved, and an important means is provided for multi-target tracking optimization. The embodiment has at least the following advantages.

[0229] The embodiment dynamically adjusts the tracking strategy through interaction with the environment. In the case of dense pedestrians, occlusion occurs, and at this time the RL agent can learn to preferentially match targets with high historical trajectory similarity, instead of relying only on the IoU of the current detection box. In a complex interactive modeling traffic scene, pedestrians interact frequently.

[0230] The embodiment reduces the dependence on the manual design of artificial rules required by traditional MOT for data association: the cost matrix of the Hungarian algorithm, and the RL automatically learns the optimal association strategy through the reward function, can replace the Hungarian algorithm, and directly outputs the association probability of the detection box and the tracking target, avoiding manual parameter tuning of the IoU threshold, motion model weight and the like; the traditional method needs to be optimized in steps, but the joint optimization of multi-task RL can simultaneously optimize the detection association, trajectory prediction, ID management and other sub-tasks. At the same time, the global optimum is realized through a hybrid action space, which includes discrete actions: selecting a detection box association target, and continuous actions: adjusting the noise parameters of the Kalman filter or the appearance feature weight;

[0231] For the reward function, the long-term tracking performance optimization reward function guides the long-term target. The traditional method optimizes frame by frame, which is easy to cause the ID identity of pedestrians to switch, such as mis-matching after temporary occlusion. Compared with the traditional method, the present application can dynamically adjust the association strategy when occlusion occurs through the reward function design of reinforcement learning, and reduce the target loss problem caused by occlusion. RL gives cumulative rewards by designing a long-term reward: ID index increment, and every 10 frames statistics ID retention rate, and a short-term reward: single-frame detection association accuracy (IoU), to encourage the agent to maintain ID identity consistency;

[0232] The embodiment adds an LSTM layer in the Critic network, which is used to capture the motion trend of the pedestrian target in the time sequence. Especially in the scene where pedestrians quickly pass through or suddenly stop, the future motion trajectory of the pedestrian can be effectively predicted, and the tracking failure caused by sudden motion can be reduced.

[0233] The embodiment processes uncertainty and sparse data robustness enhancement. In low visibility (rain and fog), sensor noise and other scenes, the tracking failure rate of traditional methods is high. However, the embodiment can improve the robustness to noise and missing data by adversarial training or domain randomization, pre-train the strategy in the simulation environment, solve the long-term loss problem of pedestrian target caused by occlusion, and guide the agent to actively explore potential associations.

[0234] In the embodiment, a distributed RL framework for multi-target cooperation and resource allocation is used to assign an independent agent (Multi-Agent RL) to each tracking target, and resources are coordinated through centralized training and decentralized execution (CTDE). In a pedestrian traffic video stream, agents can cooperate to avoid multiple trackers (pedestrian targets) competing for the same detection box. RL reduces computational overhead and optimizes computing resources through action pruning and asynchronous inference.

[0235] Through the above steps, the method of reinforcement learning can effectively face the uncertain environment in traffic monitoring and picture occlusion or missed detection, and association errors.

[0236] It should be particularly pointed out that each step in each embodiment of the above multi-target tracking method can be crossed, replaced, added, deleted and reduced. Therefore, these reasonable permutations and combinations of the multi-target tracking method should also belong to the protection scope of the present disclosure, and the protection scope of the present disclosure should not be limited to the above embodiments.

[0237] Based on the same inventive concept, the present disclosure also provides a multi-target tracking device, as follows. Since the principle of solving problems in the device embodiment is similar to that of the above-mentioned method embodiment, the implementation of the device embodiment can be referred to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0238] Figure 10 is a block diagram of a multi-target tracking device according to an exemplary embodiment. Referring to Figure 10 The multi-target tracking device 1000 provided by the embodiment of the present disclosure can include a video acquisition module 1001, a bounding box determination module 1002, a first frame state vector determination module 1003, an offset prediction module 1004, a bounding box prediction module 1005, and an object association module 1006.

[0239] The video acquisition module 1001 can be configured to acquire a target video, the target video including a plurality of images, and the plurality of images including a first image; the bounding box determination module 1002 can be configured to perform object bounding box detection on the first image by using a pre-trained detector to obtain at least one candidate object bounding box; the first frame state vector determination module 1003 can be configured to determine a first frame state vector corresponding to the first image according to each candidate object bounding box; the offset prediction module 1004 can be configured to process the first frame state vector by using a first network model to obtain a bounding box offset corresponding to each candidate object bounding box; the bounding box prediction module 1005 can be configured to determine a predicted object bounding box corresponding to the first image according to each candidate object bounding box and the bounding box offset corresponding to each candidate object bounding box; and the object association module 1006 can be configured to input the at least one predicted object bounding box into a target tracking algorithm to determine the identity of each object in the first image to perform cross-frame object association.

[0240] It should be noted that the video acquisition module 1001, the bounding box determination module 1002, the first frame state vector determination module 1003, the offset prediction module 1004, the bounding box prediction module 1005, and the object association module 1006 correspond to S202-S212 in the method embodiment, and the above modules have the same examples and application scenarios as the corresponding steps, but are not limited to the disclosure of the above method embodiments. It should be noted that the above modules as part of the device can be executed in a computer system such as a set of computer executable instructions.

[0241] In some embodiments, the multi-target tracking apparatus 1000 can further include a real bounding box determination module, an intersection-over-union determination module, and a reward evaluation value determination module.

[0242] The real bounding box determination module can be configured to acquire a real object bounding box corresponding to each object in the first image; the intersection-over-union determination module can be configured to determine the intersection-over-union between the predicted object bounding box and the real object bounding box; and the reward evaluation value determination module can be configured to determine a reward evaluation value based on the intersection-over-union, so as to update the network parameters of the first network model based on the reward evaluation value.

[0243] In some embodiments, the multi-target tracking apparatus 1000 can further include a first adjacent frame image acquisition module and a frame ratio determination module.

[0244] The first adjacent frame image acquisition module can be configured to acquire N frame images before the first frame image as the first adjacent frame images, where N is an integer greater than or equal to 1; and the frame proportion determination module can be configured to determine a frame proportion of correctly maintaining identity consistency in the first adjacent frame images according to the identity of each object in the first frame image and the identity corresponding to each object in each first adjacent frame image.

[0245] The reward evaluation value determination module can include a reward value determination submodule.

[0246] The reward value determination submodule can be configured to determine the reward evaluation value based on the intersection over union and the proportion of correctly maintaining identity consistency, so as to update the parameters of the first network model according to the reward evaluation value.

[0247] In some embodiments, the boundary box offset corresponding to each candidate object boundary box is a predicted action output by the first network model; and the multi-target tracking device 1000 can further include a second adjacent image acquisition module, a second adjacent frame state vector acquisition module, a first value evaluation value acquisition module, and a parameter update module.

[0248] The second adjacent image acquisition module can be configured to acquire k frame images before the first frame image as the second adjacent frame images, where k is an integer greater than or equal to 1; the second adjacent frame state vector acquisition module can be configured to acquire a second adjacent frame state vector corresponding to each second adjacent frame image; the first value evaluation value acquisition module can be configured to perform a prediction process on the predicted action of the first network model, the first frame state vector, and each second adjacent frame state vector by the second network model to obtain a first value evaluation value, the first value evaluation value being used to evaluate the value corresponding to the predicted action and the first frame state vector; and the parameter update module can be configured to update the parameters of the first network model and the second network model based on the first value evaluation value.

[0249] In some embodiments, the parameter update module can include a reward evaluation value acquisition submodule, a TD error determination submodule, a mean square loss determination submodule, a first parameter update submodule, and a second parameter update submodule.

[0250] The reward evaluation value acquisition submodule can be configured to acquire a reward evaluation value determined based on the output value of the first network model; the TD error determination submodule can be configured to determine a TD error corresponding to the second network model based on the reward evaluation value; the mean square loss determination submodule can be configured to determine a mean square loss based on the TD error and the first value evaluation value; the first parameter update submodule can be configured to update the parameters of the first network model based on the TD error; and the second parameter update submodule can be configured to update the parameters of the second network model by the mean square loss.

[0251] In some embodiments, the first frame state vector determination module 1003 can include an object feature determination sub-module and a splicing sub-module.

[0252] The object feature determination sub-module can be configured to determine the object feature corresponding to each candidate bounding box, and the splicing sub-module can be configured to splice the object features corresponding to the candidate bounding boxes in the first frame image according to objects to obtain the first frame state vector corresponding to the first frame image.

[0253] In some embodiments, the at least one candidate object bounding box includes a target candidate object bounding box, and the object feature determination sub-module can include an information determination unit and an object feature determination unit.

[0254] The information determination unit can be configured to determine the spatial position information, the motion speed information, the appearance feature information, and the inter-target speed difference information corresponding to the target candidate object bounding box, and the object feature determination unit can be configured to determine the object feature corresponding to the target candidate object bounding box according to the spatial position information, the motion speed information, the appearance feature information, and the inter-target speed difference information corresponding to the target candidate object bounding box.

[0255] Since the functions of the apparatus 1000 have been described in detail in the corresponding method embodiments, the present disclosure will not be repeated here.

[0256] The modules and / or sub-modules and / or units described in the embodiments of the present disclosure can be implemented in the form of software or in the form of hardware. The described modules and / or sub-modules and / or units can also be arranged in a processor. In some cases, the names of these modules and / or sub-modules and / or units do not constitute a limitation on the modules and / or sub-modules and / or units themselves.

[0257] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module or a part of a program segment that includes one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders from those shown in the drawings. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams or flowcharts, and combinations of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer program instructions.

[0258] Further, the above-described diagrams are merely schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not intended for limiting purposes. It is readily understood that the processes shown in the above-described diagrams do not indicate or limit the time sequence of the processes. In addition, it is also readily understood that the processes can be executed synchronously or asynchronously, for example, in a plurality of modules.

[0259] Figure 11 A structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. It should be noted that Figure 11 The electronic device 1100 shown is merely an example, and should not impose any limitation on the functions and usage range of the embodiments of the present disclosure.

[0260] As Figure 11 shown, the electronic device 1100 includes a central processing unit (CPU) 1101, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1102 or programs loaded from a storage section 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are also stored. The CPU 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0261] The following components are connected to the I / O interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as necessary. A removable recording medium 1111 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 1110 as necessary, so that a computer program read therefrom is installed into the storage section 1108 as necessary.

[0262] In particular, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product including a computer program carried on a computer-readable storage medium, the computer program containing computer program instructions for executing the method shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication section 1109, and / or installed from the removable recording medium 1111. When the computer program is executed by the central processing unit (CPU) 1101, the above-described functions defined in the system of the present disclosure are executed.

[0263] Note that the computer-readable storage medium shown in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carrying computer-readable computer program instructions in a baseband or as part of a carrier wave. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium that can send, propagate or transmit the program for use by or in conjunction with an instruction execution system, device or apparatus. The computer program instructions contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0264] As another aspect, the disclosure also provides a computer readable storage medium, which can be included in the device described in the above embodiments, or can exist independently without being assembled into the device. The computer readable storage medium carries one or more programs, which, when executed by the device, enable the device to implement functions including: obtaining a target video, the target video including a plurality of frames of images, wherein the plurality of frames of images include a first frame of image; performing object bounding box detection on the first frame of image by a pre-trained detector to obtain at least one candidate object bounding box; determining a first frame state vector corresponding to the first frame of image according to each candidate object bounding box; processing the first frame state vector by a first network model to obtain a bounding box offset corresponding to each candidate object bounding box; determining a predicted object bounding box corresponding to the first frame of image according to each candidate object bounding box and the bounding box offset corresponding to each candidate object bounding box; and inputting the at least one predicted object bounding box into a target tracking algorithm to determine the identity of each object in the first frame of image for cross-frame object association.

[0265] According to an aspect of the disclosure, a computer program product or computer program is provided, which includes computer program instructions stored in a computer readable storage medium. The computer program instructions are read from the computer readable storage medium and executed by a processor to implement the method provided in various optional implementation manners of the above embodiments.

[0266] From the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions of the embodiments of the disclosure can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of computer program instructions for causing an electronic device (which can be a server or a terminal device, etc.) to execute the method according to the embodiments of the disclosure.

[0267] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure as disclosed herein. The disclosure is intended to cover any variations, uses, or adaptations of the disclosure following, in general, the principles of the disclosure and including such departures from the disclosure as come within known or customary practice in the art to which the disclosure pertains. The specification and examples are to be regarded as exemplary only, and the true scope and spirit of the disclosure are indicated by the claims.

[0268] It is to be understood that the present disclosure is not limited to the detailed description, drawings, or implementation methods already shown here, but rather the present disclosure is intended to encompass various modifications and equivalent arrangements within the spirit and scope of the appended claims.

Claims

1. A multi-target tracking method, characterized by, The method comprises: obtaining a target video, wherein the target video comprises a plurality of frames of images, and the plurality of frames of images comprise a first frame of image; performing object bounding box detection on the first frame of image by using a pre-trained detector to obtain at least one candidate object bounding box; determining a first frame state vector corresponding to the first frame of image according to each candidate object bounding box; processing the first frame state vector by using a first network model to obtain a bounding box offset corresponding to each candidate object bounding box; determining a predicted object bounding box corresponding to the first frame of image according to each candidate object bounding box and the bounding box offset corresponding to each candidate object bounding box; inputting the at least one predicted object bounding box into a target tracking algorithm to determine the identity of each object in the first frame of image for cross-frame object association.

2. The method of claim 1, wherein, The method further comprises: obtaining a real object bounding box corresponding to each object in the first frame of image; determining an intersection-over-union between the predicted object bounding box and the real object bounding box; determining a reward evaluation value based on the intersection-over-union, so as to update network parameters of the first network model based on the reward evaluation value.

3. The method of claim 2, wherein, The method further comprises: obtaining N frames of image before the first frame of image as first adjacent frames of image; N is an integer greater than or equal to 1; determining a proportion of frames in which the identity is correctly maintained in the first adjacent frames of image according to the identity of each object in the first frame of image and the identity corresponding to each object in each first adjacent frame of image; wherein the determination of the reward evaluation value based on the intersection-over-union, so as to update the network parameters of the first network model based on the reward evaluation value, comprises: determining the reward evaluation value based on the intersection-over-union and the proportion of correctly maintained identity, so as to update the parameters of the first network model according to the reward evaluation value.

4. The method of claim 1, wherein, The bounding box offset corresponding to each candidate object bounding box is a predicted action output by the first network model; and the method further comprises: obtaining k frames of image before the first frame of image as second adjacent frames of image; k is an integer greater than or equal to 1; obtaining a second adjacent frame state vector corresponding to each second adjacent frame of image; performing prediction processing on the predicted action of the first network model, the first frame state vector, and each second adjacent frame state vector by using a second network model to obtain a first value evaluation value, wherein the first value evaluation value is used to evaluate the value corresponding to the predicted action and the first frame state vector; updating the parameters of the first network model and the second network model based on the first value evaluation value.

5. The method of claim 4, wherein, The updating of the parameters of the first network model and the second network model based on the first value evaluation value comprises: obtaining a reward evaluation value determined based on the output value of the first network model; determining a TD error corresponding to the second network model based on the reward evaluation value; determining a mean square loss based on the TD error and the first value evaluation value; updating the parameters of the first network model based on the TD error; and updating the parameters of the second network model based on the mean square loss. The second network model is updated in parameters by the mean square loss.

6. The method of claim 1, wherein, The first frame state vector corresponding to the first frame image is determined according to each candidate object bounding box, including: An object feature corresponding to each candidate bounding box is determined. The object features corresponding to each candidate bounding box in the first frame image are spliced by object to obtain the first frame state vector corresponding to the first frame image.

7. The method of claim 6, wherein, The at least one candidate object bounding box includes a target candidate object bounding box; wherein the object feature corresponding to each candidate bounding box is determined, including: Spatial position information, motion speed information, appearance feature information and target-to-target speed difference information corresponding to the target candidate object bounding box are determined. The object feature corresponding to the target candidate object bounding box is determined according to the spatial position information, motion speed information, appearance feature information and target-to-target speed difference information corresponding to the target candidate object bounding box.

8. A multi-target tracking device, characterized by, Including: A video acquisition module is configured to acquire a target video, the target video including a plurality of frames of images, wherein the plurality of frames of images include a first frame of image; A bounding box determination module is configured to perform object bounding box detection on the first frame of image by a pre-trained detector to obtain at least one candidate object bounding box; A first frame state vector determination module is configured to determine a first frame state vector corresponding to the first frame of image according to each candidate object bounding box; An offset prediction module is configured to process the first frame state vector by a first network model to obtain a bounding box offset corresponding to each candidate object bounding box; A bounding box prediction module is configured to determine a predicted object bounding box corresponding to the first frame of image according to each candidate object bounding box and the bounding box offset corresponding to each candidate object bounding box; An object association module is configured to input the at least one predicted object bounding box into a target tracking algorithm to determine an identity of each object in the first frame of image for cross-frame object association.

9. An electronic device, comprising: Including: A memory and a processor; The memory is configured to store computer program instructions; the processor is configured to invoke the computer program instructions stored in the memory to implement the multi-target tracking method according to any one of claims 1-7.

10. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions are executed by the processor to implement the multi-target tracking method according to any one of claims 1-7.

11. A computer program product comprising computer program instructions stored in a computer readable storage medium, characterized in that, The computer program instructions are executed by the processor to implement the method according to any one of claims 1-7.