Multi-target labeling method and device, equipment and storage medium

By setting the highest priority for accurate labels and implementing step-by-step manual review in the multi-target tracking model, the problem of low labeling accuracy in multi-target tracking models is solved, labeling accuracy is improved, and manual review time is reduced.

CN121600433APending Publication Date: 2026-03-03GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411137398.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Multi-target tracking models have limited accuracy in labeling specific targets in video data, leading to rampant misidentification and IDs, and requiring lengthy manual review.

Method used

By obtaining the initial accurate labels and bounding boxes of the target video data, and combining them with the detection boxes, the multi-target tracking model is optimized by setting the accurate labels with the highest priority and gradually performing manual review and correction.

Benefits of technology

It improves the labeling accuracy of multi-target tracking models, shortens the manual review time, and avoids misassignment and uncontrolled growth of IDs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600433A_ABST
    Figure CN121600433A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a multi-target labeling method and device, equipment and a storage medium. The method comprises the steps of obtaining target video data, wherein each target object appearing for the first time has an initial accurate label and a labeling box in a corresponding video frame; detecting each target object appearing in each video frame in the target video data to obtain a detection frame marking a detection result; the target video data, the initial accurate labels, the labeling boxes and the detection boxes are input into a multi-target tracking model together, target objects are tracked and matched through the multi-target tracking model, the highest priority is set for the accurate labels, first labeling labels corresponding to all the target objects in all the video frames are obtained, and the accurate labels comprise the initial accurate labels and the detection boxes; the first labeling labels corresponding to the same target object are the same and are initial accurate labels corresponding to the target object. According to the technical means, the technical problem that the labeling accuracy of a multi-target tracking model is limited when tracking labeling is carried out on a plurality of specific targets in the video data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a multi-target annotation method, apparatus, device, and storage medium. Background Technology

[0002] Multi-object tracking annotation is fundamental to multi-object tracking research, requiring the annotation of bounding boxes and IDs for multiple specific objects within each video frame. In related technologies, after inputting video data into a multi-object tracking model, the model can identify the same specific object in different video frames and assign the same ID to the same specific object in different video frames, thus achieving object annotation. Then, the multi-object tracking model assigns a new ID to each newly identified specific object. However, due to the limited accuracy of multi-object tracking models, misidentification is possible. For example, a previously identified specific object may be misidentified as a newly identified specific object, or vice versa. Assigning IDs to misidentified specific objects not only results in incorrect IDs but also leads to the problem of IDs growing uncontrollably. Therefore, manual review of the IDs for each specific object in each video frame is necessary to correct erroneous IDs. However, when the amount of video data requiring manual review and correction is large, it consumes significant manpower and time. Summary of the Invention

[0003] This application provides a multi-target annotation method, apparatus, device, and storage medium to solve the technical problem in the related art where the annotation accuracy of multi-target tracking models is limited when tracking and annotating multiple specific targets in video data.

[0004] Firstly, one embodiment of this application provides a multi-target annotation method, including:

[0005] Obtain the target video data and the initial accurate labels and bounding boxes of each target object that appears for the first time in the target video data in the corresponding video frame. Each target object that appears for the first time has a corresponding initial accurate label and bounding box.

[0006] Each target object appearing in each video frame of the target video data is detected to obtain a detection box for marking the detection results. Each target object detected in each video frame has a corresponding detection box.

[0007] The target video data, the initial accurate label, the bounding box, and the detection box are input together into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching, and sets the highest priority for the accurate label during tracking and matching to obtain the first label corresponding to each target object in each video frame. The accurate label includes the initial accurate label. The first labels corresponding to the same target object tracked by the multi-target tracking model are the same and are all the initial accurate labels corresponding to the target object.

[0008] The above-described method involves acquiring target video data with initial accurate labels and bounding boxes, detecting each target object within the video data to obtain detection boxes that identify the detection results, and then inputting the target video data, initial accurate labels, bounding boxes, and detection boxes into a multi-target tracking model. The multi-target tracking model tracks and matches the target objects, assigning the highest matching priority to the accurate labels during the tracking process. The model then obtains the first label for each target object, with the same first label corresponding to the same target object, and both being the initial accurate labels for that target object. This technical solution addresses the problem of limited labeling accuracy in multi-target tracking models when tracking and labeling multiple specific targets in video data in related technologies. In the multi-object tracking model, the input accurate labels are used as a reference for tracking and matching. This ensures that the labels obtained from the tracking and matching (currently the first label) are all manually set accurate labels. That is, the first label corresponding to the same target object is the accurate label marked when the target object first appears. This avoids the problem of ID (i.e., label) growing arbitrarily. In addition, each newly appearing target object will not be misidentified. That is, the multi-object tracking model does not need to identify newly appearing target objects or create new labels (because truly newly appearing target objects all have corresponding initial accurate labels). This separates detection and matching tracking (i.e., the multi-object tracking model does not need to rely on detection boxes to identify newly appearing target objects), which can improve the labeling accuracy of the multi-object tracking model.

[0009] In one embodiment of this application, after the multi-object tracking model performs target object tracking and matching and sets the highest priority for accurate labels during tracking and matching to obtain the first annotation label corresponding to each target object in each video frame, the process includes:

[0010] Obtain a first accurate label and a corresponding bounding box based on a portion of the first first label. The portion of the first label is the first label corresponding to each target object in each video frame obtained according to the first frame number interval in the target video data. The first accurate label is the accurate label obtained after reviewing and correcting the portion of the first label. The bounding box corresponding to the first accurate label is the accurate bounding box obtained after reviewing and correcting the detection box corresponding to the corresponding first label.

[0011] The target video data, the first accurate label, the initial accurate label, the annotation box, and the detection box are input together into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching, and sets the highest priority for the accurate label during tracking and matching to obtain the second annotation label corresponding to each target object in each video frame. The accurate label includes the initial accurate label and the first accurate label.

[0012] The above-mentioned method involves manually reviewing and correcting some of the labels obtained by the multi-object tracking model, and then using the multi-object tracking model again to track and process the target video data. In the process, the accurate labels are used as a reference (i.e., the highest priority). This can significantly improve the accuracy of the labels obtained by the multi-object tracking model and thus shorten the manual review time.

[0013] In one embodiment of this application, the step of tracking and matching target objects by the multi-target tracking model and setting the highest priority for accurate labels during tracking and matching to obtain the second annotation label corresponding to each target object in each video frame includes:

[0014] Obtain a second accurate label based on a portion of the second annotation labels and the corresponding annotation box of the second accurate label. The portion of the second annotation labels are the second annotation labels corresponding to each target object in each video frame obtained according to the second frame number interval in the target video data. The second accurate label is the accurate label obtained after reviewing and correcting the portion of the second annotation labels. The annotation box corresponding to the second accurate label is the accurate annotation box obtained after reviewing and correcting the detection box corresponding to the corresponding second annotation label. The second frame number is less than the first frame number.

[0015] The target video data, the second accurate label, the first accurate label, the initial accurate label, the annotation box, and the detection box are input together into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching, and sets the highest priority for the accurate label during tracking and matching to obtain the third annotation label corresponding to each target object in each video frame. The accurate label includes the initial accurate label, the first accurate label, and the second accurate label.

[0016] The above-mentioned approach involves a step-by-step process where some of the labels obtained by the multi-target tracking model are first manually reviewed and corrected. This ensures that the multi-target tracking model uses accurate labels as a reference (i.e., the highest priority) when it tracks and processes target video data again. This significantly improves the accuracy of the labels obtained by the multi-target tracking model and thus shortens the manual review time.

[0017] In one embodiment of this application, the step of tracking and matching target objects by the multi-target tracking model and setting the highest priority for accurate labels during tracking and matching includes:

[0018] The multi-target tracking model matches the bounding boxes corresponding to each accurate label in the current video frame of the target video data with each tracking trajectory in the trajectory set to obtain a set of successfully matched bounding boxes, a set of unmatched first bounding boxes, and a set of unmatched first trajectories. The trajectory set is a set of each tracking trajectory obtained during the processing of the multi-target tracking model. Each tracking trajectory corresponds to a target object and represents the position trajectory of the currently identified target object in the target video data.

[0019] The multi-target tracking model matches each bounding box in the current video frame with each detection box in the current video frame to obtain a first set of successfully matched detection boxes, a first set of unmatched detection boxes, and a second set of unmatched bounding boxes.

[0020] The multi-target tracking model matches the detection boxes in the first set of unmatched detection boxes with the tracking trajectories in the first set of unmatched trajectories to obtain the second set of successfully matched detection boxes, the second set of unmatched detection boxes, and the second set of unmatched trajectories.

[0021] The multi-target tracking model creates new tracking trajectories based on each of the unmatched bounding boxes in the first set of unmatched bounding boxes, and the new tracking trajectories use the accurate labels corresponding to the bounding boxes as labels.

[0022] The multi-target tracking model updates the corresponding matched tracking trajectory based on the set of successfully matched bounding boxes and the set of successfully matched second detection boxes, so that the updated tracking trajectory shares the same label.

[0023] The multi-target tracking model retains each tracking trajectory in the set of unmatched second trajectories;

[0024] The multi-target tracking model obtains a set of trajectories for use in the next video frame based on the new tracking trajectory, the updated tracking trajectory, and the retained tracking trajectory.

[0025] The multi-target tracking model ignores the set of successfully matched first detection boxes, the set of unmatched second annotation boxes, and the set of unmatched second detection boxes.

[0026] The multi-target tracking model updates the next video frame in the target video data to the current video frame, and then returns to perform the operation of matching the bounding box corresponding to each accurate label in the current video frame with each tracking trajectory in the trajectory set, until each video frame in the target video data has been processed by the multi-target tracking model.

[0027] As described above, the multi-object tracking model prioritizes accurate labels for matching, using them as a reference for determining annotation labels. Furthermore, each newly created tracking trajectory originates entirely from accurate labels, reducing the probability of false identification and preventing the uncontrolled growth of annotation labels, thus improving the annotation accuracy of the multi-object tracking model. Moreover, the multi-object tracking model retains every tracking trajectory generated during processing, enabling it to detect target objects even in scenarios where the target object has disappeared for a long time and then reappears.

[0028] In one embodiment of this application, it further includes:

[0029] The multi-target tracking model sets both the new tracking trajectory and the updated tracking trajectory to an active state, and sets the retained tracking trajectory to an inactive state.

[0030] As described above, the multi-target tracking model can reasonably set the state of each tracking trajectory based on the matching results of the current video frame, so that when processing subsequent video frames, the tracking trajectory can be reasonably matched with the required annotation box or detection box based on the state of the tracking trajectory.

[0031] In one embodiment of this application, each detection box has a corresponding confidence level, and the confidence level that reaches the confidence level threshold is considered high confidence level, while the confidence level that does not reach the confidence level threshold is considered low confidence level.

[0032] The step of matching the detection boxes in the first set of unmatched detection boxes with the tracking trajectories in the first set of unmatched trajectories by the multi-target tracking model to obtain the second set of successfully matched detection boxes, the second set of unmatched detection boxes, and the second set of unmatched trajectories includes:

[0033] The multi-target tracking model matches the high-confidence detection boxes in the first set of unmatched detection boxes with the active tracking trajectories in the first set of unmatched trajectories to obtain the third set of successfully matched detection boxes, the third set of unmatched detection boxes, and the third set of unmatched trajectories.

[0034] The multi-target tracking model matches the low-confidence detection boxes in the first set of unmatched detection boxes with the tracking trajectories in the third set of unmatched trajectories to obtain the fourth set of successfully matched detection boxes, the fourth set of unmatched detection boxes, and the fourth set of unmatched trajectories.

[0035] The multi-target tracking model matches the detection boxes in the third unmatched subset with the inactive tracking trajectories in the first unmatched subset to obtain a fifth successfully matched subset, a fifth unmatched subset, and a fifth unmatched subset. The third successfully matched subset, the fourth successfully matched subset, and the fifth successfully matched subset constitute the second successfully matched subset. The fourth unmatched subset and the fifth unmatched subset constitute the second unmatched subset. The fourth unmatched subset and the fifth unmatched subset constitute the second unmatched subset.

[0036] As mentioned above, for low-confidence detection boxes without accurate labels (such as occluded target objects), the multi-object tracking model can also perform tracking matching to match the tracking trajectory, thus enabling tracking matching of low-confidence detection boxes.

[0037] In one embodiment of this application, the step of matching each bounding box in the current video frame with each detection box in the current video frame by the multi-object tracking model to obtain a first set of successfully matched detection boxes, a first set of unmatched detection boxes, and a second set of unmatched bounding boxes includes:

[0038] The multi-target tracking model calculates the overlap between each bounding box and each detection box in the current video frame. If the overlap between the bounding box and the detection box reaches the overlap threshold, it is determined that the bounding box and the detection box are successfully matched. If the overlap between the bounding box and the detection box does not reach the overlap threshold, it is determined that the bounding box and the detection box are not successfully matched.

[0039] The multi-target tracking model obtains a first set of successfully matched detection boxes, a first set of unmatched detection boxes, and a second set of unmatched annotation boxes based on the successfully matched bounding boxes and detection boxes, as well as the unmatched bounding boxes and detection boxes.

[0040] As described above, deduplicating the bounding boxes and detection boxes can help multi-target tracking models avoid duplicate matching.

[0041] In one embodiment of this application, the target video data is video data obtained by shooting in a specific confined space, and the specific confined space contains multiple target objects.

[0042] As mentioned above, by leveraging the characteristics of a specific constrained space and assigning the highest priority to accurate labels, errors in multi-object tracking models can be effectively corrected in a timely manner, significantly improving the accuracy of multi-object tracking models.

[0043] Secondly, one embodiment of this application also provides a multi-target annotation device, comprising:

[0044] The first acquisition unit is used to acquire target video data and the initial accurate labels and bounding boxes of each target object that appears for the first time in the target video data in the corresponding video frame. Each target object that appears for the first time has a corresponding initial accurate label and bounding box.

[0045] The detection unit is used to detect each target object appearing in each video frame of the target video data, so as to obtain a detection box for marking the detection results. Each target object detected in each video frame has a corresponding detection box.

[0046] The first tracking unit is used to input the target video data, the initial accurate label, the annotation box, and the detection box into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching, and sets the highest priority for the accurate label during tracking and matching to obtain the first annotation label corresponding to each target object in each video frame. The accurate label includes the initial accurate label. The first annotation labels corresponding to the same target object tracked by the multi-target tracking model are the same and are all the initial accurate labels corresponding to the target object.

[0047] Thirdly, one embodiment of this application also provides a multi-target annotation device, including: one or more processors and a memory;

[0048] The memory is used to store one or more programs;

[0049] When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-target annotation method as described in the first aspect.

[0050] Fourthly, one embodiment of this application also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the multi-target annotation method as described in the first aspect.

[0051] The beneficial effects of the multi-target annotation device, equipment, and storage medium provided above can be referenced from the beneficial effects of the multi-target annotation method. Attached Figure Description

[0052] Figure 1 This is a schematic diagram of the structure of a multi-target annotation device provided in one embodiment of this application;

[0053] Figure 2 A flowchart illustrating a multi-target annotation method provided in one embodiment of this application;

[0054] Figure 3 A schematic diagram of an initial accurate label provided in one embodiment of this application;

[0055] Figure 4 A flowchart illustrating a multi-target annotation method provided in another embodiment of this application;

[0056] Figure 5 This is a schematic diagram illustrating the processing flow of a multi-target tracking model provided in one embodiment of this application;

[0057] Figure 6 A schematic diagram illustrating the manual review time provided in one embodiment of this application;

[0058] Figure 7 This is a schematic diagram of the structure of a multi-target annotation device provided in one embodiment of this application. Detailed Implementation

[0059] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and not for limiting the scope of the application. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present application are shown in the drawings, not the entire structure.

[0060] Multi-object tracking refers to the instance detection of multiple specific targets (such as multiple people or objects in video data) in video data and assigning them unique IDs, thereby obtaining the motion trajectory of each specific target. Multi-object tracking is the cornerstone of tasks such as trajectory prediction and temporal behavior analysis. Especially for multi-object tracking of the human body, it faces many challenges such as occlusion, motion blur, deformation, and interference from similar targets. Therefore, it remains a key and popular research direction in computer vision (CV).

[0061] Multi-object tracking annotation is fundamental to multi-object tracking research. Multi-object tracking annotation in dense scenes is a common application scenario. Dense scenes refer to video recordings where the shooting space is relatively limited, and the density of target objects within that space is high with low mobility. Examples include teaching and meeting scenes. In dense scenes, multi-object tracking annotation not only requires annotating the dense target bounding boxes but also considering the relationships between video frames, assigning each bounding box an ID. The bounding box can be understood as a rectangle used to identify (i.e., frame) a specific target, and the ID can be understood as the identifier of that specific target; the same specific target should have the same ID in different video frames. Taking a teaching scene as an example, multi-object tracking annotation can clearly identify the number of people in the classroom, as well as their trajectories and positions.

[0062] Multi-target tracking annotation can be performed manually, whereby humans annotate the bounding boxes and IDs of each specific target within each video frame, and then manually review the annotated bounding boxes and IDs. Since video data contains a large number of video frames, an automatic inter-frame interpolation annotation tool can be used to reduce the number of manually annotated video frames. This tool uses the manually annotated bounding boxes and IDs from preceding and following video frames to linearly interpolate specific targets in intermediate video frames. For example, after humans annotate the bounding boxes and IDs of each specific target in the first and fifth video frames respectively, the automatic interpolation tool can then annotate the bounding boxes and IDs of specific targets in the second to fourth video frames based on the aforementioned annotation results. While this saves on the number of manually annotated video frames, the annotation and review process still requires considerable time. For instance, in a classroom setting, manually annotating and reviewing 45 minutes of video data would take approximately 30 days.

[0063] In related technologies, considering the time-consuming nature of manual work in dense scenes, a classic semi-automatic annotation method was designed. This method first uses a detection model to detect each specific target within each video frame, obtaining bounding boxes to identify each target. Then, a multi-target tracking model tracks each target, assigning an ID to each target, with identical IDs assigned to the same target. Finally, the annotation results are manually reviewed. The detection model is a pre-trained neural network model that can detect multiple specific targets of the same category (e.g., the human body) in each video frame and draw corresponding bounding boxes. The multi-target tracking model performs target tracking, assigning the same ID to the same tracked target. Generally, the multi-target tracking model matches the bounding boxes in the current video frame with historical tracking trajectories (obtained from previous video frames) and updates the tracking trajectories based on the matching results. For example, if a bounding box in the current video frame matches a historical tracking trajectory, the bounding box is added to that historical trajectory, and the ID corresponding to that historical tracking trajectory is used as the bounding box's ID.

[0064] Currently, the multi-object tracking model used is ByteTrack, a pre-labeled tracker where the tracker refers to the tracked object and the pre-label refers to the assigned ID. ByteTrack is currently the best-performing model using the MOT20 multi-object tracking dataset. The implementation of ByteTrack involves matching the high-scoring bounding boxes (whose scores can also be understood as their confidence levels, generated during bounding box detection) in the current video frame with historical tracking trajectories. Based on the matching results, unmatched bounding boxes are stored in the set `Dremain`, and unmatched historical tracking trajectories are stored in the set `Tremian`. Next, the low-scoring bounding boxes in the current video frame are matched with historical tracking trajectories in the `Tremian` set. Based on the matching results, the set of unmatched historical tracking trajectories is stored in the set `Tre-remain`, and unmatched, low-scoring bounding boxes are deleted. The set of historical tracking trajectories in `Tre-remain` is placed into the set `Tcost`. Historical tracking trajectories in the set `Tcost` are deleted after a certain period (e.g., 30 frames) without a successful match. If a target bounding box in the Dream set scores high and survives for more than a few frames (e.g., 2 frames), a new historical tracking trajectory is created based on the target bounding box and a new ID is assigned.

[0065] Semi-automatic annotation methods have two time-consuming points: bounding box annotation (i.e., detection of specific targets) and ID annotation (i.e., tracking of target objects). While ByteTrack's accuracy is high, it is still limited. For example, in dense scenes, specific targets are often occluded. In a classroom setting, for instance, a teacher moving around might obscure students. In such cases, ByteTrack may misidentify students, such as misidentifying a previously present student as a newly appearing one, or misidentifying student A as student B. This can lead to arbitrary ID growth and inaccurate ID assignment. Therefore, after assigning IDs to specific targets, ByteTrack still requires manual review of the annotation results, and manually correcting erroneous IDs is the most time-consuming part of the annotation process. To shorten the manual review time, the annotated video frames can be reviewed at intervals, for example, manually reviewing video frames every 3 seconds. While this reduces the number of video frames reviewed, it still requires a significant amount of review time. For example, in a classroom setting, 45 minutes of video data requires manual review of 900 video frames. Because ByteTrack's accuracy is limited, the time required for manual review and ID correction is still approximately 9 days.

[0066] In summary, even with the best-performing ByteTrack for semi-automatic annotation in dense scenarios, manual review still takes a relatively long time.

[0067] Based on this, this application provides a multi-target annotation method. This method uses an improved multi-target tracking model based on ByteTrack. The improved multi-target tracking model can be understood as adding manually annotated IDs as input to ByteTrack and setting the highest matching priority for the manually annotated IDs. The accuracy of the improved multi-target tracking model is higher than that of ByteTrack, and the time required for manual review of the annotation results of the improved multi-target tracking model is shorter (because the amount of data that needs to be modified is reduced). Specifically, when applying the improved multi-target tracking model, each target object (i.e., a specific target) in the video data is first manually annotated. That is, the target object is manually added with a bounding box and ID for each target object that appears for the first time in the video data. Then, the manually annotated video data is input into the improved multi-target tracking model. That is, the video data received by the multi-target tracking model already contains the annotation results of each target object when it first appears. At this time, the multi-target tracking model can combine the results of manual annotation to track each target object. That is, the results of manual annotation are used to correct the annotation results of the model, and the problem of IDs growing arbitrarily can be avoided, thereby improving the accuracy of the multi-target tracking model in annotating target objects.

[0068] The multi-target annotation method provided in this application embodiment can be executed by a multi-target annotation device. This multi-target annotation device can be implemented through software and / or hardware, and can consist of two or more physical entities, or a single physical entity. Currently, multi-target annotation devices can be computers, servers, or other devices with data processing capabilities.

[0069] Figure 1 This is a schematic diagram of a multi-target annotation device provided in one embodiment of this application. (Reference) Figure 1 The multi-target annotation device includes a processor 11 and a memory 12. The processor 11 and the memory 12 can be connected via a bus or other means.

[0070] The number of processors 11 can be one or more. Figure 1 Taking a processor 11 as an example, the processor 11 may include processing units such as an application processor (AP), a graphics processing unit (GPU), and a central processing unit (CPU).

[0071] The memory 12, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the multi-target annotation device in this embodiment. The memory 12 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the multi-target annotation device. Furthermore, the memory 12 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 12 may further include remotely located memories 12 relative to the processor 11, which can be connected to the multi-target annotation device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0072] In addition, the multi-target labeling device may also include one or more components such as a display screen, a component for accessing a network (such as the Internet), a communication interface, a power supply, a speaker, a camera, and an input device, etc., but the embodiments are not limited thereto.

[0073] When a multi-object annotation device executes a multi-object annotation method, Figure 2 A flowchart of a multi-target annotation method provided in one embodiment of this application is shown below. Figure 2The multi-target annotation method includes steps 210-230:

[0074] Step 210: Obtain the target video data and the initial accurate labels and bounding boxes of each target object that appears for the first time in the target video data in the corresponding video frame. Each target object that appears for the first time has a corresponding initial accurate label and bounding box.

[0075] Target video data refers to video data that needs to be tracked and labeled for multiple targets. It consists of multiple video frames, and a video frame can be understood as a single image frame in the video data. In one embodiment, target video data is video data captured in a specific confined space, which contains multiple target objects. A specific confined space can be understood as a space with a fixed range, such as a classroom in a teaching scenario or a conference room in a meeting scenario. In addition, a fixed area designated outdoors can also be considered a specific confined space. Multiple target objects exist within the specific confined space. A target object refers to a specific target that needs to be tracked and labeled. Here, a person is used as an example to describe the target object. The movement of target objects within the specific confined space is low. For example, students in a classroom rarely move except when answering questions, and occasionally students enter and exit, while teachers move within the classroom. In a conference room, the participants do not move significantly, and the number of people entering and exiting the conference room during the meeting is also very small. Optionally, the density of target objects within the specific confined space can be high, that is, the target objects can be relatively dense. In this case, the target video data can also be considered as video data obtained in a dense scene. Optionally, the methods for capturing and acquiring the target video data are not currently limited.

[0076] In one embodiment, when acquiring target video data, initial accurate labels and bounding boxes are also acquired. The initial accurate labels and bounding boxes are stored in association with the target video data so that the initial accurate labels and bounding boxes can be read or displayed synchronously when the target video data is read or displayed.

[0077] Here, a label refers to the ID of a target object, which can consist of letters, numbers, and / or text. Each target object has a unique label, distinguishing different target objects. The encoding rules for labels are currently unrestricted. An accurate label is a label that has been manually verified as correct. An initial accurate label is a label manually set for the first appearance of a target object. It can be understood that manually setting a label signifies manual verification of its correctness. A bounding box is the bounding box of the target object corresponding to an accurate label. Each accurate label has a corresponding bounding box. Currently, the bounding box corresponding to the initial accurate label can be manually set, meaning that the bounding box and initial accurate label are manually added to the target object in the video frame.

[0078] Currently, each target object appearing for the first time in the target video data has a corresponding bounding box and an initial accurate label, meaning there is a one-to-one correspondence between the bounding box and the initial accurate label. When manually adding initial accurate labels and bounding boxes, the operator can first iterate through the target video data, adding a bounding box and setting a corresponding initial accurate label for each target object appearing for the first time. It's understood that once a target object appears repeatedly, the operator will no longer add labels or bounding boxes. For example, if a target object appears in the first video frame, the operator adds a bounding box and an initial accurate label, and then no longer checks whether the target object appears in subsequent video frames; that is, the operator will not add bounding boxes or initial accurate labels even if the target object appears in subsequent video frames. Optionally, when manually adding bounding boxes and initial accurate labels, the operator can add a rectangular bounding box as the bounding box for the target object in the video frame and add the corresponding initial accurate label around the bounding box.

[0079] Taking target video data captured in a classroom scene as an example, let's assume there's only one teacher in the target video data. The teacher's label is "Teacher," and the student's labels are defined according to their seat row and column numbers, starting from "1." During manual annotation, in the first video frame of the target video data, a bounding box and an initial accurate label are added to each target object. The bounding box is a rectangular area enclosing the human body. If the target object is a teacher, the corresponding initial accurate label is "Teacher." If the target object is a student sitting in row 1, column 1, the student's initial accurate label is "11." If the target object is a student sitting in row 1, column 2, the student's initial accurate label is "12." If the target object is a student sitting in row 2, column 1, the student's initial accurate label is "21." This method allows for manual addition of bounding boxes and initial accurate labels to each target object in the first video frame. Next, the system manually checks other video frames in the target video data for newly appearing target objects. If a new target object is found, a bounding box and an initial accurate label are added to it in the video frame where it first appears. This process continues until every target object in the target video data has a corresponding bounding box and an initial accurate label in its first appearance frame. It's understandable that if no new target object appears in a video frame, then no corresponding bounding box and initial accurate label will appear in that frame. For example, Figure 3 This is a schematic diagram of an initial accurate label provided in one embodiment of this application, with reference to... Figure 3 Each target object in the current video frame is a target object that appears for the first time. After manual setting, each target object has a corresponding label box 52 and an initial accurate label 53. The initial accurate labels are "11", "12", "13", "14", "21", "22", "23" and "24" respectively.

[0080] Optionally, in practical applications, manual annotation can be used not only to annotate target objects appearing for the first time, but also to annotate individual difficult target objects. Difficult target objects can be understood as those that are difficult to detect automatically, such as occluded target objects. In this case, the difficult objects also have initial accurate labels and bounding boxes after manual annotation. For example, if a target object has an initial accurate label and bounding box manually annotated in the video frame where it first appears, and this target object is a difficult object in a subsequent video frame, then manual annotation can be used to add another initial accurate label and bounding box for that target object in that video frame, so that the difficult object in that video frame can be accurately identified by the subsequent multi-object tracking model. In practical applications, manual annotation can also use other rules to select recurring target objects and add initial accurate labels and bounding boxes to them, so that the subsequent multi-object tracking model can refer to a larger number of accurate labels during tracking and matching, thereby improving accuracy.

[0081] It should be noted that the device used to manually add bounding boxes and initial accurate labels to the target video data can be the same device or a different device than the multi-target annotation device. When using different devices, after manually adding the bounding boxes and initial accurate labels, the bounding boxes and initial accurate labels are associated with and saved with the target video data for the multi-target annotation device to retrieve.

[0082] Step 220: Detect each target object appearing in each video frame of the target video data to obtain a detection box for marking the detection results. Each target object detected in each video frame has a corresponding detection box.

[0083] For example, after acquiring the target video data, each target object in each video frame of the target video data is detected, and a detection box is added to each detected target object in the video frame. The detection box is similar to the annotation box, both being target boxes added to the target object. The difference is that the annotation box is manually confirmed, while the detection box is automatically detected by a multi-target annotation device.

[0084] Optionally, multi-object annotation devices can achieve detection using a target object detection model. For example, when the target object is a person, the detection model can be a human detection model, used to detect each human body appearing in a video frame. The detection model can be implemented using a neural network model, which, after training, can be deployed in the multi-object annotation device for use. Currently, after inputting the target video data into the detection model, each detection box is obtained. The target video data input into the detection model may or may not contain initial accurate labels and bounding boxes, or it may contain initial accurate labels and bounding boxes where the initial labels and bounding boxes have no impact on the detection results.

[0085] In practical applications, the detection of target objects can also be achieved in other ways, and the embodiments do not limit this.

[0086] Understandably, when a target object with an added bounding box and an initial accurate label is detected, a corresponding detection box will also appear.

[0087] It should be noted that detection boxes and annotation boxes can be distinguished by multi-target annotation devices. That is, multi-target annotation devices can know whether the rectangle that encloses the target object is a detection box or an annotation box.

[0088] After the detection is completed, the detection frame is stored in association with the target video data so that the detection frame can be read or displayed synchronously when the target video data is read or displayed.

[0089] Step 230: Input the target video data, initial accurate labels, bounding boxes, and detection boxes into the multi-object tracking model. The multi-object tracking model performs target object tracking and matching, and sets the highest priority for accurate labels during tracking and matching to obtain the first label corresponding to each target object in each video frame. The accurate labels include the initial accurate labels. The first labels corresponding to the same target object tracked by the multi-object tracking model are the same and are all the initial accurate labels corresponding to the target object.

[0090] For example, a multi-object tracking model is used to achieve multi-object tracking, that is, to identify the same target object in different video frames of target video data and assign the same annotation label. Here, the annotation label refers to the label set by the multi-object tracking model for the target object in the target video data. The annotation label can also be understood as a pre-label, that is, the label predicted for the target object by the multi-object tracking model after tracking and matching the target object.

[0091] Multi-object tracking models can be implemented using deep learning networks. In one embodiment, an improved version of ByteTrack is used. This improved model prioritizes accurate labels, matching the bounding boxes corresponding to accurate labels with historical tracking trajectories first, then matching detection boxes with historical tracking trajectories that did not match any bounding boxes. This achieves target object tracking and matching, allowing for the addition of labels to targets without accurate labels. Furthermore, the creation of each historical tracking trajectory relies solely on accurate labels; the model cannot create new labels for targets, only accurate ones. Since initial accurate labels are already manually set for each target object, the model can use these labels to create corresponding historical tracking trajectories, resulting in more accurate labels for targets matched based on these historical trajectories. This also prevents arbitrary label growth and improves the overall accuracy of the multi-object tracking model.

[0092] For example, after obtaining the detection bounding boxes of each target object in the target video data, the target video data, initial accurate labels, annotation boxes, and detection bounding boxes are input together into the multi-object tracking model for processing. During processing, the accurate labels with the highest priority include the initial accurate labels, each accurate label has a corresponding annotation box, and the annotation labels added by the multi-object tracking model for the target objects in the video frame are recorded as the first annotation labels.

[0093] In one embodiment, the multi-object tracking model performs target object tracking and matching, and sets the highest priority for accurate labels during tracking and matching, including steps 231-239:

[0094] Step 231: The multi-target tracking model matches the bounding boxes corresponding to each accurate label in the current video frame of the target video data with each tracking trajectory in the trajectory set to obtain a set of successfully matched bounding boxes, a set of unmatched first bounding boxes, and a set of unmatched first trajectories. The trajectory set is a set of each tracking trajectory obtained during the processing of the multi-target tracking model. Each tracking trajectory corresponds to a target object and represents the position trajectory of the currently identified target object in the target video data.

[0095] For example, a tracking trajectory refers to the historical tracking trajectory obtained during the tracking and matching process of a multi-object tracking model, which can refer to the historical tracking trajectories used in the existing ByteTrack. A tracking trajectory can be understood as the trajectory of the target object's bounding box (currently it can be a detection box or a labeled box) in each video frame when the multi-object tracking model tracks and matches the same target object, which can reflect the position trajectory of the target object in each video frame. That is, the tracking trajectory consists of detection boxes and / or labeled boxes, and the tracking trajectory is updated according to the matching results of the multi-object tracking model.

[0096] The trajectory set is the collection of all tracking trajectories currently processed by the multi-target tracking model.

[0097] For example, the multi-target tracking model processes each video frame in the target video data sequentially from front to back, and the processing flow for each video frame is the same. In this embodiment, the processing flow of the multi-target tracking model is described using the processing of one video frame as an example.

[0098] In one embodiment, the video frame currently being processed by the multi-object tracking model is designated as the current video frame. The current video frame can be any frame of the multi-object tracking model. It is understood that when the multi-object tracking model reads the current video frame, it can simultaneously read the corresponding accurate label, bounding box (if the current video frame has a corresponding accurate label and bounding box), and detection box.

[0099] When processing the current video frame, the multi-object tracking model first finds the accurate label in the current video frame, and then determines the bounding box corresponding to the accurate label. Currently, the accurate label found is the initial accurate label.

[0100] Next, the bounding boxes corresponding to the found accurate labels are matched with each tracking trajectory in the trajectory set. Currently, the latest target boxes in each tracking trajectory are the detection boxes or labeled boxes added to the tracking trajectory after the most recent update.

[0101] When there are multiple accurate labels, each corresponding bounding box needs to be matched with the tracking trajectory. Optionally, the matching method can refer to the existing ByteTrack implementation of target box matching with historical tracking trajectories, for example, using a preset matching algorithm to match the bounding boxes and tracking trajectories. Specifically, first, predict the possible position and size of the target box in the current video frame for each tracking trajectory in the trajectory set. Then, calculate the IoU (Intersection over Union) value between each pair of the predicted target box and the bounding boxes corresponding to each accurate label to obtain the IoU loss matrix. According to the IoU loss matrix, match all tracking trajectories with the bounding boxes corresponding to each accurate label and obtain the matching results. Alternatively, other methods can be used to implement the matching. For example, if the bounding label corresponding to a certain tracking trajectory is the same as the accurate label corresponding to a certain bounding box, it can be considered that the tracking trajectory and the bounding box are successfully matched. Match all tracking trajectories with the bounding boxes corresponding to each accurate label in this way and obtain the matching results.

[0102] Currently, the matching results are: successfully matched tracking trajectories and bounding boxes (i.e., each bounding box has a corresponding matching tracking trajectory), unmatched tracking trajectories (i.e., no matching bounding boxes), and unmatched bounding boxes (i.e., no matching tracking trajectories). The multi-target tracking model adds all successfully matched bounding boxes to the set of successfully matched bounding boxes, currently denoted as M1; adds the unmatched bounding boxes to the first set of unmatched bounding boxes, currently denoted as UD1; and adds the unmatched tracking trajectories to the first set of unmatched trajectories, currently denoted as UT1.

[0103] In this context, M1 can be understood as the set of bounding boxes that successfully matched the tracking trajectory. A successful match between a bounding box and the tracking trajectory indicates that the target object marked by the bounding box and the target object tracked by the tracking trajectory are predicted to be the same target object, meaning that the target object corresponding to the tracking trajectory has been tracked in the current video frame. UD1 can be understood as the set of bounding boxes that did not successfully match the tracking trajectory. A bounding box that did not successfully match the tracking trajectory can be understood as the target object marked by the bounding box being different from all the currently identified and tracked target objects. UT1 can be understood as the set of tracking trajectories that did not successfully match the bounding boxes. A tracking trajectory that did not successfully match the bounding boxes can be understood as the target object corresponding to the tracking trajectory not being tracked in the current video frame.

[0104] Understandably, when the multi-target tracking model processes a video frame, M1, UD1, and UT1 are all empty. After all tracking trajectories and bounding boxes are matched, M1, UD1, and UT1 are filled with the corresponding data.

[0105] It's important to note that since only some video frames currently have accurate labels (meaning some frames may lack accurate labels), the multi-object tracking model may encounter situations where it cannot find an accurate label during the label-finding process; in such cases, no bounding box exists. The multi-object tracking model can consider this as a case of no successfully matched tracking trajectory and bounding box, with M1 and UD1 both empty, and UT1 containing all tracking trajectories. It's understandable that the only labels input to the multi-object tracking model are accurate labels; therefore, if the multi-object tracking model finds a label, it can consider it to have found an accurate label.

[0106] Currently, the matching of the tracking trajectory and the bounding boxes corresponding to the accurate labels can be considered as the first matching performed by the multi-object tracking model when processing the current video frame, meaning that the matching of accurate labels has the highest priority. After the first matching is completed, step 232 is executed.

[0107] Step 232: The multi-object tracking model matches each bounding box in the current video frame with each detection box in the current video frame to obtain the first detection box matching set, the first detection box not matching set, and the second bounding box not matching set.

[0108] Since step 220 detects each target object in every video frame, each detected target object has a corresponding bounding box. Therefore, for a target object appearing for the first time, it has both a detected bounding box and a manually labeled bounding box, and their positions in the corresponding video frame will highly overlap. It's understandable that after the multi-target tracking model matches the labeled bounding boxes with the tracking trajectory, it still needs to match unmatched tracking trajectories (i.e., the tracking trajectory in UT1) with the detected bounding boxes. If the detected bounding box and the labeled bounding box correspond to the same target object, then duplicate matching will occur. To avoid duplicate matching, the multi-target tracking model removes highly overlapping detected bounding boxes.

[0109] In one embodiment, when the multi-object tracking model removes detection boxes with high overlap, it needs to first determine the detection boxes and annotation boxes with high positional overlap. It is understood that detection boxes and annotation boxes serve the same purpose; for the same target object, the positions of the detection boxes and annotation boxes in the video frame should highly overlap. At this point, the multi-object tracking model can first match each detection box and each annotation box appearing in the current video frame (i.e., calculate the positional overlap) to determine the detection boxes and annotation boxes with high overlap. Based on this, step 232 specifically includes steps 2321-2322:

[0110] Step 2321: The multi-object tracking model calculates the overlap between each bounding box and each detection box in the current video frame. If the overlap between the bounding box and the detection box reaches the overlap threshold, it is determined that the bounding box and the detection box are successfully matched. If the overlap between the bounding box and the detection box does not reach the overlap threshold, it is determined that the bounding box and the detection box are not successfully matched.

[0111] For example, the multi-object tracking model performs pairwise calculations on each labeled bounding box and each detected bounding box in the current video frame to calculate the overlap between the regions containing the labeled and detected bounding boxes. The method of calculating the overlap is not currently limited; for example, it can determine the area of ​​the overlapping region between the detected and labeled bounding boxes, and then determine the ratio of this area to the sum of the areas of the detected and labeled bounding boxes. This ratio can then be used as the overlap. It is understood that the positions (i.e., pixel coordinates) and sizes of the detected and labeled bounding boxes in the video frame are determined during annotation; this step can directly use the positions and sizes to determine the areas of the detected and labeled bounding boxes.

[0112] After calculating the overlap ratio, it is determined whether the overlap ratio reaches the overlap ratio threshold. The overlap ratio threshold can be set according to the actual situation. When the overlap ratio reaches the overlap ratio threshold, the corresponding bounding box and detection box are considered to be highly overlapped, meaning they correspond to the same target object, and the bounding box and detection box are considered to have successfully matched. When the overlap ratio does not reach the overlap ratio threshold, the corresponding bounding box and detection box are considered to have low overlap, meaning they do not correspond to the same target object, and the bounding box and detection box are considered to have failed to match.

[0113] Step 2322: The multi-object tracking model obtains the first set of successfully matched detection boxes, the first set of unmatched detection boxes, and the second set of unmatched annotation boxes based on the successfully matched bounding boxes and detection boxes, as well as the unmatched bounding boxes and detection boxes.

[0114] For example, after matching bounding boxes and detection boxes, the multi-object tracking model obtains successfully matched detection boxes (i.e., each detection box has a highly overlapping bounding box), unmatched detection boxes (i.e., none of them have highly overlapping bounding boxes), and unmatched bounding boxes (i.e., none of them have highly overlapping bounding boxes). The multi-object tracking model adds all successfully matched detection boxes to the first set of successfully matched detection boxes, currently denoted as M2; adds the unmatched detection boxes to the first set of unmatched detection boxes, currently denoted as UD2; and adds the unmatched bounding boxes to the second set of unmatched bounding boxes, currently denoted as UT2.

[0115] Here, M2 can be understood as the set of detection boxes that successfully matched the labeled boxes, UD2 as the set of detection boxes that did not successfully match the labeled boxes, and UT2 as the set of labeled boxes that did not successfully match the detection boxes. The multi-object tracking model then uses UD2 to match UT1 again.

[0116] Understandably, when the multi-object tracking model processes a video frame, M2, UD2, and UT2 are all empty. After all detection boxes and annotation boxes are matched, M2, UD2, and UT2 are filled with the corresponding data.

[0117] It should be noted that since only some video frames currently have accurate labels, meaning there's a possibility that a current video frame might not have an accurate label, the multi-object tracking model may encounter situations where it cannot find an accurate label during the label-finding process; that is, no bounding boxes will be found. In such cases, the multi-object tracking model can consider that there are no successfully matched detection boxes and bounding boxes, M2 and UT2 are both empty, and UD2 contains all detection boxes from the current video frame.

[0118] Currently, the matching of the labeled bounding boxes and the detection boxes can be considered as the second matching performed by the multi-object tracking model when processing the current video frame. After the second matching is completed, step 233 is executed. In practical applications, step 232 can also be executed before step 231, or steps 231 and 232 can be executed simultaneously.

[0119] Step 233: The multi-target tracking model matches the detection boxes in the first set of unmatched detection boxes with the tracking trajectories in the first set of unmatched trajectories to obtain the second set of successfully matched detection boxes, the second set of unmatched detection boxes, and the second set of unmatched trajectories.

[0120] For example, the multi-object tracking model matches the detection boxes in UD2 with the tracking trajectories in UT1 to obtain successfully matched detection boxes and tracking trajectories, unmatched detection boxes, and unmatched tracking trajectories. Then, the multi-object tracking model adds all successfully matched detection boxes to the second set of successfully matched detection boxes, adds the unmatched detection boxes to the second set of unmatched detection boxes, and adds the unmatched tracking trajectories to the second set of unmatched tracking trajectories. The second set of successfully matched detection boxes can be understood as the set of detection boxes that successfully matched the tracking trajectory. A successful match between a detection box and a tracking trajectory indicates that the target object identified by the detection box and the target object tracked by the tracking trajectory are predicted to be the same target object, meaning that the target object corresponding to the tracking trajectory has been tracked in the current video frame. The second set of unmatched detection boxes can be understood as the set of detection boxes that did not successfully match the tracking trajectory. A detection box that did not successfully match the tracking trajectory can be understood as one whose target object is different from any of the currently identified and tracked target objects. The second set of unmatched trajectories can be understood as the set of tracking trajectories that failed to match the detection box. Tracking trajectories that failed to match the detection box can be understood as the target object corresponding to the tracking trajectory not being tracked in the current video frame.

[0121] An optional approach is to first predict the possible location and size of the bounding box in the current video frame for each tracking trajectory in UT1. If the prediction was already achieved when matching the bounding boxes and tracking trajectories, this step can directly use the prediction results. Then, the IoU (Intersection over Union) values ​​between each pair of predicted bounding boxes (the bounding boxes corresponding to each tracking trajectory in UT1) and each detection box in UD2 are calculated to obtain the IoU loss matrix. Based on the IoU loss matrix, each tracking trajectory in UT1 and each detection box in UD2 are matched to obtain the matching results.

[0122] Another option is to perform matching according to the three matching processes of the target bounding box and the historical tracking trajectory in ByteTrack. The difference is that the tracking trajectory used by the current multi-target tracking model is the tracking trajectory remaining after matching with the labeled bounding box, and the detection box used is the detection box remaining after matching with the labeled bounding box.

[0123] In one embodiment, taking the three matching processes in ByteTrack as an example, the matching process between each tracking trajectory in UT1 and each detection box in UD2 is described. In this case, step 233 may specifically include steps 2331 to 2333:

[0124] Step 2331: The multi-target tracking model matches the high-confidence detection boxes in the first unmatched detection box set with the active tracking trajectories in the first unmatched trajectory set to obtain the third successfully matched subset of detection boxes, the third unmatched subset of detection boxes, and the third unmatched subset of trajectory.

[0125] For example, in step 220, when the multi-target annotation device performs target object detection on the target video data, each detection box has a corresponding confidence level. The confidence level represents the reliability of the detection box, i.e., the reliability of the detected target object. A higher confidence level indicates a higher accuracy in detecting the target object. Taking a person as an example, a higher confidence level indicates a greater reliability that the object in the detection box is a person. For instance, if a person is obscured by another person, the confidence level corresponding to that person's detection box will be lower than the confidence level when the person is not obscured.

[0126] Currently, multi-object tracking models categorize detection boxes into high-confidence and low-confidence boxes based on their confidence levels. This distinction is made using a pre-set confidence threshold, which can be set according to specific circumstances, such as 0.5. When the confidence level reaches the threshold, it is considered high confidence, and the corresponding detection box is considered a high-confidence box. Conversely, when the confidence level does not reach the threshold, it is considered low confidence, and the corresponding detection box is considered a low-confidence box. In other words, each detection box has a corresponding confidence level, and confidence levels that reach the threshold are considered high confidence, while those that do not are considered low confidence.

[0127] For example, all tracking trajectories obtained by the multi-object tracking model can be divided into active and inactive tracking trajectories. Active tracking trajectories are those that, when the multi-object tracking model processed the previous video frame, successfully matched the corresponding bounding box or detection box and were updated based on the matched box. Inactive tracking trajectories are those that, when the multi-object tracking model processed the previous video frame, did not successfully match the corresponding bounding box or detection box; that is, inactive tracking trajectories were not updated after the previous video frame was processed. Both active and inactive tracking trajectories are stored in a trajectory set for use by the multi-object tracking model.

[0128] It should be noted that when the multi-target tracking model matches the bounding boxes and tracking trajectories, it uses all tracking trajectories, including those in active and inactive states.

[0129] In one embodiment, when the multi-target tracking model matches detection boxes and tracking trajectories, it first matches high-confidence detection boxes in UD2 with active tracking trajectories in UT1. That is, it first matches tracking trajectories that were successfully matched in the previous frame but not yet matched in the current frame with high-confidence detection boxes belonging to the target object in the current frame. During matching, the possible position and size of the target box in the current video frame for each active tracking trajectory in UT1 are predicted. If the prediction step has already been implemented during the matching of the target boxes and tracking trajectories, the prediction results can be used directly in this step. Then, the IoU (Intersection over Union) values ​​between each pair of predicted target boxes (target boxes corresponding to active tracking trajectories in UT1) and high-confidence detection boxes in UD2 are calculated to obtain the IoU loss matrix. Based on the IoU loss matrix, each active tracking trajectory in UT1 is matched with each high-confidence detection box in UD2, and the matching results are obtained.

[0130] Currently, the matching results are: successfully matched tracking trajectories and detection boxes (i.e., each successfully matched detection box has a corresponding matching tracking trajectory), unmatched tracking trajectories (i.e., no matching detection box), and unmatched detection boxes (i.e., no matching tracking trajectory). Then, the multi-object tracking model adds the successfully matched detection boxes to the third subset of successfully matched detection boxes, currently denoted as M3; adds the unmatched detection boxes to the third subset of unmatched detection boxes, currently denoted as UD3; and adds the unmatched tracking trajectories to the third subset of unmatched trajectories, currently denoted as UT3.

[0131] Here, M3 can be understood as the set of detection boxes that successfully matched the active tracking trajectory and belong to the high-confidence category; UD3 can be understood as the set of detection boxes that did not successfully match the active tracking trajectory and belong to the high-confidence category; and UT3 can be understood as the set of tracking trajectories that were active but did not successfully match the high-confidence detection boxes.

[0132] Understandably, when the multi-target tracking model processes a video frame, M3, UD3, and UT3 are all empty. After the high-confidence detection box in UD2 is matched with the active tracking trajectory in UT1, M3, UD3, and UT3 are filled with the corresponding data.

[0133] Currently, the matching of the high-confidence detection boxes in UD2 and the active tracking trajectories in UT1 can be considered as the third matching performed by the multi-object tracking model when processing the current video frame. After the third matching is completed, step 2332 is executed.

[0134] Step 2332: The multi-target tracking model matches the low-confidence detection boxes in the first unmatched detection box set with the tracking trajectories in the third unmatched trajectory set to obtain the fourth unmatched detection box subset, the fourth unmatched detection box subset, and the fourth unmatched trajectory subset.

[0135] For example, after obtaining M3, UD3, and UT3 in the multi-object tracking model, the low-confidence detection boxes in UD2 are matched with the tracking trajectories in UT3 (i.e., those in an active state but not matched with high-confidence detection boxes). In other words, tracking trajectories that were successfully matched in the previous frame but not matched with either the labeled box or the high-confidence detection box in the current frame are matched with low-confidence detection boxes belonging to the target object in the current frame. During matching, the possible position and size of the target box in the current video frame for each tracking trajectory in UT3 are first predicted. If the prediction step has already been implemented in the previous matching process, the prediction results can be used directly in this step. Then, the IoU (Intersection over Union) values ​​between each pair of predicted target boxes (target boxes corresponding to the tracking trajectories in UT3) and low-confidence detection boxes in UD2 are calculated to obtain the IoU loss matrix. Based on the IoU loss matrix, each tracking trajectory in UT3 is matched with each low-confidence detection box in UD2, and the matching results are obtained.

[0136] Currently, the matching results are: successfully matched tracking trajectories and detection boxes (i.e., each successfully matched detection box has a corresponding matching tracking trajectory), unmatched tracking trajectories (i.e., no matching detection box), and unmatched detection boxes (i.e., no matching tracking trajectory). Next, the multi-object tracking model adds the successfully matched detection boxes to the fourth subset of successfully matched detection boxes, currently denoted as M4; adds the unmatched detection boxes to the fourth subset of unmatched detection boxes, currently denoted as UD4; and adds the unmatched tracking trajectories to the fourth subset of unmatched trajectories, currently denoted as UT4.

[0137] Here, M4 can be understood as the set of detection boxes that successfully matched the active tracking trajectory and belong to the low-confidence category; UD4 can be understood as the set of detection boxes that did not successfully match the active tracking trajectory and belong to the low-confidence category; and UT4 can be understood as the set of tracking trajectories that were active but did not successfully match the low-confidence detection boxes.

[0138] Understandably, when the multi-target tracking model processes a video frame, M4, UD4, and UT4 are all empty. After the low-confidence detection box in UD2 is matched with the tracking trajectory in UT3, M4, UD4, and UT4 are filled with the corresponding data.

[0139] It should be noted that there may be cases where there are no low-confidence detection boxes in UT2. In this case, each tracking trajectory in UT3 can be considered as having failed to match a detection box. In this situation, the resulting M4 and UD4 will be empty.

[0140] Currently, the matching of the low-confidence detection boxes in UD2 and the tracking trajectories in UT3 can be considered as the fourth matching performed by the multi-object tracking model when processing the current video frame. After the fourth matching is completed, step 2333 is executed.

[0141] Step 2333: The multi-target tracking model matches the detection boxes in the third unmatched subset with the inactive tracking trajectories in the first unmatched subset to obtain the fifth successfully matched subset, the fifth unmatched subset, and the fifth unmatched subset. The third successfully matched subset, the fourth successfully matched subset, and the fifth successfully matched subset form the second successfully matched subset. The fourth unmatched subset and the fifth unmatched subset form the second unmatched subset. The fourth unmatched subset and the fifth unmatched subset form the second unmatched subset.

[0142] For example, after obtaining M4, UD4, and UT4, the multi-target tracking model uses the detection boxes in UD3 to match the inactive tracking trajectories in UT1. That is, it matches tracking trajectories that failed to match in the previous frame and have no matching bounding boxes in the current frame with high-confidence detection boxes that have no matching tracking trajectory in the current frame. During matching, the possible position and size of the target box in the current video frame for each inactive tracking trajectory in UT1 are first predicted. If the prediction step has already been implemented when matching the bounding boxes and tracking trajectories, then the prediction results can be used directly in this step. Next, the IoU (Intersection over Union) values ​​between each pair of predicted target boxes (target boxes corresponding to inactive tracking trajectories in UT1) and detection boxes in UD3 (all high-confidence detection boxes) are calculated to obtain the IoU loss matrix. Based on the IoU loss matrix, each inactive tracking trajectory in UT1 is matched with each detection box in UD3, and the matching results are obtained.

[0143] Currently, the matching results are: successfully matched tracking trajectories and detection boxes (i.e., each successfully matched detection box has a corresponding matching tracking trajectory), unmatched tracking trajectories (i.e., no matching detection box), and unmatched detection boxes (i.e., no matching tracking trajectory). Next, the multi-object tracking model adds the successfully matched detection boxes to the fifth successfully matched detection box subset, currently denoted as M5; adds the unmatched detection boxes to the fifth unmatched detection box subset, currently denoted as UD5; and adds the unmatched tracking trajectories to the fifth unmatched trajectory subset, currently denoted as UT5.

[0144] Here, M5 can be understood as the set of high-confidence detection boxes that successfully matched the tracking trajectory in the inactive state; UD5 can be understood as the set of high-confidence detection boxes that did not successfully match the tracking trajectory in the inactive state; and UT5 can be understood as the set of tracking trajectories in the inactive state that did not successfully match high-confidence detection boxes.

[0145] Understandably, when the multi-target tracking model processes a video frame, M5, UD5, and UT5 are all empty. After the detection box in UD3 is matched with the tracking trajectory in UT1 which is inactive, M5, UD5, and UT5 are filled with the corresponding data.

[0146] It should be noted that there may be cases where there are no inactive tracking trajectories in UT1. In this case, each detection box in UD3 can be considered as having failed to match an inactive tracking trajectory. In this situation, the resulting M5 and UT5 will be empty.

[0147] It should be noted that there are cases where UD2 is empty. For example, in the first video frame, each target object is manually labeled, meaning each target object has a corresponding bounding box. In this case, the detection boxes in the first video frame may have a high degree of overlap with their corresponding bounding boxes. Therefore, each detection box is added to M2, resulting in UD2 being empty. In this case, M3, M4, and M5 obtained in the previous steps are all empty, as are UD3, UD4, and UD5.

[0148] After obtaining M5, UD5, and UT5, M3, M4, and M5 form the second set of successfully matched detection boxes, in which each detection box in the set is successfully matched with the tracking trajectory. UD4 and UD5 form the second set of unmatched detection boxes, in which each detection box in the set is not matched with the tracking trajectory. UT4 and UT5 form the second set of unmatched tracking trajectories, in which each tracking trajectory is not matched with a detection box or a label box.

[0149] Currently, the matching of the detection box in UD3 and the inactive tracking trajectory in UT1 can be considered as the fifth matching performed by the multi-object tracking model when processing the current video frame. After the fifth matching is completed, step 234 is executed.

[0150] As described above, there are cases where the set used for matching is empty. In this case, the multi-object tracking model can continue matching using the empty set, but there will be no successful matching results. That is, for each video frame in the target video data, the multi-object tracking model performs the aforementioned five matching operations.

[0151] Step 234: The multi-target tracking model creates new tracking trajectories based on each of the unmatched bounding boxes in the first set of bounding boxes, and the new tracking trajectories use the accurate labels corresponding to the bounding boxes as the label labels.

[0152] For example, each bounding box in UD1 is a manually confirmed accurate target box, meaning the target object within the bounding box actually appears in the current video frame. However, the bounding box does not match a tracking trajectory, indicating that the target object may not have appeared in previous video frames. Therefore, it is highly likely that the target object is a newly appearing target object in the current video frame. Thus, in this embodiment, the multi-target tracking model creates a new tracking trajectory for each bounding box in UD1, with one bounding box corresponding to one new tracking trajectory. The position of the new tracking trajectory is the position of the bounding box in the current video frame, and the annotation label used for the new tracking trajectory is the accurate label corresponding to the bounding box (currently the initial accurate label). Furthermore, the multi-target tracking model sets the new tracking trajectory to an active state and adds it to the trajectory set.

[0153] It is understandable that this step can also be performed after step 231. That is, after obtaining UD1 in step 231, the multi-target tracking model can create a new tracking trajectory based on the labeled box in UD1 and set it to the active state.

[0154] It should be noted that if the current video frame is the first video frame in the target video data, the multi-target tracking model has not yet created a tracking trajectory, that is, the trajectory set in step 231 is empty. In this case, when matching according to step 231, each bounding box can be considered as having no matching tracking trajectory. That is, after being added to UD1, the multi-target tracking model can create a corresponding tracking trajectory for each bounding box and set it to the active state for subsequent use.

[0155] It is understandable that each tracking trajectory has a corresponding label, which is set synchronously when the tracking trajectory is created. The label is generally an accurate label because the tracking trajectory can only be created based on the label box and use the label corresponding to the label box. Furthermore, the label corresponding to the label box is an accurate label. Therefore, the label of the tracking trajectory can only be an accurate label.

[0156] Step 235: The multi-target tracking model updates the corresponding matched tracking trajectory based on the set of successfully matched bounding boxes and the set of successfully matched second detection boxes, so that the updated tracking trajectory shares the same annotation label.

[0157] For example, the bounding boxes in M1 and the detection boxes in M3, M4, and M5 all successfully matched the tracking trajectory. Therefore, the multi-object tracking model updates the matched tracking trajectory using the bounding boxes in M1, that is, updates the bounding boxes to the latest target boxes in the corresponding tracking trajectory, and makes the bounding boxes retain the labeling used by the tracking trajectory; that is, the label corresponding to the newly added bounding box is the label used by the tracking trajectory. Generally, the exact label corresponding to the bounding box is the same as the label used by the matched tracking trajectory. The multi-object tracking model updates the matched tracking trajectory using the detection boxes in M3, M4, and M5, that is, updates the detection boxes to the latest target boxes in the corresponding tracking trajectory, and makes the detection boxes retain the labeling used by the tracking trajectory; that is, the label corresponding to the newly added detection box is the label used by the tracking trajectory.

[0158] The updated tracking trajectory is set to active so that it is clear that the updated tracking trajectory was successfully matched during the processing of the next video frame.

[0159] It is understandable that updating the corresponding matched tracking trajectory based on M1 in this step can be performed after step 231. That is, after obtaining M1 in step 231, the multi-target tracking model can update the corresponding tracking trajectory based on the bounding boxes in M1 and set it to the active state. Similarly, updating the corresponding matched tracking trajectories based on M3, M4, and M5 in this step can be performed after steps 2331, 2332, and 2333, respectively. That is, after obtaining M3 in step 2331, the multi-target tracking model can update the corresponding tracking trajectory based on the detection boxes in M3 and set it to the active state; after obtaining M4 in step 2332, the multi-target tracking model can update the corresponding tracking trajectory based on the detection boxes in M4 and set it to the active state; and after obtaining M5 in step 2333, the multi-target tracking model can update the corresponding tracking trajectory based on the detection boxes in M5 and set it to the active state.

[0160] Step 236: The multi-target tracking model retains each tracking trajectory in the set of unmatched second trajectories.

[0161] For example, if the tracking trajectories in UT4 and UT5 fail to match the corresponding bounding boxes or detection boxes, then, unlike the existing ByteTrack model which deletes the inactive tracking trajectories that fail to match, the multi-target tracking model retains the tracking trajectories in UT4 and UT5 and sets the retained tracking trajectories to an inactive state, so that in the processing of the next video frame, it is clear that the retained tracking trajectories have not been matched successfully.

[0162] Understandably, in a specific confined space, the flow of target objects is relatively small, and the number of target objects is limited. Therefore, target objects may disappear for a long time and then reappear. For example, students in a classroom may leave the classroom for a period of time and then return, or attendees in a meeting room may leave the meeting room and then return. Therefore, for a specific confined space, the multi-target tracking model always retains every tracking trajectory generated during the processing (regardless of whether it matches a bounding box or a detection box).

[0163] It is understandable that in this step, retaining UT4 can be performed after step 2332. That is, after obtaining UT4 in step 2332, the multi-target tracking model can retain the tracking trajectory in UT4 and set it to an inactive state. In this step, retaining UT5 ​​can be performed after step 2333. That is, after obtaining UT5 ​​in step 2333, the multi-target tracking model can retain the tracking trajectory in UT5 and set it to an inactive state.

[0164] Step 237: The multi-target tracking model obtains a set of trajectories for use in the next video frame based on the new tracking trajectory, the updated tracking trajectory, and the retained tracking trajectory.

[0165] For example, the multi-target tracking model updates all currently obtained tracking trajectories (including new tracking trajectories obtained based on UD1, tracking trajectories updated based on M1, M3, M4 and M5, and tracking trajectories retained based on UT4 and UT5) into the trajectory set for use in the next video frame.

[0166] Step 238: The multi-object tracking model ignores the set of first detection boxes that are successfully matched, the set of second annotation boxes that are not successfully matched, and the set of second detection boxes that are not successfully matched.

[0167] For example, the multi-object tracking model ignores M2, UT2, UD4, and UD5. Ignoring M2 can be considered as removing detection boxes with high overlap with the bounding boxes. Ignoring UT2 can be considered as ignoring bounding boxes with low overlap with the detection boxes. It's understood that each bounding box is matched with the tracking trajectory, therefore, no additional matching processing is needed for the bounding boxes in UT2. Ignoring UD4 and UD5 can be considered as ignoring detection boxes that failed to match the tracking trajectory. Failure to match the tracking trajectory indicates that the corresponding detection box may be a false detection or not being tracked; therefore, it is ignored.

[0168] Step 239: The multi-target tracking model updates the next video frame in the target video data to the current video frame, and returns to perform the operation of matching the bounding box corresponding to each accurate label in the current video frame with each tracking trajectory in the trajectory set, until each video frame in the target video data has been processed by the multi-target tracking model.

[0169] For example, after steps 231-238 are completed, it can be considered that the processing of the current video frame is finished. Then, the multi-object tracking model determines whether the current video frame is the last frame of the target video data. If not, the multi-object tracking model obtains the next video frame after the current video frame and updates it to the current video frame, then returns to step 231 to continue processing the current video frame. If it is the last frame, it is determined that the processing of the target video data is finished. At this point, the annotation label corresponding to each target box (which may be a bounding box or a detection box) in each tracking trajectory can be considered as the first annotation label obtained based on the multi-object tracking model. It can be understood that since the annotation labels of the tracking trajectory are all accurate labels, and the currently used accurate label is the initial accurate label, the first annotation label is equal to the initial accurate label.

[0170] Afterwards, the target video data with the first annotation label can be manually reviewed. This is understandable because, since the first annotation label is the initial accurate label, the multi-object tracking model does not have the right to create new labels. Therefore, the accuracy of the first annotation label obtained by the multi-object tracking model is relatively high, reducing the number of first annotation labels that need to be modified during manual review and shortening the manual review time.

[0171] The above-described method involves acquiring target video data with initial accurate labels and bounding boxes, detecting each target object within the video data to obtain detection boxes that identify the detection results, and then inputting the target video data, initial accurate labels, bounding boxes, and detection boxes into a multi-target tracking model. The multi-target tracking model tracks and matches the target objects, assigning the highest matching priority to the accurate labels during the tracking process. The model then obtains the first label for each target object, with the same first label corresponding to the same target object, and both being the initial accurate labels for that target object. This technical solution addresses the problem of limited labeling accuracy in multi-target tracking models when tracking and labeling multiple specific targets in video data in related technologies. Multi-object tracking models use accurate input labels as a reference for tracking and matching, ensuring that the matched labels (currently the first label) are accurate. This means the first label for the same target object is the accurate label it was marked with when it first appeared, preventing the uncontrolled growth of IDs (labels). Furthermore, each newly appearing target object is not misidentified; the multi-object tracking model does not need to identify new targets or create new labels (because truly new targets already have corresponding initial accurate labels). This separation of detection and matching tracking (meaning the multi-object tracking model does not rely on bounding boxes to identify and track new targets) improves its labeling accuracy. Multi-object tracking models prioritize accurate labels for matching, using them as a reference for label determination. Each newly created tracking trajectory originates entirely from accurate labels, reducing the probability of misidentification and preventing the uncontrolled growth of labels, thus improving the accuracy. Moreover, the multi-object tracking model retains every tracking trajectory generated during processing, enabling it to detect targets even after they have disappeared for a long time and then reappear. Furthermore, the multi-object tracking model can rationally set the state of each tracking trajectory based on the matching results of the current video frame. This facilitates the matching of tracking trajectories with the required bounding boxes or detection boxes when processing subsequent video frames, taking into account the state of the tracking trajectories. Moreover, the multi-object tracking model can also perform tracking matching for low-confidence detection boxes without accurate labels (such as occluded objects), thus achieving tracking matching for low-confidence detection boxes. Deduplicating the bounding boxes and detection boxes helps the multi-object tracking model avoid redundant matching. The aforementioned method, especially in specific constrained spaces, can leverage the characteristics of such spaces to assign the highest priority to accurate labels, effectively using accurate labels to correct errors in the multi-object tracking model in a timely manner, significantly improving the accuracy of the multi-object tracking model.

[0172] Figure 4A flowchart of a multi-target annotation method provided for another embodiment of this application. Figure 4 The multi-target annotation method shown is in Figure 2 Based on the multi-target annotation method shown, manual review and correction of the first annotation label in the interval video frame is added, and the target video data is reprocessed based on the manual review and correction results using the multi-target tracking model to further improve the accuracy of the annotation labels output by the multi-target tracking model.

[0173] refer to Figure 4 The multi-target annotation method includes steps 310-370:

[0174] Step 310: Obtain the target video data and the initial accurate labels and bounding boxes of each target object that appears for the first time in the target video data in the corresponding video frame. Each target object that appears for the first time has a corresponding initial accurate label and bounding box.

[0175] Step 320: Detect each target object appearing in each video frame of the target video data to obtain a detection box for marking the detection results. Each target object detected in each video frame has a corresponding detection box.

[0176] Step 330: Input the target video data, initial accurate labels, bounding boxes, and detection boxes into the multi-object tracking model. The multi-object tracking model performs target object tracking and matching, and sets the highest priority for accurate labels during tracking and matching to obtain the first label corresponding to each target object in each video frame. The accurate labels include the initial accurate labels. The first labels corresponding to the same target object tracked by the multi-object tracking model are the same and are all the initial accurate labels corresponding to the target object.

[0177] Step 340: Obtain the first accurate label and the corresponding annotation box based on the partial first annotation label. The partial first annotation label is the first annotation label corresponding to each target object in each video frame obtained according to the first frame number interval in the target video data. The first accurate label is the accurate label obtained after reviewing and correcting the partial first annotation label. The annotation box corresponding to the first accurate label is the accurate annotation box obtained after reviewing and correcting the detection box corresponding to the corresponding first annotation label.

[0178] For example, after obtaining the first annotation label, a portion of the first annotation labels in the target video data are manually reviewed. If the first annotation label is accurate, it is retained; if it is inaccurate, it is modified to become an accurate label. The standard for manual review is whether the initial accurate label set for the target object corresponding to the first annotation label matches the first annotation label. If they match, the first annotation label is accurate; if they do not match, the first annotation label is inaccurate, and the first annotation label is modified to be accurate by the manual reviewer.

[0179] The first annotation label for manual review is the first annotation label in the interval video frames. In one embodiment, during manual review, the first annotation label of each video frame at intervals of the first frame number is reviewed. The first frame number can be set according to actual conditions. For example, if the target video data has 900 frames and the first frame number is 50, then manual review will be performed on the first annotation label of every 50 frames in the target video data. In this case, only 18 video frames need to be reviewed manually, resulting in a smaller review workload.

[0180] After manual review, the first labeled tag (i.e., the tag confirmed accurate by humans) is considered the accurate tag. Currently, the accurate tag obtained after manual review of the first labeled tag is recorded as the first accurate tag. It is understandable that there may be overlap between the first accurate tag obtained through manual review and the initial accurate tag. For example, the first video frame may have an initial accurate tag, and this initial accurate tag, after being used as the first labeled tag, may also be manually reviewed to obtain the first accurate tag. In principle, both tags belong to the same tag, so in practical applications, only one type of tag needs to be retained. For instance, if the first labeled tag in the first video frame is the initial accurate tag during manual review, then manual review of the initial accurate tag is unnecessary.

[0181] For example, after the first accurate label is obtained through manual review, the detection box corresponding to the first accurate label can be used as the annotation box corresponding to the first accurate label. The annotation box corresponding to the first accurate label is the accurate annotation box obtained after manual review of the detection box corresponding to the first annotation label.

[0182] Next, the first accurate label and its corresponding bounding box, the initial accurate label and its corresponding bounding box, and the target video data are manually saved. Unapproved first-label labels are deleted, leaving only the detection boxes corresponding to the first label. Afterward, the multi-target labeling device can acquire the target video data, the initial accurate label and its corresponding bounding box, the first accurate label and its corresponding bounding box, and the detection boxes.

[0183] Step 350: Input the target video data, the first accurate label, the initial accurate label, the bounding box, and the detection box into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching and sets the highest priority for the accurate label during tracking and matching to obtain the second label corresponding to each target object in each video frame. The accurate label includes the initial accurate label and the first accurate label.

[0184] For example, the processing procedure of the multi-object tracking model in this step is the same as that in step 330. The difference is that in this step, the accurate labels used by the multi-object tracking model include both initial accurate labels and first accurate labels, meaning there are more accurate labels, which ensures a higher accuracy rate for the labels output by the multi-object tracking model. For instance, if the multi-object tracking model misidentifies a target object when processing a video frame, making the first label of that target object inaccurate (i.e., the first label of that target object is the initial accurate label of another target object), then when other video frames following that video frame are processed by the multi-object tracking model, the misidentified matching results will be used, affecting the labeling accuracy of the multi-object tracking model. In this case, manual review can correct the erroneous first label in the video frame to obtain the first accurate label. Therefore, when the multi-object tracking model processes the target video frame again, the existence of the first accurate label can avoid misidentification of the target object, that is, it will not identify the target object as another target object. This avoids the impact of misidentification on the tracking and matching of other video frames after the target video frame, thus improving the labeling accuracy of the multi-object tracking model.

[0185] Currently, the labels output by the multi-object tracking model are denoted as the second labels. The accuracy of the second labels is higher than that of the first labels. This means that the second labels corresponding to the same target object tracked by the multi-object tracking model are identical and are all the initial accurate labels for the target object.

[0186] It should be noted that for multi-target tracking models, there is no need to distinguish between initial accurate labels and first accurate labels; any label input into the multi-target tracking model will be considered an accurate label. Currently, the use of initial accurate labels and first accurate labels is only to differentiate between manually set labels and labels that have undergone initial manual review.

[0187] Optionally, the detection boxes used in the multi-object tracking model processing are still the detection boxes output in step 320. Since the multi-object tracking model has already matched the labeled boxes with the detection boxes in step 330, highly overlapping detection boxes may not appear when the labeled boxes and detection boxes are matched in this processing of the multi-object tracking model.

[0188] After obtaining the second labels based on the multi-target tracking model, the entire set of target video data and the second labels can be manually reviewed. In one embodiment, to further shorten the time for manual full-scale review, the second labels can be manually reviewed and corrected at intervals between video frames, and then the target object is matched and tracked again based on the accurate labels corrected by manual review, thereby improving the accuracy of the labels output by the multi-target tracking model. In this case, steps 360 and 370 can also be included after step 350:

[0189] Step 360: Obtain the second accurate label and the corresponding bounding box of the second accurate label based on the partial second label. The partial second label is the second label corresponding to each target object in each video frame obtained according to the second frame number interval in the target video data. The second accurate label is the accurate label obtained after reviewing and correcting the partial second label. The bounding box corresponding to the second accurate label is the accurate bounding box obtained after reviewing and correcting the detection box corresponding to the corresponding second label. The second frame number is less than the first frame number.

[0190] For example, after obtaining the second annotation labels, a portion of the second annotation labels in the target video data are manually reviewed. The manual review of the second annotation labels is the same as the manual review of the first annotation labels, except that the second annotation labels are reviewed at intervals equal to the second frame number. The second frame number can be set according to actual conditions, and it is less than the first frame number. In one embodiment, the first frame number is even, and the second frame number is half of the first frame number. In this case, when reviewing the second annotation labels in the video frames at intervals equal to the second frame number, some video frames have already been reviewed when the first annotation labels were reviewed. Therefore, the manual review only needs to review the second annotation labels in the video frames that have not yet been reviewed. For example, if the target video data has 900 frames, with 50 first frames and 25 second frames, then the manual review will review the second annotation labels in every 25 frames of the target video data, ignoring the video frames reviewed in the first review. In this case, the manual review only needs to review 18 video frames, resulting in a smaller review workload.

[0191] Once the manual review is completed, the second label that has been reviewed (i.e., the label that has been confirmed to be accurate by humans) can be considered an accurate label. Currently, the accurate label obtained after the manual review of the second label is recorded as the second accurate label.

[0192] For example, after the second accurate label is obtained through manual review, the detection box corresponding to the second accurate label can be used as the annotation box corresponding to the second accurate label. That is, the annotation box corresponding to the second accurate label is the accurate annotation box obtained after manual review of the detection box corresponding to the second annotation label.

[0193] Next, the second accurate label and its corresponding bounding box, the first accurate label and its corresponding bounding box, the initial accurate label and its corresponding bounding box, and the target video data are manually saved. Unapproved second accurate labels (excluding unapproved first and initial accurate labels) are deleted, leaving only the detection boxes corresponding to the second accurate labels. Afterward, the multi-target annotation device can acquire the target video data, the initial accurate labels and their corresponding bounding boxes, the first accurate labels and their corresponding bounding boxes, and the second accurate labels and their corresponding detection boxes.

[0194] Step 370: Input the target video data, second accurate label, first accurate label, initial accurate label, bounding box, and detection box into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching, and sets the highest priority for the accurate label during tracking and matching to obtain the third label corresponding to each target object in each video frame. The accurate label includes the initial accurate label, the first accurate label, and the second accurate label.

[0195] For example, the processing procedure for the multi-target tracking model in this step is the same as that in steps 330 and 350. The difference is that in this step, the accurate labels used by the multi-target tracking model include an initial accurate label, a first accurate label, and a second accurate label, meaning there are more accurate labels, which ensures a higher accuracy rate for the labels output by the multi-target tracking model. Currently, the labels output by the multi-target tracking model are denoted as the third label. The accuracy rate of the third label is higher than that of the second and first labels.

[0196] It should be noted that for multi-target tracking models, there is no need to distinguish between initial accurate labels, first accurate labels, and second accurate labels. Any label input into the multi-target tracking model will be considered an accurate label. Currently, the use of initial accurate labels, first accurate labels, and second accurate labels is only to differentiate between manually set labels and labels that have undergone two manual reviews, facilitating a better understanding of the technical solution.

[0197] Generally, two manual frame reviews at intervals are sufficient to achieve a high accuracy rate, thus allowing for shorter review times in subsequent manual reviews. This means that the time required for a full manual review of the target video data and the third annotation label (i.e., reviewing all video frames) is significantly shorter than in related technologies. In practical applications, a third manual review at intervals can also be performed to ensure even more accurate annotation labels output by the multi-object tracking model.

[0198] Understandably, in the process of using semi-automatic annotation methods in related technologies, ID error switching requires manual correction, which accounts for a significant portion of the manual review workload. ID error switching can be understood as the occlusion caused by the movement of the target object, leading to the interchange of IDs for different specific targets during ByteTrack processing, such as identifying target object A as target object B. In this case, manual correction of the interchanged IDs is required, consuming a considerable amount of manual time. In this case, the multi-target annotation method in this embodiment obtains accurate labels by manually correcting the annotation labels output by the multi-target tracking model multiple times. These accurate labels, along with the target video data, are then input into the multi-target tracking model. This allows the multi-target tracking model to use the accurate labels as a reference, reducing the probability of ID error switching, i.e., reducing the probability of misidentification, and improving the annotation accuracy of the multi-target tracking model.

[0199] The above-mentioned approach, employing a step-by-step method, involves manually reviewing and correcting some of the labels obtained by the multi-object tracking model. This ensures that the accurate labels are used as a reference (i.e., the highest priority) when the multi-object tracking model tracks target video data again, significantly improving the accuracy of the labels obtained by the multi-object tracking model. Generally, after 2 to 3 rounds of manual review of some labels, the efficiency of the full manual review will increase by about 3 times.

[0200] The multi-target annotation method provided in this application embodiment is illustrated below. In this example, the target object is a person, the target video data is video data shot in a classroom scene, the duration of the target video data is 45 minutes, and the total number of frames is 900.

[0201] First, manual annotation involves adding bounding boxes and initial accurate labels to each target object appearing for the first time in the first and subsequent video frames of the target video data. Then, a multi-target annotation device acquires the target video data and the initial accurate labels and bounding boxes added for the first appearance of the target objects. Next, the multi-target annotation device detects the target objects appearing in each video frame of the target video data to obtain detection boxes identifying the detected target objects and the corresponding confidence scores for each detection box. A confidence score greater than or equal to 0.5 is considered high confidence, and a confidence score less than 0.5 is considered low confidence.

[0202] Subsequently, the multi-target annotation device inputs the target video data, initial accurate labels, annotation boxes, and detection boxes into the multi-target tracking model. Figure 5 This is a schematic diagram illustrating the processing flow of a multi-target tracking model provided in one embodiment of this application. (Reference) Figure 5 In the process of multi-object tracking model processing, the target bounding box in the currently processed video frame ( Figure 5 The term "Candidates" includes the label boxes corresponding to the exact labels. Figure 5 (referred to as Annotations) and the detection boxes generated during automatic device detection ( Figure 5 (Refered as Dets). First, the multi-target tracking model performs its first (i.e., Figure 5 Matching the number 1 in the data, specifically matching the Annotations with the set of tracks ( Figure 5 The tracking trajectories in the Stracks are matched to obtain M1 and UD1 (i.e., ...). Figure 5 U_D1) and UT1 (i.e. Figure 5 (U_T1). Afterwards, the multi-target tracking model performs a second (i.e., Figure 5 Matching the number 2 in the code involves matching Annotations with Dets to obtain M2 and UD2 (i.e., ...). Figure 5 U_D2) and UT2 (i.e. Figure 5 (U_T2). After that, the matching tracking strategy of ByteTrack can be referenced, and the multi-object tracking model will perform a third (i.e., Figure 5 The number 3 in the middle), the fourth time (i.e. Figure 5 The number 4) and the fifth (i.e. Figure 5 The third match is specifically using the detection box corresponding to the high confidence level in UD2 (the number 5 in the original text). Figure 5 The high score is denoted as high_score in the middle and the active tracking trajectory in UT1 is ( Figure 5 The matching is performed using the term activated_stracks to obtain M3 and UD3 (i.e., ...). Figure 5 U_D3) and UT3 (i.e. Figure 5 In U_T3), the fourth match specifically uses the detection box corresponding to the low confidence level in UD2 ( Figure 5 The low score (denoted as low_score) is matched with the tracking trajectory in UT3 to obtain M4 and UT4 (i.e., Figure 5 U_T4) and UD4 (i.e. Figure 5 In U_D4), the fifth match specifically uses the detection box in UD3 and the tracking trajectory in U1 that is inactive. Figure 5 The code is denoted as non_activated_stracks and used for matching to obtain M5 and UT5 (i.e., ...). Figure 5U_T5) and UD5 (i.e. Figure 5 (U_D5).

[0203] Among them, the multi-target tracking model ignores (i.e.) Figure 5 In the Ignore) M2, UT2, UD4, and UD5 multi-object tracking models, a new tracking trajectory is created for each bounding box in UD1 and set to the active state (i.e., Figure 5 The tracking trajectory is set to the initial accurate label corresponding to the bounding box, and the new (activated) label is used for the tracking trajectory. The multi-object tracking model updates the matched tracking trajectory based on the detection boxes in M5, M4, and M3, and the annotation boxes in M1, and sets it to the active state (i.e., the new (activated) label). Figure 5 The Update(activated) function in the multi-target tracking model retains the tracking trajectories from UT4 and UT5 and sets them to an inactive state (i.e., Update(activated)). Figure 5 The Remain (non-activated) track is added to Stats for use when processing the next video frame. New tracking tracks, updated tracking tracks, and retained tracking tracks are all added to Stats.

[0204] After the multi-target tracking model has processed all video frames in the target video data, it can obtain the first label of each target object. The target objects identified by the target boxes (located in different video frames) on a tracking trajectory obtained by the multi-target tracking model are the same target objects matched by the multi-target tracking model. These target objects share the same first label, and the first label is the initial accurate label corresponding to the creation of the tracking trajectory.

[0205] Next, the first annotation labels of the video frames in the target video data are manually reviewed at 50-frame intervals to obtain the first accurate label after review and correction. The detection box corresponding to the first accurate label is then used as the annotation box. It can be understood that the first annotation label obtained from the initial accurate label does not require further manual review and is still used as the initial accurate label.

[0206] Subsequently, the multi-object annotation device inputs the target video data, initial manual labels, first accurate labels, annotation boxes, and detection boxes into the multi-object tracking model, which then continues to track the target video data according to the target's initial labels, first accurate labels, annotation boxes, and detection boxes. Figure 5 The process is followed to obtain the second label.

[0207] Next, the second annotation labels of the video frames in the target video data are manually reviewed at 25-frame intervals to obtain the revised second accurate labels, and the detection boxes corresponding to the second accurate labels are used as annotation boxes. It can be understood that the second annotation labels obtained from the initial accurate labels and the first accurate labels do not need to be manually reviewed again and are still used as the initial accurate labels and the first accurate labels.

[0208] Subsequently, the multi-object annotation device inputs the target video data, initial accurate labels, first accurate labels, second accurate labels, annotation boxes, and detection boxes into the multi-object tracking model, which then continues to track the target video data according to the target's location. Figure 5 The process is followed to obtain the third label.

[0209] Next, the third annotation labels in each video frame of the target video data are reviewed manually to obtain the final label of the target object.

[0210] Figure 6 A schematic diagram illustrating the manual review time provided in one embodiment of this application, with reference to... Figure 6 The paper demonstrates that manually annotating 900 video frames using existing purely manual methods, even with the aid of automatic inter-frame interpolation tools, requires a 30-day review period. Using existing semi-automatic annotation methods, manual review of the same 900 frames requires 9 days. However, employing the multi-target annotation method described in the previous embodiment, with two separate manual reviews at intervals, followed by a full manual review of all video frames, reduces the review time to 3 days, significantly shortening the manual review period.

[0211] One embodiment of this application also provides a multi-target annotation device. Figure 7 This is a schematic diagram of a multi-target annotation device provided in one embodiment of this application, with reference to... Figure 7 The multi-target labeling device includes a first acquisition unit 401, a detection unit 402, and a first tracking unit 403.

[0212] The system includes a first acquisition unit 401, which acquires target video data and the initial accurate labels and bounding boxes of each target object appearing for the first time in the target video data in the corresponding video frames. Each target object appearing for the first time has a corresponding initial accurate label and bounding box. A detection unit 402 is used to detect each target object appearing in each video frame of the target video data to obtain detection boxes for marking the detection results. Each target object detected in each video frame has a corresponding detection box. A first tracking unit 403 is used to input the target video data, the initial accurate labels, the bounding boxes, and the detection boxes into a multi-target tracking model. The multi-target tracking model performs target object tracking and matching and sets the highest priority for accurate labels during tracking and matching to obtain the first label corresponding to each target object in each video frame. The accurate label includes the initial accurate label. The first labels corresponding to the same target object tracked by the multi-target tracking model are the same and are all the initial accurate labels corresponding to the target object.

[0213] In one embodiment of this application, it further includes: a second acquisition unit, configured to, after the multi-target tracking model performs target object tracking and matching and sets the highest priority for accurate labels during tracking and matching to obtain a first annotation label corresponding to each target object in each video frame, acquire a first accurate label obtained based on a portion of the first annotation labels and an annotation box corresponding to the first accurate label, wherein the portion of the first annotation labels are the first annotation labels corresponding to each target object in each video frame obtained at a first frame interval in the target video data, the first accurate label is an accurate label obtained after reviewing and correcting the portion of the first annotation labels, and the annotation box corresponding to the first accurate label is an accurate annotation box obtained after reviewing and correcting the detection box corresponding to the corresponding first annotation label; and a second tracking unit, configured to input the target video data, the first accurate label, the initial accurate label, the annotation box, and the detection box together into the multi-target tracking model, wherein the multi-target tracking model performs target object tracking and matching and sets the highest priority for accurate labels during tracking and matching to obtain a second annotation label corresponding to each target object in each video frame, wherein the accurate label includes the initial accurate label and the first accurate label.

[0214] In one embodiment of this application, it further includes: a third acquisition unit, configured to, after the multi-target tracking model performs target object tracking and matching and sets the highest priority for accurate labels during tracking and matching to obtain second annotation labels corresponding to each target object in each video frame, acquire a second accurate label obtained based on a portion of the second annotation labels and an annotation box corresponding to the second accurate label, wherein the portion of the second annotation labels are second annotation labels corresponding to each target object in each video frame obtained at second frame intervals in the target video data, the second accurate label is an accurate label obtained after reviewing and correcting the portion of the second annotation labels, and the annotation box corresponding to the second accurate label is an accurate annotation box obtained after reviewing and correcting the detection box corresponding to the corresponding second annotation label, wherein the second frame number is less than the first frame number; and a third tracking unit, configured to input the target video data, the second accurate label, the first accurate label, the initial accurate label, the annotation box, and the detection box together into the multi-target tracking model, wherein the multi-target tracking model performs target object tracking and matching and sets the highest priority for accurate labels during tracking and matching to obtain a third annotation label corresponding to each target object in each video frame, wherein the accurate label includes the initial accurate label, the first accurate label, and the second accurate label.

[0215] In one embodiment of this application, the first tracking unit, the second tracking unit, and the third tracking unit may each include: an input subunit, used to input the target video data, accurate labels, the bounding boxes, and the detection boxes together into a multi-target tracking model; and a first matching subunit, used by the multi-target tracking model to match the bounding boxes corresponding to each accurate label in the current video frame of the target video data with each tracking trajectory in the trajectory set, to obtain a set of successfully matched bounding boxes, a set of unmatched first bounding boxes, and a set of unmatched first trajectories, wherein the trajectory set is a set of each tracking trajectory obtained during the processing of the multi-target tracking model, and each tracking trajectory corresponds to... A target object, representing the position trajectory of the currently identified target object in the target video data; a second matching subunit, used by the multi-target tracking model to match each bounding box in the current video frame with each detection box in the current video frame, to obtain a first set of successfully matched detection boxes, a first set of unmatched detection boxes, and a second set of unmatched bounding boxes; a third matching subunit, used by the multi-target tracking model to match the detection boxes in the first set of unmatched detection boxes with the tracking trajectories in the first set of unmatched trajectories, to obtain a second set of successfully matched detection boxes, a second set of unmatched detection boxes, and a second trajectory. The multi-target tracking model is configured with the following sub-units: a set of unmatched bounding boxes; a trajectory creation sub-unit, used by the multi-target tracking model to create new tracking trajectories based on each bounding box in the first set of unmatched bounding boxes, with the new tracking trajectories using the accurate labels corresponding to the bounding boxes as their labels; a trajectory update sub-unit, used by the multi-target tracking model to update the corresponding matched tracking trajectories based on the set of successfully matched bounding boxes and the second set of successfully matched detection boxes, so that the updated tracking trajectories share the same labels; a trajectory retention sub-unit, used by the multi-target tracking model to retain each tracking trajectory in the second set of unmatched trajectory; and a set determination sub-unit, used by the multi-target tracking model to determine the set of unmatched bounding boxes. Based on the new tracking trajectory, the updated tracking trajectory, and the retained tracking trajectory, a trajectory set for use in the next video frame is obtained; a set ignoring subunit is used by the multi-target tracking model to ignore the first detection box matching success set, the second annotation box not matching success set, and the second detection box not matching success set; a video frame updating subunit is used by the multi-target tracking model to update the next video frame in the target video data to the current video frame, and return to perform the operation of matching the annotation box corresponding to each accurate label in the current video frame with each tracking trajectory in the trajectory set, until each video frame in the target video data has been processed by the multi-target tracking model.

[0216] In one embodiment of this application, it further includes: a trajectory state setting unit, used by the multi-target tracking model to set both the new tracking trajectory and the updated tracking trajectory to an active state, and to set the retained tracking trajectory to an inactive state.

[0217] In one embodiment of this application, each detection box has a corresponding confidence level, and the confidence level reaching the confidence level threshold is considered high confidence, while the confidence level not reaching the confidence level threshold is considered low confidence. The third matching subunit includes: a fourth matching grandchild unit, used by the multi-target tracking model to match the detection boxes with high confidence in the first unmatched detection box set with the tracking trajectories in the active state in the first unmatched trajectory set, to obtain a third detection box matching successful subset, a third detection box unmatched subset, and a third trajectory unmatched subset; a fifth matching grandchild unit, used by the multi-target tracking model to match the detection boxes with low confidence in the first unmatched detection box set with the tracking trajectories in the third unmatched trajectory set, to obtain a fourth detection box matching successful subset, a fourth detection box unmatched subset, and a fourth trajectory unmatched subset; a sixth matching grandchild unit, used by the... The multi-target tracking model matches the detection boxes in the third unmatched subset with the inactive tracking trajectories in the first unmatched subset to obtain a fifth successfully matched subset, a fifth unmatched subset, and a fifth unmatched subset. The third successfully matched subset, the fourth successfully matched subset, and the fifth successfully matched subset constitute the second successfully matched subset. The fourth unmatched subset and the fifth unmatched subset constitute the second unmatched subset. The fourth unmatched subset and the fifth unmatched subset constitute the second unmatched subset.

[0218] In one embodiment of this application, the second matching subunit includes: an overlap determination subunit, configured to calculate the overlap between each bounding box and each detection box in the current video frame by the multi-object tracking model; if the overlap between the bounding box and the detection box reaches an overlap threshold, the bounding box and the detection box are determined to be successfully matched; if the overlap between the bounding box and the detection box does not reach the overlap threshold, the bounding box and the detection box are determined to be unmatched; and a matching confirmation subunit, configured to obtain a first set of successfully matched detection boxes, a first set of unmatched detection boxes, and a second set of unmatched bounding boxes by the multi-object tracking model based on the successfully matched bounding boxes and detection boxes and the unmatched bounding boxes and detection boxes.

[0219] In one embodiment of this application, the target video data is video data obtained by shooting in a specific confined space, and the specific confined space contains multiple target objects.

[0220] The multi-target annotation device provided in this application embodiment is included in the multi-target annotation device and can be used to execute the multi-target annotation method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0221] It is worth noting that in the embodiments of the multi-target annotation device described above, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this application.

[0222] One embodiment of this application also provides a multi-target annotation device, see reference. Figure 1 The multi-target annotation device includes a processor 11 and a memory 12. The processor 11 and memory 12 can be connected via a bus or other means. The memory 12 stores one or more programs; when one or more programs are executed by one or more processors 11, the processors 11 implement the multi-target annotation method described in any of the foregoing embodiments. For details regarding each component, please refer to the foregoing description.

[0223] The aforementioned multi-target annotation device is used to execute arbitrary multi-target annotation methods and has corresponding functions and beneficial effects. For specific details not described here, please refer to the relevant descriptions of the aforementioned multi-target annotation methods.

[0224] One embodiment of this application also provides a storage medium containing computer-executable instructions, which, when executed by a processor, are used to perform relevant operations in the multi-target annotation method provided in any embodiment of this application, and have corresponding functions and beneficial effects.

[0225] Those skilled in the art will understand that embodiments of this application may provide methods, systems, or computer program products.

[0226] Therefore, this application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing module of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing module of the computer or other programmable data processing apparatus, produce implementations of the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0227] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0228] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0229] Note that the above are merely preferred embodiments and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this application, the scope of which is determined by the scope of the appended claims.

Claims

1. A multi-target annotation method, characterized in that, include: Obtain the target video data and the initial accurate labels and bounding boxes of each target object that appears for the first time in the target video data in the corresponding video frame. Each target object that appears for the first time has a corresponding initial accurate label and bounding box. Each target object appearing in each video frame of the target video data is detected to obtain a detection box for marking the detection results. Each target object detected in each video frame has a corresponding detection box. The target video data, the initial accurate label, the bounding box, and the detection box are input together into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching, and sets the highest priority for the accurate label during tracking and matching to obtain the first label corresponding to each target object in each video frame. The accurate label includes the initial accurate label. The first labels corresponding to the same target object tracked by the multi-target tracking model are the same and are all the initial accurate labels corresponding to the target object.

2. The multi-target annotation method according to claim 1, characterized in that, The step of tracking and matching target objects using the multi-target tracking model and setting the highest priority for accurate labels during tracking and matching to obtain the first labeled label corresponding to each target object in each video frame includes: Obtain a first accurate label and a corresponding bounding box based on a portion of the first first label. The portion of the first label is the first label corresponding to each target object in each video frame obtained according to the first frame number interval in the target video data. The first accurate label is the accurate label obtained after reviewing and correcting the portion of the first label. The bounding box corresponding to the first accurate label is the accurate bounding box obtained after reviewing and correcting the detection box corresponding to the corresponding first label. The target video data, the first accurate label, the initial accurate label, the annotation box, and the detection box are input together into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching, and sets the highest priority for the accurate label during tracking and matching to obtain the second annotation label corresponding to each target object in each video frame. The accurate label includes the initial accurate label and the first accurate label.

3. The multi-target annotation method according to claim 2, characterized in that, The step of tracking and matching target objects using the multi-object tracking model and setting the highest priority for accurate labels during tracking and matching to obtain the second annotation label corresponding to each target object in each video frame includes: Obtain a second accurate label based on a portion of the second annotation labels and the corresponding annotation box of the second accurate label. The portion of the second annotation labels are the second annotation labels corresponding to each target object in each video frame obtained according to the second frame number interval in the target video data. The second accurate label is the accurate label obtained after reviewing and correcting the portion of the second annotation labels. The annotation box corresponding to the second accurate label is the accurate annotation box obtained after reviewing and correcting the detection box corresponding to the corresponding second annotation label. The second frame number is less than the first frame number. The target video data, the second accurate label, the first accurate label, the initial accurate label, the annotation box, and the detection box are input together into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching, and sets the highest priority for the accurate label during tracking and matching to obtain the third annotation label corresponding to each target object in each video frame. The accurate label includes the initial accurate label, the first accurate label, and the second accurate label.

4. The multi-target annotation method according to any one of claims 1-3, characterized in that, The step of tracking and matching target objects using the multi-target tracking model and setting the highest priority for accurate labels during tracking and matching includes: The multi-target tracking model matches the bounding boxes corresponding to each accurate label in the current video frame of the target video data with each tracking trajectory in the trajectory set to obtain a set of successfully matched bounding boxes, a set of unmatched first bounding boxes, and a set of unmatched first trajectories. The trajectory set is a set of each tracking trajectory obtained during the processing of the multi-target tracking model. Each tracking trajectory corresponds to a target object and represents the position trajectory of the currently identified target object in the target video data. The multi-target tracking model matches each bounding box in the current video frame with each detection box in the current video frame to obtain a first set of successfully matched detection boxes, a first set of unmatched detection boxes, and a second set of unmatched bounding boxes. The multi-target tracking model matches the detection boxes in the first set of unmatched detection boxes with the tracking trajectories in the first set of unmatched trajectories to obtain the second set of successfully matched detection boxes, the second set of unmatched detection boxes, and the second set of unmatched trajectories. The multi-target tracking model creates new tracking trajectories based on each of the unmatched bounding boxes in the first set of unmatched bounding boxes, and the new tracking trajectories use the accurate labels corresponding to the bounding boxes as labels. The multi-target tracking model updates the corresponding matched tracking trajectory based on the set of successfully matched bounding boxes and the set of successfully matched second detection boxes, so that the updated tracking trajectory shares the same label. The multi-target tracking model retains each tracking trajectory in the set of unmatched second trajectories; The multi-target tracking model obtains a set of trajectories for use in the next video frame based on the new tracking trajectory, the updated tracking trajectory, and the retained tracking trajectory. The multi-target tracking model ignores the set of successfully matched first detection boxes, the set of unmatched second annotation boxes, and the set of unmatched second detection boxes. The multi-target tracking model updates the next video frame in the target video data to the current video frame, and then returns to perform the operation of matching the bounding box corresponding to each accurate label in the current video frame with each tracking trajectory in the trajectory set, until each video frame in the target video data has been processed by the multi-target tracking model.

5. The multi-target annotation method according to claim 4, characterized in that, Also includes: The multi-target tracking model sets both the new tracking trajectory and the updated tracking trajectory to an active state, and sets the retained tracking trajectory to an inactive state.

6. The multi-target annotation method according to claim 5, characterized in that, Each detection box has a corresponding confidence level, and a confidence level that reaches the confidence level threshold is considered high confidence, while a confidence level that does not reach the confidence level threshold is considered low confidence. The step of matching the detection boxes in the first set of unmatched detection boxes with the tracking trajectories in the first set of unmatched trajectories by the multi-target tracking model to obtain the second set of successfully matched detection boxes, the second set of unmatched detection boxes, and the second set of unmatched trajectories includes: The multi-target tracking model matches the high-confidence detection boxes in the first set of unmatched detection boxes with the active tracking trajectories in the first set of unmatched trajectories to obtain the third set of successfully matched detection boxes, the third set of unmatched detection boxes, and the third set of unmatched trajectories. The multi-target tracking model matches the low-confidence detection boxes in the first set of unmatched detection boxes with the tracking trajectories in the third set of unmatched trajectories to obtain the fourth set of successfully matched detection boxes, the fourth set of unmatched detection boxes, and the fourth set of unmatched trajectories. The multi-target tracking model matches the detection boxes in the third unmatched subset with the inactive tracking trajectories in the first unmatched subset to obtain a fifth successfully matched subset, a fifth unmatched subset, and a fifth unmatched subset. The third successfully matched subset, the fourth successfully matched subset, and the fifth successfully matched subset constitute the second successfully matched subset. The fourth unmatched subset and the fifth unmatched subset constitute the second unmatched subset. The fourth unmatched subset and the fifth unmatched subset constitute the second unmatched subset.

7. The multi-target annotation method according to claim 4, characterized in that, The step of matching each bounding box in the current video frame with each detection box in the current video frame by the multi-object tracking model to obtain a first set of successfully matched detection boxes, a first set of unmatched detection boxes, and a second set of unmatched bounding boxes includes: The multi-target tracking model calculates the overlap between each bounding box and each detection box in the current video frame. If the overlap between the bounding box and the detection box reaches the overlap threshold, it is determined that the bounding box and the detection box are successfully matched. If the overlap between the bounding box and the detection box does not reach the overlap threshold, it is determined that the bounding box and the detection box are not successfully matched. The multi-target tracking model obtains a first set of successfully matched detection boxes, a first set of unmatched detection boxes, and a second set of unmatched annotation boxes based on the successfully matched bounding boxes and detection boxes, as well as the unmatched bounding boxes and detection boxes.

8. The multi-target annotation method according to claim 1, characterized in that, The target video data is video data obtained by shooting in a specific confined space, which contains multiple target objects.

9. A multi-target annotation device, characterized in that, include: The first acquisition unit is used to acquire target video data and the initial accurate labels and bounding boxes of each target object that appears for the first time in the target video data in the corresponding video frame. Each target object that appears for the first time has a corresponding initial accurate label and bounding box. The detection unit is used to detect each target object appearing in each video frame of the target video data, so as to obtain a detection box for marking the detection results. Each target object detected in each video frame has a corresponding detection box. The first tracking unit is used to input the target video data, the initial accurate label, the annotation box, and the detection box into the multi-target tracking model. The multi-target tracking model performs target object tracking and matching, and sets the highest priority for the accurate label during tracking and matching to obtain the first annotation label corresponding to each target object in each video frame. The accurate label includes the initial accurate label. The first annotation labels corresponding to the same target object tracked by the multi-target tracking model are the same and are all the initial accurate labels corresponding to the target object.

10. A multi-target annotation device, characterized in that, include: One or more processors and memory; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-target annotation method as described in any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the multi-target annotation method as described in any one of claims 1-8.