A video-based instance tracking method, device, equipment and storage medium

By building an end-to-end instance detection framework, using the backbone network and instance segmentation network to directly detect and track instance targets in video frames, the problem of inefficient instance tracking caused by relying on post-processing methods in the prior art is solved, and efficient instance tracking is achieved.

CN113822134BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110813442.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-19
Publication Date
2025-06-27
Estimated Expiration
2041-07-19

AI Technical Summary

Technical Problem

The existing instance segmentation algorithm relies on post-processing methods such as non-maximum suppression (NMS) during the instance tracking process, resulting in low instance tracking efficiency and difficulty in achieving end-to-end inference.

Method used

By building an end-to-end instance detection framework, the backbone network and instance segmentation network are used to directly detect instance targets in video frames, and track them through instance query vectors, avoiding dependence on post-processing methods such as NMS.

Benefits of technology

The video-based instance tracking efficiency is improved, so that instance detection and tracking can be performed without relying on post-processing methods, significantly improving the inference speed and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113822134B_ABST
    Figure CN113822134B_ABST
Patent Text Reader

Abstract

The present application discloses a video-based instance tracking method implemented by using artificial intelligence technology, including: obtaining a target feature map through a backbone network based on a target video frame in a video to be detected; obtaining N regions of interest (ROIs) of bounding boxes from the target feature map according to N instance bounding boxes; obtaining N first detection results through an instance segmentation network based on N instance query vectors and N bounding box ROIs; determining at least one instance similarity according to the N first detection results and M second detection results; and determining an instance tracking result of the target video frame according to the at least one instance similarity. The present application also provides a device, a device and a storage medium. By constructing an end-to-end instance detection framework, the present application realizes instance detection without relying on post-processing methods such as non-maximum suppression, and further tracks instance targets based on instance identifiers, thereby improving the efficiency of video-based instance tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a video-based instance tracking method, apparatus, device, and storage medium. Background Art

[0002] Instance segmentation is a crucial preprocessing for image recognition and computer vision, and is widely used in various fields. For example, instance segmentation can be used for tasks such as object recognition, object detection, and object tracking. In the object detection task, it is necessary to not only detect the category of the object in the image, but also detect the bounding box of the object.

[0003] Currently, instance segmentation algorithms generally follow the process of "detect first and then segment", that is, through object detection based on prior boxes, the instances of interest in the video are detected and segmented. Specifically, when screening positive samples during the training process, it is necessary to match based on the intersection over union between the prior box and the ground truth box.

[0004] However, since the prior box follows the principle of "one-to-many (i.e., one ground truth box corresponds to multiple prior boxes)" during the training process, in the test stage, it is necessary to rely on post-processing methods such as Non-Maximum Suppression (NMS) to reduce duplicate instance predictions, and it is difficult to perform end-to-end inference, resulting in low instance tracking efficiency. Summary of the Invention

[0005] Embodiments of this application provide a video-based instance tracking method, apparatus, device, and storage medium. This application constructs an end-to-end instance detection framework, realizes instance detection without relying on post-processing methods such as NMS, and then tracks the instance target based on the instance identifier, thereby improving the efficiency of video-based instance tracking.

[0006] In view of this, on the one hand, this application provides a video-based instance tracking method, including:

[0007] Based on the target video frame in the video to be detected, obtain a target feature map through a backbone network, where the target video frame is the T-th video frame in the video to be detected, and T is an integer greater than 1;

[0008] According to N instance bounding boxes, obtain N regions of interest (ROIs) of the bounding boxes from the target feature map, where each instance bounding box is used to extract a corresponding bounding box ROI, and N is an integer greater than or equal to 1;

[0009] Based on N instance query vectors and N bounding box ROIs, obtain N first detection results through an instance segmentation network, where each first detection result includes a first class probability value, a first instance bounding box, and a first instance embedding vector;

[0010] According to the N first detection results and M second detection results, determine at least one instance similarity, where each second detection result includes a second class probability value, a second instance bounding box, and a second instance embedding vector, and the M second detection results are obtained according to the first (T - 1) video frames in the video to be detected, each second detection result corresponds to an instance identifier, and M is an integer greater than or equal to 1;

[0011] According to at least one instance similarity, determine the instance tracking result of the target video frame, where the instance tracking result includes at least one instance identifier, and the same instance identifier represents the same instance in the video to be detected.

[0012] On the other hand, this application provides an instance tracking device, including:

[0013] An acquisition module, configured to obtain a target feature map through a backbone network based on the target video frame in the video to be detected, where the target video frame is the T-th video frame in the video to be detected, and T is an integer greater than 1;

[0014] The acquisition module is further configured to obtain N bounding box regions of interest (ROIs) from the target feature map according to the N instance bounding boxes, where each instance bounding box is used to extract a corresponding bounding box ROI, and N is an integer greater than or equal to 1;

[0015] The acquisition module is further configured to obtain N first detection results through an instance segmentation network based on the N instance query vectors and the N bounding box ROIs, where each first detection result includes a first class probability value, a first instance bounding box, and a first instance embedding vector;

[0016] A determination module, configured to determine at least one instance similarity according to the N first detection results and the M second detection results, where each second detection result includes a second class probability value, a second instance bounding box, and a second instance embedding vector, the M second detection results are obtained according to the first (T - 1) video frames in the video to be detected, each second detection result corresponds to an instance identifier, and M is an integer greater than or equal to 1;

[0017] The determination module is further configured to determine the instance tracking result of the target video frame according to at least one instance similarity, where the instance tracking result includes at least one instance identifier, and the same instance identifier represents the same instance in the video to be detected.

[0018] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0019] An acquisition module, specifically configured to perform dot multiplication on each of the N bounding box ROIs in the feature dimension by using N instance query vectors to obtain N enhanced bounding box ROIs;

[0020] Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results;

[0021] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0022] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

[0023] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0024] An acquisition module, specifically configured to obtain at least one set of bounding box dynamic parameters through a fully connected layer based on the N instance query vectors;

[0025] Use at least one set of bounding box dynamic parameters to perform dot multiplication on each of the N bounding box ROIs in the feature dimension to obtain N enhanced bounding box ROIs;

[0026] Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results;

[0027] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0028] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

[0029] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0030] The obtaining module is further configured to obtain N mask ROIs from the target video frame according to the N instance bounding boxes, where each instance bounding box is further configured to extract a corresponding mask ROI.

[0031] The obtaining module is specifically configured to obtain N first detection results through an instance segmentation network based on N instance query vectors, N bounding box ROIs, and N mask ROIs, where each first detection result further includes a first instance foreground mask.

[0032] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0033] The obtaining module is specifically configured to perform dot multiplication on each of the N bounding box ROIs in the N bounding box ROIs in the feature dimension by using the N instance query vectors to obtain N enhanced bounding box ROIs;

[0034] Perform dot multiplication on each of the N mask ROIs in the N mask ROIs in the feature dimension by using the N instance query vectors to obtain N enhanced mask ROIs;

[0035] Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results;

[0036] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0037] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results;

[0038] Based on the N enhanced mask ROIs, obtain N first instance foreground masks through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

[0039] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0040] The obtaining module is specifically configured to obtain at least one set of bounding box dynamic parameters and at least one set of mask dynamic parameters through a fully connected layer based on the N instance query vectors, where each set of mask dynamic parameters includes N mask dynamic sub-parameters.

[0041] Using at least one set of bounding box dynamic parameters, perform a dot product on each of the N bounding box ROIs in the feature dimension to obtain N enhanced bounding box ROIs;

[0042] Using at least one set of mask dynamic parameters, perform a dot product on each of the N mask ROIs in the feature dimension to obtain N enhanced mask ROIs;

[0043] Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results;

[0044] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0045] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results;

[0046] Based on the N enhanced mask ROIs, obtain N first instance foreground masks through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

[0047] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0048] The determination module is specifically configured to sort the N first detection results in descending order according to the first category probability values to obtain the sorted N first detection results;

[0049] Select the first K first detection results from the sorted N first detection results, where K is an integer greater than or equal to 1 and less than or equal to N;

[0050] According to the K first detection results and the M second detection results, determine the instance similarity between each first detection result and each second detection result to obtain (K * M) instance similarities.

[0051] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0052] The determination module is specifically configured to determine the instance embedding vector similarity between each first detection result and each second detection result according to the first instance embedding vector included in each first detection result and the second instance embedding vector included in each second detection result;

[0053] Determine the spatial similarity between each first detection result and each second detection result according to the first instance bounding box included in each first detection result and the second instance bounding box included in each second detection result;

[0054] Determine the class similarity between each first detection result and each second detection result according to the first class probability value included in each first detection result and the second class probability value included in each second detection result;

[0055] Determine the instance similarity between each first detection result and each second detection result according to the instance embedding vector similarity, spatial similarity, class similarity between each first detection result and each second detection result, and the first class probability value included in each first detection result.

[0056] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, K is less than or equal to M;

[0057] The determination module is specifically configured to construct a mapping relationship between K first detection results and M second detection results based on the bipartite graph matching algorithm according to (K * M) instance similarities, and obtain K mapping relationships;

[0058] If the instance similarities corresponding to P mapping relationships among the K mapping relationships are less than or equal to the instance similarity threshold, then delete the P mapping relationships from the K mapping relationships to obtain (K - P) mapping relationships, where P is an integer greater than or equal to 1 and less than or equal to K;

[0059] Determine the instance tracking result of the target video frame according to the second detection result corresponding to each mapping relationship among the (K - P) mapping relationships, where each second detection result corresponds to an instance identifier;

[0060] The determination module is further configured to use the first detection results corresponding to the P mapping relationships as the second detection results to obtain (M + P) second detection results.

[0061] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, K is greater than or equal to M;

[0062] The determination module is specifically configured to construct a mapping relationship between K first detection results and M second detection results based on the bipartite graph matching algorithm according to (K * M) instance similarities, and obtain M mapping relationships;

[0063] If the instance similarity corresponding to Q mapping relationships among the M mapping relationships is less than or equal to the instance similarity threshold, then delete the Q mapping relationships from the M mapping relationships to obtain (M - Q) mapping relationships, where Q is an integer greater than or equal to 1 and less than or equal to M;

[0064] Determine the instance tracking result of the target video frame according to the second detection result corresponding to each mapping relationship among the (M - Q) mapping relationships, where each second detection result corresponds to an instance identifier;

[0065] The determining module is further configured to use the Q mapping relationships and the (K - M) first detection results as the second detection results to obtain (Q + K) second detection results.

[0066] In a possible design, in another implementation manner on the other hand of the embodiments of the present application, the instance tracking device further includes a training module;

[0067] The obtaining module is further configured to obtain a sample feature map through the to-be-trained backbone network based on the first sample video frames in the to-be-trained video, where the first sample video frames have labeled categories, labeled instance bounding boxes, and labeled instance identifiers;

[0068] The obtaining module is further configured to obtain N predicted bounding box ROIs from the sample feature map according to N to-be-trained instance bounding boxes, where each to-be-trained instance bounding box is used to extract a corresponding predicted bounding box ROI;

[0069] The obtaining module is further configured to obtain N first prediction results through the to-be-trained instance segmentation network based on the N to-be-trained instance query vectors and the N predicted bounding box ROIs, where each first prediction result includes a first predicted class probability value, a first predicted instance bounding box, and a first predicted instance embedding vector;

[0070] The determining module is further configured to determine at least one predicted instance similarity according to the N first prediction results and the N second prediction results, where each second prediction result includes a second predicted class probability value, a second predicted instance bounding box, and a second predicted instance embedding vector, the N second prediction results are from the second sample video frames in the to-be-trained video, and each second prediction result corresponds to an instance identifier;

[0071] The determining module is further configured to determine the predicted instance tracking result of the first sample video frame according to at least one predicted instance similarity;

[0072] A training module, configured to update the parameters of the backbone network to be trained, N bounding boxes of instances to be trained, N query vectors of instances to be trained, and the instance segmentation network to be trained through a loss function according to the prediction instance tracking result, the labeled instance identifier, N first prediction results, the labeled category, and the labeled instance bounding box.

[0073] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0074] An acquisition module is further configured to obtain a sample feature map through the backbone network to be trained based on the first sample video frame in the video to be trained, where the first sample video frame has a labeled category, a labeled instance bounding box, a labeled instance identifier, and a labeled instance foreground mask;

[0075] The acquisition module is further configured to obtain N predicted bounding box ROIs from the sample feature map according to the N bounding boxes of instances to be trained, where each bounding box of an instance to be trained is used to extract a corresponding predicted bounding box ROI;

[0076] The acquisition module is further configured to obtain N predicted mask ROIs from the sample feature map according to the N bounding boxes of instances to be trained, where each bounding box of an instance to be trained is further used to extract a corresponding predicted mask ROI;

[0077] The acquisition module is further configured to obtain N first prediction results through the instance segmentation network to be trained based on the N query vectors of instances to be trained, the N predicted bounding box ROIs, and the N predicted mask ROIs, where each first prediction result includes a first predicted category probability value, a first predicted instance bounding box, a first predicted instance embedding vector, and a predicted instance foreground mask;

[0078] A determination module is further configured to determine at least one predicted instance similarity according to the N first prediction results and the N second prediction results, where each second prediction result includes a second predicted category probability value, a second predicted instance bounding box, and a second predicted instance embedding vector, and the N second prediction results are from the second sample video frame in the video to be trained, and each second prediction result corresponds to an instance identifier;

[0079] The determination module is further configured to determine the prediction instance tracking result of the first sample video frame according to at least one predicted instance similarity;

[0080] The training module is further configured to update the parameters of the backbone network to be trained, the N bounding boxes of instances to be trained, the N query vectors of instances to be trained, and the instance segmentation network to be trained through a loss function according to the prediction instance tracking result, the labeled instance identifier, the N first prediction results, the labeled category, the labeled instance bounding box, and the labeled instance foreground mask.

[0081] On the other hand, the present application provides a computer device, including: a memory, a processor, and a bus system;

[0082] The memory is used to store programs;

[0083] The processor is used to execute the programs in the memory, and the processor is used to execute the methods of the above aspects according to the instructions in the program code;

[0084] The bus system is used to connect the memory and the processor, so that the memory and the processor can communicate.

[0085] On the other hand, the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a computer, the computer is enabled to execute the methods of the above aspects.

[0086] In another aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.

[0087] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:

[0088] In the embodiments of the present application, a video-based instance tracking method is provided. Based on the target video frame in the video to be detected, the target feature map is obtained through the backbone network, and then according to N instance bounding boxes, N regions of interest (ROIs) of the bounding boxes are obtained from the target feature map. Then, based on the N instance query vectors and the N bounding box ROIs, N first detection results are obtained through the instance segmentation network. Based on this, at least one instance similarity is determined according to the N first detection results and the M second detection results. Finally, the instance tracking result of the target video frame can be determined according to at least one instance similarity. The instance tracking result includes at least one instance identifier, and the same instance identifier represents the same instance in the video to be detected. Through the above method, in the process of video-based instance tracking, the instance query vector can be used to directly detect the instances of interest in the video frame, and then the instances in the video frame are matched with the instances in the previous video frame in the video for similarity. Finally, the instance tracking of the video frame is realized according to the matching situation. The present application constructs an end-to-end instance detection framework, realizes instance detection without relying on post-processing methods such as NMS, and then tracks the instance target based on the instance identifier, thereby improving the efficiency of video-based instance tracking. Description of the Drawings

[0089] Figure 1It is an environmental schematic diagram of the instance tracking system in the embodiment of this application;

[0090] Figure 2 It is a schematic diagram of an end-to-end video instance segmentation framework based on instance query in the embodiment of this application;

[0091] Figure 3 It is a schematic diagram of the process of the video-based instance tracking method in the embodiment of this application;

[0092] Figure 4 It is a schematic diagram of the instance tracking result for the target video frame in the embodiment of this application;

[0093] Figure 5 It is another schematic diagram of the instance tracking result for the target video frame in the embodiment of this application;

[0094] Figure 6 It is a schematic diagram of implementing dynamic convolution in the embodiment of this application;

[0095] Figure 7 It is a schematic diagram of the bounding box ROI and the mask ROI in the embodiment of this application;

[0096] Figure 8 It is another schematic diagram of the instance tracking result for the target video frame in the embodiment of this application;

[0097] Figure 9 It is another schematic diagram of the instance tracking result for the target video frame in the embodiment of this application;

[0098] Figure 10 It is a schematic diagram of implementing dynamic convolution in the embodiment of this application;

[0099] Figure 11 It is a schematic diagram of implementing instance matching based on the bipartite graph matching algorithm in the embodiment of this application;

[0100] Figure 12 It is another schematic diagram of implementing instance matching based on the bipartite graph matching algorithm in the embodiment of this application;

[0101] Figure 13 It is a schematic diagram of the instance tracking device in the embodiment of this application;

[0102] Figure 14 It is a schematic diagram of the structure of the computer device in the embodiment of this application. Detailed implementation manners

[0103] The embodiments of the present application provide a video-based instance tracking method, device, equipment and storage medium. By constructing an end-to-end instance detection framework, the present application realizes instance detection without relying on post-processing methods such as NMS, and then tracks instance targets based on instance identifiers, thereby improving the efficiency of video-based instance tracking.

[0104] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0105] With the continuous development of computer technology, image processing technology based on artificial intelligence has become increasingly mature. In the field of image recognition, neural networks usually need to learn image features to detect instance targets in images. Instance targets can be human faces, animals, objects, landscapes, buildings, subtitles, etc. Detection of instance targets usually requires a rough detection of the position of the instance target first to determine the features of the candidate rectangular box as the input for fine detection. Then, the specific position of the instance target can be determined through fine detection, and the category of the instance target can be classified to detect the instance target in the image according to the category and specific position of the instance target. In the field of video recognition, instance targets in each video frame of the video can be recognized, and then the recognized instance targets can be compared with the instance targets in other previous video frames to generate an instance tracking result for the video frame. Several applications of video instance detection and segmentation will be introduced below. In actual applications, other specific scenarios may also be involved, which will not be enumerated here.

[0106] I. Subtitle detection;

[0107] Based on video instance tracking, detect and track the target instances (such as subtitles) in the video, then extract the target instances with the same instance identifier, and use optical character recognition (OCR) technology to recognize the subtitle target instances.

[0108] II. Autopilot;

[0109] Based on video instance tracking, detect and track target instances in the video (such as obstacles like vehicles and pedestrians). For target instances with the same instance identifier, if the target instance is on the lane, control the vehicle to avoid it; if the target instance is not on the lane, control the vehicle to continue driving along the trajectory.

[0110] III. Replace the scene;

[0111] Based on video instance tracking, segment and track the target instance (such as Person A) in Video A, and then use the segmented target instance as the foreground and place it in Video B, thus realizing the replacement of the background.

[0112] In order to achieve more efficient instance tracking in the above scenarios, this application proposes a video-based instance tracking method, which is applied to Figure 1 the instance tracking system shown in the figure. As shown in the figure, the instance tracking system includes a terminal device, or the instance tracking system includes a terminal device and a server. Among them, the client is deployed on the terminal device. The client can run on the terminal device in the form of a browser, or can also run on the terminal device in the form of an independent application (APP), etc. For the specific presentation form of the client, no limitation is made here. The server involved in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a laptop computer, a handheld computer, a personal computer, a smart TV, a smart watch, a vehicle-mounted device, a wearable device, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and no limitation is made here in this application. The number of servers and terminal devices is also not limited. The solution provided by this application can be completed independently by the terminal device, or can be completed independently by the server, or can also be completed by the cooperation of the terminal device and the server. Regarding this, no specific limitation is made in this application.

[0113] Exemplarily, in one case, the user selects a video to be detected. Then, the terminal device can call the local network model to perform instance recognition on the video, and thus output and display the instance tracking results of each video frame.

[0114] Exemplarily, in another case, the user selects a video to be detected and uploads the video to the server. Then, the server can call the local network model to perform instance recognition on the video. Thus, the instance tracking results of each video frame are output, and the instance tracking results are fed back to the terminal device, and the terminal device displays the instance tracking results of each video frame.

[0115] It should be noted that the process of detecting and segmenting video frames involves computer vision technology (CV) and machine learning (ML). Among them, computer vision technology is a science that studies how to enable machines to "see". Further speaking, it refers to using cameras and computers to replace human eyes to perform machine vision such as object recognition, tracking, and measurement on targets, and further performing graphic processing to make the computer process into images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0116] Machine learning is a multi-disciplinary cross-cutting discipline that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0117] Both computer vision technology and machine learning belong to artificial intelligence (AI) technology. Among them, artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0118] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0119] The following will introduce the Query Video Instance Segmentation (QueryVIS) framework used in the training process. Please refer to Figure 2 , Figure 2 This is a schematic diagram of an end-to-end video instance segmentation framework based on instance query in an embodiment of this application. As shown in the figure, in the training stage, the QueryVIS framework uses sample pairs (i.e., reference video frames and auxiliary video frames) sampled from the same video to be trained as input. Then, the backbone network to be trained extracts features from the reference video frames and auxiliary video frames, and then extracts the Region of Interest (ROI) features from the extracted features. The instance query vector directly enhances the ROI features and then inputs them into the task heads of different tasks to perform discrimination on relevant tasks.

[0120] Combined with the above introduction, the following will describe the video-based instance tracking method in this application. Please refer to Figure 3 An embodiment of the instance tracking method in an embodiment of this application includes:

[0121] 110. Based on the target video frame in the video to be detected, obtain the target feature map through the backbone network, where the target video frame is the T-th video frame in the video to be detected, and T is an integer greater than 1;

[0122] In one or more embodiments, the instance tracking device obtains the target video frame from the video to be detected, where the target video frame is the T-th frame in the video to be detected. Then, the target video frame is input into the backbone network, and thus the target feature map corresponding to the target video frame is obtained. Among them, the backbone network belongs to a feature extraction network, and its network parameters can be initialized with the weights pre-trained on the MS-COCO instance segmentation dataset, and the dataset used is the COCO example segmentation dataset.

[0123] It should be noted that the backbone network can adopt a residual network (ResNet), such as ResNet50, ResNet101, or ResNeXt101, etc. Optionally, the backbone network can adopt a hierarchical vision transformer using shifted windows (SwinTransformer), such as SwinTransformer Tiny, SwinTransformer Small, SwinTransformer Bas, or SwinTransformer Large, etc.

[0124] It should be noted that the instance tracking device can be deployed on a terminal device, or the instance tracking device can be deployed on a server, or the instance tracking device can be deployed on an instance tracking system composed of a terminal device and a server, which is not limited here.

[0125] 120. Obtain N regions of interest (ROIs) of bounding boxes from the target feature map according to N instance bounding boxes, where each instance bounding box is used to extract a corresponding bounding box ROI, and N is an integer greater than or equal to 1.

[0126] In one or more embodiments, the instance tracking device uses N trained instance bounding boxes to extract corresponding N bounding box ROIs from the target feature map, where ROI refers to the key area that the algorithm focuses on. Each instance bounding box is used to extract a corresponding bounding box ROI, and each bounding box ROI usually has a fixed resolution. For example, each bounding box ROI is 7×7 in size.

[0127] It should be noted that N can be an integer greater than or equal to 1. In this application, N can be set to 300. Optionally, N can also be set to 100 or other values, which is not limited here.

[0128] 130. Based on N instance query vectors and N bounding box ROIs, obtain N first detection results through an instance segmentation network, where each first detection result includes a first class probability value, a first instance bounding box, and a first instance embedding vector.

[0129] In one or more embodiments, there is a one-to-one correspondence between the instance bounding box and the instance query vector. Based on this, the instance tracking device performs convolution processing on the bounding box ROI using the instance query vector, and then inputs the convolved bounding box ROI into the instance segmentation network. The instance segmentation network respectively outputs the first detection results corresponding to each convolved bounding box ROI, that is, N first detection results are obtained. Each first detection result includes a first class probability value, a first instance bounding box, and a first instance embedding vector. The first class probability value is the maximum probability value in the probability distribution. The first instance bounding box is represented as (x1, y1, x2, y2), where x1 and y1 can be the coordinates of the upper left vertex of the first instance bounding box, and x2 and y2 can be the coordinates of the lower right vertex of the first instance bounding box. The first instance embedding vector can be represented as a 1*256 vector.

[0130] 140. Determine at least one instance similarity according to the N first detection results and the M second detection results, where each second detection result includes a second class probability value, a second instance bounding box, and a second instance embedding vector, and the M second detection results are obtained according to the first (T - 1) video frames in the video to be detected. Each second detection result corresponds to an instance identifier, and M is an integer greater than or equal to 1;

[0131] In one or more embodiments, the instance tracking device matches the N first detection results with the M second detection results, where the M second detection results are detected from the first (T - 1) video frames in the video to be detected. Assume T is 100, then based on steps 110 to 130, M second detection results can be detected from 99 video frames. Each second detection result includes a second class probability value, a second instance bounding box, and a second instance embedding vector. Thus, the instance similarity between the first detection result and the second detection result can be calculated.

[0132] 150. Determine the instance tracking result of the target video frame according to at least one instance similarity, where the instance tracking result includes at least one instance identifier, and the same instance identifier represents the same instance in the video to be detected.

[0133] In one or more embodiments, after obtaining at least one instance similarity, the instance tracking device can use a matching algorithm (for example, the nearest neighbor matching algorithm or the bipartite graph matching algorithm, etc.) to match the first detection results and the second detection results. Assume that the first detection result A matches the second detection result B successfully, and the first detection result C matches the second detection result D successfully (that is, the tracking of the instance target is achieved). Thus, it can be determined that the instance tracking result of the target video frame includes two instance identifiers, that is, the instance identifier corresponding to the second detection result B and the instance identifier corresponding to the second detection result D.

[0134] In an embodiment of the present application, a video-based instance tracking method is provided. Through the above method, in the process of video-based instance tracking, the instance query vector can be used to directly detect the instances of interest in the video frame, and then the instances in the video frame are matched with the instances in the previous video frame in the video for similarity. Finally, the instance tracking of the video frame is realized according to the matching situation. The present application constructs an end-to-end instance detection framework, realizes instance detection without relying on post-processing methods such as NMS, and then tracks the instance target based on the instance identifier, thereby improving the efficiency of video-based instance tracking.

[0135] Optionally, Figure 3 Based on the corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, based on N instance query vectors and N bounding box ROIs, N first detection results are obtained through an instance segmentation network, which may specifically include:

[0136] Using N instance query vectors, perform dot product on each of the N bounding box ROIs in the feature dimension to obtain N enhanced bounding box ROIs;

[0137] Based on the N enhanced bounding box ROIs, obtain N first class probability values through the class discrimination network included in the instance segmentation network, where the N first class probability values are included in the N first detection results;

[0138] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0139] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

[0140] In one or more embodiments, a method for realizing instance target tracking based on an instance segmentation network is introduced. For each bounding box ROI, the multi-head attention mechanism can also be used to extract features from the entire bounding box ROI, and better model effects can be obtained under a large range of receptive fields.

[0141] Specifically, for easy understanding, please refer to Figure 4 , Figure 4This is a schematic diagram of an instance tracking result for a target video frame in an embodiment of the present application. As shown in the figure, the target video frame is input into the backbone network to extract the target feature map. Then, N instance bounding boxes are used to perform an alignment operation on the ROI. For example, the RoIAlign operation is used to extract the bounding box ROI, and thus N bounding box ROIs are obtained. The instance query vector is used to perform a dot product on the bounding box ROI in the feature dimension to obtain an enhanced bounding box ROI. Since there is a one-to-one correspondence between the instance query vector and the instance bounding box, N enhanced bounding box ROIs can be obtained.

[0142] It should be noted that the instance segmentation network includes a class discrimination network, a bounding box regression network, and an embedding vector network. Based on this, the N enhanced bounding box ROIs are input into the class discrimination network, and the class discrimination network outputs the first class probability value corresponding to each enhanced bounding box ROI. The N enhanced bounding box ROIs are input into the bounding box regression network, and the bounding box regression network outputs the first instance bounding box corresponding to each enhanced bounding box ROI. The N enhanced bounding box ROIs are input into the embedding vector network, and the embedding vector network outputs the first instance embedding vector corresponding to each enhanced bounding box ROI.

[0143] Based on this, the N first detection results predicted for the target video frame are matched with the M second detection results predicted for the previous (T - 1) video frames, and the instance tracking result of the target video frame is generated based on the matching result. In addition, the instance tracking result can also be displayed. For example, it is displayed that the instance type is "car", the position of the instance bounding box, and the instance identifier is "101". It should be noted that the same instance identifier represents the same instance in the video to be detected.

[0144] Secondly, in the embodiment of the present application, a method for implementing instance object tracking based on an instance segmentation network is provided. Through the above method, the detection and tracking of instance objects in a video can be achieved without extracting the instance foreground mask, and an end-to-end video instance segmentation framework is designed, thereby simplifying the complexity of the video instance segmentation model and facilitating the improvement of the model inference speed.

[0145] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiment of the present application based on the corresponding various embodiments, based on the N instance query vectors and the N bounding box ROIs, N first detection results are obtained through the instance segmentation network, which may specifically include:

[0146] Based on the N instance query vectors, at least one set of bounding box dynamic parameters is obtained through a fully connected layer;

[0147] Using at least one set of bounding box dynamic parameters, perform a dot product on each of the N bounding box ROIs in the feature dimension to obtain N enhanced bounding box ROIs;

[0148] Based on the N enhanced bounding box ROIs, obtain N first-class probability values through the class discrimination network included in the instance segmentation network, where the N first-class probability values are included in the N first detection results;

[0149] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0150] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

[0151] In one or more embodiments, a method for implementing instance object tracking based on an instance segmentation network is introduced. For each bounding box ROI, a multi-head attention mechanism can also be used to extract features of the entire bounding box ROI, and better model effects can be obtained under a large range of receptive fields.

[0152] Specifically, for ease of understanding, please refer to Figure 5 , Figure 5 FIG. is another schematic diagram of an instance tracking result for a target video frame in an embodiment of the present application. As shown in the figure, the target video frame is input into the backbone network to extract the target feature map. Then, N instance bounding boxes are used to perform an alignment operation on the ROI. For example, the RoIAlign operation is used to extract the bounding box ROI, and thus N bounding box ROIs are obtained. An instance query vector is used to generate dynamic parameters, and then the dynamic parameters are used to perform a dot product on the bounding box ROI in the feature dimension to obtain an enhanced bounding box ROI. Since there is a one-to-one correspondence between the instance query vector and the instance bounding box, N enhanced bounding box ROIs can be obtained.

[0153] It should be noted that the instance segmentation network includes a class discrimination network, a bounding box regression network, and an embedding vector network. Based on this, the N enhanced bounding box ROIs are input into the class discrimination network, and the class discrimination network outputs the first-class probability value corresponding to each enhanced bounding box ROI. The N enhanced bounding box ROIs are input into the bounding box regression network, and the bounding box regression network outputs the first instance bounding box corresponding to each enhanced bounding box ROI. The N enhanced bounding box ROIs are input into the embedding vector network, and the embedding vector network outputs the first instance embedding vector corresponding to each enhanced bounding box ROI.

[0154] Based on this, N first detection results predicted from the target video frame are matched with M second detection results predicted from the previous (T - 1) video frames, and an instance tracking result of the target video frame is generated based on the matching result. In addition, the instance tracking result can also be displayed.

[0155] Next, in conjunction with Figure 6 , the process of dynamic convolution will be introduced. Here, a bounding box ROI will be used as an example for introduction. It can be understood that a similar method is adopted for all N bounding box ROIs, and details will not be elaborated here. Please refer to Figure 6 , Figure 6 FIG.

[0156] It should be noted that in practical applications, 1 dynamic convolution or more than 2 dynamic convolutions can also be performed. In this application, 2 dynamic convolutions are taken as an example for introduction. However, this should not be construed as a limitation to this application.

[0157] Secondly, in the embodiments of the present application, a method for implementing instance object tracking based on an instance segmentation network is provided. Through the above method, the detection and tracking of instance objects in a video can be realized without extracting an instance foreground mask, and an end-to-end video instance segmentation framework is designed, thereby simplifying the complexity of the video instance segmentation model and facilitating the improvement of the model inference speed. At the same time, after ROI feature extraction, dynamic convolution is performed between the instance query vector and the ROI feature, thereby generating enhanced instance features, which is beneficial to obtaining better model effects.

[0158] Optionally, based on the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, it may further include:

[0159] According to N instance bounding boxes, N masked ROIs are obtained from the target video frame, where each instance bounding box is also used to extract a corresponding masked ROI;

[0160] Based on N instance query vectors and N bounding box ROIs, N first detection results are obtained through an instance segmentation network, which may specifically include:

[0161] Based on N instance query vectors, N bounding box ROIs, and N mask ROIs, N first detection results are obtained through an instance segmentation network, where each first detection result further includes a first instance foreground mask.

[0162] In one or more embodiments, a method for extracting mask ROIs for tracking is introduced. The instance tracking device uses N instance bounding boxes and can also extract N mask ROIs from the target video frame, that is, each instance bounding box can not only extract a corresponding bounding box ROI but also provide a corresponding mask ROI. Among them, the size of the mask ROI is larger than that of the bounding box ROI.

[0163] Specifically, N instance query vectors are used to perform convolution on N bounding box ROIs and N mask ROIs respectively, and based on the convolution results, instance segmentation is further performed to obtain N first instance foreground masks. For ease of understanding, please refer to Figure 7 , Figure 7 which is a schematic diagram of the bounding box ROI and the mask ROI in the embodiments of the present application. As shown in Figure (A) of Figure 7 , taking the instance target as "car" as an example, its corresponding bounding box ROI is the minimum bounding box for the instance target. As shown in Figure (B) of Figure 7 , taking the instance target as "car" as an example, its corresponding mask ROI is the foreground segmentation result for the instance target.

[0164] Secondly, in the embodiments of the present application, a method for extracting mask ROIs for tracking is provided. Through the above method, the instance targets in the video can be further segmented, which improves the single-stage instance segmentation network and extends it to the field of video instance segmentation, not only reducing the number of model parameters but also improving the segmentation fineness.

[0165] Optionally, on the basis of the corresponding embodiments above Figure 3 , in another optional embodiment provided by the embodiments of the present application, based on N instance query vectors, N bounding box ROIs, and N mask ROIs, N first detection results are obtained through an instance segmentation network, which may specifically include:

[0166] Using N instance query vectors to perform dot multiplication on each of the N bounding box ROIs in the feature dimension to obtain N enhanced bounding box ROIs;

[0167] Using N instance query vectors, perform a dot product on each of the N masked ROIs in the feature dimension to obtain N enhanced masked ROIs;

[0168] Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results;

[0169] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0170] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results;

[0171] Based on the N enhanced masked ROIs, obtain N first instance foreground masks through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

[0172] In one or more embodiments, a method for implementing instance object tracking based on an instance segmentation network is introduced. For each bounding box ROI and each masked ROI, a multi-head attention mechanism can also be used to extract features from the entire bounding box ROI and the entire masked ROI respectively. In a large receptive field, better model performance can be achieved.

[0173] Specifically, for ease of understanding, please refer to Figure 8 , Figure 8 is another schematic diagram of an instance tracking result for the target video frame in the embodiments of the present application. As shown in the figure, the target video frame is input into the backbone network to extract the target feature map. Then, N instance bounding boxes are used to perform an alignment operation on the ROI. For example, the RoIAlign operation is used to extract the bounding box ROI and the masked ROI, thereby obtaining N bounding box ROIs and N masked ROIs. A dot product is performed on the bounding box ROI in the feature dimension using the instance query vector to obtain the enhanced bounding box ROI. Similarly, a dot product is performed on the masked ROI in the feature dimension using the instance query vector to obtain the enhanced masked ROI. Since the instance query vector and the instance bounding box have a one-to-one correspondence, N enhanced bounding box ROIs and N enhanced bounding box ROIs are obtained.

[0174] It should be noted that the instance segmentation network includes a class discrimination network, a bounding box regression network, an embedding vector network, and a mask generation network. Based on this, N enhanced bounding box ROIs are input into the class discrimination network, and the class discrimination network outputs the first class probability value corresponding to each enhanced bounding box ROI. The N enhanced bounding box ROIs are input into the bounding box regression network, and the bounding box regression network outputs the first instance bounding box corresponding to each enhanced bounding box ROI. The N enhanced bounding box ROIs are input into the embedding vector network, and the embedding vector network outputs the first instance embedding vector corresponding to each enhanced bounding box ROI. The N enhanced mask ROIs are input into the mask generation network, and the mask generation network outputs the first instance foreground mask corresponding to each enhanced mask ROI.

[0175] Based on this, the N first detection results predicted from the target video frame are matched with the M second detection results predicted from the previous (T - 1) video frames, and an instance tracking result of the target video frame is generated based on the matching result. In addition, the instance tracking result can also be displayed.

[0176] Again, in the embodiments of the present application, a method for implementing instance target tracking based on an instance segmentation network is provided. Through the above method, the detection and tracking of instance targets in a video can be achieved when extracting the instance foreground mask, and an end-to-end video instance segmentation framework is designed, thereby simplifying the complexity of the video instance segmentation model and facilitating the improvement of the model inference speed.

[0177] Optionally, on the basis of the above Figure 3 corresponding respective embodiments, in another optional embodiment provided by the embodiments of the present application, based on N instance query vectors, N bounding box ROIs, and N mask ROIs, N first detection results are obtained through the instance segmentation network, which may specifically include:

[0178] Based on the N instance query vectors, at least one set of bounding box dynamic parameters and at least one set of mask dynamic parameters are obtained through a fully connected layer, and each set of mask dynamic parameters includes N mask dynamic sub-parameters;

[0179] Using at least one set of bounding box dynamic parameters, each bounding box ROI among the N bounding box ROIs is dot-multiplied in the feature dimension to obtain N enhanced bounding box ROIs;

[0180] Using at least one set of mask dynamic parameters, each mask ROI among the N mask ROIs is dot-multiplied in the feature dimension to obtain N enhanced mask ROIs;

[0181] Based on the N enhanced bounding box ROIs, N first class probability values are obtained through the class discrimination network included in the instance segmentation network, where the N first class probability values are included in the N first detection results;

[0182] Based on N enhanced bounding box ROIs, N first instance bounding boxes are obtained through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0183] Based on N enhanced bounding box ROIs, N first instance embedding vectors are obtained through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results;

[0184] Based on N enhanced mask ROIs, N first instance foreground masks are obtained through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

[0185] In one or more embodiments, a method for implementing instance object tracking based on an instance segmentation network is introduced. For each bounding box ROI and each mask ROI, a multi-head attention mechanism can also be used to extract features from the entire bounding box ROI and the entire mask ROI respectively. In a large range of receptive fields, better model effects can be achieved.

[0186] Specifically, for ease of understanding, please refer to Figure 9 , Figure 9 which is another schematic diagram of the instance tracking result for the target video frame in the embodiment of the present application. As shown in the figure, the target video frame is input into the backbone network to extract the target feature map. Then, N instance bounding boxes are used to perform an alignment operation on the ROI. For example, the RoIAlign operation is used to extract the bounding box ROI and the mask ROI, thereby obtaining N bounding box ROIs and N mask ROIs. An instance query vector is used to generate a set of dynamic parameters, and then the set of dynamic parameters is used to perform a dot product on the bounding box ROI in the feature dimension to obtain the enhanced bounding box ROI. Since there is a one-to-one correspondence between the instance query vector and the instance bounding box, N enhanced bounding box ROIs can be obtained. Similarly, an instance query vector is used to generate another set of dynamic parameters, and then the set of dynamic parameters is used to perform a dot product on the mask ROI in the feature dimension to obtain the enhanced mask ROI. Since there is a one-to-one correspondence between the instance query vector and the instance bounding box, N enhanced mask ROIs can be obtained.

[0187] It should be noted that the instance segmentation network includes a class discrimination network, a bounding box regression network, an embedding vector network, and a mask generation network. Based on this, N enhanced bounding box ROIs are input into the class discrimination network, and the class discrimination network outputs the first class probability value corresponding to each enhanced bounding box ROI. The N enhanced bounding box ROIs are input into the bounding box regression network, and the bounding box regression network outputs the first instance bounding box corresponding to each enhanced bounding box ROI. The N enhanced bounding box ROIs are input into the embedding vector network, and the embedding vector network outputs the first instance embedding vector corresponding to each enhanced bounding box ROI. The N enhanced mask ROIs are input into the mask generation network, and the mask generation network outputs the first instance foreground mask corresponding to each enhanced mask ROI.

[0188] Based on this, the N first detection results predicted from the target video frame are matched with the M second detection results predicted from the previous (T - 1) video frames, and the instance tracking result of the target video frame is generated based on the matching result. In addition, the instance tracking result can also be displayed.

[0189] The following will be combined with Figure 10 , to introduce the process of dynamic convolution. The following will take a bounding box ROI and a mask ROI as an example for introduction. It can be understood that similar methods are used for processing N bounding box ROIs and N mask ROIs, which will not be elaborated here. Please refer to Figure 10 , Figure 10 is a schematic diagram for implementing dynamic convolution in an embodiment of this application. As shown in the figure, after the bounding box ROI is extracted, the instance query vector performs dynamic convolution with the features of the bounding box ROI. Taking two dynamic convolutions as an example, the instance query vector is input into the fully connected layer 1, and the fully connected layer 1 outputs two sets of bounding box dynamic parameters. Among them, one set of bounding box dynamic parameters includes N dynamic parameters A, and the other set of bounding box dynamic parameters includes N dynamic parameters B. Based on this, the bounding box ROI is multiplied point - by - point (i.e., a 1 * 1 convolution operation) using the dynamic parameter A, and then the dynamically convolved bounding box ROI is multiplied point - by - point (i.e., a 1 * 1 convolution operation) using the dynamic parameter B, thereby obtaining the enhanced bounding box ROI.

[0190] Similarly, after the mask ROI is extracted, the instance query vector performs dynamic convolution with the features of the mask ROI. Taking two dynamic convolutions as an example, the instance query vector is input into the fully connected layer 2, and the fully connected layer 2 outputs two sets of bounding box dynamic parameters. Among them, one set of bounding box dynamic parameters includes N dynamic parameters C, and the other set of bounding box dynamic parameters includes N dynamic parameters D. Based on this, the mask ROI is multiplied point - by - point (i.e., a 1 * 1 convolution operation) using the dynamic parameter C, and then the dynamically convolved mask ROI is multiplied point - by - point (i.e., a 1 * 1 convolution operation) using the dynamic parameter D, thereby obtaining the enhanced mask ROI.

[0191] It should be noted that in practical applications, 1 dynamic convolution can also be performed, or more than 2 dynamic convolutions can be performed. In this application, 2 dynamic convolutions are taken as an example for introduction. However, this should not be construed as a limitation to this application.

[0192] Furthermore, in the embodiments of this application, a method for implementing instance object tracking based on an instance segmentation network is provided. Through the above method, the detection and tracking of instance objects in a video can be realized when extracting the instance foreground mask, and an end-to-end video instance segmentation framework is designed, thereby simplifying the complexity of the video instance segmentation model and facilitating the improvement of the model inference speed. At the same time, after ROI feature extraction, dynamic convolution is performed between the instance query vector and the ROI feature to generate enhanced instance features, which is conducive to achieving better model effects.

[0193] Optionally, based on the corresponding respective embodiments above, in another optional embodiment provided by the embodiments of this application, according to N first detection results and M second detection results, at least one instance similarity is determined, which may specifically include: Figure 3 Sort the N first detection results in descending order according to the first category probability value to obtain the sorted N first detection results;

[0194] Select the first K first detection results from the sorted N first detection results, where K is an integer greater than or equal to 1 and less than or equal to N;

[0195] According to the K first detection results and the M second detection results, determine the instance similarity between each first detection result and each second detection result to obtain (K * M) instance similarities.

[0196] In one or more embodiments, a method for screening out K first detection results for instance matching is introduced. To improve the instance matching effect, an online instance connection method can also be adopted, and at the same time, bidirectional softmax and instance similarity are used to further improve the effect of online instance connection. Since the calculated (K * M) instance similarities are represented in matrix form, bidirectional softmax can be used to normalize the (K * M) instance similarities.

[0197]

[0198] ​Specifically, for each video frame in the video to be detected, the corresponding detection results (including class probability values, instance bounding boxes, and instance embedding vectors) can be output through QueryVIS, and then the detection results of each frame are stored in the candidate pool. Moreover, when the target video frame is detected, there are already M second detection results in the candidate pool. Therefore, it is necessary to calculate the similarity between the K first detection results for the target video frame and the M second detection results already in the candidate pool. Next, how to select K first detection results from N first detection results will be introduced.

[0199] Exemplarily, assume that N is 300 and K is 10. Based on this, for the first class probability value included in each of the 300 first detection results, the N first detection results are re-sorted in descending order of the first class probability value, thereby obtaining the sorted N first detection results. Then, select the top K first detection results arranged at the forefront from the sorted N first detection results. For example, obtain 10 first detection results. Finally, calculate the pairwise similarity between the K first detection results and the M second detection results, so as to obtain (K * M) instance similarities. Among them, the (K * M) instance similarities can be presented in the form of a similarity matrix.

[0200] Secondly, in the embodiments of the present application, a method for selecting K first detection results for instance matching is provided. Through the above method, considering that matching all N first detection results will not only consume more resources but also lead to a lower efficiency of instance target tracking. Therefore, the top K first detection results with the largest class probability values are selected from the N first detection results, and the K first detection results are used for subsequent matching, thereby improving the matching efficiency and the matching accuracy at the same time.

[0201] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, according to the K first detection results and the M second detection results, determine the instance similarity between each first detection result and each second detection result, which may specifically include:

[0202] Determine the instance embedding vector similarity between each first detection result and each second detection result according to the first instance embedding vector included in each first detection result and the second instance embedding vector included in each second detection result;

[0203] Determine the spatial similarity between each first detection result and each second detection result according to the first instance bounding box included in each first detection result and the second instance bounding box included in each second detection result;

[0204] Determine the class similarity between each first detection result and each second detection result according to the first class probability value included in each first detection result and the second class probability value included in each second detection result;

[0205] Determine the instance similarity between each first detection result and each second detection result according to the instance embedding vector similarity, spatial similarity, class similarity between each first detection result and each second detection result, and the first class probability value included in each first detection result.

[0206] In one or more embodiments, a method for calculating instance similarity is introduced. As can be seen from the foregoing embodiments, based on K first detection results and M second detection results, (K * M) instance similarities can be calculated. For the sake of convenience of introduction, hereinafter, the example of calculating the instance similarity between one first detection result and one second detection result will be used for illustration, and the instance similarities between other detection results are also calculated in a similar manner, which will not be elaborated here.

[0207] Specifically, the present application can calculate the instance similarity in the following manner:

[0208] Similarity = A * B * C * D; Equation 1

[0209] Among them, Similarity represents the instance similarity between two detection results, A represents the instance embedding vector similarity between two detection results, B represents the spatial similarity between two detection results, C represents the class similarity between two detection results, and D represents the first class probability value included in the first detection result.

[0210] Next, the methods for calculating the instance embedding vector similarity, spatial similarity, and class similarity will be introduced in sequence.

[0211] I. Instance embedding vector similarity;

[0212] The cosine similarity can be used to calculate the instance embedding vector similarity between the first instance embedding vector and the second instance embedding vector. Alternatively, the inner product of the first instance embedding vector and the second instance embedding vector can be used as the instance embedding vector similarity. Or, other similarity algorithms can be used to calculate the instance embedding vector similarity.

[0213] II. Spatial similarity;

[0214] The spatial correlation is the intersection over union (IOU) value of the first instance bounding box and the second instance bounding box.

[0215] III. Class similarity;

[0216] The class similarity can ensure that the instance connection process is only carried out between instances of the same class. That is, the first class can be determined according to the first class probability value. For example, the first class probability value is 0.9, and the first class corresponding to 0.9 is "puppy". The second class can be determined according to the second class probability value. For example, the second class probability value is 0.8, and the second class corresponding to 0.8 is "kitten". Then the first class and the second class are different. Therefore, the class similarity between the two is "0".

[0217] For another example, the first class probability value is 0.9, and the first class corresponding to 0.9 is "puppy". The second class can be determined according to the second class probability value. For example, the second class probability value is 0.8, and the second class corresponding to 0.8 is "puppy". Then the first class and the second class are the same. Therefore, the class similarity between the two is "1".

[0218] Furthermore, in the embodiments of the present application, a method for calculating instance similarity is provided. Through the above method, the instance similarity between two detection results can be directly calculated by using the product of the instance embedding vector similarity, the spatial similarity, the class similarity, and the class probability value. Therefore, there is no need to train other weight parameters, which not only reduces the computational complexity but also can reduce the amount of parameter training. At the same time, the (K * M) instance similarities can be normalized by bidirectional softmax to ensure that the value range of the probability is between 0 and 1, thereby improving the calculation accuracy and accelerating the convergence speed.

[0219] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, K is less than or equal to M;

[0220] Determine the instance tracking result of the target video frame according to at least one instance similarity, which may specifically include:

[0221] Based on the bipartite graph matching algorithm, according to the (K * M) instance similarities, construct the mapping relationship between the K first detection results and the M second detection results to obtain K mapping relationships;

[0222] If there are P mapping relationships among the K mapping relationships whose corresponding instance similarities are less than or equal to the instance similarity threshold, then delete the P mapping relationships from the K mapping relationships to obtain (K - P) mapping relationships, where P is an integer greater than or equal to 1 and less than or equal to K;

[0223] Determine the instance tracking result of the target video frame according to each second detection result corresponding to each mapping relationship in the (K - P) mapping relationships, where each second detection result corresponds to an instance identifier;

[0224] It may further include:

[0225] Use the first detection results corresponding to P mapping relationships as the second detection results, obtaining (M + P) second detection results.

[0226] In one or more embodiments, a method for matching K first detection results with M second detection results is introduced. As can be seen from the foregoing embodiments, the value of K can be an integer less than or equal to the value of M, and an instance similarity is calculated for each first detection result and the second detection result. Based on the (K * M) instance similarities, an optimal matching result can be calculated using the bipartite graph matching algorithm, and under the optimal matching result, the sum of the instance similarities corresponding to the matching results is the smallest. Based on this, mapping relationships between K first detection results and M second detection results can be constructed, thereby obtaining K mapping relationships.

[0227] Specifically, by way of example, for ease of understanding, please refer to Figure 11 , Figure 11 FIG. 11 is a schematic diagram of implementing instance matching based on the bipartite graph matching algorithm in an embodiment of the present application. As shown in the figure, assume that the K first detection results include the first detection result A, the first detection result B, and the first detection result C, and the M second detection results include the second detection result V, the second detection result W, the second detection result X, the second detection result Y, and the second detection result Z. Assume that the K mapping relationships obtained after matching are shown in Table 1.

[0228] Table 1

[0229] Mapping relationship Instance similarity A-X 0.9 B-V 0.3 C-Y 0.8

[0230] It can be seen that when K is 3 and M is 5, 3 mapping relationships can be obtained, and each mapping relationship has a corresponding instance similarity. Assume that the instance similarity threshold is 0.4. Then, only the instance similarity corresponding to the "B - Y" mapping relationship is less than the instance similarity threshold (i.e., P is 1). Therefore, this mapping relationship will be excluded from the K mapping relationships, and the remaining 2 mapping relationships (i.e., the remaining K - P mapping relationships) will be obtained.

[0231] Based on this, use the instance identifier corresponding to the second detection result X as the instance identifier corresponding to the first detection result A, and use the instance identifier corresponding to the second detection result Y as the instance identifier corresponding to the first detection result C. Thus, the instance tracking result of the target video frame is obtained.

[0232] Further, since the P mapping relationships fail to match successfully, the first detection results corresponding to the P mapping relationships can be added to the candidate pool as the second detection results for subsequent matching. For example, the first detection result B is used as the second detection result B and added to the candidate pool. Therefore, the candidate pool includes the original M second detection results and the P newly added second detection results, and thus, (M + P) second detection results are obtained.

[0233] It should be noted that the bipartite graph matching algorithm adopted in this application can be the Hungarian algorithm, or the maximum matching algorithm of the bipartite graph. Or the perfect matching algorithm, etc., which is not limited here.

[0234] Again, in the embodiments of the present application, a method for matching K first detection results with M second detection results is provided. Through the above method, when K is less than or equal to M, the bipartite graph matching algorithm is used to preferentially match the K first detection results, so as to obtain a result with better overall matching degree, thereby improving the accuracy and consistency of instance object tracking.

[0235] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding respective embodiments, K is greater than or equal to M;

[0236] Determining the instance tracking result of the target video frame according to at least one instance similarity may specifically include:

[0237] Based on the bipartite graph matching algorithm, according to the (K * M) instance similarities, construct the mapping relationships between the K first detection results and the M second detection results to obtain M mapping relationships;

[0238] If there are Q mapping relationships among the M mapping relationships whose corresponding instance similarities are less than or equal to the instance similarity threshold, then delete the Q mapping relationships from the M mapping relationships to obtain (M - Q) mapping relationships, where Q is an integer greater than or equal to 1 and less than or equal to M;

[0239] Determine the instance tracking result of the target video frame according to the second detection result corresponding to each mapping relationship among the (M - Q) mapping relationships, where each second detection result corresponds to an instance identifier;

[0240] It may further include:

[0241] Use the Q mapping relationships and the (K - M) first detection results as the second detection results to obtain (Q + K) second detection results.

[0242] In one or more embodiments, a method for matching K first detection results with M second detection results is introduced. As can be seen from the foregoing embodiments, the value of K can be an integer greater than or equal to the value of M, and an instance similarity is calculated for each first detection result and the second detection result. Based on the (K * M) instance similarities, an optimal matching result can be calculated based on the bipartite graph matching algorithm, and under the optimal matching result, the sum of the instance similarities corresponding to the matching result is the smallest. Based on this, a mapping relationship between the K first detection results and the M second detection results can be constructed, and M mapping relationships are obtained therefrom.

[0243] Specifically, by way of example, for ease of understanding, please refer to Figure 12 , Figure 12 FIG. is another schematic diagram of implementing instance matching based on the bipartite graph matching algorithm in the embodiments of the present application. As shown in the figure, it is assumed that the K first detection results include the first detection result A, the first detection result B, the first detection result C, the first detection result D, and the first detection result E, and the M second detection results include the second detection result X, the second detection result Y, and the second detection result Z. It is assumed that the M mapping relationships obtained after matching are shown in Table 2.

[0244] Table 2

[0245] Mapping relationship Instance similarity A-Z 0.7 B-Y 0.9 D-X 0.1

[0246] It can be seen that in the case where K is 5 and M is 3, 3 mapping relationships can be obtained, and each mapping relationship has a corresponding instance similarity. Assuming that the instance similarity threshold is 0.4, then only the instance similarity corresponding to the "D - X" mapping relationship is less than the instance similarity threshold (i.e., Q is 1). Therefore, this mapping relationship will be excluded from the M mapping relationships, and the remaining 2 mapping relationships are obtained (i.e., the remaining M - Q mapping relationships).

[0247] Based on this, the instance identifier corresponding to the second detection result Z is used as the instance identifier corresponding to the first detection result A, and the instance identifier corresponding to the second detection result Y is used as the instance identifier corresponding to the first detection result B. Thus, the instance tracking result of the target video frame is obtained.

[0248] Further, since the Q mapping relationships fail to match successfully, the first detection results corresponding to the Q mapping relationships can be added to the candidate pool as the second detection results for subsequent matching. For example, the first detection result D is used as the second detection result B and added to the candidate pool. In addition, the first detection results C and E also fail to match successfully. Therefore, the remaining (K - M) first detection results can be added to the candidate pool as the second detection results for subsequent matching. Based on this, the candidate pool includes the original M second detection results, Q newly added second detection results, and (K - M) newly added second detection results. Thus, (Q + K) (i.e., M + Q + K - M) second detection results are obtained.

[0249] It should be noted that the bipartite graph matching algorithm adopted in this application can be the Hungarian algorithm, or the maximum matching algorithm of the bipartite graph, or the perfect matching algorithm, etc., which is not limited here.

[0250] Again, in the embodiments of this application, a method for matching K first detection results with M second detection results is provided. Through the above method, when K is greater than or equal to M, the bipartite graph matching algorithm is used to preferentially match the M second detection results, so as to obtain better overall matching results, thereby improving the accuracy and consistency of instance object tracking.

[0251] Optionally, based on the corresponding embodiments above Figure 3 In another optional embodiment provided by the embodiments of this application, it may further include:

[0252] Based on the first sample video frames in the video to be trained, sample feature maps are obtained through the backbone network to be trained, where the first sample video frames have labeled categories, labeled instance bounding boxes, and labeled instance identifiers;

[0253] According to N bounding boxes of instances to be trained, N predicted bounding boxes ROI are obtained from the sample feature maps, where each bounding box of an instance to be trained is used to extract a corresponding predicted bounding box ROI;

[0254] Based on N query vectors of instances to be trained and N predicted bounding boxes ROI, N first prediction results are obtained through the instance segmentation network to be trained, where each first prediction result includes a first predicted class probability value, a first predicted instance bounding box, and a first predicted instance embedding vector;

[0255] Determine at least one predicted instance similarity according to N first prediction results and N second prediction results, where each second prediction result includes a second predicted class probability value, a second predicted instance bounding box, and a second predicted instance embedding vector, and each second prediction result corresponds to an instance identifier. The N second prediction results are from second sample video frames in the video to be trained.

[0256] Determine the predicted instance tracking result of the first sample video frame according to at least one predicted instance similarity.

[0257] According to the predicted instance tracking result, the annotated instance identifier, the N first prediction results, the annotated class, and the annotated instance bounding box, update the parameters of the backbone network to be trained, the N instance bounding boxes to be trained, the N instance query vectors to be trained, and the instance segmentation network to be trained through a loss function.

[0258] In one or more embodiments, a method for training a backbone network, instance bounding boxes, instance query vectors, and an instance segmentation network is introduced. Based on this, the QueryVIS to be trained includes the backbone network to be trained, N instance bounding boxes to be trained, N instance query vectors to be trained, and the instance segmentation network to be trained. During training, sample video frames from the same video to be trained can be used. Then, each sample video frame can be scaled to a certain size. For example, the short side of the sample video frame is scaled to (320, 800) pixels, and the long side of the sample video frame is scaled to 1333. In addition, Adaptive Moment Estimation Weight Decay (AdamW) can be used as an optimizer to train on 8 graphics cards, where the number of frames used for one gradient descent is 32. During the training process, 12 full iterations are performed on the video instance segmentation dataset, and the initial learning rate is set to 0.000025. The learning rate is divided by 10 after the 8th and 11th full iterations respectively.

[0259] Specifically, for QueryVIS, there are N instance query vectors. When performing sample connection and loss function calculation, one instance query vector is responsible for one ground truth bounding box (i.e., the annotated instance bounding box). Based on this, in the training stage, after the instance query vector performs bounding box ROI prediction, by calculating the loss with the ground truth bounding box one by one, a two-dimensional loss matrix between the prediction of the instance query vector and the ground truth bounding box is obtained. Then, the bipartite graph matching algorithm is used to perform one-to-one matching between the instance query vector and the ground truth bounding box, and the matching result between the instance query and the ground truth bounding box is obtained. The instance query vector that matches the ground truth bounding box is responsible for predicting the class of the ground truth bounding box (i.e., the first predicted class probability value), the bounding box (i.e., the first predicted instance bounding box), and the instance embedding vector (i.e., the first predicted instance embedding vector). That is, the loss function is the instance-level loss between the prediction of this instance query vector and its corresponding ground truth bounding box. It should be noted that for the instance query vector that does not match the ground truth bounding box, the corresponding loss function only includes the classification loss where the ground truth bounding box is a negative sample.

[0260] The training process of QueryVIS is specifically as follows: Using the first sample video frame in the video to be trained as the input of the backbone network to be trained, the backbone network to be trained is used to extract features to obtain the sample feature map. Then, using N ground truth instance bounding boxes, N predicted bounding box ROIs are extracted from the sample feature map. Next, using N instance query vectors to be trained, dynamic convolution operations or ordinary convolution operations are performed on the N predicted bounding box ROIs to generate N enhanced predicted bounding box ROIs. The N enhanced predicted bounding box ROIs are used as the input of the instance segmentation network to be trained, and the instance segmentation network to be trained outputs N first predicted class probability values, N first predicted instance bounding boxes, and N first predicted instance embedding vectors.

[0261] It should be noted that since the training uses paired sample video frames, therefore, similar processing also needs to be performed on the second sample video frame to obtain N second predicted results. The second sample video frame comes from any frame in the video to be trained. Then, the predicted instance similarity between the first predicted result and the second predicted result can be calculated using the above formula 1, and at least one predicted instance similarity is obtained. Then, the predicted instance tracking result can be determined from the first sample video frame, that is, the instance identifier that the predicted instance tracking result can predict.

[0262] Design a loss function. Exemplarily, the Focal loss function can be adopted to calculate the loss value between N first predicted class probability values and the true annotation class. And use the number of positive samples as a scaling factor to scale the calculated loss function. Exemplarily, the generalized IoU (GIOU) loss function and the L1 loss function can be used to calculate the loss value between N first predicted instance bounding boxes and the true annotation instance bounding boxes. Exemplarily, classification loss can be adopted to calculate the loss value between the instance identifiers included in the predicted instance tracking result and the true annotation instance identifiers.

[0263] Sum or weighted sum the above loss values to obtain a comprehensive loss value. Based on the comprehensive loss value, use backpropagation and gradient descent algorithms to train QueryVIS, that is, update the parameters of the backbone network to be trained, N instance bounding boxes to be trained, N instance query vectors to be trained, and the instance segmentation network to be trained.

[0264] Secondly, in the embodiments of the present application, a method for training a backbone network, instance bounding boxes, instance query vectors, and instance segmentation network is provided. Through the above method, "one-to-one" sampled samples are connected with the true values to detect and segment instances of interest, and complete end-to-end training and inference can be performed.

[0265] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to the above respective embodiments, it may further include:

[0266] Based on the first sample video frame in the video to be trained, obtain a sample feature map through the backbone network to be trained, where the first sample video frame has an annotation class, an annotation instance bounding box, an annotation instance identifier, and an annotation instance foreground mask;

[0267] According to N instance bounding boxes to be trained, obtain N predicted bounding box ROIs from the sample feature map, where each instance bounding box to be trained is used to extract a corresponding predicted bounding box ROI;

[0268] According to N instance bounding boxes to be trained, obtain N predicted mask ROIs from the sample feature map, where each instance bounding box to be trained is also used to extract a corresponding predicted mask ROI;

[0269] Based on N instance query vectors to be trained, N predicted bounding box ROIs, and N predicted mask ROIs, obtain N first predicted results through the instance segmentation network to be trained, where each first predicted result includes a first predicted class probability value, a first predicted instance bounding box, a first predicted instance embedding vector, and a predicted instance foreground mask;

[0270] Determine at least one prediction instance similarity based on N first prediction results and N second prediction results, where each second prediction result includes a second predicted class probability value, a second predicted instance bounding box, and a second predicted instance embedding vector. The N second prediction results are from second sample video frames in the video to be trained, and each second prediction result corresponds to an instance identifier.

[0271] Determine the prediction instance tracking result of the first sample video frame based on at least one prediction instance similarity.

[0272] Update the parameters of the backbone network to be trained, N instance bounding boxes to be trained, N instance query vectors to be trained, and the instance segmentation network to be trained through a loss function according to the prediction instance tracking result, the labeled instance identifier, N first prediction results, the labeled class, the labeled instance bounding box, and the labeled instance foreground mask.

[0273] In one or more embodiments, a method for training a backbone network, instance bounding boxes, instance query vectors, and an instance segmentation network is introduced. The training related parameters are as described in the foregoing embodiments, so details are not repeated here.

[0274] Specifically, similar to the previous embodiment, for QueryVIS, there are N instance query vectors. When performing sample connection and loss function calculation, one instance query vector is responsible for one ground truth bounding box (i.e., the labeled instance bounding box). During the training phase, the instance query vectors that match the ground truth bounding boxes are also responsible for predicting the instance mask of the ground truth bounding box (i.e., the predicted instance foreground mask). It should be noted that other training processes are similar to those described in the foregoing embodiments, so details are not repeated here.

[0275] The training process of QueryVIS is specifically as follows: Use the first sample video frame in the video to be trained as the input of the backbone network to be trained, and use the backbone network to be trained to extract features to obtain a sample feature map. Then, use N instance bounding boxes to be trained to extract N predicted bounding box ROIs and N predicted mask ROIs from the sample feature map. Next, use N instance query vectors to be trained to perform dynamic convolution operations or ordinary convolution operations on the N predicted bounding box ROIs and N predicted mask ROIs to generate enhanced N predicted bounding box ROIs and enhanced N predicted mask ROIs. Use the enhanced N predicted bounding box ROIs and enhanced N predicted mask ROIs as the input of the instance segmentation network to be trained, and the instance segmentation network to be trained outputs N first predicted class probability values, N first predicted instance bounding boxes, N first predicted instance embedding vectors, and N predicted instance foreground masks.

[0276] It should be noted that since the training uses paired sample video frames, the second sample video frame also needs to be processed similarly to obtain N second prediction results. The second sample video frame is from any frame in the video to be trained. Then, the prediction instance similarity between the first prediction result and the second prediction result can be calculated using the above formula (1), thereby obtaining at least one prediction instance similarity. Then, the prediction instance tracking result can be determined from the first sample video frame, that is, the instance identifier that can be predicted by the prediction instance tracking result.

[0277] Design a loss function. Exemplarily, the Focal loss function can be used to calculate the loss value between the N first prediction class probabilities and the true annotation class. And use the number of positive samples as a scaling factor to scale the calculated loss function. Exemplarily, the generalized IoU (GIOU) loss function and the L1 loss function can be used to calculate the loss value between the N first prediction instance bounding boxes and the true annotation instance bounding boxes. Exemplarily, the classification loss can be used to calculate the loss value between the instance identifier included in the prediction instance tracking result and the true annotation instance identifier. Exemplarily, the DICE loss function and the L1 loss function can be used to calculate the loss value between the N prediction instance foreground masks and the true annotation instance foreground masks.

[0278] Perform operations such as summing or weighted summing on the above loss values to obtain a comprehensive loss value. Based on the comprehensive loss value, use backpropagation and gradient descent algorithms to train QueryVIS, that is, update the parameters of the backbone network to be trained, N instance bounding boxes to be trained, N instance query vectors to be trained, and the instance segmentation network to be trained.

[0279] Secondly, in the embodiments of the present application, a method for training a backbone network, instance bounding boxes, instance query vectors, and an instance segmentation network is provided. Through the above method, the "one-to-one" sampling samples are connected with the true values to detect and segment the instances of interest, and full end-to-end training and inference can be performed.

[0280] Based on the instance tracking method provided by the present application, instance detection, segmentation, and tracking of the input video can be accurately and quickly performed. Leading results have been achieved on multiple open-source datasets, where these open-source datasets include the 2019 dataset of YouTube video instance segmentation (YouTube-VIS(2019)) and the 2021 dataset of YouTube video instance segmentation (YouTube-VIS(2021)). Specifically, refer to Table 3, which shows the results of system-level comparison on the YouTube-VIS(2019) dataset.

[0281] Table 3

[0282]

[0283] Among them, the Mean Average Precision (mAP) is an evaluation metric for video instance segmentation algorithms. It can be seen that this application proposes a VIS method based on instance query to construct an end-to-end and post-processing-free video instance segmentation model. At the same time, by constructing a unified tracking head, the number of manual parameters in the tracking task head is reduced, and at the same time, good tracking performance is achieved for different tracking tasks. On the publicly available dataset YouTube-VIS validation set, this method outperforms the current state-of-the-art video instance segmentation algorithms in terms of both speed and accuracy.

[0284] The instance tracking device in this application will be described in detail below. Please refer to Figure 13 , Figure 13 which is a schematic diagram of an embodiment of the instance tracking device in an embodiment of this application. The instance tracking device 20 includes:

[0285] An acquisition module 210, configured to obtain a target feature map through a backbone network based on a target video frame in a video to be detected, where the target video frame is the T-th video frame in the video to be detected, and T is an integer greater than 1;

[0286] The acquisition module 210 is further configured to obtain N regions of interest (ROIs) of bounding boxes from the target feature map according to N instance bounding boxes, where each instance bounding box is used to extract a corresponding bounding box ROI, and N is an integer greater than or equal to 1;

[0287] The acquisition module 210 is further configured to obtain N first detection results through an instance segmentation network based on N instance query vectors and N bounding box ROIs, where each first detection result includes a first class probability value, a first instance bounding box, and a first instance embedding vector;

[0288] A determination module 220, configured to determine at least one instance similarity according to N first detection results and M second detection results, where each second detection result includes a second class probability value, a second instance bounding box, and a second instance embedding vector, and the M second detection results are obtained according to the first (T-1) video frames in the video to be detected, and each second detection result corresponds to an instance identifier, and M is an integer greater than or equal to 1;

[0289] The determination module 220 is further configured to determine an instance tracking result of the target video frame according to at least one instance similarity, where the instance tracking result includes at least one instance identifier, and the same instance identifier represents the same instance in the video to be detected.

[0290] In an embodiment of the present application, an instance tracking device is provided. By using the above device, during video-based instance tracking, the instance query vector can be used to directly detect the instances of interest in the video frame, and then the instances in the video frame are matched with the instances in the previous video frame in the video for similarity. Finally, the instance tracking of the video frame is achieved according to the matching situation. The present application realizes instance detection without relying on post-processing methods such as NMS by constructing an end-to-end instance detection framework, and then tracks the instance target based on the instance identifier, thereby improving the efficiency of video-based instance tracking.

[0291] Optionally, based on the corresponding embodiment above, in another embodiment of the instance tracking device 20 provided in the embodiment of the present application, Figure 13 The obtaining module 210 is specifically configured to perform a dot product on each of the N bounding box ROIs in the feature dimension by using N instance query vectors to obtain N enhanced bounding box ROIs;

[0292] Based on the N enhanced bounding box ROIs, N first category probability values are obtained through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results;

[0293] Based on the N enhanced bounding box ROIs, N first instance bounding boxes are obtained through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0294] Based on the N enhanced bounding box ROIs, N first instance embedding vectors are obtained through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

[0295] In an embodiment of the present application, an instance tracking device is provided. By using the above device, the detection and tracking of instance targets in a video can be realized without extracting the instance foreground mask, and an end-to-end video instance segmentation framework is designed, thereby simplifying the complexity of the video instance segmentation model and being beneficial to improving the speed of model inference.

[0296] Optionally, based on the corresponding embodiment above, in another embodiment of the instance tracking device 20 provided in the embodiment of the present application,

[0297] The obtaining module 210 is specifically configured to obtain at least one set of bounding box dynamic parameters through a fully connected layer based on the N instance query vectors; Figure 13 Based on the N enhanced bounding box ROIs, N first instance bounding boxes are obtained through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0298] Based on the N enhanced bounding box ROIs, N first instance embedding vectors are obtained through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

[0299] Using at least one set of bounding box dynamic parameters, perform a dot product on each of the N bounding box ROIs in the feature dimension to obtain N enhanced bounding box ROIs;

[0300] Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results;

[0301] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0302] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

[0303] In the embodiments of the present application, an instance tracking device is provided. By using the above device, the detection and tracking of instance targets in a video can be achieved without extracting the instance foreground mask, and an end-to-end video instance segmentation framework is designed, thereby simplifying the complexity of the video instance segmentation model and facilitating the improvement of the model inference speed. At the same time, after the ROI feature extraction, dynamic convolution is performed between the instance query vector and the ROI feature, thereby generating enhanced instance features, which is conducive to obtaining better model effects.

[0304] Optionally, based on the above Figure 13 In another embodiment of the instance tracking device 20 provided in the embodiments of the present application, on the basis of the corresponding embodiment,

[0305] The acquisition module 210 is further configured to obtain N mask ROIs from the target video frame according to the N instance bounding boxes, where each instance bounding box is further configured to extract a corresponding mask ROI;

[0306] The acquisition module 210 is specifically configured to obtain N first detection results through the instance segmentation network based on the N instance query vectors, the N bounding box ROIs, and the N mask ROIs, where each first detection result further includes a first instance foreground mask.

[0307] In the embodiments of the present application, an instance tracking device is provided. By using the above device, the instance targets in the video can be further segmented, the improvement of the single-stage instance segmentation network is realized, and it is extended to the field of video instance segmentation, which not only reduces the number of model parameters, but also can improve the segmentation fineness.

[0308] Optionally, based on the above Figure 13Based on the corresponding embodiment, in another embodiment of the instance tracking device 20 provided in the embodiments of the present application,

[0309] An acquisition module 210, specifically configured to perform dot multiplication on each of the N bounding box ROIs in the feature dimension by using N instance query vectors to obtain N enhanced bounding box ROIs;

[0310] Perform dot multiplication on each of the N mask ROIs in the feature dimension by using N instance query vectors to obtain N enhanced mask ROIs;

[0311] Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results;

[0312] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0313] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results;

[0314] Based on the N enhanced mask ROIs, obtain N first instance foreground masks through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

[0315] In the embodiments of the present application, an instance tracking device is provided. By using the above device, the detection and tracking of instance targets in a video can be realized when extracting an instance foreground mask, and an end-to-end video instance segmentation framework is designed, thereby simplifying the complexity of the video instance segmentation model and being beneficial to improving the inference speed of the model.

[0316] Optionally, based on the corresponding embodiment, in another embodiment of the instance tracking device 20 provided in the embodiments of the present application, Figure 13 Based on the corresponding embodiment, in another embodiment of the instance tracking device 20 provided in the embodiments of the present application,

[0317] An acquisition module 210, specifically configured to obtain at least one set of bounding box dynamic parameters and at least one set of mask dynamic parameters through a fully connected layer based on N instance query vectors, and each set of mask dynamic parameters includes N mask dynamic sub-parameters;

[0318] Perform dot multiplication on each of the N bounding box ROIs in the feature dimension by using at least one set of bounding box dynamic parameters to obtain N enhanced bounding box ROIs;

[0319] Using at least one set of mask dynamic parameters, perform dot product on each of the N masked ROIs in the feature dimension to obtain N enhanced masked ROIs;

[0320] Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results;

[0321] Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results;

[0322] Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results;

[0323] Based on the N enhanced masked ROIs, obtain N first instance foreground masks through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

[0324] In the embodiments of the present application, an instance tracking device is provided. By using the above device, the detection and tracking of instance targets in a video can be achieved when extracting the instance foreground mask, and an end-to-end video instance segmentation framework is designed, thereby simplifying the complexity of the video instance segmentation model and facilitating the improvement of the model inference speed. At the same time, after ROI feature extraction, dynamic convolution is performed between the instance query vector and the ROI feature, thereby generating enhanced instance features, which is beneficial to obtaining better model effects.

[0325] Optionally, based on the above Figure 13 In another embodiment of the instance tracking device 20 provided in the embodiments of the present application, on the basis of the corresponding embodiment

[0326] The determination module 220 is specifically configured to sort the N first detection results in descending order according to the first category probability values to obtain the sorted N first detection results;

[0327] Select the first K first detection results from the sorted N first detection results, where K is an integer greater than or equal to 1 and less than or equal to N;

[0328] According to the K first detection results and the M second detection results, determine the instance similarity between each first detection result and each second detection result to obtain (K * M) instance similarities.

[0329] In an embodiment of the present application, an instance tracking device is provided. By using the above device, considering that matching all N first detection results not only consumes a lot of resources but also leads to low efficiency in instance target tracking. Therefore, the top K first detection results with the largest class probability values are selected from the N first detection results, and the K first detection results are used for subsequent matching, thereby improving the matching efficiency and the matching accuracy at the same time.

[0330] Optionally, based on the corresponding embodiment above, in another embodiment of the instance tracking device 20 provided in the embodiment of the present application, Figure 13 the determining module 220 is specifically configured to determine the instance embedding vector similarity between each first detection result and each second detection result according to the first instance embedding vector included in each first detection result and the second instance embedding vector included in each second detection result;

[0331] Determine the spatial similarity between each first detection result and each second detection result according to the first instance bounding box included in each first detection result and the second instance bounding box included in each second detection result;

[0332] Determine the class similarity between each first detection result and each second detection result according to the first class probability value included in each first detection result and the second class probability value included in each second detection result;

[0333] Determine the instance similarity between each first detection result and each second detection result according to the instance embedding vector similarity, spatial similarity, class similarity between each first detection result and each second detection result, and the first class probability value included in each first detection result.

[0334] In an embodiment of the present application, an instance tracking device is provided. By using the above device, the instance similarity between two detection results can be directly calculated by using the product of the instance embedding vector similarity, spatial similarity, class similarity, and class probability value. Therefore, there is no need to train other weight parameters, which not only reduces the computational complexity but also reduces the amount of parameter training. At the same time, the (K * M) instance similarities can be normalized by bidirectional softmax to ensure that the value range of the probability is between 0 and 1, thereby improving the calculation accuracy and accelerating the convergence speed.

[0335] Optionally, based on the corresponding embodiment above, in another embodiment of the instance tracking device 20 provided in the embodiment of the present application, K is less than or equal to M;

[0336] Optionally, based on the corresponding embodiment above, in another embodiment of the instance tracking device 20 provided in the embodiment of the present application, Figure 13 the determining module 220 is specifically configured to determine the instance embedding vector similarity between each first detection result and each second detection result according to the first instance embedding vector included in each first detection result and the second instance embedding vector included in each second detection result;

[0337] A determination module 220, specifically configured to construct a mapping relationship between K first detection results and M second detection results based on a bipartite graph matching algorithm according to (K*M) instance similarities, obtaining K mapping relationships;

[0338] If there are P mapping relationships among the K mapping relationships whose corresponding instance similarities are less than or equal to the instance similarity threshold, then delete the P mapping relationships from the K mapping relationships, obtaining (K - P) mapping relationships, where P is an integer greater than or equal to 1 and less than or equal to K;

[0339] Determine the instance tracking result of the target video frame according to each second detection result corresponding to each mapping relationship among the (K - P) mapping relationships, where each second detection result corresponds to an instance identifier;

[0340] The determination module is further configured to use the first detection results corresponding to the P mapping relationships as second detection results, obtaining (M + P) second detection results.

[0341] In an embodiment of the present application, an instance tracking device is provided. By using the above device, when K is less than or equal to M, the bipartite graph matching algorithm is used to preferentially match the K first detection results, thereby obtaining a result with a better overall matching degree. Thus, the accuracy and consistency of instance target tracking are improved.

[0342] Optionally, on the basis of the above Figure 13 corresponding embodiment, in another embodiment of the instance tracking device 20 provided in the embodiment of the present application, K is greater than or equal to M;

[0343] A determination module 220, specifically configured to construct a mapping relationship between K first detection results and M second detection results based on a bipartite graph matching algorithm according to (K*M) instance similarities, obtaining M mapping relationships;

[0344] If there are Q mapping relationships among the M mapping relationships whose corresponding instance similarities are less than or equal to the instance similarity threshold, then delete the Q mapping relationships from the M mapping relationships, obtaining (M - Q) mapping relationships, where Q is an integer greater than or equal to 1 and less than or equal to M;

[0345] Determine the instance tracking result of the target video frame according to each second detection result corresponding to each mapping relationship among the (M - Q) mapping relationships, where each second detection result corresponds to an instance identifier;

[0346] The determination module is further configured to use the Q mapping relationships and the (K - M) first detection results as second detection results, obtaining (Q + K) second detection results.

[0347] In an embodiment of the present application, an instance tracking device is provided. By using the above device, when K is greater than or equal to M, the bipartite graph matching algorithm is used to preferentially match M second detection results, so as to obtain a result with better overall matching degree. Thus, the accuracy and consistency of instance target tracking are improved.

[0348] Optionally, on the basis of the corresponding embodiment above, Figure 13 In another embodiment of the instance tracking device 20 provided in the embodiment of the present application, the instance tracking device 20 further includes a training module 203;

[0349] The obtaining module 210 is further configured to obtain a sample feature map through the backbone network to be trained based on the first sample video frame in the video to be trained, where the first sample video frame has an annotated category, an annotated instance bounding box, and an annotated instance identifier;

[0350] The obtaining module 210 is further configured to obtain N predicted bounding box ROIs from the sample feature map according to N instance bounding boxes to be trained, where each instance bounding box to be trained is used to extract a corresponding predicted bounding box ROI;

[0351] The obtaining module 210 is further configured to obtain N first prediction results through the instance segmentation network to be trained based on N instance query vectors to be trained and N predicted bounding box ROIs, where each first prediction result includes a first predicted category probability value, a first predicted instance bounding box, and a first predicted instance embedding vector;

[0352] The determining module 220 is further configured to determine at least one predicted instance similarity according to the N first prediction results and the N second prediction results, where each second prediction result includes a second predicted category probability value, a second predicted instance bounding box, and a second predicted instance embedding vector, and the N second prediction results are from the second sample video frames in the video to be trained, and each second prediction result corresponds to an instance identifier;

[0353] The determining module 220 is further configured to determine the predicted instance tracking result of the first sample video frame according to at least one predicted instance similarity;

[0354] The training module 230 is configured to update the parameters of the backbone network to be trained, N instance bounding boxes to be trained, N instance query vectors to be trained, and the instance segmentation network to be trained through a loss function according to the predicted instance tracking result, the annotated instance identifier, the N first prediction results, the annotated category, and the annotated instance bounding box.

[0355] In an embodiment of the present application, an instance tracking device is provided. By using the above device, a "one-to-one" connection is established between the sampling samples and the true values, so as to detect and segment the instances of interest, and full end-to-end training and inference can be performed.

[0356] Optionally, based on the corresponding embodiment above, in another embodiment of the instance tracking device 20 provided in the embodiment of the present application, Figure 13 the obtaining module 210 is further configured to obtain a sample feature map through a backbone network to be trained based on the first sample video frame in the video to be trained, where the first sample video frame has an annotation category, an annotation instance bounding box, an annotation instance identifier, and an annotation instance foreground mask;

[0357] the obtaining module 210 is further configured to obtain N predicted bounding box ROIs from the sample feature map according to N instance bounding boxes to be trained, where each instance bounding box to be trained is used to extract a corresponding predicted bounding box ROI;

[0358] the obtaining module 210 is further configured to obtain N predicted mask ROIs from the sample feature map according to N instance bounding boxes to be trained, where each instance bounding box to be trained is also used to extract a corresponding predicted mask ROI;

[0359] the obtaining module 210 is further configured to obtain N first prediction results through the instance segmentation network to be trained based on N instance query vectors to be trained, N predicted bounding box ROIs, and N predicted mask ROIs, where each first prediction result includes a first predicted class probability value, a first predicted instance bounding box, a first predicted instance embedding vector, and a predicted instance foreground mask;

[0360] the determining module 220 is further configured to determine at least one predicted instance similarity according to the N first prediction results and the N second prediction results, where each second prediction result includes a second predicted class probability value, a second predicted instance bounding box, and a second predicted instance embedding vector, and the N second prediction results are from the second sample video frame in the video to be trained, and each second prediction result corresponds to an instance identifier;

[0361] the determining module 220 is further configured to determine the predicted instance tracking result of the first sample video frame according to at least one predicted instance similarity;

[0362] the training module 230 is further configured to update the parameters of the backbone network to be trained, the N instance bounding boxes to be trained, the N instance query vectors to be trained, and the instance segmentation network to be trained through a loss function according to the predicted instance tracking result, the annotation instance identifier, the N first prediction results, the annotation category, the annotation instance bounding box, and the annotation instance foreground mask.

[0363] the training module 230 is further configured to update the parameters of the backbone network to be trained, the N instance bounding boxes to be trained, the N instance query vectors to be trained, and the instance segmentation network to be trained through a loss function according to the predicted instance tracking result, the annotation instance identifier, the N first prediction results, the annotation category, the annotation instance bounding box, and the annotation instance foreground mask.

[0364] In an embodiment of the present application, an instance tracking device is provided. By using the above device, a "one-to-one" sampling sample is connected to the true value, so as to detect and segment the instance of interest, and full end-to-end training and inference can be performed.

[0365] Figure 14 It is a schematic structural diagram of a computer device 30 according to an embodiment of the present application. The computer device 30 may include an input device 310, an output device 320, a processor 330, and a memory 340. The output device in the embodiment of the present application may be a display device.

[0366] The memory 340 may include a read-only memory and a random access memory, and provide instructions and data to the processor 330. A part of the memory 340 may further include a non-volatile random access memory (Non-Volatile Random Access Memory, NVRAM).

[0367] The memory 340 stores the following elements, executable modules or data structures, or subsets thereof, or extended sets thereof:

[0368] Operation instructions: including various operation instructions for implementing various operations.

[0369] Operating system: including various system programs for implementing various basic services and processing hardware-based tasks.

[0370] The processor 330 controls the operation of the computer device 30. The processor 330 may also be referred to as a central processing unit (Central Processing Unit, CPU). The memory 340 may include a read-only memory and a random access memory, and provide instructions and data to the processor 330. A part of the memory 340 may further include NVRAM. In a specific application, the various components of the computer device 30 are coupled together through a bus system 350. The bus system 350 may include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clarity, all kinds of buses are labeled as the bus system 350 in the figure.

[0371] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 330. The processor 330 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 330 or the instructions in the form of software. The above-mentioned processor 330 may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 340, and the processor 330 reads the information in the memory 340 and combines its hardware to complete the steps of the above method.

[0372] Figure 14 For the relevant description, reference can be made to Figure 3 the relevant description and effects in the method part for understanding, and details will not be elaborated here.

[0373] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0374] The embodiments of the present application also provide a computer program product including a program. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0375] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0376] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0377] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0378] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0379] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0380] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present application.

Claims

1. A video-based instance tracking method, characterized in that, Including: Based on a target video frame in a video to be detected, obtaining a target feature map through a backbone network, where the target video frame is the T-th video frame in the video to be detected, and T is an integer greater than 1; According to N instance bounding boxes, obtaining N region of interests (ROIs) of bounding boxes from the target feature map, where each instance bounding box is used to extract a corresponding bounding box ROI, N is an integer greater than or equal to 1, and the bounding box ROI has a fixed resolution; Based on N instance query vectors and the N bounding box ROIs, obtaining N first detection results through an instance segmentation network, where there is a one-to-one correspondence between the instance bounding boxes and the instance query vectors, performing convolutional processing on the bounding box ROIs using the instance query vectors, and inputting the convolved bounding box ROIs into the instance segmentation network to obtain the corresponding first detection results. Each first detection result includes a first class probability value, a first instance bounding box, and a first instance embedding vector; Determining at least one instance similarity according to the N first detection results and M second detection results, including: sorting the N first detection results in descending order according to the first class probability values to obtain the sorted N first detection results; selecting the first K first detection results from the sorted N first detection results, where K is an integer greater than or equal to 1 and less than or equal to N; determining the instance similarity between each first detection result and each second detection result according to the K first detection results and the M second detection results to obtain (K * M) instance similarities. Each second detection result includes a second class probability value, a second instance bounding box, and a second instance embedding vector. The M second detection results are obtained according to the first (T - 1) video frames in the video to be detected, and each second detection result corresponds to an instance identifier. M is an integer greater than or equal to 1; Determining an instance tracking result of the target video frame according to the at least one instance similarity, where the instance tracking result includes at least one instance identifier, and the same instance identifier represents the same instance in the video to be detected.

2. The example tracking method according to claim 1, wherein The obtaining N first detection results through the instance segmentation network based on the N instance query vectors and the N bounding box ROIs includes: Performing dot multiplication on each of the N bounding box ROIs in the N bounding box ROIs in the feature dimension using the N instance query vectors to obtain N enhanced bounding box ROIs; Based on the N enhanced bounding box ROIs, obtaining N first class probability values through a class discriminant network included in the instance segmentation network, where the N first class probability values are included in the N first detection results; Based on the N enhanced bounding box ROIs, obtaining N first instance bounding boxes through a bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results; Based on the N enhanced bounding box ROIs, N first instance embedding vectors are obtained through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

3. The example tracking method according to claim 1, wherein The obtaining of the N first detection results through the instance segmentation network based on the N instance query vectors and the N bounding box ROIs includes: Based on the N instance query vectors, at least one set of bounding box dynamic parameters is obtained through a fully connected layer; Using the at least one set of bounding box dynamic parameters, each of the N bounding box ROIs is multiplied pointwise in the feature dimension to obtain N enhanced bounding box ROIs; Based on the N enhanced bounding box ROIs, N first class probability values are obtained through the class discrimination network included in the instance segmentation network, where the N first class probability values are included in the N first detection results; Based on the N enhanced bounding box ROIs, N first instance bounding boxes are obtained through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results; Based on the N enhanced bounding box ROIs, N first instance embedding vectors are obtained through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

4. The instance tracking method according to claim 1, wherein The method further includes: According to the N instance bounding boxes, N mask ROIs are obtained from the target video frame, where each instance bounding box is further used to extract a corresponding mask ROI; The obtaining of the N first detection results through the instance segmentation network based on the N instance query vectors and the N bounding box ROIs includes: Based on the N instance query vectors, the N bounding box ROIs, and the N mask ROIs, N first detection results are obtained through the instance segmentation network, where each first detection result further includes a first instance foreground mask.

5. The instance tracking method according to claim 4, wherein The obtaining of the N first detection results through the instance segmentation network based on the N instance query vectors, the N bounding box ROIs, and the N mask ROIs includes: Using the N instance query vectors, each of the N bounding box ROIs is multiplied pointwise in the feature dimension to obtain N enhanced bounding box ROIs; Using the N instance query vectors, each of the N mask ROIs is multiplied pointwise in the feature dimension to obtain N enhanced mask ROIs; Based on the N enhanced bounding box ROIs, N first class probability values are obtained through the class discrimination network included in the instance segmentation network, where the N first class probability values are included in the N first detection results; Based on the N enhanced bounding box ROIs, N first instance bounding boxes are obtained through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results; Based on the N enhanced bounding box ROIs, N first instance embedding vectors are obtained through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results; Based on the N enhanced mask ROIs, N first instance foreground masks are obtained through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

6. The instance tracking method according to claim 4, wherein The obtaining of the N first detection results through the instance segmentation network based on the N instance query vectors, the N bounding box ROIs, and the N mask ROIs includes: Based on the N instance query vectors, at least one set of bounding box dynamic parameters and at least one set of mask dynamic parameters are obtained through a fully connected layer, and each set of mask dynamic parameters includes N mask dynamic sub-parameters; Using the at least one set of bounding box dynamic parameters, each of the N bounding box ROIs is multiplied pointwise in the feature dimension to obtain N enhanced bounding box ROIs; Using the at least one set of mask dynamic parameters, each of the N mask ROIs is multiplied pointwise in the feature dimension to obtain N enhanced mask ROIs; Based on the N enhanced bounding box ROIs, N first class probability values are obtained through the class discrimination network included in the instance segmentation network, where the N first class probability values are included in the N first detection results; Based on the N enhanced bounding box ROIs, N first instance bounding boxes are obtained through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results; Based on the N enhanced bounding box ROIs, N first instance embedding vectors are obtained through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results; Based on the N enhanced mask ROIs, N first instance foreground masks are obtained through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

7. The instance tracking method according to claim 1, wherein The determining of the instance similarity between each first detection result and each second detection result according to the K first detection results and the M second detection results includes: Determining the instance embedding vector similarity between each first detection result and each second detection result according to the first instance embedding vector included in each first detection result and the second instance embedding vector included in each second detection result; Determining the spatial similarity between each first detection result and each second detection result according to the first instance bounding box included in each first detection result and the second instance bounding box included in each second detection result; Determining the class similarity between each first detection result and each second detection result according to the first class probability value included in each first detection result and the second class probability value included in each second detection result; Determine the instance similarity between each of the first detection results and each of the second detection results based on the instance embedding vector similarity, the spatial similarity, the category similarity between each of the first detection results and each of the second detection results, and the first category probability value included in each of the first detection results.

8. The example tracking method according to claim 1, wherein The K is less than or equal to the M; The determining the instance tracking result of the target video frame according to the at least one instance similarity includes: Based on the bipartite graph matching algorithm, construct a mapping relationship between the K first detection results and the M second detection results according to the (K * M) instance similarities, and obtain K mapping relationships; If there are P mapping relationships among the K mapping relationships whose corresponding instance similarities are less than or equal to the instance similarity threshold, then delete the P mapping relationships from the K mapping relationships to obtain (K - P) mapping relationships, where the P is an integer greater than or equal to 1 and less than or equal to the K; Determine the instance tracking result of the target video frame according to the second detection result corresponding to each mapping relationship among the (K - P) mapping relationships, where each second detection result corresponds to an instance identifier; The method further includes: Use the first detection results corresponding to the P mapping relationships as the second detection results to obtain (M + P) second detection results.

9. The instance tracking method according to claim 1, wherein The K is greater than or equal to the M; The determining the instance tracking result of the target video frame according to the at least one instance similarity includes: Based on the bipartite graph matching algorithm, construct a mapping relationship between the K first detection results and the M second detection results according to the (K * M) instance similarities, and obtain M mapping relationships; If there are Q mapping relationships among the M mapping relationships whose corresponding instance similarities are less than or equal to the instance similarity threshold, then delete the Q mapping relationships from the M mapping relationships to obtain (M - Q) mapping relationships, where the Q is an integer greater than or equal to 1 and less than or equal to the M; Determine the instance tracking result of the target video frame according to the second detection result corresponding to each mapping relationship among the (M - Q) mapping relationships, where each second detection result corresponds to an instance identifier; The method further includes: Use the Q mapping relationships and the (K - M) first detection results as the second detection results to obtain (Q + K) second detection results.

10. The instance tracking method according to claim 1, wherein The method further includes: Based on the first sample video frames in the video to be trained, obtain the sample feature maps through the backbone network to be trained, where the first sample video frames have labeled categories, labeled instance bounding boxes, and labeled instance identifiers; According to N to-be-trained instance bounding boxes, obtain N predicted bounding box ROIs from the sample feature maps, where each to-be-trained instance bounding box is used to extract a corresponding predicted bounding box ROI; Based on N query vectors of training instances and the N predicted bounding box ROIs, obtain N first prediction results through the training instance segmentation network, where each first prediction result includes a first predicted class probability value, a first predicted instance bounding box, and a first predicted instance embedding vector; Determine at least one predicted instance similarity according to the N first prediction results and N second prediction results, where each second prediction result includes a second predicted class probability value, a second predicted instance bounding box, and a second predicted instance embedding vector, the N second prediction results are from the second sample video frames in the training video, and each second prediction result corresponds to an instance identifier; Determine the predicted instance tracking result of the first sample video frame according to the at least one predicted instance similarity; According to the predicted instance tracking result, the labeled instance identifier, the N first prediction results, the labeled class, and the labeled instance bounding box, update the parameters of the training backbone network, the N training instance bounding boxes, the N training instance query vectors, and the training instance segmentation network through a loss function.

11. The instance tracking method according to claim 4, wherein The method further includes: Based on the first sample video frame in the training video, obtain a sample feature map through the training backbone network, where the first sample video frame has a labeled class, a labeled instance bounding box, a labeled instance identifier, and a labeled instance foreground mask; Obtain N predicted bounding box ROIs from the sample feature map according to the N training instance bounding boxes, where each training instance bounding box is used to extract a corresponding predicted bounding box ROI; Obtain N predicted mask ROIs from the sample feature map according to the N training instance bounding boxes, where each training instance bounding box is also used to extract a corresponding predicted mask ROI; Based on the N training instance query vectors, the N predicted bounding box ROIs, and the N predicted mask ROIs, obtain the N first prediction results through the training instance segmentation network, where each first prediction result includes a first predicted class probability value, a first predicted instance bounding box, a first predicted instance embedding vector, and a predicted instance foreground mask; Determine at least one predicted instance similarity according to the N first prediction results and N second prediction results, where each second prediction result includes a second predicted class probability value, a second predicted instance bounding box, and a second predicted instance embedding vector, the N second prediction results are from the second sample video frames in the training video, and each second prediction result corresponds to an instance identifier; Determine the predicted instance tracking result of the first sample video frame according to the at least one predicted instance similarity; Based on the predicted instance tracking results, the labeled instance identifiers, the N first prediction results, the labeled categories, the labeled instance bounding boxes, and the labeled instance foreground masks, update the parameters of the backbone network to be trained, the N instance bounding boxes to be trained, the N instance query vectors to be trained, and the instance segmentation network to be trained through a loss function.

12. An instance tracking device, characterized in that, Including: An acquisition module, configured to obtain a target feature map through a backbone network based on a target video frame in a video to be detected, where the target video frame is the T-th video frame in the video to be detected, and T is an integer greater than 1; The acquisition module is further configured to obtain N region of interests (ROIs) of bounding boxes from the target feature map according to N instance bounding boxes, where each instance bounding box is used to extract a corresponding bounding box ROI, N is an integer greater than or equal to 1, and the bounding box ROI has a fixed resolution; The acquisition module is further configured to obtain N first detection results through an instance segmentation network based on N instance query vectors and the N bounding box ROIs, where there is a one-to-one correspondence between the instance bounding boxes and the instance query vectors, perform convolutional processing on the bounding box ROIs using the instance query vectors, and input the convolved bounding box ROIs into the instance segmentation network to obtain the corresponding first detection results. Each first detection result includes a first category probability value, a first instance bounding box, and a first instance embedding vector; A determination module, configured to determine at least one instance similarity according to the N first detection results and M second detection results, including: sorting the N first detection results in descending order according to the first category probability values to obtain the sorted N first detection results; selecting the first K first detection results from the sorted N first detection results, where K is an integer greater than or equal to 1 and less than or equal to N; determining the instance similarity between each first detection result and each second detection result according to the K first detection results and the M second detection results to obtain (K * M) instance similarities. Each second detection result includes a second category probability value, a second instance bounding box, and a second instance embedding vector. The M second detection results are obtained according to the first (T - 1) video frames in the video to be detected, and each second detection result corresponds to an instance identifier. M is an integer greater than or equal to 1; The determination module is further configured to determine the instance tracking result of the target video frame according to the at least one instance similarity, where the instance tracking result includes at least one instance identifier, and the same instance identifier represents the same instance in the video to be detected.

13. The device according to claim 12, characterized in that, The acquisition module is specifically configured to: Perform dot product on each of the N bounding box ROIs in the N bounding box ROIs in the feature dimension using the N instance query vectors to obtain N enhanced bounding box ROIs; Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results; Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results; Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

14. The device according to claim 12, characterized in that, The obtaining module is specifically configured to: Based on the N instance query vectors, obtain at least one set of bounding box dynamic parameters through a fully connected layer; Using the at least one set of bounding box dynamic parameters, perform dot multiplication on each of the N bounding box ROIs in the feature dimension to obtain N enhanced bounding box ROIs; Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results; Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results; Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results.

15. The device according to claim 12, characterized in that, The obtaining module is further configured to obtain N mask ROIs from the target video frame according to the N instance bounding boxes, where each instance bounding box is further used to extract a corresponding mask ROI; The obtaining module is specifically configured to obtain N first detection results through an instance segmentation network based on the N instance query vectors, the N bounding box ROIs, and the N mask ROIs, where each first detection result further includes a first instance foreground mask.

16. The device according to claim 15, wherein The obtaining module is specifically configured to: Using the N instance query vectors, perform dot multiplication on each of the N bounding box ROIs in the feature dimension to obtain N enhanced bounding box ROIs; Using the N instance query vectors, perform dot multiplication on each of the N mask ROIs in the feature dimension to obtain N enhanced mask ROIs; Based on the N enhanced bounding box ROIs, obtain N first category probability values through the category discrimination network included in the instance segmentation network, where the N first category probability values are included in the N first detection results; Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results. Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results; Based on the N enhanced mask ROIs, obtain N first instance foreground masks through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

17. The device according to claim 15, characterized in that, The obtaining module is specifically configured to: Based on the N instance query vectors, obtain at least one set of bounding box dynamic parameters and at least one set of mask dynamic parameters through a fully connected layer, and each set of mask dynamic parameters includes N mask dynamic sub-parameters; Use the at least one set of bounding box dynamic parameters to perform a dot product on each of the N bounding box ROIs in the feature dimension to obtain N enhanced bounding box ROIs; Use the at least one set of mask dynamic parameters to perform a dot product on each of the N mask ROIs in the feature dimension to obtain N enhanced mask ROIs; Based on the N enhanced bounding box ROIs, obtain N first class probability values through the class discrimination network included in the instance segmentation network, where the N first class probability values are included in the N first detection results; Based on the N enhanced bounding box ROIs, obtain N first instance bounding boxes through the bounding box regression network included in the instance segmentation network, where the N first instance bounding boxes are included in the N first detection results; Based on the N enhanced bounding box ROIs, obtain N first instance embedding vectors through the embedding vector network included in the instance segmentation network, where the N first instance embedding vectors are included in the N first detection results; Based on the N enhanced mask ROIs, obtain N first instance foreground masks through the mask generation network included in the instance segmentation network, where the N first instance foreground masks are included in the N first detection results.

18. The device according to claim 12, wherein The determining module is specifically configured to: Determine the instance embedding vector similarity between each first detection result and each second detection result according to the first instance embedding vector included in each first detection result and the second instance embedding vector included in each second detection result; Determine the spatial similarity between each first detection result and each second detection result according to the first instance bounding box included in each first detection result and the second instance bounding box included in each second detection result; Determine the class similarity between each first detection result and each second detection result according to the first class probability value included in each first detection result and the second class probability value included in each second detection result; Determine the instance similarity between each of the first detection results and each of the second detection results according to the instance embedding vector similarity, the spatial similarity, the category similarity between each of the first detection results and each of the second detection results, and the first category probability value included in each of the first detection results.

19. A computer device, characterized in that, Comprising: A memory, a processor, and a bus system; Wherein, the memory is used for storing programs; The processor is used for executing the programs in the memory, and the processor is used for executing the instance tracking method according to any one of claims 1 to 11 according to the instructions in the program code; The bus system is used for connecting the memory and the processor, so that the memory and the processor communicate with each other.

20. A computer-readable storage medium, comprising instructions that, when running on a computer, cause the computer to execute the instance tracking method according to any one of claims 1 to 11.

21. A computer program product, characterized in that, The computer program product comprises a program that, when running on a computer, causes the computer to execute the instance tracking method according to any one of claims 1 to 11.