A video labeling method, device, apparatus and storage medium
By matching and segmenting keyframes in the video, the problem that target tracking algorithms cannot handle situations where the target object does not appear in the first frame is solved, thus achieving automation and efficiency improvement in video annotation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NETEASE LINGDONG (HANGZHOU) TECHNOLOGY CO LTD
- Filing Date
- 2024-11-01
- Publication Date
- 2026-05-01
AI Technical Summary
Existing target tracking algorithms cannot perform target tracking processing on video data where the target object does not appear in the first frame, resulting in reduced efficiency and accuracy of the automated video annotation process.
By using keyframe matching, keyframes containing the target object are automatically identified from the video to be labeled. The keyframes are then used as video segmentation points to segment the video to be labeled, and target tracking is performed on the image sequences located before and after the keyframes.
It automates video annotation, improving the efficiency and accuracy of video annotation and avoiding the need for separate judgment on whether the first frame contains the target object.
Smart Images

Figure CN119484933B_ABST
Abstract
Description
A video annotation method, apparatus, device, and storage medium Technical Field
[0001] This application relates to the field of image processing technology, and more specifically, to a video annotation method, apparatus, device, and storage medium. Background Technology
[0002] In scenarios such as excavator operation and loader material handling, it is often necessary to record the working process of the vehicles in the scene on video. However, in addition to the target objects related to the video shooting task, such as the working vehicles and the materials being shoveled, the video often contains background elements unrelated to the video shooting task, such as trees and warehouses. Therefore, in order to improve the data quality of the video data, users need to label the above-mentioned target objects contained in the video data so that they can more intuitively and quickly locate the labeled target objects from the labeled video.
[0003] Currently, target tracking algorithms can be used to track specified objects in videos and mark their positions within each frame. However, since these algorithms typically only work on image sequences where the target object appears in the first frame and the images were captured consecutively, they cannot track video data where the target object does not appear in the first frame. This hinders the automation of video annotation, resulting in reduced efficiency and accuracy. Summary of the Invention
[0004] In view of this, this application provides a video annotation method, apparatus, device, and storage medium. By using keyframe matching, it automatically determines the keyframes containing the target object from the video to be annotated, and uses the keyframes as video segmentation points to segment the video to be annotated. It then performs target tracking processing on two sets of image sequences in the video to be annotated, one before the keyframe and the other after the keyframe. Thus, without changing the working principle of the target tracking algorithm, the video annotation method does not need to separately determine whether the first frame of the video to be annotated contains the target object, thereby automating the video annotation process and effectively improving the efficiency and accuracy of video annotation.
[0005] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.
[0006] In a first aspect, embodiments of this application provide a video annotation method, the video annotation method comprising:
[0007] Based on the standard keyframes of the target object, image frames that match the standard keyframes are determined from the video to be labeled as keyframes of the target object in the video to be labeled; wherein, the standard keyframes include the location marker information of the target object;
[0008] The standard keyframe, the keyframe, and the first image sequence are combined to obtain a first sub-video; the standard keyframe and the keyframe are combined to obtain a second sub-video; wherein, the first image sequence represents the image sequence obtained by reversing the image frames in the video to be labeled that are located before the keyframe; the second image sequence represents the image sequence composed of the image frames in the video to be labeled that are located after the keyframe.
[0009] Using the target object as the tracking target, target tracking processing is performed on the first sub-video and the second sub-video respectively to obtain the target tracking processing results corresponding to the first sub-video and the second sub-video respectively;
[0010] Based on the order of image frames in the video to be labeled, the target tracking processing results corresponding to the first sub-video and the second sub-video are recombined to obtain the target tracking result of the target object in the video to be labeled.
[0011] Secondly, embodiments of this application provide a video annotation device, the video annotation device comprising:
[0012] The matching module is used to determine, based on the standard keyframes of the target object, image frames that match the standard keyframes from the video to be labeled as keyframes of the target object in the video to be labeled; wherein, the standard keyframes include the location marker information of the target object;
[0013] A grouping module is used to combine the standard keyframe, the keyframe and a first image sequence to obtain a first sub-video, and to combine the standard keyframe and a second image sequence to obtain a second sub-video; wherein, the first image sequence represents the image sequence obtained by reversing the image frames in the video to be labeled that are located before the keyframe; the second image sequence represents the image sequence composed of the image frames in the video to be labeled that are located after the keyframe;
[0014] The target tracking module is used to perform target tracking processing on the first sub-video and the second sub-video respectively, taking the target object as the tracking target, to obtain the target tracking processing results corresponding to the first sub-video and the second sub-video respectively;
[0015] The recombination module is used to recombine the target tracking processing results corresponding to the first sub-video and the second sub-video respectively according to the order of image frames in the video to be labeled, so as to obtain the target tracking result of the target object in the video to be labeled.
[0016] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the video annotation method described above.
[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the video annotation method described above.
[0018] The technical solutions provided by the embodiments of this application may include the following beneficial effects:
[0019] This application provides a video annotation method, apparatus, device, and storage medium. By using keyframe matching, it automatically determines keyframes containing target objects from the video to be annotated, and uses these keyframes as video segmentation points to segment the video. It then performs target tracking processing on two sets of image sequences located before and after the keyframes in the video to be annotated. Thus, without changing the working principle of the target tracking algorithm, the video annotation method does not need to separately determine whether the first frame of the video to be annotated contains the target object, thereby automating video annotation and effectively improving the efficiency and accuracy of video annotation. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 shows a flowchart of a video annotation method provided in an embodiment of this application;
[0022] Figure 2 shows a schematic diagram of the structure of a first sub-video and a second sub-video provided in an embodiment of this application;
[0023] Figure 3 shows a flowchart illustrating a method for determining standard keyframes according to an embodiment of this application;
[0024] Figure 4 shows a flowchart of a method for deduplicating target tracking results of multiple target objects in a video to be labeled, provided by an embodiment of this application.
[0025] Figure 5 shows a schematic diagram of a video annotation device provided in an embodiment of this application;
[0026] Figure 6 is a schematic diagram of the structure of an electronic device 600 provided in an embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0028] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0029] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0030] Currently, target tracking algorithms can be used to track specified objects in videos and mark their positions within each frame. However, since these algorithms typically only work on image sequences where the target object appears in the first frame and the images were captured consecutively, they cannot track video data where the target object does not appear in the first frame. This hinders the automation of video annotation, resulting in reduced efficiency and accuracy.
[0031] Based on this, embodiments of this application provide a video annotation method, apparatus, device, and storage medium. By using keyframe matching, keyframes containing target objects are automatically determined from the video to be annotated. The video to be annotated is then segmented using these keyframes as video segmentation points. Target tracking processing is performed on two sets of image sequences in the video to be annotated, one before the keyframe and the other after the keyframe. Thus, without changing the working principle of the target tracking algorithm, the video annotation method does not need to separately determine whether the first frame of the video to be annotated contains the target object, achieving automation of video annotation and effectively improving the efficiency and accuracy of video annotation.
[0032] In one embodiment of this application, a video annotation method can run on a terminal device or a server. The terminal device can be a local terminal device. When the video annotation method runs on a server, it can be implemented and executed based on a cloud interaction system, which includes a server and client devices (i.e., terminal devices).
[0033] To facilitate understanding of the embodiments of this application, a video annotation method, apparatus, device, and storage medium provided in the embodiments of this application will be described in detail below.
[0034] Referring to Figure 1, which illustrates a flowchart of a video annotation method provided in an embodiment of this application, the video annotation method includes steps S101-S104; specifically:
[0035] S101, based on the standard keyframes of the target object, determine the image frames that match the standard keyframes from the video to be labeled as the keyframes of the target object in the video to be labeled.
[0036] S102, the standard keyframe and the keyframe are combined with the first image sequence to obtain the first sub-video, and the standard keyframe and the keyframe are combined with the second image sequence to obtain the second sub-video.
[0037] S103, using the target object as the tracking target, target tracking processing is performed on the first sub-video and the second sub-video respectively to obtain the target tracking processing results corresponding to the first sub-video and the second sub-video respectively.
[0038] S104, based on the order of the image frames in the video to be labeled, the target tracking processing results corresponding to the first sub-video and the second sub-video are recombined to obtain the target tracking result of the target object in the video to be labeled.
[0039] The video annotation method provided in this application automatically determines the keyframes containing the target object from the video to be annotated by keyframe matching, and uses the keyframes as video segmentation points to segment the video to be annotated. Target tracking processing is then performed on two sets of image sequences located before and after the keyframes in the video to be annotated. Thus, without changing the working principle of the target tracking algorithm, the video annotation method does not need to separately determine whether the first frame of the video to be annotated contains the target object, achieving automation of video annotation and effectively improving the efficiency and accuracy of video annotation.
[0040] The following is an exemplary description of each step in the video annotation method provided in the embodiments of this application:
[0041] S101, based on the standard keyframes of the target object, determine the image frames that match the standard keyframes from the video to be labeled as the keyframes of the target object in the video to be labeled.
[0042] Here, the target object refers to the entity object that needs to be labeled in the video to be labeled. The specific type of entity object represented by the target object can be determined according to the video shooting scene of the video to be labeled. The target object can represent only one entity object, or it can include multiple entity objects of different categories. This application embodiment does not limit the specific type and number of entity objects represented by the target object.
[0043] For example, taking the video shooting scene of the video to be labeled as an excavator operation scene as an example, the above-mentioned target object can be the excavator body (including the boom and bucket), or the excavator body, the material to be loaded (such as a pile of soil, a pile of materials, etc.), the truck transporting the material, etc.
[0044] Here, the above-mentioned keyframe annotation includes the location marker information of the target object. That is, the above-mentioned standard keyframe can be obtained by marking the target object in the image frame containing the target object.
[0045] It should be noted that the aforementioned standard keyframes may or may not come from the video to be labeled. For example, based on the video shooting scene of the video to be labeled, multiple videos from that shooting scene can be obtained. By selecting the image frame containing the target object from the multiple videos and marking the target object contained in the image frame, the aforementioned standard image frame can be obtained. The standard image frame obtained in this way is applicable to the video labeling of target objects in all videos to be labeled taken in the same video shooting scene, thereby improving the efficiency of video labeling.
[0046] Here, considering that the aforementioned standard keyframes may not belong to the video to be labeled, when performing step S101, the image frame with the highest degree of matching with the aforementioned standard keyframes can be determined from the video to be labeled by feature matching as the keyframe of the target object in the video to be labeled (i.e., the image frame that matches the aforementioned standard keyframes).
[0047] Specifically, as an optional embodiment, step S101 can be performed using ORB (Oriented Fast and Rotated BRIEF, a feature detection and description algorithm) feature matching. In the ORB feature matching method, for each image frame in the video to be labeled, the ORB algorithm can first detect key points in the image frame, and then calculate a feature vector (i.e., the ORB feature corresponding to the key point) for each key point, thus obtaining multiple ORB features corresponding to the image frame. Similarly, multiple ORB features corresponding to the standard key frame can also be obtained through the ORB algorithm. When the Euclidean distance between ORB features is less than a preset threshold, it can be determined that two ORB features match. The more ORB features that match, the higher the degree of matching between the image frame and the standard key frame. Thus, the image frame with the highest degree of matching with the standard key frame can be determined from the video to be labeled as the key frame of the target object in the video to be labeled.
[0048] It should be noted that the feature matching method that can be used when performing step S101 is not unique. This application embodiment does not limit the specific implementation method of the above feature matching.
[0049] S102, the standard keyframe and the keyframe are combined with the first image sequence to obtain the first sub-video, and the standard keyframe and the keyframe are combined with the second image sequence to obtain the second sub-video.
[0050] Here, the first image sequence mentioned above represents the image sequence obtained by reversing the image frames in the video to be labeled that are located before the keyframe.
[0051] Specifically, Figure 2 shows a schematic diagram of the structure of a first sub-video and a second sub-video provided in an embodiment of this application. As shown in Figure 2, the video to be labeled is segmented using the keyframe determined in step S101 as the video segmentation point. Multiple image frames in the video to be labeled that are located before the keyframe can be determined (starting from the first frame in the video to be labeled shown in Figure 2, up to the frame before the keyframe). Considering that the target tracking algorithm can usually only perform target tracking processing on image sequences in which the target object appears in the first frame and the shooting time is continuous, when performing step S102, after placing the standard keyframe containing the target object and the keyframe in front, it is also necessary to reverse the multiple image frames in the video to be labeled that are located before the keyframe to obtain a first image sequence whose shooting time is continuous with the keyframe. Thus, according to the arrangement order of the standard keyframe, the keyframe and the first image sequence, the first sub-video (i.e., the video that can be target tracked by the target tracking algorithm) composed of the standard keyframe, the keyframe and the first image sequence is obtained.
[0052] Here, the second image sequence described above represents an image sequence consisting of image frames located after the keyframe in the video to be annotated.
[0053] Specifically, as shown in Figure 2, after segmenting the video to be labeled using the keyframes determined in step S101 as video segmentation points, multiple image frames located after the keyframes in the video to be labeled (starting from the frame after the keyframe until the last frame in the video to be labeled) can also be determined. Considering that target tracking algorithms can usually only perform target tracking processing on image sequences where the target object appears in the first frame and the shooting time is continuous, when executing step S102, after placing the standard keyframe containing the target object and the keyframe at the front, the multiple image frames located after the keyframe that are continuous with the shooting time of the keyframe can be directly used as the second image sequence. Thus, according to the arrangement order of the standard keyframe, the keyframe, and the second image sequence, the second sub-video (i.e., the video that can be target tracked by the target tracking algorithm) composed of the standard keyframe, the keyframe, and the second image sequence is obtained.
[0054] S103, using the target object as the tracking target, target tracking processing is performed on the first sub-video and the second sub-video respectively to obtain the target tracking processing results corresponding to the first sub-video and the second sub-video respectively.
[0055] Here, when performing target tracking processing on the target objects contained in the first sub-video using the target tracking algorithm, based on the position marking information of the target objects contained in the standard keyframes, the target tracking algorithm can use the standard keyframes located at the foremost position in the first sub-video as prompt information related to the tracking target (equivalent to clearly identifying the target objects marked in the standard keyframes as the tracking targets). Based on this prompt information, the algorithm performs target tracking processing on the keyframes and the first image sequence that are captured in consecutive time in the first sub-video, and obtains the target tracking processing result corresponding to the first sub-video (i.e., the first sub-video marked with the position information of the target objects in each image frame of the first sub-video).
[0056] Here, when performing target tracking processing on the aforementioned target object contained in the second sub-video using the target tracking algorithm, the target tracking algorithm can also use the aforementioned standard keyframe located at the foremost position in the second sub-video as cue information related to the tracking target. Based on this cue information, target tracking processing is performed on the aforementioned keyframes and the aforementioned second image sequence that are captured in consecutive time in the second sub-video, to obtain the target tracking processing result corresponding to the second sub-video (i.e., the second sub-video marked with the position information of the target object in each image frame of the second sub-video).
[0057] It should be noted that the above target tracking algorithm can be the TrackAnything algorithm, or other target tracking algorithms that can only perform target tracking processing on image sequences in which the target object appears in the first frame and the shooting time is continuous. The embodiments of this application do not limit the specific algorithm represented by the above target tracking algorithm.
[0058] S104, based on the order of the image frames in the video to be labeled, the target tracking processing results corresponding to the first sub-video and the second sub-video are recombined to obtain the target tracking result of the target object in the video to be labeled.
[0059] Here, referring to the content of steps S102-S103 above, it can be seen that since the first image sequence in the first sub-video is the image sequence obtained by reversing the image frames in the video to be labeled that are located before the key frame, after obtaining the target tracking processing results corresponding to the first sub-video and the second sub-video respectively, it is also necessary to adjust the arrangement order of each image frame in the target tracking processing results corresponding to the first sub-video and the second sub-video respectively according to the sorting of the image frames in the video to be labeled, so as to obtain the labeled video with the same sorting as the video to be labeled and containing the position marking information of the target object as the target tracking result of the target object in the video to be labeled.
[0060] In this embodiment of the application, in conjunction with the above description, step S104 can be performed according to the method shown in steps a1-a4 below, specifically:
[0061] Step a1: Obtain the target tracking processing result of the first image sequence from the target tracking processing result corresponding to the first sub-video.
[0062] Here, referring to Figure 2, since the video to be labeled does not contain the above-mentioned standard keyframes, and the above-mentioned keyframes appear in both the first sub-video and the second sub-video, the target tracking processing result of the first image sequence (i.e., the first image sequence marked with the position information of the target object in each image frame of the first image sequence) can be obtained from the target tracking processing result corresponding to the first sub-video.
[0063] Step a2: Obtain the target tracking processing result of the second image sequence from the target tracking processing result corresponding to the second sub-video.
[0064] Here, referring to Figure 2, since the video to be labeled does not contain the above-mentioned standard keyframes, and the above-mentioned keyframes appear in both the first sub-video and the second sub-video, the target tracking processing result of the second image sequence (i.e., the second image sequence marked with the position information of the target object in each image frame of the second image sequence) can be obtained from the target tracking processing result corresponding to the second sub-video.
[0065] Step a3: Obtain the target tracking processing result of the key frame from the target tracking processing result corresponding to the first sub-video or the second sub-video.
[0066] Here, referring to Figure 2, since the keyframes appear in both the first sub-video and the second sub-video, the target tracking processing results of the keyframes (i.e., keyframes marked with the position information of the target object in the keyframes) can be obtained from the target tracking processing results corresponding to the first sub-video, or the target tracking processing results of the keyframes can be obtained from the target tracking processing results corresponding to the second sub-video, as long as there is no duplicate acquisition.
[0067] Step a4: Based on the order of the image frames in the video to be labeled, adjust the order of the image frames in the target tracking processing results of the first image sequence, the target tracking processing results of the key frames, and the target tracking processing results of the second image sequence to obtain the target tracking result of the target object in the video to be labeled.
[0068] Here, referring to Figure 2, based on the order of the image frames in the video to be labeled, the labeled image frames in the target tracking processing result of the first image sequence can be reversed to obtain the target tracking processing result of the first image sequence after reversal. Then, according to the order of the target tracking processing result of the first image sequence after reversal, the target tracking processing result of the key frame, and the target tracking processing result of the second image sequence, the labeled video composed of the above three target tracking processing results is obtained (i.e., the target tracking result of the target object in the video to be labeled).
[0069] It should be noted that, regardless of whether there is one or multiple target objects, each target object (that is, each entity object included in the target object) in the video to be annotated can be annotated according to the annotation method shown in steps S101-S104. The repetitive parts will not be repeated here.
[0070] The specific implementation process of each of the above steps in the embodiments of this application will be described in detail below:
[0071] Regarding the method for determining the standard keyframe in step S101 above, in an optional embodiment, Figure 3 shows a flowchart of a method for determining the standard keyframe provided by an embodiment of this application. As shown in Figure 3, before executing step S101, the method includes steps S301-S302; specifically:
[0072] S301, based on the video shooting scene of the video to be labeled, determine the target image frame containing the target object from multiple original videos corresponding to the video shooting scene.
[0073] Here, the original video mentioned above can be a historical video taken in the video shooting scene. For example, if the video shooting scene is an excavator working scene, the original video mentioned above can be a historical video of multiple excavation operations of the excavator taken within a certain time period.
[0074] Specifically, when there are multiple image frames containing the target object in the aforementioned multiple original videos, as an optional embodiment, the image frame containing the target object and displaying the target object most clearly and completely can be determined from the multiple original videos as the target image frame.
[0075] It should be noted that, in the embodiments of this application, it is only necessary to ensure that the target image frame contains the target object. The embodiments of this application do not impose any limitations on the specific selection criteria for the target image frame.
[0076] S302, the target image frame is input into the image segmentation model, the target object contained in the target image frame is labeled by the image segmentation model, and the labeled target image frame is output as the standard keyframe of the target object.
[0077] Here, the target image frame is input into the image segmentation model. The image segmentation model can use the specified target object as a segmentation cue, predict local image regions in the target image frame that belong to the same entity category as the segmentation cue, and generate a segmentation mask corresponding to each predicted local region as the segmentation mask corresponding to the target object (that is, to label the target object contained in the target image frame); where the segmentation mask corresponding to the target object is the position marking information of the target object in the target image frame.
[0078] It should be noted that the above image segmentation models include, but are not limited to, the SAM model (Segment Anything model), the Mask-RCNN model, etc.; the specific model structure of the above image segmentation models is not limited in any way in the embodiments of this application.
[0079] Based on the video annotation method shown in steps S101-S104 above, considering that when there are multiple target objects to be annotated in the video to be annotated (i.e., the target objects include multiple entity objects of different categories), since multiple target objects may appear in the same image frame, there may be an overlap of the position markers (i.e. the above segmentation mask) of multiple target objects in the same image frame.
[0080] Here, to solve the above problems, in an optional implementation, Figure 4 shows a flowchart of a method for deduplicating target tracking results of multiple target objects in a video to be labeled, provided by an embodiment of this application. As shown in Figure 4, after executing step S104, the method includes steps S401-S402; specifically:
[0081] S401, when the target object includes multiple entity objects of different categories, based on the target tracking results corresponding to the multiple entity objects in the video to be labeled, determine whether there are abnormal image frames with overlapping position markers among the target tracking results of the multiple entity objects.
[0082] Here, when there are multiple target objects (i.e., the target objects include multiple entity objects of different categories), since the target tracking processing is performed on multiple different entity objects appearing in the same video to be labeled, the target tracking results corresponding to the multiple entity objects in the video to be labeled (i.e., the segmentation mask of each entity object in each image frame in the video to be labeled) can be used to determine if the segmentation masks (i.e., mask markers) of multiple (two or more) entity objects overlap in the same image frame. Then, the image frame where the segmentation masks overlap is the above-mentioned abnormal image frame, and the overlapping segmentation mask is the local image region where the position markers overlap in the abnormal image frame.
[0083] S401, if it is determined that the abnormal image frame exists, segmentation prediction is performed on the local image region where the position markers overlap in the abnormal image frame, and the entity objects to which different pixel units in the local image region belong are determined according to the segmentation prediction result.
[0084] Here, the method for segmentation prediction of local image regions (i.e. overlapping segmentation masks) with overlapping position markers in abnormal image frames is not unique, and the embodiments of this application do not limit it in any way.
[0085] Specifically, based on the fact that the aforementioned local image region (i.e., the overlapping segmentation mask) contains multiple sub-image regions enclosed by region boundaries (i.e., in the overlapping segmentation mask, multiple local regions enclosed by line segments often appear), in the first optional embodiment, for the multiple sub-image regions contained in the aforementioned local image region, according to the region boundary enclosing the sub-image region, the entity object whose position marker is connected to the region boundary is determined as the segmentation prediction result of the sub-image region.
[0086] For example, taking a target object including an excavator body and a pile of excavated material as an example, if the segmentation masks of the excavator body and the pile of material overlap in the same image frame, the image frame can be determined to be the aforementioned abnormal image frame. The segmentation mask of the overlapping part is determined to be the local image region of the aforementioned position mark overlap. At this time, for each local region (i.e., each sub-image region) enclosed by line segments in the segmentation mask of the overlapping part, it can be determined whether the region boundary (i.e., the line segment enclosing the local region) of the local region is connected to the segmentation mask of the excavator body or the segmentation mask of the pile of material. If the region boundary of the local region is connected to the segmentation mask of the excavator body, it can be determined that the local region belongs to the segmentation mask of the excavator body. If the region boundary of the local region is connected to the segmentation mask of the pile of material, it can be determined that the local region belongs to the segmentation mask of the pile of material. If there is no connection, the local region can be removed and the next local region can be determined.
[0087] Specifically, in the second optional implementation, the above-mentioned local image region (i.e., the overlapping segmentation mask) can be input into the image segmentation model by means of an image segmentation model, and the entity objects contained in the local image region can be segmented and predicted by the image segmentation model to obtain the segmentation prediction result of the local image region.
[0088] Here, the aforementioned local image regions (i.e., overlapping segmentation masks) are input into the image segmentation model. Based on the known information that multiple overlapping entity objects are involved in the aforementioned local image regions, the image segmentation model can use the aforementioned known information as segmentation cues. It can predict image regions from the local image regions that belong to the same entity category as each entity object involved in the aforementioned segmentation cues, and generate a segmentation mask corresponding to each predicted image region as the segmentation prediction result for that image region.
[0089] Specifically, as another optional embodiment, the segmentation prediction results of the local image regions with overlapping location markers can be comprehensively determined by combining the two optional implementation methods of the above-mentioned segmentation prediction, according to the methods described in steps b1-b4 below:
[0090] Step b1: For the multiple sub-image regions contained in the local image region, based on the region boundary that surrounds the sub-image region, determine the entity object whose position marker is connected to the boundary of the region as the segmentation prediction result of the sub-image region.
[0091] Here, the specific implementation of step b1 is the same as the first optional implementation described above, and the repetitions will not be repeated here.
[0092] Step b2: Input the local image region into the image segmentation model, and use the image segmentation model to segment and predict the entity objects contained in the local image region to obtain the segmentation prediction result of the local image region.
[0093] Here, the specific implementation of step b2 is the same as the second optional implementation described above, and the repetitions will not be repeated here.
[0094] Step b3: If the segmentation prediction result of the sub-image region matches the segmentation prediction result of the local image region, then determine the entity object to which the sub-image region belongs based on the segmentation prediction result of the sub-image region.
[0095] Here, if the segmentation prediction result of the sub-image region is consistent with (i.e., matches) the segmentation prediction result of the local image region, then the entity object to which the sub-image region belongs can be directly determined as the segmentation prediction result of the sub-image region.
[0096] Step b4: If the segmentation prediction result of the sub-image region does not match the segmentation prediction result of the local image region, then remove the sub-image region from the local image region.
[0097] Here, if the segmentation prediction result of the sub-image region is inconsistent with the segmentation prediction result of the local image region (i.e., mismatch), the sub-image region is removed from the local image region (equivalent to removing the sub-image region whose entity object is uncertain).
[0098] Based on the video annotation method provided in this application embodiment, keyframe matching is used to automatically determine keyframes containing target objects from the video to be annotated. The keyframes are then used as video segmentation points to segment the video to be annotated. Target tracking processing is performed on two sets of image sequences located before and after the keyframes in the video to be annotated. Thus, without changing the working principle of the target tracking algorithm, the video annotation method does not need to separately determine whether the first frame of the video to be annotated contains the target object, achieving automation of video annotation and effectively improving the efficiency and accuracy of video annotation.
[0099] Based on the same inventive concept, this application also provides a video annotation device corresponding to the above-mentioned video annotation method. Since the principle of solving the problem by the video annotation device in the embodiments of this application is similar to that of the above-mentioned video annotation method in the embodiments of this application, the implementation of the video annotation device can refer to the implementation of the above-mentioned video annotation method, and the repeated parts will not be described again.
[0100] Referring to Figure 5, which shows a schematic diagram of a video annotation device provided in an embodiment of this application, the video annotation device includes:
[0101] The matching module 501 is used to determine, based on the standard keyframes of the target object, an image frame that matches the standard keyframe from the video to be labeled as the keyframe of the target object in the video to be labeled; wherein, the standard keyframe includes the location marker information of the target object;
[0102] Grouping module 502 is used to combine the standard keyframe, the keyframe and a first image sequence to obtain a first sub-video, and to combine the standard keyframe and a second image sequence to obtain a second sub-video; wherein, the first image sequence represents the image sequence obtained by reversing the image frames in the video to be labeled that are located before the keyframe; the second image sequence represents the image sequence composed of the image frames in the video to be labeled that are located after the keyframe;
[0103] The target tracking module 503 is used to perform target tracking processing on the first sub-video and the second sub-video respectively, taking the target object as the tracking target, to obtain the target tracking processing results corresponding to the first sub-video and the second sub-video respectively;
[0104] The recombination module 504 is used to recombine the target tracking processing results corresponding to the first sub-video and the second sub-video respectively according to the sorting of image frames in the video to be labeled, so as to obtain the target tracking result of the target object in the video to be labeled.
[0105] In an optional embodiment, the video annotation device further includes: a standard keyframe determination module, wherein the standard keyframe determination module is used to:
[0106] Based on the video shooting scene of the video to be labeled, a target image frame containing the target object is determined from multiple original videos corresponding to the video shooting scene;
[0107] The target image frame is input into the image segmentation model, and the target object contained in the target image frame is labeled by the image segmentation model. The labeled target image frame is then output as the standard keyframe of the target object.
[0108] In an optional implementation, when recombining the target tracking processing results corresponding to the first sub-video and the second sub-video respectively according to the order of image frames in the video to be labeled, the recombining module 504 is used to:
[0109] Obtain the target tracking processing result of the first image sequence from the target tracking processing result corresponding to the first sub-video;
[0110] Obtain the target tracking processing result of the second image sequence from the target tracking processing result corresponding to the second sub-video;
[0111] Obtain the target tracking processing result of the key frame from the target tracking processing result corresponding to the first sub-video or the second sub-video;
[0112] Based on the order of the image frames in the video to be labeled, the order of the image frames in the target tracking processing results of the first image sequence, the target tracking processing results of the key frames, and the target tracking processing results of the second image sequence is adjusted to obtain the target tracking result of the target object in the video to be labeled.
[0113] In one optional embodiment, the video annotation device further includes a deduplication module, wherein the deduplication module is used to:
[0114] When the target object includes multiple entity objects of different categories, based on the target tracking results corresponding to the multiple entity objects in the video to be labeled, it is determined whether there are abnormal image frames with overlapping position markers among the target tracking results of the multiple entity objects.
[0115] If the abnormal image frame is determined to exist, the local image region with overlapping position markers in the abnormal image frame is segmented and predicted, and the entity objects to which different pixel units in the local image region belong are determined based on the segmentation and prediction results.
[0116] In an optional implementation, when performing segmentation prediction on local image regions with overlapping location markers in the abnormal image frame, the deduplication module is used to:
[0117] For multiple sub-image regions contained in the local image region, based on the region boundary that surrounds the sub-image region, the entity object whose position marker is connected to the boundary of the region is determined as the segmentation prediction result of the sub-image region.
[0118] In an optional implementation, when performing segmentation prediction on local image regions with overlapping location markers in the abnormal image frame, the deduplication module is used to:
[0119] The local image region is input into the image segmentation model, and the image segmentation model is used to segment and predict the entity objects contained in the local image region to obtain the segmentation prediction result of the local image region.
[0120] In an optional implementation, when performing segmentation prediction on local image regions with overlapping location markers in the abnormal image frame, the deduplication module is used to:
[0121] For multiple sub-image regions contained in the local image region, based on the region boundary that surrounds the sub-image region, the entity object whose position marker is connected to the boundary of the region is determined as the segmentation prediction result of the sub-image region;
[0122] The local image region is input into the image segmentation model, and the image segmentation model is used to segment and predict the entity objects contained in the local image region to obtain the segmentation prediction result of the local image region.
[0123] If the segmentation prediction result of the sub-image region matches the segmentation prediction result of the local image region, then the entity object to which the sub-image region belongs is determined based on the segmentation prediction result of the sub-image region.
[0124] If the segmentation prediction result of the sub-image region does not match the segmentation prediction result of the local image region, then the sub-image region is removed from the local image region.
[0125] Based on the video annotation apparatus provided in this application embodiment, the keyframe containing the target object is automatically determined from the video to be annotated by keyframe matching. The keyframe is used as the video segmentation point to segment the video to be annotated. Target tracking processing is then performed on the two sets of image sequences located before and after the keyframe in the video to be annotated. Thus, without changing the working principle of the target tracking algorithm, the video annotation method does not need to separately determine whether the first frame of the video to be annotated contains the target object, thereby achieving automation of video annotation and effectively improving the efficiency and accuracy of video annotation.
[0126] Based on the same inventive concept, this application also provides an electronic device corresponding to the above-mentioned video annotation method. Since the principle of solving the problem by the electronic device in the embodiments of this application is similar to that of the above-mentioned video annotation method in the embodiments of this application, the implementation of the electronic device can refer to the implementation of the above-mentioned video annotation method, and the repeated parts will not be described again.
[0127] Figure 6 is a schematic diagram of the structure of an electronic device 600 provided in an embodiment of this application, including: a processor 601, a memory 602, and a bus 603. The memory 602 stores machine-readable instructions executable by the processor 601. When the electronic device runs a video annotation method as described in the embodiment, the processor 601 communicates with the memory 602 through the bus 603. The processor 601 executes the machine-readable instructions, wherein the processor 601 executes the following steps when executing the machine-readable instructions:
[0128] Based on the standard keyframes of the target object, image frames that match the standard keyframes are determined from the video to be labeled as keyframes of the target object in the video to be labeled; wherein, the standard keyframes include the location marker information of the target object;
[0129] The standard keyframe, the keyframe, and the first image sequence are combined to obtain a first sub-video; the standard keyframe and the keyframe are combined to obtain a second sub-video; wherein, the first image sequence represents the image sequence obtained by reversing the image frames in the video to be labeled that are located before the keyframe; the second image sequence represents the image sequence composed of the image frames in the video to be labeled that are located after the keyframe.
[0130] Using the target object as the tracking target, target tracking processing is performed on the first sub-video and the second sub-video respectively to obtain the target tracking processing results corresponding to the first sub-video and the second sub-video respectively;
[0131] Based on the order of image frames in the video to be labeled, the target tracking processing results corresponding to the first sub-video and the second sub-video are recombined to obtain the target tracking result of the target object in the video to be labeled.
[0132] In an alternative implementation, the processor 601 is further configured to:
[0133] Based on the video shooting scene of the video to be labeled, a target image frame containing the target object is determined from multiple original videos corresponding to the video shooting scene;
[0134] The target image frame is input into the image segmentation model, and the target object contained in the target image frame is labeled by the image segmentation model. The labeled target image frame is then output as the standard keyframe of the target object.
[0135] In an optional implementation, when recombining the target tracking processing results corresponding to the first sub-video and the second sub-video respectively according to the order of image frames in the video to be labeled, the processor 601 is used to:
[0136] Obtain the target tracking processing result of the first image sequence from the target tracking processing result corresponding to the first sub-video;
[0137] Obtain the target tracking processing result of the second image sequence from the target tracking processing result corresponding to the second sub-video;
[0138] Obtain the target tracking processing result of the key frame from the target tracking processing result corresponding to the first sub-video or the second sub-video;
[0139] Based on the order of the image frames in the video to be labeled, the order of the image frames in the target tracking processing results of the first image sequence, the target tracking processing results of the key frames, and the target tracking processing results of the second image sequence is adjusted to obtain the target tracking result of the target object in the video to be labeled.
[0140] In an alternative implementation, the processor 601 is further configured to:
[0141] When the target object includes multiple entity objects of different categories, based on the target tracking results corresponding to the multiple entity objects in the video to be labeled, it is determined whether there are abnormal image frames with overlapping position markers among the target tracking results of the multiple entity objects.
[0142] If the abnormal image frame is determined to exist, the local image region with overlapping position markers in the abnormal image frame is segmented and predicted, and the entity objects to which different pixel units in the local image region belong are determined based on the segmentation and prediction results.
[0143] In an optional implementation, when performing segmentation prediction on local image regions with overlapping location markers in the anomalous image frame, the processor 601 is configured to:
[0144] For multiple sub-image regions contained in the local image region, based on the region boundary that surrounds the sub-image region, the entity object whose position marker is connected to the boundary of the region is determined as the segmentation prediction result of the sub-image region.
[0145] In an optional implementation, when performing segmentation prediction on local image regions with overlapping location markers in the anomalous image frame, the processor 601 is configured to:
[0146] The local image region is input into the image segmentation model, and the image segmentation model is used to segment and predict the entity objects contained in the local image region to obtain the segmentation prediction result of the local image region.
[0147] In an optional implementation, when performing segmentation prediction on local image regions with overlapping location markers in the anomalous image frame, the processor 601 is configured to:
[0148] For multiple sub-image regions contained in the local image region, based on the region boundary that surrounds the sub-image region, the entity object whose position marker is connected to the boundary of the region is determined as the segmentation prediction result of the sub-image region;
[0149] The local image region is input into the image segmentation model, and the image segmentation model is used to segment and predict the entity objects contained in the local image region to obtain the segmentation prediction result of the local image region.
[0150] If the segmentation prediction result of the sub-image region matches the segmentation prediction result of the local image region, then the entity object to which the sub-image region belongs is determined based on the segmentation prediction result of the sub-image region.
[0151] If the segmentation prediction result of the sub-image region does not match the segmentation prediction result of the local image region, then the sub-image region is removed from the local image region.
[0152] The electronic device provided in this application automatically determines keyframes containing target objects from the video to be labeled by keyframe matching, and uses the keyframes as video segmentation points to segment the video to be labeled. Target tracking processing is then performed on two sets of image sequences in the video to be labeled, one before the keyframe and the other after the keyframe. This eliminates the need to separately determine whether the first frame of the video to be labeled contains the target object without changing the working principle of the target tracking algorithm, thus automating video labeling and effectively improving the efficiency and accuracy of video labeling.
[0153] Based on the same inventive concept, embodiments of this application also provide a computer-readable storage medium storing a computer program, which is executed by a processor, wherein the processor performs the following steps:
[0154] Based on the standard keyframes of the target object, image frames that match the standard keyframes are determined from the video to be labeled as keyframes of the target object in the video to be labeled; wherein, the standard keyframes include the location marker information of the target object;
[0155] The standard keyframe, the keyframe, and the first image sequence are combined to obtain a first sub-video; the standard keyframe and the keyframe are combined to obtain a second sub-video; wherein, the first image sequence represents the image sequence obtained by reversing the image frames in the video to be labeled that are located before the keyframe; the second image sequence represents the image sequence composed of the image frames in the video to be labeled that are located after the keyframe.
[0156] Using the target object as the tracking target, target tracking processing is performed on the first sub-video and the second sub-video respectively to obtain the target tracking processing results corresponding to the first sub-video and the second sub-video respectively;
[0157] Based on the order of image frames in the video to be labeled, the target tracking processing results corresponding to the first sub-video and the second sub-video are recombined to obtain the target tracking result of the target object in the video to be labeled.
[0158] In one alternative implementation, the processor is further configured to:
[0159] Based on the video shooting scene of the video to be labeled, a target image frame containing the target object is determined from multiple original videos corresponding to the video shooting scene;
[0160] The target image frame is input into the image segmentation model, and the target object contained in the target image frame is labeled by the image segmentation model. The labeled target image frame is then output as the standard keyframe of the target object.
[0161] In an optional implementation, when recombining the target tracking processing results corresponding to the first sub-video and the second sub-video respectively according to the order of image frames in the video to be labeled, the processor is configured to:
[0162] Obtain the target tracking processing result of the first image sequence from the target tracking processing result corresponding to the first sub-video;
[0163] Obtain the target tracking processing result of the second image sequence from the target tracking processing result corresponding to the second sub-video;
[0164] Obtain the target tracking processing result of the key frame from the target tracking processing result corresponding to the first sub-video or the second sub-video;
[0165] Based on the order of the image frames in the video to be labeled, the order of the image frames in the target tracking processing results of the first image sequence, the target tracking processing results of the key frames, and the target tracking processing results of the second image sequence is adjusted to obtain the target tracking result of the target object in the video to be labeled.
[0166] In one alternative implementation, the processor is further configured to:
[0167] When the target object includes multiple entity objects of different categories, based on the target tracking results corresponding to the multiple entity objects in the video to be labeled, it is determined whether there are abnormal image frames with overlapping position markers among the target tracking results of the multiple entity objects.
[0168] If the abnormal image frame is determined to exist, the local image region with overlapping position markers in the abnormal image frame is segmented and predicted, and the entity objects to which different pixel units in the local image region belong are determined based on the segmentation and prediction results.
[0169] In one optional implementation, when performing segmentation prediction on local image regions with overlapping location markers in the anomalous image frame, the processor is configured to:
[0170] For multiple sub-image regions contained in the local image region, based on the region boundary that surrounds the sub-image region, the entity object whose position marker is connected to the boundary of the region is determined as the segmentation prediction result of the sub-image region.
[0171] In one optional implementation, when performing segmentation prediction on local image regions with overlapping location markers in the anomalous image frame, the processor is configured to:
[0172] The local image region is input into the image segmentation model, and the image segmentation model is used to segment and predict the entity objects contained in the local image region to obtain the segmentation prediction result of the local image region.
[0173] In one optional implementation, when performing segmentation prediction on local image regions with overlapping location markers in the anomalous image frame, the processor is configured to:
[0174] For multiple sub-image regions contained in the local image region, based on the region boundary that surrounds the sub-image region, the entity object whose position marker is connected to the boundary of the region is determined as the segmentation prediction result of the sub-image region;
[0175] The local image region is input into the image segmentation model, and the image segmentation model is used to segment and predict the entity objects contained in the local image region to obtain the segmentation prediction result of the local image region.
[0176] If the segmentation prediction result of the sub-image region matches the segmentation prediction result of the local image region, then the entity object to which the sub-image region belongs is determined based on the segmentation prediction result of the sub-image region.
[0177] If the segmentation prediction result of the sub-image region does not match the segmentation prediction result of the local image region, then the sub-image region is removed from the local image region.
[0178] The computer-readable storage medium provided in this application embodiment automatically determines keyframes containing target objects from the video to be labeled by keyframe matching. The keyframes are then used as video segmentation points to segment the video to be labeled. Target tracking processing is performed on two sets of image sequences located before and after the keyframes in the video to be labeled. Thus, without changing the working principle of the target tracking algorithm, the video labeling method does not need to separately determine whether the first frame of the video to be labeled contains the target object, achieving automation of video labeling and effectively improving the efficiency and accuracy of video labeling.
[0179] In this embodiment, the computer-readable storage medium can also execute other machine-readable instructions when the processor runs, to perform the video annotation methods described in other embodiments. For details on the specific video annotation method steps and principles, please refer to the description of the method-side embodiments, which will not be repeated here.
[0180] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0181] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0182] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0183] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0184] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0185] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A video annotation method, characterized in that, The video annotation method includes: determining image frames matching the standard keyframes of the target object from the video to be annotated as keyframes of the target object in the video to be annotated; wherein the standard keyframes include the location marker information of the target object; combining the standard keyframes, the keyframes and a first image sequence to obtain a first sub-video, and combining the standard keyframes, the keyframes and a second image sequence to obtain a second sub-video; wherein the first image sequence represents an image sequence obtained by reversing the image frames in the video to be annotated that are located before the keyframes; the second image sequence represents an image sequence composed of image frames in the video to be annotated that are located after the keyframes; using the target object as the tracking target, performing target tracking processing on the first sub-video and the second sub-video respectively to obtain target tracking processing results corresponding to the first sub-video and the second sub-video respectively; and recombining the target tracking processing results corresponding to the first sub-video and the second sub-video respectively according to the order of the image frames in the video to be annotated to obtain the target tracking result of the target object in the video to be annotated.
2. The video annotation method according to claim 1, characterized in that, The video annotation method further includes: determining a target image frame containing the target object from multiple original videos corresponding to the video shooting scene of the video to be annotated; inputting the target image frame into an image segmentation model, annotating the target object contained in the target image frame through the image segmentation model, and outputting the annotated target image frame as the standard keyframe of the target object.
3. The video annotation method according to claim 1, characterized in that, The step of reorganizing the target tracking processing results corresponding to the first sub-video and the second sub-video according to the order of image frames in the video to be labeled includes: obtaining the target tracking processing results of the first image sequence from the target tracking processing results corresponding to the first sub-video; obtaining the target tracking processing results of the second image sequence from the target tracking processing results corresponding to the second sub-video; obtaining the target tracking processing results of the keyframes from the target tracking processing results corresponding to the first sub-video or the second sub-video; and adjusting the order of image frames in the target tracking processing results of the first image sequence, the target tracking processing results of the keyframes, and the target tracking processing results of the second image sequence according to the order of image frames in the video to be labeled, to obtain the target tracking results of the target object in the video to be labeled.
4. The video annotation method according to claim 1, characterized in that, The video annotation method further includes: when the target object includes multiple entity objects of different categories, determining whether there are abnormal image frames with overlapping position markers among the target tracking results of the multiple entity objects in the video to be annotated, based on the target tracking results corresponding to the multiple entity objects respectively; if it is determined that there are abnormal image frames, segmenting and predicting the local image region with overlapping position markers in the abnormal image frame, and determining the entity objects to which different pixel units in the local image region belong based on the segmentation and prediction results.
5. The video annotation method according to claim 4, characterized in that, The segmentation prediction of local image regions with overlapping location markers in the abnormal image frame includes: for multiple sub-image regions contained in the local image region, determining the entity object whose location marker is connected to the boundary of the region that encloses the sub-image region as the segmentation prediction result of the sub-image region based on the region boundary that encloses the sub-image region.
6. The video annotation method according to claim 4, characterized in that, The step of segmenting and predicting local image regions with overlapping location markers in the abnormal image frame further includes: inputting the local image regions into an image segmentation model, and using the image segmentation model to segment and predict entity objects contained in the local image regions to obtain the segmentation prediction result of the local image regions.
7. The video annotation method according to claim 4, characterized in that, The segmentation prediction of local image regions with overlapping location markers in the abnormal image frame further includes: for multiple sub-image regions contained in the local image region, determining the entity object whose location marker is connected to the boundary of the region that encloses the sub-image region as the segmentation prediction result of the sub-image region based on the region boundary enclosing the sub-image region; inputting the local image region into an image segmentation model, and performing segmentation prediction on the entity object contained in the local image region through the image segmentation model to obtain the segmentation prediction result of the local image region; if the segmentation prediction result of the sub-image region matches the segmentation prediction result of the local image region, then determining the entity object to which the sub-image region belongs based on the segmentation prediction result of the sub-image region; if the segmentation prediction result of the sub-image region does not match the segmentation prediction result of the local image region, then removing the sub-image region from the local image region.
8. A video annotation device, characterized in that, The video annotation device includes: a matching module, configured to determine, based on a standard keyframe of the target object, an image frame matching the standard keyframe from the video to be annotated as the keyframe of the target object in the video to be annotated; wherein the standard keyframe includes the location marker information of the target object; a grouping module, configured to combine the standard keyframe, the keyframe and a first image sequence to obtain a first sub-video, and combine the standard keyframe, the keyframe and a second image sequence to obtain a second sub-video; wherein the first image sequence represents an image sequence obtained by reversing the image frames in the video to be annotated that are located before the keyframe; the second image sequence represents an image sequence composed of image frames in the video to be annotated that are located after the keyframe; a target tracking module, configured to use the target object as the tracking target, perform target tracking processing on the first sub-video and the second sub-video respectively, and obtain target tracking processing results corresponding to the first sub-video and the second sub-video respectively; and a recombination module, configured to recombine the target tracking processing results corresponding to the first sub-video and the second sub-video respectively according to the order of the image frames in the video to be annotated, and obtain the target tracking result of the target object in the video to be annotated.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the video annotation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the video annotation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Video semi-automatic target labeling method integrating target detection and tracking
CN110929560A
Video target detection data labeling method and device, equipment and storage medium
CN116935267A