Video labeling method and device, storage medium, electronic equipment and program product

By synchronous playback and labeling in multiple viewing videos using the timeline, the problem of inefficient viewing video labeling in the prior art is solved, and efficient and accurate consistent labeling is achieved.

CN120343313APending Publication Date: 2025-07-18BEIJING HORIZON INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510637741.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing video annotation scheme requires users to label videos from each perspective separately, resulting in repetitive work and inefficiency, making it difficult to ensure the consistency and accuracy of video annotation results from different perspectives.

Method used

Synchronous playback is achieved by using the time axis in a video synchronously collected in multiple different perspectives, and the video clip is determined based on the first frame information and the second frame information of the target time axis, and the target element attribute information in the video clip is automatically marked.

Benefits of technology

It improves the labeling efficiency, accuracy and consistency of videos on multiple perspectives, and reduces the user's operation complexity and probability of errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343313A_ABST
    Figure CN120343313A_ABST
Patent Text Reader

Abstract

The invention discloses a video annotation method and device, a storage medium, electronic equipment and a program product, and the method comprises the steps: responding to an annotation instruction triggered by a target time axis corresponding to a target element contained in a plurality of videos, which is received through a visual interaction interface, in a process of synchronously playing the videos of the same scene, which are synchronously collected from a plurality of different visual angles; based on first frame information and second frame information corresponding to a target time axis, video clips corresponding to the first frame information and the second frame information are determined in multiple videos, and the labeling instruction comprises attribute information of a target element; according to the method, the attribute information is marked for the target element of each frame of image in the video clip to obtain the marked video clip, and the time axis is utilized to realize simultaneous marking of the videos synchronously acquired at multiple different visual angles, so that the efficiency, accuracy and consistency of marking the videos at multiple visual angles can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to data annotation technologies in the field of intelligent driving, and particularly to a video annotation method, apparatus, storage medium, electronic device, and program product. Background Art

[0002] During the driving process of an intelligent vehicle, in order to obtain accurate information about the vehicle's surrounding environment, visual sensors with different perspectives are set on the intelligent vehicle to collect videos, and the elements in the videos collected by the visual sensors with different perspectives are annotated to provide accurate basis for downstream tasks such as target behavior prediction, target trajectory prediction, and vehicle driving path planning. Under the related technologies, the existing video annotation solutions require users to annotate the videos of each perspective separately, resulting in repetitive labor and low efficiency. Moreover, the annotation results of the videos of different perspectives need to be manually synchronized by users, further increasing the labor and error probability of users, and it is difficult to ensure the consistency and accuracy of the annotation results of the videos of multiple perspectives. Summary of the Invention

[0003] To solve the above technical problems, embodiments of the present disclosure provide a video annotation method, apparatus, storage medium, electronic device, and program product, which utilize the time axis to simultaneously annotate videos synchronously collected from multiple different perspectives, helping to improve the efficiency, accuracy, and consistency of annotating videos of multiple perspectives.

[0004] According to one aspect of the embodiments of the present disclosure, there is provided a video annotation method, including: during the synchronous playback of videos of the same scene synchronously collected from multiple different perspectives, in response to a annotation instruction triggered by a target time axis corresponding to a target element included in the multiple videos received through a visual interaction interface, based on the first frame information and the second frame information corresponding to the target time axis, determining video segments corresponding to the first frame information and the second frame information in the multiple videos, where the annotation instruction includes the attribute information of the target element; annotating the attribute information of the target element for each frame image in the video segment to obtain the annotated video segment.

[0005] According to another aspect of the embodiments of the present disclosure, a video annotation device is provided, including: a visual interaction interface module, configured to synchronously play videos of the same scene collected from multiple different perspectives, and interact with a user; a video segment determination module, configured to, during the process of synchronously playing videos of the same scene collected from multiple different perspectives, in response to a marking instruction triggered by a target timeline corresponding to a target element included in multiple videos received through the visual interaction interface, determine video segments corresponding to the first frame information and the second frame information in the multiple videos based on the first frame information and the second frame information corresponding to the target timeline, where the marking instruction includes attribute information of the target element; and a marking module, configured to mark the attribute information on the target element of each frame image in the video segment to obtain marking information of the target element in the video segment.

[0006] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The storage medium stores a computer program, and when the computer program instructions are executed, the above-mentioned video annotation method is implemented.

[0007] According to yet another aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes: a memory, configured to store a computer program product; and a processor, configured to execute the computer program product stored in the memory, and when the computer program product is executed, the above-mentioned video annotation method is implemented.

[0008] According to yet another aspect of the embodiments of the present disclosure, a computer program product is provided, including computer program instructions, and when the computer program instructions are executed by a processor, the above-mentioned video annotation method is implemented.

[0009] Based on the above embodiments of the present disclosure, during the process of synchronously playing videos of the same scene collected from multiple different perspectives, if a marking instruction triggered by a target timeline corresponding to a target element included in multiple videos is received through the visual interaction interface, video segments corresponding to the first frame information and the second frame information can be determined in the multiple videos based on the first frame information and the second frame information corresponding to the target timeline, and the attribute information is marked on the target element of each frame image in the video segment to obtain the marked video segment. Thus, the technical solution provided by the present disclosure realizes the simultaneous annotation of videos collected from multiple different perspectives synchronously by using the timeline, which helps to improve the efficiency, accuracy, and consistency of annotating videos from multiple perspectives; in addition, based on the first frame information and the second frame information corresponding to the target timeline, batch annotation of each frame image in the video segment determined by the first frame information and the second frame information is realized, further improving the annotation efficiency of the video. Description of the Drawings

[0010] Figure 1It is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure;

[0011] Figure 2 It is a schematic flowchart of a video annotation method provided by an exemplary embodiment of the present disclosure;

[0012] Figure 3 It is a schematic diagram of video segment annotation provided by an exemplary embodiment of the present disclosure;

[0013] Figure 4 It is a schematic diagram of an interface for synchronously playing multiple videos provided by an exemplary embodiment of the present disclosure;

[0014] Figure 5 It is a schematic diagram of an interface for synchronously playing multiple videos provided by another exemplary embodiment of the present disclosure;

[0015] Figure 6 It is a schematic flowchart of configuring a timeline in a video annotation method provided by an exemplary embodiment of the present disclosure;

[0016] Figure 7 It is a schematic flowchart of a video annotation method provided by another exemplary embodiment of the present disclosure;

[0017] Figure 8 It is a schematic flowchart of verifying video annotation information provided by an exemplary embodiment of the present disclosure;

[0018] Figure 9 It is a schematic structural diagram of a video annotation device provided by an exemplary embodiment of the present disclosure;

[0019] Figure 10 It is a schematic structural diagram of a video annotation device provided by another exemplary embodiment of the present disclosure. Detailed Description of the Invention

[0020] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.

[0021] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.

[0022] Application Overview

[0023] In the field of autonomous driving, it is very important to annotate elements in video or image sequences collected by visual sensors from different perspectives, providing accurate bases for downstream tasks such as target behavior prediction, target trajectory prediction, and vehicle driving path planning.

[0024] Existing video annotation schemes require users to annotate each video or image sequence collected from each perspective separately, resulting in repetitive labor and low efficiency. Moreover, it is difficult to ensure the consistency of the annotation results of videos or image sequences from different perspectives. In addition, in existing video annotation schemes, it is difficult to synchronously play videos from different perspectives, and there is a lack of comparative playback of videos or image sequences from different perspectives, further increasing the error probability of users annotating videos or image sequences from different perspectives and making it difficult to ensure the consistency and accuracy of the annotation results of videos or image sequences from multiple perspectives.

[0025] In order to improve the annotation efficiency, annotation accuracy, and consistency of video annotation information from multiple perspectives, the inventors have proposed the technical solution of the present disclosure.

[0026] Exemplary Device

[0027] Figure 1 An electronic device to which the video annotation method according to the embodiments of the present disclosure can be applied is shown.

[0028] As Figure 1 shown, the electronic device includes at least one processor 11 and a memory 12.

[0029] The processor 11 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 14 to perform desired functions.

[0030] The memory 12 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 11 can run one or more computer program instructions to implement the video annotation methods of the various embodiments of the present disclosure above and / or other desired functions. The video annotation method includes: during the synchronous playback of videos of the same scene synchronously collected from multiple different perspectives, in response to a marking instruction triggered by a target time axis corresponding to a target element included in multiple videos received through a visual interaction interface, based on the first frame information and the second frame information corresponding to the target time axis, determining video segments corresponding to the first frame information and the second frame information in the multiple videos, where the marking instruction includes the attribute information of the target element; marking the attribute information of the target element for each frame image in the video segment to obtain the marked video segment.

[0031] In one example, the electronic device may further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0032] The input device 13 may further include, for example, a keyboard, a mouse, and the like.

[0033] The output device 14 may output various information to the outside, and it may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and the like.

[0034] Of course, for simplicity, Figure 1 only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, and the like are omitted. In addition, according to specific application scenarios, the electronic device 100 may further include any other appropriate components.

[0035] It should be noted that the electronic device in the technical solution of the present disclosure may be various electronic devices, including but not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like.

[0036] Exemplary Method

[0037] Figure 2 is a schematic flowchart of a video annotation method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to Figure 1 an electronic device or a video annotation device, as Figure 2 shown, and includes the following steps:

[0038] Step 201, during the synchronous playback process of videos of the same scene collected from multiple different perspectives synchronously, in response to a marking instruction triggered by a target time axis corresponding to a target element included in multiple videos received through a visual interaction interface, based on the first frame information and the second frame information corresponding to the target time axis, determine video segments corresponding to the first frame information and the second frame information in multiple videos, and the marking instruction includes the attribute information of the target element.

[0039] Among them, the perspective is used to indicate the orientation for video or image acquisition of the scene. Different perspectives in the embodiments of the present disclosure may include a front view perspective, a left front view perspective, a right front view perspective, a rear view perspective, a left rear view perspective, a right rear view perspective, and the like.

[0040] In the embodiments of the present disclosure, videos of the same scene collected synchronously from different perspectives can be obtained by synchronously collecting through multiple image acquisition devices provided on a vehicle. For example, video data of the vehicle's surrounding environment can be synchronously collected through cameras with a front view perspective, a left front view perspective, a right front view perspective, a rear view perspective, a left rear view perspective, and a right rear view perspective. The present disclosure does not limit the number, installation position, and viewing angle range of the image acquisition devices. The present disclosure also does not limit the type of the image acquisition devices, which can be a monocular camera, a binocular camera, a depth camera, etc.

[0041] Among them, synchronous acquisition means that the acquisition times of the images by multiple image acquisition devices are completely synchronous, and time synchronization can be achieved through devices connected by a network. For example, multiple image acquisition devices can exchange synchronization information with the master clock within a specific period, and use the timestamp information in the sending and receiving times to achieve the synchronization of video acquisition for the same scene by multiple image acquisition devices.

[0042] In specific implementation, when there is an asynchronous situation among multiple videos collected from multiple perspectives, time alignment of each video can be performed by methods such as interpolation.

[0043] Among them, synchronous playback means a playback mode in which multiple videos start playing simultaneously, and can be paused or continued simultaneously through a switching button or other interaction methods.

[0044] In the embodiments of the present disclosure, the target elements may include, but are not limited to, any one or more of the following: left turn signal lights, right turn signal lights, straight-ahead signal lights, traffic signs, lane lines, vehicles, pedestrians, and other static or dynamic elements that need to be marked. The attribute information of the target elements is used to indicate the form, state, movement direction, etc. of the target elements. The attribute information of different types of target elements is different. For example, if the target element is a left turn signal light, the attribute information of the target element may include: the color of the left turn signal light, such as the attribute information of the left turn signal light is red, or the attribute information of the left turn signal light is green; if the target element is a vehicle, the attribute information of the target element may include: the door opening and closing state, the headlight usage state, the relative pose relationship with the vehicle itself, etc.

[0045] Among them, the time axis refers to a timeline used to mark the target elements in the video according to the video segmentation method. The length of the time axis is the same as the time length of the multiple videos to be marked. Each time axis is used to mark a target element. The time axis can be adjusted by the user according to the actual marking requirements and can be segmented based on the attribute change state of the target element. Different time axes can be sequentially displayed in the marking area of the visual interaction interface according to the configured order.

[0046] Among them, the annotation instruction is used to indicate the instruction for annotating the target element in the video. The annotation instruction includes the attribute information of the target element, as well as the first frame information and the second frame information corresponding to the target timeline to be annotated.

[0047] In the implementation manner of the present disclosure, since each timeline is a timeline corresponding to the video, the first frame information and the second frame information corresponding to the timeline can be used to represent the frame numbers of the start frame and the end frame of the video segment in the video respectively.

[0048] In some optional implementation manners, the annotation instruction may further include: the identification information of the target element, which is used to uniquely identify the target element and may be the name of the target element. The attribute information of the target element is used to indicate the form, state, movement direction, etc. of the target element. For example, if the target element is a left-turn signal light, the attribute information of the target element may include the color of the left-turn signal light, such as the attribute information of the left-turn signal light is red, or the color of the left-turn signal light is green, etc.; if the target element is a vehicle, the attribute information of the target element may include: the driving direction of the vehicle, the door opening and closing state, the headlight usage state, the relative pose relationship with the vehicle itself, etc.

[0049] Among them, the type of the target element has been configured in the configuration information of the timeline when setting the timeline. When annotating, it is only necessary to annotate the attribute information such as the form, state, movement direction, etc. of the target element in different image frames.

[0050] Step 202: Annotate the attribute information of the target element in each frame of the video segment to obtain the annotated video segment.

[0051] See Figure 3 , which schematically shows the implementation of the annotation of the straight-ahead lights in each video segment through the timeline corresponding to the target element "straight-ahead light", and the method of annotating the straight-ahead lights in the video segment indicated by the start frame 377 and the end frame 566 through the annotation information setting window ([[]] Figure 3 the window indicated by label 31 in [[])], and by setting the color information and time information, the annotation attribute value "red" of the straight-ahead lights in each frame of the image between the start frame 377 and the end frame 566 can be realized, and the annotated video segment (frame 377 - frame 566) with the straight-ahead lights annotated is obtained.

[0052] In this implementation manner, through the first frame information and the second frame information corresponding to the timeline, video segments are cut out from multiple videos. The video segment is a video segment with the first frame as the start frame and the second frame as the end frame), and then by annotating the video segment, the effect of annotating each frame of the image in the video segment is achieved.

[0053] The video annotation method provided in this embodiment, during the synchronous playback of videos of the same scene collected from multiple different perspectives, if an annotation instruction triggered by a target time axis corresponding to a target element included in multiple videos is received through a visual interaction interface, video segments corresponding to the first frame information and the second frame information can be determined in multiple videos based on the first frame information and the second frame information corresponding to the target time axis, and attribute information of the target element can be annotated for each frame image in the video segments to obtain the annotated video segments. Thus, the technical solution provided in this disclosure uses the time axis to simultaneously annotate videos collected from multiple different perspectives, improving the annotation efficiency, annotation accuracy, and consistency of video annotation for multiple perspectives; in addition, based on the first frame information and the second frame information corresponding to the target time axis, batch annotation of each frame image in the video segments is realized, further improving the video annotation efficiency.

[0054] In some alternative implementation manners, in order to synchronously play videos of the same scene collected from multiple different perspectives, for any video to be annotated, video decoding processing can be performed to obtain an image sequence corresponding to any video to be annotated; the image sequence corresponding to any video to be annotated is rendered on the interface canvas of the corresponding visual interaction interface at the original video frame rate, and the rendering progress of the image sequences in different interface canvases is controlled by a playback component, so that synchronous playback of each video on the visual interaction interface can be realized.

[0055] When synchronously playing videos of the same scene collected from multiple different perspectives through a visual interaction interface, the videos of the same scene collected from multiple different perspectives can be synchronously played in multiple video playback areas of the visual interaction interface in a grid mode, such as Figure 4 shown, playing videos from each perspective in a grid mode helps to improve the viewing effect of users watching multiple videos simultaneously, and can determine, based on the synchronous playback of videos from each perspective, a video with a better display effect for the target element, which helps to improve the video annotation accuracy.

[0056] Such as Figure 5As shown, multiple videos collected from multiple different perspectives can also be synchronously played in multiple video playback areas of the visual interaction interface in window mode. Among them, the multiple video playback areas include a main playback area (the area indicated by label 51) and a secondary playback area (the area indicated by label 52). In the main playback area, different perspective videos can be switched and played based on a video switching instruction triggered by the user. The secondary playback area plays videos other than the video played in the main playback area. By playing videos from each perspective in window mode, it is possible to play the video collected from the best perspective in the main playback area, improving the viewing effect of the user watching the video. While playing videos from other perspectives in the secondary playback area helps to make up for the poor display effect of the target element in some frames of the video played in the main playback area with the videos from other perspectives played in the secondary playback area.

[0057] In the embodiments of the present disclosure, the grid mode and the window mode can be switched based on a mode switching instruction triggered by the user to improve the accuracy of video annotation.

[0058] Figure 6 It is a schematic flowchart of configuring the timeline in the video annotation method provided by an exemplary embodiment of the present disclosure. In this embodiment, an example of how to configure the timeline is used for exemplary illustration. As Figure 6 shown, it includes the following steps:

[0059] Step 601, receive a timeline configuration instruction triggered through the visual interaction interface.

[0060] Among them, the timeline configuration instruction is an instruction used to indicate setting corresponding timelines for each target element to be annotated. Before annotating the video, the user can determine the target elements that may need to be annotated according to the annotation requirements. For example, in the field of autonomous driving, it may be necessary to identify and annotate traffic lights, vehicles parked by the roadside, vehicles entering the ramp, etc. Then, it can be determined that the target elements to be annotated include traffic lights, road signs, parked vehicles, ramps, etc. Furthermore, a configuration instruction for configuring the timelines corresponding to traffic lights, road signs, parked vehicles, and ramps can be triggered through the visual interaction interface.

[0061] In some optional embodiments, multiple timelines can be preset in the video annotation tool (software). The user selects the timelines that may be needed for this annotation from the multiple timelines, and then the timeline configuration instruction can be triggered. For example, timelines associated with elements such as pedestrians, vehicles, traffic lights (including straight-ahead lights), ramps, etc. are preset in the video annotation tool (software). When starting to perform an annotation task, the required timelines can be selected from the multiple preset timelines according to the annotation requirements of this time.

[0062] In some other alternative embodiments, the user can trigger a timeline configuration instruction for a target element to be marked through the timeline configuration menu or panel of the video annotation tool (software), and create a corresponding timeline.

[0063] In specific implementation, according to the requirements of this video annotation, it can be determined which elements in the video may need to be marked. For example, currently, it is necessary to mark the driving scenario of a vehicle, and the elements to be marked include elements such as the vehicle and lane lines. Then, the timeline information of the vehicle and lane lines needs to be configured.

[0064] Step 602: Based on the timeline configuration instruction, determine at least one timeline displayed in the annotation area of the visual interaction interface, and the at least one timeline is used to implement the operation of marking target elements in multiple videos.

[0065] In this embodiment, the timeline displayed in the annotation area of the visual interaction interface is the timeline configured by the timeline configuration instruction. For example, for a traffic signal light, 6 timelines are configured, namely the "straight-ahead light" timeline, the "inferred straight-ahead state" timeline, the "left-turn light" timeline, the "inferred left-turn state" timeline, the "right-turn light" timeline, and the "inferred right-turn state" timeline. Then, these 6 configured timelines can be correspondingly displayed in the annotation area.

[0066] Among them, although only the three target elements of the "straight-ahead light", "left-turn light", and "right-turn light" are marked above, since the attribute information of the three target elements may not be directly determined from the video and needs to be inferred by the annotator based on some elements in the video, therefore, additional timelines for inferring the corresponding states are set for the three target elements: the "inferred left-turn light state" timeline, the "inferred straight-ahead state" timeline, and the "right-turn light" timeline, which helps to mark the target elements better and more comprehensively.

[0067] From the above description, it can be seen that in the technical solution of the present disclosure, the configuration of the timeline is for better marking of target elements. One target element can correspond to one timeline or two timelines. If it corresponds to two timelines, the basis for setting the attribute information of the target element for the two timelines can be different.

[0068] Exemplarily, for the annotation triggered by the time axis of the "straight-ahead light", the generated annotation information is the true straight-ahead light annotation information obtained by observing the target element of the "straight-ahead light", while for the annotation triggered by the time axis of the "inference state of the straight-ahead light", the generated annotation information is the inferred straight-ahead light annotation information obtained by inferring the target element of the "straight-ahead light". Comparing the two, the accuracy of the annotation information triggered by the time axis of the "straight-ahead light" is higher. However, the annotation triggered by the time axis of the "inference state of the straight-ahead light" can assist in determining the situation where the state of the straight-ahead light cannot be accurately observed, which helps to improve the comprehensiveness and robustness of the annotation of the target element.

[0069] Step 603, display at least one time axis in the annotation area of the visual interaction interface.

[0070] The time axis configuration method provided in this embodiment, by configuring corresponding time axes for each target element to be annotated in multiple videos based on video annotation requirements, helps to implement the annotation of target elements in videos by means of video segmentation based on the time axis, improving the annotation efficiency. In addition, flexibly setting the corresponding time axis according to the annotation requirements helps to more accurately control the content display in the annotation area of the visual interaction interface, which can not only meet the annotation requirements of video annotation but also avoid displaying unnecessary content.

[0071] Figure 7 is a schematic flowchart of a video annotation method provided by another exemplary embodiment of the present disclosure; as Figure 7 shown, this embodiment includes the following steps:

[0072] Step 701, in response to receiving an annotation operation triggered by the target time axis corresponding to the target element, display an annotation information setting window in the annotation area of the visual interaction interface.

[0073] In this embodiment, the annotation operation is an operation for the user to trigger video segmentation and implement annotation through the time axis. Specifically, the user can trigger the annotation operation by clicking on the time axis, or trigger the annotation operation through the annotation button on the visual interaction interface.

[0074] Among them, the annotation information setting window is Figure 3 the window indicated by label 31 in

[0075] Step 702, receive the attribute information input by the user through the annotation information setting window.

[0076] Exemplarily, refer to Figure 3 , through the annotation information setting window in the visual interaction interface ( Figure 3For the window indicated by reference numeral 31, a method for a user to annotate the straight-ahead lights in the video clip indicated by the start frame 377 (as the first frame information) and the end frame 566 (as the second frame information). By setting the color information and time information, the annotation attribute value "red" of the straight-ahead lights for each frame image between the start frame 377 and the end frame 566 can be achieved, and the annotated video clip (frames 377 - 566) with the straight-ahead lights annotation completed is obtained.

[0077] In some embodiments, based on the playback progress of the video when the user triggers the annotation operation, the corresponding "time" information can be automatically filled in the annotation information setting window. Specifically, based on the current playback frame of multiple images to be annotated when the annotation operation is received and the historical playback frame of multiple images to be annotated when the previous annotation operation was triggered, the video clip can be determined, where the frame adjacent to the current playback frame before it is the end frame of the video clip, and the historical playback frame is the start frame of the video clip.

[0078] In some embodiments, the time information in the attribute information can be automatically segmented based on the attribute change state of the target element. For example, when the straight-ahead light changes from green to yellow, the time axis is segmented at the time point when it changes to yellow, thereby correspondingly realizing video segmentation for the frame image in the video when the green light changes to yellow. The frame image when the green light changes to yellow belongs to the latter video clip, and the frame image immediately before the frame image when the green light changes to yellow belongs to the former video clip.

[0079] Step 703, in response to receiving the confirmation instruction that the attribute information setting is completed, confirm that the annotation instruction is received.

[0080] In some embodiments, after the input of the attribute information is completed through the annotation information setting window, the confirmation instruction for setting completion can be triggered by a button in the visual interaction interface, whereby the electronic device can receive the annotation instruction for annotating the target element in the video clip (such as the video clip corresponding to frames 377 - 566 above).

[0081] Step 704, based on the first frame information and the second frame information corresponding to the target time axis, determine the video clip corresponding to the first frame information and the second frame information in multiple videos. The annotation instruction includes the attribute information of the target element.

[0082] Step 705, annotate the attribute information of the target element for each frame image in the video clip to obtain the annotated video clip.

[0083] In some embodiments, the implementation manners of steps 704 - 705 can refer to Figure 2 the descriptions of steps 201 - 202 in the illustrated embodiments, which will not be elaborated here.

[0084] The video annotation method provided in this embodiment displays an annotation information setting window in the annotation area of the visual interaction interface by triggering an annotation operation for the target time axis corresponding to the target element. Thus, the attribute information input by the user through the annotation information setting window can be received, and after receiving the confirmation instruction that the attribute information setting is completed, it is confirmed that the annotation instruction is received, so as to realize the annotation of videos simultaneously collected from multiple different perspectives. Based on the annotation operation triggered by the user, this embodiment can automatically generate the start frame (the first frame information) and the end frame (the second frame information) of the video segment, avoiding the complexity of the user manually inputting the start frame and the end frame of the video segment and reducing the error probability.

[0085] After completing the annotation of the target elements of multiple videos, the annotation information corresponding to the target elements of multiple videos can be processed. Figure 8 It is a schematic flowchart of verifying video annotation information provided by an exemplary embodiment of the present disclosure, as Figure 8 shown, including the following steps:

[0086] Step 801, in response to receiving a confirmation instruction that the annotation of the target elements of multiple videos is completed through the visual interaction interface, verify the annotation information corresponding to the target elements of multiple videos.

[0087] Among them, when annotating a video, multiple target elements in the video are usually annotated. After completing the annotation of each target element, the electronic device can receive a confirmation instruction triggered by the user that the annotation is completed. After receiving the confirmation instruction triggered by the user that the annotation is completed, the electronic device will simultaneously receive the annotation information corresponding to this video annotation operation reported by the video annotation tool (software). The annotation information includes the annotation information corresponding to each target element in the videos of multiple perspectives. For example, if this video annotation operation annotates the straight-ahead light, left-turn light, and right-turn light in the videos of multiple perspectives, the annotation information includes the attribute information of the straight-ahead light, left-turn light, and right-turn light in each frame of the videos of multiple perspectives. Since the annotation of the target element is implemented in the way of video segmentation, the attribute information of the target element in multiple consecutive frames of images may be the same. Therefore, the annotation information can also be stored and recorded in the form of video segments.

[0088] During specific verification, the integrity verification of the annotation information for each type of target element can be performed separately. For example, to verify whether the attribute information of the target element "straight-ahead light" is annotated in each frame of the video from multiple perspectives. If the attribute information of the target element "straight-ahead light" is not set for the video segment corresponding to frames 566 - 788 in the video, then there is a problem of missing annotation for the target element, indicating that the verification fails. In some other embodiments, the annotation information of any two or more target elements can also be cross-verified to check whether there are conflicts in the annotation information of different target elements. For example, for the same video segment, the attribute value of the target element "left-turn light" is red, while the attribute value of the target element "inferred left-turn light state" is green. Since "inferred left-turn light state" is the left-turn light state inferred by the user based on certain elements in the video, when their attribute values are different within the same video segment, there is a conflict in their annotation information.

[0089] Exemplarily, according to the video from the left-front perspective, it is determined that in the video segment of frames 377 - 566, the straight-ahead light is red. In the video from the front perspective, a bus completely blocks the left-turn light. However, based on the video of the bus turning left, it can be determined that the vehicle can turn left, and further, the value of the color information in the attribute information of "inferred left-turn light state" can be determined to be green. But within the same time period (the video segment of frames 377 - 566), the attribute value of "left-turn light" and the attribute value of "inferred left-turn light state" should be the same. If they are not the same, it can be inferred that there is a conflict.

[0090] In some optional embodiments, the annotation information corresponding to each annotated video segment can be traversed, and it can be determined whether there is no annotation information on the video segment for any target element. For example, for the "straight-ahead light", in video segment 3 (frames 233 - 376), the attribute information of the "straight-ahead light" is not set, indicating that there is a problem of missing annotation for the "straight-ahead light" annotation information; by traversing the annotation information of related elements, it is determined whether there are conflicts in the annotation information.

[0091] Furthermore, if the verification result indicates successful verification, the annotation information of the video can be saved as a file in a standard format, such as JSON or XML format; if the verification result indicates failed verification, a prompt message can be generated through step 802.

[0092] Step 802, in response to the verification result indicating failed verification, generate a prompt message, and the prompt message carries the annotation information for which the verification fails.

[0093] In the embodiments of the present disclosure, if there are omissions or conflicts in the annotation information, a prompt message can be generated.

[0094] In some embodiments, the prompt information may be text-based prompt information. For example, a window may pop up and display a text prompt message "Missing label for [405-566] frames in the straight-ahead inference state". In other embodiments, the prompt information may also be voice-based prompt information. For example, a voice message "Missing label for [405-566] frames in the straight-ahead inference state" may be played. In still other embodiments, both text-based prompt information and voice-based prompt information may be output simultaneously. In the embodiments of the present disclosure, the prompt information may be other information that can indicate the failure of the annotation information verification, and the form of the prompt information is not specifically limited.

[0095] Among them, the annotation information with verification failure is used to indicate the missing or conflicting annotation information.

[0096] Step 803: Output prompt information so that the user can modify the annotation information with verification failure based on the prompt information.

[0097] Step 804: Obtain the video segment corresponding to the annotation information with verification failure carried by the prompt information.

[0098] In some embodiments, the first frame information and the second frame information of the video segment are carried in the annotation information with verification failure in the prompt information. By analyzing the prompt information, the corresponding video segment can be obtained.

[0099] Step 805: Play back the video segment corresponding to the annotation information with verification failure.

[0100] In some alternative embodiments, a hyperlink to the video segment corresponding to the annotation information with verification failure may be embedded in the prompt information. Thus, the user can click on the prompt information to play back the video segment corresponding to the annotation information with verification failure again.

[0101] In some embodiments, when playing back the video segment corresponding to the annotation information with verification failure, the video segments of multiple videos from different perspectives can be played back synchronously.

[0102] Step 806: In response to receiving an annotation information modification operation sent through the annotation area of the visual interaction interface, perform a modification operation on the annotation information with verification failure based on the annotation information modification operation to obtain the annotation information of the video segment corresponding to the annotation information with verification failure. The modification operation includes modifying at least one of the first frame information, the second frame information, and the attribute information.

[0103] Among them, the annotation information modification operation is an operation triggered by the user through the timeline for modifying the annotation information. Specifically, the user can trigger the annotation information operation by clicking on the timeline, or trigger the annotation information modification operation through the modification button on the visual interaction interface.

[0104] In this embodiment, when making modifications, the first-frame information and the second-frame information of the video segment corresponding to the annotation information can be modified. For example, if the video segment "inference straight-ahead state [405 - 566] frames" is missed in annotation, and it is determined from the played-back video that the attribute value of "inference straight-ahead state" for the video segment of frames [405 - 521] is green light, and the attribute value of "inference straight-ahead state" for the video segment of frames [522 - 566] is green light flashing, then the video segment of [405 - 566] frames can be divided into two video segments (the video segment of [405 - 521] frames and the video segment of [522 - 566] frames), and corresponding attribute information can be set for the two video segments respectively.

[0105] The method for verifying and modifying video annotation information provided in this embodiment can detect and prompt the user about missing, misannotated, or conflicting data by performing a comprehensive verification when the target elements of multiple videos are all annotated, and can play back the video segments that fail the verification to assist the user in making precise modifications, which helps to ensure that incorrect annotation information does not flow through, and significantly improves the quality and reliability of the annotation information.

[0106] In another alternative implementation, after the prompt information is output through step 803, a video playback operation triggered by the user through the video playback area of the visual interaction interface can also be received; and based on the video playback operation, multiple annotated videos are played back; in response to receiving an annotation information modification operation sent through the annotation area of the visual interaction interface, based on the annotation information modification operation, a modification operation is performed on the annotation information that fails the verification to obtain the annotation information of the video segment corresponding to the annotation information that fails the verification, and the modification operation includes modifying at least one of the first-frame information, the second-frame information, and the attribute information.

[0107] In this implementation, if there are many annotation information that fail the verification indicated in the prompt information, for example, there are more than 10 annotation information that fail the verification, then a playback operation can be directly performed on multiple annotated videos, and after receiving an annotation information modification operation during the playback process, based on the annotation information modification operation, the annotation information that fails the verification is modified. This implementation helps to uniformly modify the annotation information of the annotated videos based on the playback of the annotated videos when there are many annotation information that fail the verification, and improves the reliability of modifying the annotation information.

[0108] Exemplary Apparatus

[0109] Figure 9 is a schematic structural diagram of a video annotation device provided by an exemplary embodiment of the present disclosure; as Figure 9 shown, the device includes:

[0110] A display module 91, configured to synchronously play videos of the same scene collected from multiple different perspectives, and to interact with a user;

[0111] A determination module 92, configured to, during the synchronous playback of videos of the same scene collected from multiple different perspectives, in response to a labeling instruction triggered by a target timeline corresponding to a target element included in multiple videos received through a visual interaction interface, determine video segments corresponding to the first frame information and the second frame information in the multiple videos based on the first frame information and the second frame information corresponding to the target timeline, where the labeling instruction includes attribute information of the target element;

[0112] A labeling module 93, configured to label the attribute information of the target element in each frame image of the video segment to obtain the labeling information of the target element in the video segment

[0113] Figure 10 It is a schematic structural diagram of a video labeling device provided by another exemplary embodiment of the present disclosure. As Figure 10 shown, on the basis of the embodiment shown in Figure 9 In some embodiments, the display module 91 may include:

[0114] A first playback sub-module 911, configured to synchronously play videos of the same scene collected from multiple different perspectives in multiple video playback areas of the visual interaction interface in a grid mode; or,

[0115] A second playback sub-module 912, configured to synchronously play videos collected from multiple different perspectives in multiple video playback areas of the visual interaction interface in a window mode, where the multiple video playback areas include a main playback area and a secondary playback area, and different perspective videos can be switched and played in the main playback area based on a video switching instruction triggered by a user, and other videos except the video played in the main playback area are played in the secondary playback area.

[0116] In some embodiments, the video labeling device may further include:

[0117] A first receiving module 94, configured to receive a timeline configuration instruction triggered through the visual interaction interface;

[0118] A timeline determination module 95, configured to determine at least one timeline displayed in a labeling area of the visual interaction interface based on the timeline configuration instruction, where the at least one timeline is used to implement an operation of labeling a target element in multiple videos;

[0119] A first display module 96, configured to display at least one timeline in the labeling area of the visual interaction interface.

[0120] In some embodiments, the video annotation device may further include:

[0121] A second display module 97, configured to display an annotation information setting window in an annotation area of the visual interaction interface in response to receiving an annotation operation triggered by a target timeline corresponding to a target element;

[0122] A second receiving module 98, configured to receive attribute information input by a user through the annotation information setting window;

[0123] A confirmation module 99, configured to confirm receiving an annotation instruction in response to receiving a confirmation instruction indicating that the setting of the attribute information is completed.

[0124] In some embodiments, the video annotation device may further include:

[0125] A verification module 11, configured to verify annotation information corresponding to target elements of multiple videos in response to receiving a confirmation instruction indicating that the annotation of the target elements of multiple videos is completed through the visual interaction interface;

[0126] A generation module 12, configured to generate a prompt message carrying the annotation information for which the verification fails in response to the verification result indicating that the verification fails;

[0127] An output module 13, configured to output the prompt message so that the user can modify the annotation information for which the verification fails based on the prompt message.

[0128] In some embodiments, the video annotation device may further include:

[0129] An acquisition module 14, configured to acquire a video segment corresponding to the annotation information for which the verification fails carried in the prompt message;

[0130] A first playback module 15, configured to playback the video segment corresponding to the annotation information for which the verification fails;

[0131] A first modification module 16, configured to perform a modification operation on the annotation information for which the verification fails based on an annotation information modification operation received through the annotation area of the visual interaction interface, so as to obtain the annotation information of the video segment corresponding to the annotation information for which the verification fails, and the modification operation includes modifying at least one of the first frame information, the second frame information, and the attribute information.

[0132] In some embodiments, the device includes:

[0133] A third receiving module 17, configured to receive a video playback operation triggered through a video playback area of the visual interaction interface;

[0134] A second playback module 18 for playing back multiple annotated videos based on a video playback operation;

[0135] A second modification module 19 for, in response to receiving an annotation information modification operation sent through an annotation area of a visual interaction interface, performing a modification operation on the annotation information that fails verification based on the annotation information modification operation to obtain the annotation information of the video segment corresponding to the annotation information that fails verification, where the modification operation includes modifying at least one of first frame information, second frame information, and attribute information.

[0136] The exemplary embodiments of the present apparatus correspond to the above - mentioned exemplary method part, and the relevant content can be mutually referred to and cited. The beneficial technical effects corresponding to the exemplary embodiments of the present apparatus can be seen in the corresponding beneficial technical effects of the above - mentioned exemplary method part, and will not be elaborated here.

[0137] Exemplary Computer Program Product and Computer Readable Storage Medium

[0138] In addition to the above - mentioned method and device, the embodiments of the present disclosure can also provide a computer program product, including computer program instructions, which when run by a processor cause the processor to execute the steps in the video annotation of various embodiments of the present disclosure described in the above - mentioned "Exemplary Method" part.

[0139] The computer program product can be written in any combination of one or more programming languages for the program code to execute the operations of the embodiments of the present disclosure. The programming languages include object - oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0140] Furthermore, the embodiments of the present disclosure can also be a computer - readable storage medium, on which computer program instructions are stored, which when run by a processor cause the processor to execute the steps in the video annotation of various embodiments of the present disclosure described in the above - mentioned "Exemplary Method" part.

[0141] A computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium includes, for example but not limited to, a system, device or component of electricity, magnetism, optics, electromagnetic, infrared ray, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0142] The basic principles of the present disclosure have been described in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for illustrative and easy-to-understand purposes, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0143] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and modifications.

Claims

1. A video annotation method, comprising: During the synchronous playback of videos of the same scene collected from multiple different perspectives, in response to a annotation instruction triggered by a target timeline corresponding to a target element included in multiple videos received through a visual interaction interface, based on the first frame information and the second frame information corresponding to the target timeline, determining video segments corresponding to the first frame information and the second frame information in the multiple videos, wherein the annotation instruction includes attribute information of the target element; Annotating the target element of each frame image in the video segment with the attribute information to obtain an annotated video segment.

2. The method according to claim 1, wherein the synchronous playback of videos of the same scene collected from multiple different perspectives comprises: Synchronously playing the videos of the same scene collected from multiple different perspectives in multiple video playback areas of the visual interaction interface in a grid mode; Or, Synchronously playing the videos of the same scene collected from multiple different perspectives in multiple video playback areas of the visual interaction interface in a window mode, wherein the multiple video playback areas include a main playback area and secondary playback areas, and in the main playback area, videos of different perspectives can be switched and played based on a video switching instruction triggered by a user, and the secondary playback areas play other videos except the videos played in the main playback area.

3. The method according to claim 1, wherein Before the synchronous playback of videos of the same scene collected from multiple different perspectives, it further comprises: Receiving a timeline configuration instruction triggered through the visual interaction interface; Based on the timeline configuration instruction, determining at least one timeline displayed in an annotation area of the visual interaction interface, wherein the at least one timeline is used to implement the operation of annotating the target element in the multiple videos; Displaying the at least one timeline in the annotation area of the visual interaction interface.

4. The method according to claim 1, wherein, The receiving of the annotation instruction triggered by the target timeline corresponding to the target element included in multiple videos through the visual interaction interface comprises: In response to receiving an annotation operation triggered by the target timeline corresponding to the target element, displaying an annotation information setting window in the annotation area of the visual interaction interface; Receiving the attribute information input by the user through the annotation information setting window; In response to receiving a confirmation instruction indicating that the setting of the attribute information is completed, confirming the receipt of the annotation instruction.

5. The method according to any one of claims 1-4, wherein, After obtaining the annotated video segment, it further comprises: In response to a confirmation instruction received through the visual interaction interface indicating that the annotation of the target elements of the multiple videos is completed, verifying the annotation information corresponding to the target elements of the multiple videos; In response to the verification result indicating a verification failure, generating a prompt message, wherein the prompt message carries the annotation information for which the verification fails; Outputting the prompt message so that the user can modify the annotation information for which the verification fails based on the prompt message.

6. The method according to claim 5, wherein, After outputting the prompt message, it further comprises: Obtaining the video segment corresponding to the annotation information for which the verification fails carried by the prompt message; Playing back the video segment corresponding to the annotation information for which the verification fails. In response to receiving an annotation information modification operation sent through the annotation area of the visual interaction interface, based on the annotation information modification operation, perform a modification operation on the annotation information that fails the verification, to obtain the annotation information of the video segment corresponding to the annotation information that fails the verification, where the modification operation includes modifying at least one of the first frame information, the second frame information, and the attribute information.

7. The method according to claim 5, wherein After outputting the prompt information, it further includes: Receiving a video playback operation triggered through the video playback area of the visual interaction interface; Based on the video playback operation, perform playback on multiple annotated videos; In response to receiving an annotation information modification operation sent through the annotation area of the visual interaction interface, based on the annotation information modification operation, perform a modification operation on the annotation information that fails the verification, to obtain the annotation information of the video segment corresponding to the annotation information that fails the verification, where the modification operation includes modifying at least one of the first frame information, the second frame information, and the attribute information.

8. A video annotation device, comprising: A visual interaction interface module, configured to synchronously play videos of the same scene collected from multiple different perspectives, and interact with a user; A video segment determination module, configured to, during the process of synchronously playing videos of the same scene collected from multiple different perspectives, in response to receiving an annotation instruction triggered by a target time axis corresponding to a target element included in multiple videos through the visual interaction interface, based on the first frame information and the second frame information corresponding to the target time axis, determine a video segment corresponding to the first frame information and the second frame information in the multiple videos, where the annotation instruction includes the attribute information of the target element; An annotation module, configured to annotate the attribute information to the target element of each frame image in the video segment, to obtain the annotation information of the target element in the video segment.

9. A computer-readable storage medium, where the storage medium stores computer program instructions, and when the computer program instructions are executed, the method according to any one of claims 1-7 above is implemented.

10. An electronic device, where the electronic device includes: A memory, configured to store a computer program product; A processor, configured to execute the computer program product stored in the memory, and when the computer program product is executed, the method according to any one of claims 1-7 above is implemented.

11. A computer program product, comprising computer program instructions, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1-7 above is implemented.