Action labeling method, device, apparatus and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF AUTOMATION CHINESE ACAD OF SCI
- Filing Date
- 2023-03-13
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]但这种标注方式较慢,并且直接将一秒内的其他帧抛弃对少样本动作是一种浪费
[0033]The action annotation method, apparatus, device, and storage medium provided by this invention, by cropping a first video from a video to be processed, extracting the first video frame of the first video at preset frame intervals, determining the detection box annotation results of each target organism in each video frame of the video to be processed (excluding the first video frame) based on the detection box annotation results of each target organism in the first video frame, and then matching the action recognition results of each target organism in the video to be processed with the detection box annotation results of each target organism to obtain the action annotation results of each target organism in the video to be processed. This not only preserves the video frames of each action sample in the video to be processed, but also reduces the time spent on manual annotation of detection boxes and actions in the prior art, thereby improving the speed of action annotation.
Smart Images

Figure CN116486477B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an action annotation method, apparatus, device, and storage medium. Background Technology
[0002] Since human and animal movements are continuous processes, recording them often relies on video annotation. For videos featuring only one action, there's no need to distinguish the actor performing the action; simply recording the action in segments from start to finish is sufficient. However, when multiple actions appear in a video, we need to differentiate the actors performing each action. This approach is difficult to implement.
[0003] The Atomic Visual Actions (AVA) dataset is a dataset for annotating people and animals in a group. It uses a method of annotating one frame every second, and records the location information (detection box) and action information of each action in the video with the action performer as the center.
[0004] However, this annotation method is slow, and discarding other frames within a second is a waste for few-sample actions. Summary of the Invention
[0005] To address the problems existing in the prior art, embodiments of the present invention provide an action annotation method, apparatus, device, and storage medium.
[0006] This invention provides an action annotation method, comprising:
[0007] At least one first video is identified; the first video is a portion of the video to be processed; the video to be processed contains multiple target organisms;
[0008] Extract the first video frame of the first video according to a preset frame interval;
[0009] Based on the detection box annotation results of each target organism in the first video frame, the detection box annotation results of each target organism in each video frame other than the first video frame of the video to be processed are determined.
[0010] The action recognition result of each target organism in the video to be processed is matched with the detection box annotation result of each target organism in each video frame of the video to be processed to obtain the action annotation result of each target organism in the video to be processed.
[0011] According to an action annotation method provided by the present invention, the step of determining the detection box annotation results of each target organism in each video frame other than the first video frame of the video to be processed based on the detection box annotation results of each target organism in the first video frame includes:
[0012] Based on the bounding box annotation results of each target organism in the first video frame, a detector for each target organism is constructed;
[0013] Based on the detector, the detection box annotation results for each target organism in each video frame of the video to be processed, excluding the first video frame, are determined.
[0014] According to an action annotation method provided by the present invention, the step of determining the bounding box annotation result of each target organism in each video frame other than the first video frame of the video to be processed based on the detector includes:
[0015] Based on the detector, the automatic annotation result of the detection box for each target organism in the second video frame is obtained; the second video frame is a portion of the video to be processed excluding the first video frame.
[0016] The detection bounding box correction annotation results for the automatic annotation results of the detection bounding boxes of each target organism in the second video frame are determined;
[0017] The detector is trained based on the corrected annotation results of the detection box to obtain the updated detector;
[0018] Based on the updated detector, the automatic annotation result of the detection box for each target organism in the third video frame is obtained; the third video frame is the video frame other than the first video frame and the second video frame in the video to be processed;
[0019] Based on the automatic annotation results of the detection boxes of each target organism in the third video frame and the corrected annotation results of the detection boxes of each target organism in the second video frame, the annotation results of the detection boxes of each target organism in each video frame of the video to be processed, excluding the first video frame, are determined.
[0020] According to an action annotation method provided by the present invention, the second video frame is a video frame of the first video other than the first video frame.
[0021] According to an action annotation method provided by the present invention, the action recognition result of each target organism includes the action of each target organism and the time period in which the action of the target organism occurs.
[0022] According to an action annotation method provided by the present invention, the step of matching the action recognition result of each target creature in the video to be processed with the detection box annotation result of each target creature in each video frame of the video to be processed to obtain the action annotation result of each target creature in the video to be processed includes:
[0023] Determine the video frames corresponding to the time periods in which the actions of each target organism occur in the video to be processed;
[0024] The action of each target organism in the video frame corresponding to the time period in which the action of each target organism occurs is matched with the detection box annotation result of each target organism to obtain the action annotation result of each target organism in the video to be processed.
[0025] The present invention also provides an action annotation device, comprising:
[0026] A video processing unit is configured to determine at least one first video; the first video is a portion of a video to be processed; the video to be processed contains multiple target organisms;
[0027] The frame extraction unit is used to extract the first video frame of the first video according to a preset frame interval.
[0028] The detection box annotation unit is used to determine the detection box annotation results of each target organism in each video frame other than the first video frame of the video to be processed, based on the detection box annotation results of each target organism in the first video frame.
[0029] The matching unit is used to match the action recognition result of each target organism in the video to be processed with the detection box annotation result of each target organism in each video frame of the video to be processed, so as to obtain the action annotation result of each target organism in the video to be processed.
[0030] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the action annotation methods described above.
[0031] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the action annotation method as described above.
[0032] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the action annotation methods described above.
[0033] The action annotation method, apparatus, device, and storage medium provided by this invention, by cropping a first video from a video to be processed, extracting the first video frame of the first video at preset frame intervals, determining the detection box annotation results of each target organism in each video frame of the video to be processed (excluding the first video frame) based on the detection box annotation results of each target organism in the first video frame, and then matching the action recognition results of each target organism in the video to be processed with the detection box annotation results of each target organism to obtain the action annotation results of each target organism in the video to be processed. This not only preserves the video frames of each action sample in the video to be processed, but also reduces the time spent on manual annotation of detection boxes and actions in the prior art, thereby improving the speed of action annotation. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0035] Figure 1 A flowchart illustrating the action annotation method provided by this invention;
[0036] Figure 2 A structural diagram of the motion annotation system provided by this invention;
[0037] Figure 3 A flowchart illustrating the workflow of the location marking module provided by this invention;
[0038] Figure 4 This is a schematic diagram of the action annotation device provided by the present invention;
[0039] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0041] The action annotation method provided by this invention is executed by a processing device that can receive user input and has certain computing capabilities.
[0042] The following is combined with Figure 1 The action annotation method of the present invention is described using a computer device as an example. Figure 1 This is a flowchart illustrating the action annotation method provided by the present invention, as shown below. Figure 1 As shown, the method includes the following steps:
[0043] Step 100: Determine at least one first video; the first video is a portion of the video to be processed; the video to be processed contains multiple target organisms.
[0044] Specifically, the video to be processed can be either manually filmed or filmed by a fixed-position camera, and the subjects of the filming include multiple target organisms. In this embodiment of the invention, it is necessary to annotate the actions and positions of the multiple organisms in the video to be processed. For example, the actions and positions of a group of people, or the actions and positions of a group of animals, can be annotated.
[0045] The computer device can first determine one or more first videos. The first videos can be portions of videos cropped from the videos to be processed. In practical applications, the first videos can be video clips selected by the user from the videos to be processed, including relatively complete footage of the target creature and / or relatively rare movements of the target creature.
[0046] Step 101: Extract the first video frame of the first video according to the preset frame interval.
[0047] Specifically, after determining one or more first videos, the computer device can perform frame-by-frame processing on the first videos and extract the first video frames according to a preset frame interval. The preset frame interval can be set according to the actual situation. For example, if the first video is divided into 30 frames per second, one frame can be extracted every 10 frames, that is, 3 frames can be extracted per second as the first video frames.
[0048] Step 102: Based on the bounding box annotation results of each target organism in the first video frame, determine the bounding box annotation results of each target organism in each video frame other than the first video frame of the video to be processed.
[0049] Specifically, after extracting the first video frame, the computer device can determine the detection box annotation result for each target organism in the first video frame. The detection box annotation result for each target organism in the first video frame can be a detection box obtained by the user manually annotating the position of the target organism.
[0050] Then, the computer device can determine the bounding box annotation results of each target organism in each video frame other than the first video frame of the video to be processed based on the bounding box annotation results of each target organism in the first video frame.
[0051] The first video is cropped from the video to be processed, and the first video frame of the first video is extracted for manual annotation to obtain the detection box annotation results. Then, the detection box annotation results of other video frames in the video to be processed are determined based on the detection box annotation results in the first video frame. Compared with manually annotating one frame every second in the video to be processed, the annotation workload is reduced.
[0052] Step 103: Match the action recognition result of each target creature in the video to be processed with the detection box annotation result of each target creature in each video frame of the video to be processed to obtain the action annotation result of each target creature in the video to be processed.
[0053] Specifically, the motion recognition results of each target organism in the video to be processed can be determined by the user and input into the computer device.
[0054] After determining the bounding box annotations for each target organism in each video frame of the video to be processed, the computer device can match the motion recognition results of each target organism with the bounding box annotations for each target organism in each video frame, thereby obtaining the motion annotation results for each target organism in the video to be processed. The motion annotation results for each target organism in the video to be processed are the location information (bounding box) and motion information of each target organism in the video to be processed.
[0055] The action annotation method provided by this invention involves cropping a first video from a video to be processed, extracting the first video frame of the first video at preset frame intervals, determining the detection box annotation results of each target organism in each video frame of the video to be processed (excluding the first video frame) based on the detection box annotation results of each target organism in the first video frame, and then matching the action recognition results of each target organism in the video to be processed with the detection box annotation results of each target organism to obtain the action annotation results of each target organism in the video to be processed. This method not only preserves the video frames of each action sample in the video to be processed, but also reduces the time spent on manual annotation of detection boxes and actions in the prior art, thereby improving the speed of action annotation.
[0056] Optionally, based on the bounding box annotation results of each target organism in the first video frame, the bounding box annotation results of each target organism in each video frame other than the first video frame of the video to be processed are determined, including:
[0057] Based on the bounding box annotations of each target organism in the first video frame, a detector is constructed for each target organism.
[0058] Based on the detector, the bounding box annotation results for each target organism in each video frame other than the first video frame of the video to be processed are determined.
[0059] Specifically, after determining the bounding box annotations for each target organism in the first video frame, the computer device can construct a detector for each target organism based on these bounding box annotations. The detector can be constructed by training an object detection model, such as the YOLOv5 model.
[0060] After constructing the detector for each target organism, the bounding box annotation results for each target organism in each video frame except the first video frame can be determined based on the detector for each target organism.
[0061] By extracting the first video frame from the cropped first video and then constructing a detector based on the bounding box annotation results of each target organism in the first video frame, a detector with higher detection accuracy can be obtained. The bounding box annotation results of each target organism in other video frames determined by this detector will be more consistent with the real situation.
[0062] Optionally, based on the detector, the bounding box annotation results for each target organism in each video frame other than the first video frame of the video to be processed are determined, including:
[0063] Based on the detector, the automatic annotation results of the detection box for each target organism in the second video frame are obtained; the second video frame is a portion of the video to be processed excluding the first video frame.
[0064] The detection bounding box of each target organism in the second video frame is automatically labeled, and the labeling results are corrected.
[0065] The detector is trained based on the corrected annotation results of the detection boxes to obtain the updated detector;
[0066] Based on the updated detector, the automatic annotation results of the detection box for each target organism in the third video frame are obtained; the third video frame is the video frame other than the first and second video frames of the video to be processed.
[0067] Based on the automatic annotation results of the detection bounding boxes of each target organism in the third video frame and the corrected annotation results of the detection bounding boxes of each target organism in the second video frame, the annotation results of the detection bounding boxes of each target organism in each video frame of the video to be processed, excluding the first video frame, are determined.
[0068] Specifically, after the computer device constructs a detector for each target organism, it can automatically label the detection box of each target organism in the second video frame (i.e., the detection box predicted by the detector) based on the detector of each target organism.
[0069] The second video frame can be a portion of the video to be processed, excluding the first video frame. For example, a portion of video frames can be manually or randomly selected from the video to be processed, and the computer device can automatically label the detection boxes for each organism in this portion of the video frames based on the detectors for each target organism.
[0070] Optionally, the second video frame can be a video frame of the first video other than the first video frame.
[0071] Then, the user can manually correct the automatic bounding box annotation results of the second video frame to obtain the corrected bounding box annotation results of the automatic bounding box annotation results of the second video frame (i.e., the manually annotated bounding boxes determined by the user), and input the corrected bounding box annotation results into the computer device.
[0072] After receiving the correction annotation results of the detection boxes from the user, the computer device can train the detector based on the correction annotation results of the detection boxes and obtain the updated detector.
[0073] Then, based on the updated detector, the computer device can obtain the automatic annotation results of the detection boxes for each target organism in the third video frame. The third video frame is a video frame in the video to be processed, excluding the first and second video frames.
[0074] If the second video frame is a video frame other than the first video frame, then the third video frame is a video frame of a video segment other than the first video in the video to be processed.
[0075] After obtaining the automatic annotation results of the detection boxes for each target organism in the third video frame and the corrected annotation results of the detection boxes for each target organism in the second video frame, the computer device can determine the annotation results of the detection boxes for each target organism in the video frames other than the first video frame of the video to be processed.
[0076] By determining the detection bounding box correction annotation results of the second video frame and training the detector based on the detection bounding box correction annotation results, the detector can be further optimized, the detection accuracy of each detector can be improved, and the detection bounding box annotation results of each target organism in other video frames of the video to be processed can be determined, thus ensuring the accuracy of the detection bounding box annotation results of the target organisms in other video frames.
[0077] Optionally, the action recognition result for each target organism may include the action of each target organism and the time period during which the action occurred.
[0078] For example, the target organisms include monkeys wearing yellow collars and monkeys wearing red collars. The action recognition results for the monkeys wearing yellow collars could be: 0-4 seconds, walking; 4-10 seconds, sitting. The action recognition results for the monkeys wearing red collars could be: 0-5 seconds, crawling; 5-10 seconds, lying down. It should be understood that although the time intervals in the above examples are accurate to the second, in reality, the unit of measurement for time intervals can be accurate to the millisecond, etc.
[0079] Optionally, the action recognition result of each target organism in the video to be processed is matched with the detection box annotation result of each target organism in each video frame of the video to be processed to obtain the action annotation result of each target organism in the video to be processed, including:
[0080] Determine the video frames corresponding to the time periods in which the actions of each target organism occur in the video to be processed;
[0081] Match the action of each target organism in the video frame corresponding to the time period in which the action of each target organism occurs with the detection box annotation result of each target organism to obtain the action annotation result of each target organism in the video to be processed.
[0082] Specifically, after determining the bounding box annotation results for each target organism in each video frame of the video to be processed, and determining the action recognition results for each organism in the video to be processed, the computer device can determine the video frame corresponding to the time period in which the action of each target organism in the video to be processed occurs, based on the time period in which the action occurs and the number of frames in the video to be processed.
[0083] For example, if the video to be processed is divided into 30 frames per second, the action recognition results for the monkey wearing a yellow collar can be: 0-4s, walking; 4-10s, sitting. Therefore, it can be determined that the video frames corresponding to the time periods when the monkey wearing the yellow collar walks are frames 1-120; and the video frames corresponding to the time periods when it sits are frames 121-300.
[0084] After determining the video frame corresponding to the time period in which the action of each target creature in the video to be processed occurs, the computer device can match the action of each target creature in the video frame corresponding to the time period in which the action of each target creature occurs with the detection box annotation result of each target creature, and obtain the action annotation result of each target creature in the video to be processed. That is, the detection box annotation result of each target creature in each frame is matched with the action of the target creature.
[0085] By converting the time period of each target organism's action in the action recognition results into corresponding video frames, and then matching the action of each target organism in the video frames corresponding to the time period of each target organism's action with the detection box annotation results of each target organism, the workload of action annotation is reduced compared to annotating the target organism's actions frame by frame.
[0086] The following section further illustrates the action annotation method provided by this invention through an action annotation system in a specific application scenario. Figure 2 A structural diagram of the action annotation system provided by the present invention is shown below. Figure 2 As shown, the system is used to observe the movements and behaviors of monkeys and may include the following modules: video processing module, detector construction module, location labeling module, action recognition module, detection box and action matching module.
[0087] The video processing module is used to trim videos to obtain the desired video segments.
[0088] For example, a PyQt5 interface can be designed for video cropping and frame splitting. The video cropping interface can be used to filter videos with a large number of monkeys or those exhibiting rare movements, while the video frame splitting interface can divide the cropped video into image sequences, dividing each second of video into 30 frames. Compared to command-line cropping and frame splitting, designing an interface makes these two functions much simpler and faster.
[0089] The detector building module is used to annotate detection boxes every preset number of frames in the image sequence and use the annotated information to train the detector.
[0090] For example, the detector building module can extract one frame every 10 frames to annotate the detection boxes and use this information to train a detector for each monkey, with one detector for each monkey. For example, if there are 5 monkeys in the video to be processed, namely a monkey with a yellow collar, a monkey with a green collar, a monkey with a red collar, a monkey with a black collar, and a monkey with a white collar, there are a total of 5 detectors.
[0091] The location annotation module is used to manually correct the bounding boxes automatically annotated by the detector, thereby ensuring that the mean average precision (MAP) of the detection boxes is close to 1, and further optimize the detector with the corrected data for the annotation of the next batch of data.
[0092] For example, the location annotation module provides automatic annotation and manual correction interfaces. The automatic annotation interface uses a trained detector to automatically annotate other frames. The manual correction interface is used to correct the automatic annotations, ensuring that errors are corrected and achieving annotation results of the same quality as manual annotation. The manually corrected data can then be used to train the detector, thereby optimizing it. After changing the detector, the time spent on manual correction is further reduced, thus further improving annotation efficiency.
[0093] Figure 3 A flowchart of the location marking module provided by this invention is shown below. Figure 3 As shown, the labeled video frames in n video segments (i.e., videos where one frame is extracted every 10 frames and bounding boxes are labeled) are used to train the detector. The trained detector is then used to automatically label the unlabeled video frames in the n video segments. A manual correction interface is used to correct the automatically labeled frames. The manually corrected data can then be used to train the detector, thereby optimizing it. The optimized detector is then used to label the bounding boxes in other unlabeled videos, resulting in automatic labeling results.
[0094] The action recognition module is used to identify and label the actions of each monkey from the beginning to the end of the video, similar to the actions of a single monkey.
[0095] For example, the action recognition module can record the time intervals of each monkey's actions sequentially. For instance, a monkey with a yellow collar: 0-4 seconds of walking, 4-10 seconds of sitting, and so on. A monkey with a red collar: 0-5 seconds of crawling, 5-10 seconds of lying down, and so on.
[0096] The detection box and action matching module is used to match the detection boxes obtained by the location annotation module with the actions recorded by the action recognition module, so that each monkey has a detection box and an action category in each frame.
[0097] For example, after all data is automatically labeled and manually modified, each monkey in the dataset has a bounding box. The action recognition module records the actions of each monkey. The bounding box and action matching module can convert the action recording time information into frames, so that each monkey has action information in each frame. Then, the bounding box and action information are fused to achieve the matching of bounding boxes and actions.
[0098] The motion annotation device provided by the present invention will be described below. The motion annotation device described below and the motion annotation method described above can be referred to in correspondence.
[0099] Figure 4 This is a schematic diagram of the action annotation device provided by the present invention, as shown below. Figure 4 As shown, the device includes the following units:
[0100] The video processing unit 400 is used to determine at least one first video; the first video is a portion of the video to be processed; the video to be processed contains multiple target organisms;
[0101] The frame extraction unit 410 is used to extract the first video frame of the first video according to a preset frame interval.
[0102] The detection box annotation unit 420 is used to determine the detection box annotation results of each target organism in each video frame other than the first video frame of the video to be processed based on the detection box annotation results of each target organism in the first video frame.
[0103] The matching unit 430 is used to match the action recognition result of each target organism with the detection box annotation result of each target organism in each video frame of the video to be processed, so as to obtain the action annotation result of each target organism in the video to be processed.
[0104] Optionally, based on the bounding box annotation results of each target organism in the first video frame, the bounding box annotation results of each target organism in each video frame other than the first video frame of the video to be processed are determined, including:
[0105] Based on the bounding box annotations of each target organism in the first video frame, a detector is constructed for each target organism.
[0106] Based on the detector, the bounding box annotation results for each target organism in each video frame other than the first video frame of the video to be processed are determined.
[0107] Optionally, based on the detector, the bounding box annotation results for each target organism in each video frame other than the first video frame of the video to be processed are determined, including:
[0108] Based on the detector, the automatic annotation results of the detection box for each target organism in the second video frame are obtained; the second video frame is a portion of the video to be processed excluding the first video frame.
[0109] The detection bounding box of each target organism in the second video frame is automatically labeled, and the labeling results are corrected.
[0110] The detector is trained based on the corrected annotation results of the detection boxes to obtain the updated detector;
[0111] Based on the updated detector, the automatic annotation results of the detection box for each target organism in the third video frame are obtained; the third video frame is the video frame other than the first and second video frames of the video to be processed.
[0112] Based on the automatic annotation results of the detection bounding boxes of each target organism in the third video frame and the corrected annotation results of the detection bounding boxes of each target organism in the second video frame, the annotation results of the detection bounding boxes of each target organism in each video frame of the video to be processed, excluding the first video frame, are determined.
[0113] Optionally, the second video frame is a video frame of the first video other than the first video frame.
[0114] Optionally, the action recognition result for each target organism includes the action of each target organism and the time period in which the action occurred.
[0115] Optionally, the action recognition result of each target organism in the video to be processed is matched with the detection box annotation result of each target organism in each video frame of the video to be processed to obtain the action annotation result of each target organism in the video to be processed, including:
[0116] Determine the video frames corresponding to the time periods in which the actions of each target organism occur in the video to be processed;
[0117] Match the action of each target organism in the video frame corresponding to the time period in which the action of each target organism occurs with the detection box annotation result of each target organism to obtain the action annotation result of each target organism in the video to be processed.
[0118] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 5 As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute any of the action annotation methods provided in the above embodiments.
[0119] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0120] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute any of the action annotation methods provided in the above embodiments.
[0121] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform any of the action annotation methods provided in the above embodiments.
[0122] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0123] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An action annotation method, characterized in that, include: Identify at least one first video; The first video is a portion of the video to be processed that has been cropped out. The video to be processed contains multiple target organisms; Extract the first video frame of the first video according to a preset frame interval; Based on the detection box annotation results of each target organism in the first video frame, the detection box annotation results of each target organism in each video frame other than the first video frame of the video to be processed are determined. The action recognition result of each target creature in the video to be processed is matched with the detection box annotation result of each target creature in each video frame of the video to be processed to obtain the action annotation result of each target creature in the video to be processed. The action recognition result for each target organism includes the action of each target organism and the time period in which the action occurred; The step of matching the action recognition result of each target creature in the video to be processed with the detection box annotation result of each target creature in each video frame of the video to be processed to obtain the action annotation result of each target creature in the video to be processed includes: Determine the video frames corresponding to the time periods in which the actions of each target organism occur in the video to be processed; The action of each target organism in the video frame corresponding to the time period in which the action of each target organism occurs is matched with the detection box annotation result of each target organism to obtain the action annotation result of each target organism in the video to be processed.
2. The action annotation method according to claim 1, characterized in that, The step of determining the detection bounding box annotation results for each target organism in each video frame other than the first video frame, based on the detection bounding box annotation results for each target organism in the first video frame, includes: Based on the bounding box annotation results of each target organism in the first video frame, a detector for each target organism is constructed; Based on the detector, the detection box annotation results for each target organism in each video frame of the video to be processed, excluding the first video frame, are determined.
3. The action annotation method according to claim 2, characterized in that, The step of determining the bounding box annotation results for each target organism in each video frame other than the first video frame of the video to be processed based on the detector includes: Based on the detector, the automatic annotation result of the detection box for each target organism in the second video frame is obtained; the second video frame is a portion of the video to be processed excluding the first video frame. The detection bounding box correction annotation results for the automatic annotation results of the detection bounding boxes of each target organism in the second video frame are determined; The detector is trained based on the corrected annotation results of the detection box to obtain the updated detector; Based on the updated detector, the automatic annotation result of the detection box for each target organism in the third video frame is obtained; the third video frame is the video frame other than the first video frame and the second video frame in the video to be processed; Based on the automatic annotation results of the detection boxes of each target organism in the third video frame and the corrected annotation results of the detection boxes of each target organism in the second video frame, the annotation results of the detection boxes of each target organism in each video frame of the video to be processed, excluding the first video frame, are determined.
4. The action annotation method according to claim 3, characterized in that, The second video frame is a video frame of the first video other than the first video frame.
5. An action annotation device, characterized in that, include: A video processing unit is used to determine at least one first video; The first video is a portion of the video to be processed that has been cropped out. The video to be processed contains multiple target organisms; The frame extraction unit is used to extract the first video frame of the first video according to a preset frame interval. The detection box annotation unit is used to determine the detection box annotation results of each target organism in each video frame other than the first video frame of the video to be processed, based on the detection box annotation results of each target organism in the first video frame. The matching unit is used to match the action recognition result of each target organism in the video to be processed with the detection box annotation result of each target organism in each video frame of the video to be processed, so as to obtain the action annotation result of each target organism in the video to be processed. The action recognition result for each target organism includes the action of each target organism and the time period in which the action occurred; The matching unit is specifically used for: Determine the video frames corresponding to the time periods in which the actions of each target organism occur in the video to be processed; The action of each target organism in the video frame corresponding to the time period in which the action of each target organism occurs is matched with the detection box annotation result of each target organism to obtain the action annotation result of each target organism in the video to be processed.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the action annotation method as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the action annotation method as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the action annotation method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Video annotation method, system and equipment
CN114117128A
Livestock climbing behavior labeling method and device, electronic equipment and storage medium
CN115661717A