Automatic labeling method, automatic labeling device and computer readable storage medium

Through the automatic labeling method, the target labeling area of ​​video generated by the labeling network is used and classified, and the problems of low labeling efficiency and high cost caused by human dependence in the prior art are solved, and the efficiency and accuracy of automatic labeling of videos are achieved.

CN114218434BActive Publication Date: 2025-05-16ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111320677.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-09
Publication Date
2025-05-16
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

Existing video labeling technology relies heavily on human resources, with low labeling efficiency and high cost. Manual labeling can easily lead to confusion in the definition of action start and end points, which requires review to further reduce efficiency.

Method used

The automatic labeling method is adopted to obtain the video to be labeled, and feature extraction is performed. The labeling generation network is used to generate candidate labeling areas, and the target labeling area is generated through correction processing. Finally, the content in the target labeling area is classified and processed to obtain category information.

Benefits of technology

It realizes the accuracy and efficiency of automatic video labeling, reduces manpower participation, improves the accuracy and efficiency of automatic labeling, and adapts to the annotation of different categories of videos, enhancing robustness and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218434B_ABST
    Figure CN114218434B_ABST
Patent Text Reader

Abstract

This application discloses an automatic annotation method, an automatic annotation device, and a computer-readable storage medium. The method includes: acquiring a first video to be annotated, the first video to be annotated including content to be annotated; performing feature extraction processing on the first video to be annotated to obtain first feature information; processing the first feature information using an annotation generation network to generate at least one candidate annotation region, and performing correction processing on the candidate annotation regions to generate a target annotation region, the target annotation region including the video between the start time and the end time corresponding to the content to be annotated; and classifying the content to be annotated in the target annotation region to obtain category information of the content to be annotated. Through the above methods, this application can improve the accuracy and efficiency of automatic video annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of deep learning technology, and in particular to an automatic labeling method, an automatic labeling device, and a computer-readable storage medium. Background Art

[0002] Currently, the number of videos is experiencing explosive growth, which has led to an increasing demand for video understanding technology. Rapidly locating the segment of interest in any video of indefinite length and identifying the category of the segment is of great significance for applications such as video recommendation, retrieval, and retraining of video understanding. However, some solutions in related technologies rely heavily on human resources, with low labeling efficiency and high labeling costs. In addition, manual labeling will also cause confusion in the definition of the start and end points of the same action in different videos, and the labeled videos need to be reviewed, further reducing the labeling efficiency. Summary of the invention

[0003] The present application provides an automatic labeling method, an automatic labeling device and a computer-readable storage medium, which can improve the accuracy and efficiency of automatic labeling of videos.

[0004] In order to solve the above technical problems, the technical solution adopted in the present application is: to provide an automatic annotation method, the method comprising: obtaining a first video to be annotated, the first video to be annotated including content to be annotated; performing feature extraction processing on the first video to be annotated to obtain first feature information; using an annotation generation network to process the first feature information to generate at least one candidate annotation area, and performing correction processing on the candidate annotation area to generate a target annotation area, the target annotation area including a video between a starting time and an end time corresponding to the content to be annotated; classifying the content to be annotated in the target annotation area to obtain category information of the content to be annotated.

[0005] In order to solve the above technical problems, another technical solution adopted in the present application is: to provide an automatic labeling device, which includes a memory and a processor connected to each other, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the automatic labeling method in the above technical solution.

[0006] To solve the above technical problems, another technical solution adopted in the present application is: providing a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, it is used to implement the automatic labeling method in the above technical solution.

[0007] Through the above scheme, the beneficial effects of the present application are: first, a first video to be annotated containing content to be annotated is obtained; then, features in the first video to be annotated are extracted to generate first feature information; then, the first feature information is preliminarily processed by using an annotation generation network to generate at least one candidate annotation area, and the candidate annotation area is corrected to generate a final target annotation area, which includes a video between a starting time corresponding to the content to be annotated and an end time corresponding to the content to be annotated; then, the content to be annotated in the target annotation area is classified to obtain category information of the content to be annotated; because the annotation generation network is used to detect the starting time corresponding to the content to be annotated and the end time corresponding to the content to be annotated, it is possible to select a target annotation area related to the content to be annotated from a video, that is, to extract a video clip of interest, and it is also possible to identify the category corresponding to the video clip, thereby realizing automatic annotation of the video, reducing human participation, and improving the accuracy and efficiency of automatic annotation; moreover, it is also possible to adapt to the annotation of videos of different categories, enhance the robustness of the annotation of videos of different categories and indefinite lengths, and improve the versatility and reproducibility of any automatic annotation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:

[0009] Figure 1 It is a flowchart of an embodiment of an automatic annotation method provided by the present application;

[0010] Figure 2 It is a schematic diagram of the starting time and the ending time of the target annotation segment provided by the present application;

[0011] Figure 3 is a flow chart of another embodiment of the automatic annotation method provided by the present application;

[0012] Figure 4 It is a schematic diagram of the structure of the classification network provided by this application;

[0013] Figure 5 It is a structural schematic diagram of feature interpolation provided by this application;

[0014] Figure 6 It is a structural diagram of the annotation generation network provided by this application;

[0015] Figure 7 yes Figure 3 A schematic diagram of the flow chart of step 35 in the embodiment shown;

[0016] Figure 8 It is a structural diagram of the feature enhancement module provided by this application;

[0017] Fig. 9 is a schematic diagram of the structure of the scoring grid provided by this application;

[0018] Fig.10 It is a schematic diagram of the structure of the dense2sparse unit provided in this application;

[0019] Fig.11 It is a structural schematic diagram of an embodiment of an automatic labeling device provided by the present application;

[0020] Fig.12 It is a structural schematic diagram of an embodiment of a computer-readable storage medium provided by the present application. DETAILED DESCRIPTION

[0021] The present application is further described in detail below in conjunction with the accompanying drawings and examples. It is particularly noted that the following examples are only used to illustrate the present application, but are not intended to limit the scope of the present application. Similarly, the following examples are only some embodiments of the present application rather than all embodiments, and all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0022] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0023] It should be noted that the terms "first", "second" and "third" in this application are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first", "second" and "third" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices.

[0024] See also Figure 1, Figure 1 : is a flow chart of an embodiment of an automatic annotation method provided by the present application, the method comprising:

[0025] Step 11: Obtain the first video to be labeled.

[0026] The first video to be labeled is a video that needs to be labeled. The first video to be labeled can be obtained from a video database, or the current monitoring scene can be shot to generate the first video to be labeled. Specifically, the first video to be labeled includes content to be labeled, and the content to be labeled includes behavior or event. When the content to be labeled is behavior, the category information may include long jump, swimming, walking, chopping trees, playing ball, makeup (for example, applying lipstick or eyeliner) or running, etc.; when the content to be labeled is an event, the category information may be related to abnormal events, for example, the category information may be traffic accidents, driving in the wrong direction, objects thrown on the road or objects thrown from high altitude, etc.

[0027] Step 12: Perform feature extraction processing on the first video to be labeled to obtain first feature information.

[0028] A feature extraction method is used to extract features in a first video to be labeled, and first feature information is generated. Specifically, a feature extraction network is used to process the first video to be labeled. The feature extraction network can be a deep learning network, such as: a time-sensitive network (TSN), a channel-separated convolutional network (CSN), a SlowFast network, a two-dimensional convolution (2DConv) or a video-Swin-Transformer structure. Feature extraction refers to transforming an image sequence group in the first video to be labeled into a long vector with a fixed dimension through a deep learning network. Using the long vector to represent the entire video can reduce the computational complexity of subsequently generating a target annotation area.

[0029] Step 13: Use a label generation network to process the first feature information to generate at least one candidate labeling area.

[0030] A pre-trained annotation generation network is used to perform position annotation processing on the first feature information to obtain a candidate annotation set, which includes at least one candidate annotation area, and the candidate annotation area includes a predicted starting time and an estimated end time. The predicted starting time is the estimated time when the content to be annotated in the first video to be annotated first appears, and the estimated end time is the estimated time when the content to be annotated in the first video to be annotated last appears. For example, taking the first video to be annotated with a duration of 60s and the content to be annotated being event A as an example, assuming that event A appears in the first video to be annotated for a period of 30s to 50s, after the annotation generation network processes the first video to be annotated, a candidate annotation set is obtained, which includes 3 candidate annotation areas: 1 to 20s, 21s to 40s, and 41s to 60s.

[0031] Step 14: Use the annotation generation network to correct the candidate annotation area and generate the target annotation area.

[0032] There are a large number of redundant annotation areas in the candidate annotation set, and due to the fixedness of the grid parameters in the annotation generation network, there are errors in the boundary positioning of the annotation area, resulting in poor accuracy in positioning the boundary of the annotation area; in the embodiment of the present application, after obtaining the candidate annotation set, the annotation generation network can continue to be used to process the candidate annotation set to correct the candidate annotation area, find the segment in the first video to be annotated where the content to be annotated is most likely to appear, so as to generate a target annotation area, which includes the video between the starting time corresponding to the content to be annotated in the first video to be annotated and the end time corresponding to the content to be annotated, the starting time is the time when the content to be annotated most likely appears in the first video to be annotated for the first time, and the end time is the time when the content to be annotated appears in the first video to be annotated for the last time.

[0033] Step 15: Classify the content to be marked in the target marking area to obtain category information of the content to be marked.

[0034] The target annotation area is a video clip from the starting time to the end time in the first video to be annotated. After obtaining the target annotation area, a classification network can be used to classify the target annotation area to generate category information of the content to be annotated in the target annotation area. The classification network can be a deep learning network. For example, taking the content to be annotated as behavior as an example, the target annotation area can be classified to obtain category information of whether the event is long jump, makeup or playing ball. Figure 2 As shown, after the above processing operation, the indefinite length video generates a target annotation area with a starting time of 32.6 seconds and an end time of 37.2 seconds. The target annotation area is sent to the pre-trained classification network, and the "long jump" label output by the classification network is used as the final category annotation result.

[0035] It can be understood that in other embodiments, the frame number of the image in the first video to be labeled can also be used as the starting point and end point of the labeling. For example, assuming that the number of frames of a first video to be labeled is 30 frames, the image frame in which the content to be labeled first appears in the first video to be labeled is the 5th frame, and the image frame in which the content to be labeled last appears is the 20th frame, then the video segment composed of the images of frames 5 to 20 is the target labeling area.

[0036] This embodiment actively learns the category features of videos of indefinite length and requiring annotation through a deep learning network, outputs the category annotation results of any indefinite length video, and further realizes the automation of video annotation, so as to generate more intuitive, more information-intensive, and more valuable event / action video clips, which is convenient for subsequent video review, video recommendation, and retraining of the video understanding network, and can be applied to video understanding and video analysis, etc. Moreover, since automatic annotation is realized, human intervention can be reduced, and the accuracy and efficiency of automatic annotation can be improved. In addition, it can also adapt to the annotation of videos of different categories, enhance the robustness of the annotation of videos of different and indefinite lengths, and improve the versatility and replicability of any automatic annotation tasks.

[0037] See also Figure 3 , Figure 3 : is a flow chart of another embodiment of the automatic annotation method provided by the present application, the method comprising:

[0038] Step 31: Train the classification network to obtain a trained classification network.

[0039] The training data is used to train the classification network. The first purpose of training the classification network is to extract features from the video clips and provide a model for feature extraction. The second purpose is to set category labels for the target annotation areas generated later. Therefore, the performance of the classification network directly affects the quality of the subsequent video clip annotation. Specifically, the training steps of the classification network are as follows:

[0040] 1) Select video clip samples with equal time intervals and to be labeled to form training data.

[0041] The training data includes multiple video clip samples. When selecting video clip samples, all categories to be labeled must be covered, and the video clip samples should accurately cover the starting point (i.e., the starting point moment) of the video clip samples of that category.

[0042] 2) Send the video clip samples to the classification network for training.

[0043] The structure of the classification network is as follows Figure 4As shown, the classification network includes a first normalization module, an extraction module and a calculation module. The first normalization module is a video spatiotemporal sequence normalization unit, which is used to normalize the input short video along the space and time dimensions. The purpose of time dimension normalization is to ensure that videos of different lengths can maintain the uniformity of the time dimension when they are sent to the network. Time dimension normalization can be achieved by nonlinear interpolation of adjacent multi-frame images and interval sampling, etc., which is not specifically limited in this embodiment; spatial dimension normalization is to first perform linear interpolation on the single-frame image after time sequence normalization, and then standardize the image with the same size after linear interpolation (that is, first subtract the pixel of each frame image in the video from the pixel mean, and then divide it by the variance). The feature vector extracted by the extraction module is used to replace the input short video (that is, the video clip sample). This embodiment does not specifically limit the classification network used, as long as the classification function can be achieved.

[0044] In a specific embodiment, the training of the classification network is as follows:

[0045] The training data also includes category labels corresponding to the video clip samples. A video clip sample is selected from the training data, and the first normalization module is used to normalize the video clip sample to obtain the normalized video clip sample; the extraction module is used to extract features from the normalized video clip sample to obtain sample feature information; the calculation module is used to classify the sample feature information to obtain a sample classification result, and the calculation module can be a fully connected layer; based on the sample classification result and the category label, the current loss value is calculated; based on the current loss value or the current number of training times, it is determined whether the classification network meets the preset training end condition; if the classification network does not meet the preset training end condition, the step of selecting a video clip sample from the training data is returned until the classification network meets the preset training end condition.

[0046] Furthermore, the preset stopping conditions include: loss value convergence, that is, the difference between the previous loss value and the current loss value is less than the set value; determining whether the current loss value is less than the preset loss value, the preset loss value is a preset loss threshold, if the current loss value is less than the preset loss value, it is determined that the preset stopping condition is met; the number of training times reaches the set value (for example: training 10,000 times); or the accuracy obtained when testing using the test set reaches the set conditions (for example: exceeding the preset accuracy), etc.

[0047] In another specific embodiment, see Figure 3 The classification network also includes a second normalization module, which is used to normalize the results output by the calculation module to obtain category information of the video clip sample; specifically, the second normalization module uses a softmax function to process the results output by the calculation module to generate a probability value that the content to be labeled in the video clip sample belongs to each category.

[0048] Step 32: Segment the first video to be labeled to obtain multiple video segments.

[0049] For the indefinite length video to be labeled, it can be first split into multiple video clips with equal time intervals, so that they can be sent to the trained classification network in step 31 at equal intervals to extract features.

[0050] Step 33: Use a classification network to process the video segment to obtain feature information of the second segment.

[0051] After the video segments are acquired, the video segments may be input into a classification network so that the classification network performs feature extraction processing on the video segments respectively to generate corresponding second segment feature information, where the second segment feature information is feature information of the video segments.

[0052] Step 34: Normalize all the second segment feature information to obtain the first feature information.

[0053] Considering that a complete video of indefinite length may be divided into features of arbitrary length, the entire video features corresponding to the first video to be labeled (i.e., all the feature information of the second segment) are sent to the feature linear interpolation unit for normalization, so that the lengths of the features corresponding to the first videos to be labeled of different lengths are consistent.

[0054] In a specific embodiment, in order to conveniently describe the cutting process and feature normalization process of the indefinite length video to be labeled, three variables num_clips, clip_len, and frame_interval are defined to describe the solution. num_clips is defined as the number of segments after the entire video is cut, clip_len is defined as the number of frames selected for each video segment, and frame_interval is defined as the interval between frames of each video segment. Figure 5 For example, the entire video is cut into M segments, and any segment is denoted as Clip-i (1≤i≤M). The classification network extracts the features of the image frames in each Clip-i, and the length of the output features is len. After equal-interval segmentation and fixed-interval feature extraction, the video can be replaced by a matrix vector of M×len. For different videos to be annotated, the length of the features of the entire video needs to be normalized to L×len. The normalization operation includes linear interpolation between adjacent features and standardization of each element. L represents the size of the regional grid generated subsequently. L is a value set according to experience or application requirements, for example: it is the average number of segments into which all the entire videos are divided.

[0055] It can be understood that when using the classification network for feature extraction, the calculation result before the softmax operation can be directly used, or the feature extraction result output by the extraction module can be used.

[0056] After obtaining the first feature information, the first feature information is input into the annotation generation network to generate a target annotation area; specifically, the annotation generation network includes a first annotation generation network and a second annotation generation network, the input of the first annotation generation network is the video feature normalized to a fixed size (i.e., the first feature information), and the output of the first annotation generation network includes a candidate annotation set; the second annotation generation network generates a target annotation area based on the candidate annotation set output by the first annotation generation network and the first feature information, which is described in detail below.

[0057] Step 35: Use the first annotation generation network to process the first feature information to generate a candidate annotation set.

[0058] like Figure 6 As shown, the first annotation generation network includes a feature enhancement module, a first estimation module, a second estimation module and a generation module, such as Figure 7 As shown, the following steps are used to generate a candidate annotation set:

[0059] Step 41: Use a feature enhancement module to enhance the first feature information to generate second feature information.

[0060] The feature enhancement module is a temporal feature enhancement unit, which is used to simultaneously encode the local and overall temporal information of the normalized features, learn the variable temporal features before and after the area to be labeled, and enhance the feature association of the semantic information of the previous and next frames. Compared with the method of matching based on the similarity between frames, it has higher detection accuracy and robustness for videos of different lengths and easy to confuse.

[0061] In a specific embodiment, the feature enhancement module includes a coding module and an enhancement module. The coding module may be a local-global temporal feature encoder (LGTE), and the LGTE unit is a pre-feature reorganization processing module for the first feature information, such as Figure 8 shown.

[0062] The encoding module is used to encode the first feature information to obtain the third feature information; the enhancement module is used to enhance the third feature information to obtain the second feature information; specifically, the enhancement module includes an enhancement unit and a fusion unit, the enhancement unit is used to enhance the third feature information to obtain the fourth feature information; the fusion unit is used to fuse the fourth feature information to obtain the second feature information, the enhancement unit can be a graph convolutional network with expanded balance theory (GCNEXT), and the fusion unit can be a one-dimensional convolution (CONV-1D), such as Figure 8 shown.

[0063] Furthermore, the enhancement unit includes a temporal enhancement unit and a spatial enhancement unit, which are respectively used to enhance the semantic information of previous and next frames and to aggregate the associated features of different video clips; specifically, the temporal enhancement unit is used to perform temporal enhancement processing on the third feature information to obtain temporal feature information (i.e., the temporal enhancement result); the spatial enhancement unit is used to perform spatial enhancement processing on the third feature information to obtain spatial feature information (i.e., the spatial enhancement result); the fusion unit is used to fuse the temporal feature information, the spatial feature information and the third feature information to obtain the second feature information.

[0064] Understandably, if Figure 8 As shown, the GCNeXt unit and the CONV-1D unit are combined modules and can be repeated N times to further enhance the feature fusion effect of video clips of different time lengths and categories. The specific value of N can be adjusted according to actual application requirements.

[0065] Step 42: Use the first estimation module to estimate the second feature information to obtain first score information.

[0066] The first score information includes multiple regional probabilities, and a regional grid can be created. The regional grid includes multiple grids, and the horizontal coordinate and vertical coordinate of each grid are the predicted starting time and the predicted end time respectively; the first estimation module is used to calculate the probability that the video clip corresponding to each grid is the target marked area to obtain the regional probability.

[0067] In a specific embodiment, taking the content to be annotated as a specific action (such as long jump) as an example, the regional grid is a dense candidate square grid with a side length of L. By defining the size L of the regional grid, a total of (L*L / 2) potential action clips are generated. The size L of the regional grid determines the length of the detectable action, and the length of the shortest detectable video clip is δ (δ=total video length / L); specifically, each row of the regional grid corresponds to the starting point of the action clip, and each column of the grid corresponds to the end point of the action clip. Considering the characteristic that the starting point is prior to the end point, the lower half of the regional grid is invalid.

[0068] Furthermore, the input of the first estimation module is the enhanced temporal features, and the output is the probability value of whether it is the action clip to be labeled after Sigmoid processing; for any valid action clip in the regional grid, the first estimation module is supervisedly trained to learn the features of the video sequence falling into the corresponding grid through the start time label and the end time label marked in advance by humans.

[0069] Step 43: Use a second estimation module to estimate the second feature information to obtain second score information.

[0070] The second estimation module is similar to the first estimation module. It outputs a probability value by learning the video features output by the feature enhancement module. The probability value is the probability value processed by the Sigmoid function. Specifically, the second score information includes a first probability value and a second probability value. The first probability value is the probability that the starting point prediction moment is the starting point moment, and the second probability value is the probability that the end point prediction moment is the end point moment. The time difference between the starting point prediction moment and the end point prediction moment is δ.

[0071] Step 44: Use a generation module to fuse the first score information and the second score information to obtain score information.

[0072] The generation module outputs a square score grid with a side length of L by combining the output results of the first estimation module and the second estimation module, that is, generates score information, which includes multiple score values; specifically, the combination is implemented in the form of Hadamard product, that is, multiplying the regional probability with the corresponding first probability value and the second probability value to obtain the score value. For example, Fig. 9 As shown, the value of the horizontal axis is the predicted time of the starting point, and the value of the vertical axis is the predicted time of the end point. Assuming that the first probability value, the second probability value and the regional probability at (x0, y0) are P1, P2 and P3 respectively, the score value at (x0, y0) is P1×P2×P3.

[0073] Step 45: Generate a candidate annotation set based on the score information.

[0074] Determine whether the score value in the score information is greater than a preset value; if the score value is greater than the preset value, the horizontal and vertical coordinates of the grid corresponding to the score value form a candidate marking area; or perform non-maximum suppression processing on all score values ​​to obtain a candidate marking area. The method of non-maximum suppression processing is the same as that of the related art and will not be repeated here.

[0075] Step 36: Use the second annotation generation network to correct the candidate annotation area to generate a target annotation area.

[0076] The generation module can generate a large number of candidate annotation areas, which improves the recall rate of potential target annotation areas. However, due to the fixed grid scale of the module, the start and end intervals of the candidate annotation areas (including the predicted starting time and the predicted end time) become very rigid. Therefore, the second annotation generation network is used to process the candidate annotation areas; specifically, the second annotation generation network is used to encode the first feature information to obtain the encoded information; the second annotation generation network is used to process the encoded information and the candidate annotation set to obtain the target annotation area; further, the second annotation generation network uses a part of the candidate annotation areas output by the generation module (that is, the candidate annotation areas corresponding to the grids whose predicted starting time is less than the predicted end time) as anchor boxes, and learns the corresponding features of the anchor box start (that is, starting time), end (that is, end time) and start and end intervals through the dense2sparse unit, so as to more accurately output the target annotation area.

[0077] In a specific embodiment, the dense2sparse unit uses a cascade trainer to fine-tune the candidate annotation region, and its structure is as follows: Fig.10 As shown, the candidate annotation set is sorted in descending order according to the score value, non-maximum suppression is performed on the candidate annotation set sorted by the score, and the first K candidate annotation areas are selected first, which can reduce the dense candidate annotation set to K sparse candidate annotation areas while increasing the number of candidate annotation areas with different intersection over union (IOU) ratios to form a recommended set of segments to be annotated. In order to ensure the generalization ability of dense2sparse, this scheme fixes the above L to 1000; at the same time, since fixing L to 1000 requires linear interpolation, in order to eliminate the influence of linear interpolation on the feature information of each video segment, this scheme performs linear interpolation on the feature information with a feature length greater than 1000, and fills the feature information with a feature length less than 1000 with zeros.

[0078] When training the dense2sparse unit, samples of corresponding quality are selected for training according to a specific IOU threshold, such as Fig.10As shown in the figure, the IOU threshold can be set to 0.5 in the H1 stage; then the IOU threshold of 0.5 is fine-tuned and sent to the H2 stage, and the fine-tuning result output by the H2 stage is sent to the H3 stage. Furthermore, different IOU thresholds can be set at different stages (for example, the rule is set that the higher the stage, the larger the IOU threshold), and each stage is cascaded to gradually improve the accuracy of detecting the target annotation area; for example, the IOU threshold of the H2 stage can be set to 0.6, and the IOU threshold of the H3 stage can be set to 0.7.

[0079] This embodiment adopts a dense2sparse unit, which can reduce the sparsity of the candidate annotation set while increasing the number of candidate annotation areas with different intersection-over-union ratios by performing non-maximum suppression on a dense set of candidate annotations arranged in descending order according to the score values. Moreover, a multi-level cascaded trainer is designed to achieve accurate search of video clips, which can effectively improve the accuracy of finding the clips that need to be labeled in a video.

[0080] Step 37: Classify the target annotation area through a classification network to obtain category information of the content to be annotated.

[0081] After obtaining the target annotated segment generated in step 36, it is necessary to further obtain the category information of the target annotated segment in order to output the final automatic annotation result of the video, which includes the category information and the target annotated area.

[0082] In a specific embodiment, Figure 4 As shown, the target annotation area is input into the classification network to generate corresponding category information; specifically, the first normalization module in the classification network is used to normalize the target annotation area to obtain a video processing segment; the extraction module in the classification network is used to perform feature extraction on the video processing segment to obtain first segment feature information; the calculation module in the classification network is used to classify the first segment feature information to obtain category information.

[0083] Furthermore, the calculation module may be used to classify the first segment feature information to obtain a classification result; and then the second normalization module may be used to normalize the classification result to obtain category information.

[0084] In summary, this embodiment proposes a deep learning-based automatic video labeling method. By using video clips of different categories to train classifiers and labeling generation networks, the category and start and end intervals of the video to be labeled can be automatically output. Except for the manual collection of label materials and training of the network in the early stage, no additional manpower is required, which helps to reduce labeling costs and improve labeling efficiency. Moreover, since the network has generalization capabilities, it can improve the versatility and reproducibility of any automatic labeling tasks.

[0085] See also Fig.11 , Fig.11 It is a structural diagram of an embodiment of an automatic labeling device provided in the present application. The automatic labeling device 110 includes a memory 111 and a processor 112 connected to each other. The memory 111 is used to store a computer program. When the computer program is executed by the processor 112, it is used to implement the automatic labeling method in the above embodiment.

[0086] See also Fig.12 , Fig.12 It is a structural diagram of an embodiment of a computer-readable storage medium provided in the present application. The computer-readable storage medium 120 is used to store a computer program 121. When the computer program 121 is executed by a processor, it is used to implement the automatic labeling method in the above embodiment.

[0087] The computer-readable storage medium 120 may be a server, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which may store program codes.

[0088] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only illustrative, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0089] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0090] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0091] The above descriptions are merely embodiments of the present application and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An automatic labeling method, characterized in that: include: Acquire a first video to be annotated, where the first video to be annotated includes content to be annotated; Performing feature extraction processing on the first video to be labeled to obtain first feature information; The first feature information is processed by a first annotation generation network to generate a candidate annotation set, wherein the candidate annotation set includes a candidate annotation area; wherein the first annotation generation network includes a feature enhancement module, a first estimation module, a second estimation module and a generation module, wherein the first feature information is enhanced by the feature enhancement module to generate second feature information; the second feature information is estimated by the first estimation module to obtain first score information; the second feature information is estimated by the second estimation module to obtain second score information; the first score information and the second score information are fused by the generation module to obtain score information; and the candidate annotation set is generated based on the score information; Using a second annotation generation network to correct the candidate annotation area to generate a target annotation area, wherein the target annotation area includes a video between a starting time and an end time corresponding to the content to be annotated; Classify the content to be marked in the target marking area to obtain category information of the content to be marked.

2. The automatic marking method according to claim 1, characterized in that: The step of classifying the content to be marked in the target marking area includes: Inputting the target marked area into a classification network; Using the first normalization module in the classification network to perform normalization processing on the target marked area to obtain a video processing segment; Using the extraction module in the classification network to perform feature extraction processing on the video processing segment to obtain first segment feature information; The first segment feature information is classified using a calculation module in the classification network to obtain the category information.

3. The automatic marking method according to claim 2, characterized in that: The classification network further includes a second normalization module. The step of using the calculation module in the classification network to classify the first segment feature information to obtain the category information includes: Using the calculation module to classify the first segment feature information to obtain a classification result; The second normalization module is used to normalize the classification result to obtain the category information.

4. The automatic marking method according to claim 2, characterized in that: The step of performing feature extraction processing on the first video to be labeled to obtain first feature information includes: Segmenting the first video to be labeled to obtain multiple video segments; Processing the video segment using the classification network to obtain second segment feature information, where the second segment feature information is feature information of the video segment; Normalize all the second segment feature information to obtain the first feature information.

5. The automatic marking method according to claim 1, characterized in that: The candidate annotation area includes a predicted starting point time and a predicted end point time, the predicted starting point time is an estimated time when the content to be annotated in the first video to be annotated first appears, and the predicted end point time is an estimated time when the content to be annotated in the first video to be annotated last appears, the first score information includes multiple regional probabilities, and the step of using the first estimation module to estimate the second feature information to obtain the first score information includes: Creating a regional grid, wherein the regional grid includes a plurality of grids, wherein the horizontal coordinate and the vertical coordinate of the grid are respectively the predicted start time and the predicted end time; The first estimation module is used to calculate the probability that the video segment corresponding to each grid is the target marked area to obtain the area probability.

6. The automatic marking method according to claim 5, characterized in that: The second score information includes a first probability value and a second probability value, the first probability value is the probability that the predicted starting point moment is the starting point moment, the second probability value is the probability that the predicted end point moment is the end point moment, the score information includes a plurality of score values, and the step of using the generation module to fuse the first score information and the second score information to obtain the score information includes: The area probability is multiplied by the corresponding first probability value and the second probability value to obtain the score value.

7. The automatic marking method according to claim 6, characterized in that: The step of generating the candidate annotation set based on the score information includes: Determine whether the score value in the score information is greater than a preset value; if so, the horizontal coordinate and the vertical coordinate of the grid corresponding to the score value form the candidate marking area; or Non-maximum suppression processing is performed on all the score values ​​to obtain the candidate marked areas.

8. The automatic marking method according to claim 1, characterized in that: The feature enhancement module includes an encoding module and an enhancement module. The step of using the feature enhancement module to enhance the first feature information to generate second feature information includes: Using the encoding module to encode the first feature information to obtain third feature information; The enhancement module is used to perform enhancement processing on the third characteristic information to obtain the second characteristic information.

9. The automatic marking method according to claim 8, characterized in that: The enhancement module includes an enhancement unit and a fusion unit. The step of using the enhancement module to enhance the third feature information to obtain the second feature information includes: Using the enhancement unit to perform enhancement processing on the third feature information to obtain fourth feature information; The fusion unit is used to perform fusion processing on the fourth feature information to obtain the second feature information.

10. The automatic marking method according to claim 9, characterized in that: The enhancement unit includes a timing enhancement unit and a spatial enhancement unit, and the method includes: Using the timing enhancement unit to perform timing enhancement processing on the third feature information to obtain timing feature information; Using the spatial enhancement unit to perform spatial enhancement processing on the third feature information to obtain spatial feature information; The fusion unit is used to fuse the temporal feature information, the spatial feature information and the third feature information to obtain the second feature information.

11. The automatic marking method according to claim 1, characterized in that: The step of using the second annotation generation network to correct the candidate annotation area to generate the target annotation area includes: encoding the first feature information using the second annotation generation network to obtain encoded information; The second annotation generation network is used to process the encoding information and the candidate annotation set to obtain the target annotation area.

12. An automatic labeling device, characterized in that: The invention comprises a memory and a processor connected to each other, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the automatic marking method according to any one of claims 1 to 11.

13. A computer-readable storage medium for storing a computer program, characterized in that: When the computer program is executed by a processor, it is used to implement the automatic marking method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Video annotation method and device, and storage medium

    CN110996138A