Media resource labeling method, device and equipment and storage medium

By using the annotation information of similar frames in media resource annotation to annotate non-sampled frames, the problem of high labor costs in existing technologies is solved, and an efficient annotation process is achieved.

CN113989703BActive Publication Date: 2025-12-09BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111210643.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-12-09
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

Existing text detection and recognition algorithms require a large number of manual annotations for training samples, resulting in high labor costs and affecting annotation speed.

Method used

By obtaining the initial annotation information of the sampled media resource frames of the media resource to be annotated, similar frames are identified and assimilated. The annotation information of adjacent frames is used to annotate non-sampled frames, reducing the need for manual annotation.

Benefits of technology

It reduced the cost of manual annotation and improved the efficiency of media resource annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989703B_ABST
    Figure CN113989703B_ABST
Patent Text Reader

Abstract

The present disclosure provides a media resource labeling method and device, equipment and a storage medium, and relates to the technical field of data processing, to solve the problem of high labor cost of media resource labeling. The method comprises: obtaining initial labeling information of at least two sample media resource frames in a to-be-labeled media resource; determining similar sample media resource frames from the at least two sample media resource frames for each sample media resource frame; performing assimilation processing on the initial labeling information of the objects in each sample media resource frame and the initial labeling information of the objects in the respective corresponding similar sample media resource frames to obtain target labeling information of the objects in each sample media resource frame; determining the labeling information of the objects in the non-sampling media resource between each two adjacent sample media resource frames based on the target labeling information of the objects in each sample media resource frame; and taking all the target labeling information and the labeling information of the objects in the non-sampling media resource as the labeling information of the objects in the to-be-labeled media resource.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a media resource labeling method and device, equipment and a storage medium. BACKGROUND

[0002] In order to quickly obtain the text information in the media resource, a text detection and recognition algorithm based on deep learning is usually used. However, the existing text detection and recognition algorithm is not mature enough and needs a large amount of training before being put into use.

[0003] Currently, the training samples of the text detection and recognition algorithm are obtained by frame-by-frame analysis of the media resource. If a large number of training samples are to be obtained, a large amount of manual cost will be consumed. Therefore, how to improve the text labeling speed of the media resource is crucial. SUMMARY

[0004] The present disclosure provides a media resource labeling method, device, equipment and storage medium to at least solve the problem that a large amount of manual cost is consumed in the prior art for media resource labeling.

[0005] The technical solution of the present disclosure is as follows: According to a first aspect of the present disclosure, a media resource labeling method is provided, which comprises: obtaining initial labeling information of at least two sample media resource frames in a to-be-labeled media resource, the initial labeling information being used to label objects in the sample media resource frames; for each sample media resource frame, determining a similar sample media resource frame from the at least two sample media resource frames; performing assimilation processing on the initial labeling information of the objects in each sample media resource frame and the initial labeling information of the objects in the respective corresponding similar sample media resource frame, to obtain target labeling information of the objects in each sample media resource frame; based on the target labeling information of the objects in each sample media resource frame, determining labeling information of objects in a non-sample media resource between every two adjacent sample media resource frames; the non-sample media resource is a media resource other than the sample media resource frames in the to-be-labeled media resource; and taking all the target labeling information and the labeling information of the objects in the non-sample media resource as the labeling information of the objects in the to-be-labeled media resource.

[0006] Optionally, when the target object exists in both of the two adjacent sampling media resource frames, based on the target annotation information of the object in each sampling media resource frame, the annotation information of the object in the non-sampling media resource between every two adjacent sampling media resource frames is determined, including: determining the frame number of all non-sampling media resource frames in the non-sampling media resource; determining the position information of the target object in each non-sampling media resource frame according to the frame number, the position information in the target annotation information of the target object in the front sampling media resource frame and the position information in the target annotation information of the target object in the rear sampling media resource frame; the front sampling media resource frame and the rear sampling media resource frame are two adjacent sampling media resource frames; annotating the target object in each non-sampling media resource frame according to the parameter information in the target annotation information of the target object in the front sampling media resource frame to obtain the parameter information of the target object in each non-sampling media resource frame; taking the position information and the parameter information of the target object in each non-sampling media resource frame as the annotation information of the target object in the non-sampling media resource.

[0007] Optionally, when the target object exists in the front sampling media resource frame and does not exist in the rear sampling media resource frame, based on the target annotation information of the object in each sampling media resource frame, the annotation information of the object in the non-sampling media resource between every two adjacent sampling media resource frames is determined, including: when the front sampling media resource frame is the first media resource frame, supplementing the annotation information of the target object in each non-sampling media resource frame in the non-sampling media resource according to the target annotation information of the target object in the front sampling media resource frame to obtain the annotation information of the target object in the non-sampling media resource; the front sampling media resource frame and the rear sampling media resource frame are two adjacent sampling media resource frames.

[0008] Optionally, when the target object appears in the front-sampling media resource frame and does not appear in the rear-sampling media resource frame, based on the target annotation information of the object in each sampling media resource frame, the annotation information of the object in the non-sampling media resource between every two adjacent sampling media resource frames is determined, including: when the current sampling media resource frame is not the first media resource frame, the position information of the target object in the target non-sampling media resource frame including the target object is obtained; the target non-sampling media resource frame is any media resource frame in the non-sampling media resource, and the front-sampling media resource frame and the rear-sampling media resource frame are two adjacent sampling media resource frames; when the position information of the target object in the target non-sampling media resource frame and the position information in the target annotation information of the target object in the front-sampling media resource frame are inconsistent, the frame number of all non-sampling media resource frames in the non-sampling media resource is determined; the position information of the target object in the front-sampling media resource frame, the position information of the target object in the target non-sampling media resource frame and the frame number are linearly processed to obtain the position information of the target object in each non-sampling media resource frame; based on the parameter information in the target annotation information of the target object in the front-sampling media resource frame, the parameter information of the target object in each non-sampling media resource frame is supplemented to obtain the parameter information of the target object in each non-sampling media resource frame; and the position information and the parameter information of the target object in each non-sampling media resource frame are taken as the annotation information of the target object in the non-sampling media resource.

[0009] Optionally, when the target object does not appear in the front-sampling media resource frame and appears in the rear-sampling media resource frame, based on the target annotation information of the object in each sampling media resource frame, the annotation information of the object in the non-sampling media resource between every two adjacent sampling media resource frames is determined, including: the position information of the target object in the target non-sampling media resource frame including the target object is obtained; the target non-sampling media resource frame is any media resource frame in the non-sampling media resource, and the front-sampling media resource frame and the rear-sampling media resource frame are two adjacent sampling media resource frames; when the position information of the target object in the target non-sampling media resource frame and the position information of the target object in the rear-sampling media resource frame are consistent, the target annotation information of the target object in the rear-sampling media resource frame is taken as the annotation information of the target object in each non-sampling media resource frame.

[0010] Optionally, the position information of the target object in the target non-sampled media resource frame is obtained, and the method further includes: when the position information of the target object in the target non-sampled media resource frame is inconsistent with the position information of the target object in the post-sampled media resource frame, determining the number of frames of all non-sampled media resource frames in the non-sampled media resource; performing linear processing on the position information of the target object in the post-sampled media resource frame, the position information of the target object in the target non-sampled media resource frame, and the number of frames to obtain the position information of the target object in each non-sampled media resource frame; supplementing the parameter information of the target object in each non-sampled media resource frame based on the parameter information in the target annotation information of the target object in the post-sampled media resource frame to obtain the parameter information of the target object in each non-sampled media resource frame; and taking the position information and the parameter information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resource.

[0011] Optionally, the similarity between the object in each sampled media resource frame and the object in the corresponding similar sampled media resource frame is greater than a threshold value; the similarity is the similarity between the annotation information of the object in each sampled media resource frame and the annotation information of the object in the similar sampled media resource frame, and / or the similarity between the position information of the object in each sampled media resource frame and the position information of the object in the similar sampled media resource frame.

[0012] Optionally, before taking all the target annotation information and the annotation information of the object in the non-sampled media resource as the annotation information of the object in the media resource to be annotated, the method further includes: obtaining the identification annotation information of the object in at least two sampled media resource frames; in a case where the initial annotation information of the object in the at least two sampled media resource frames is consistent with the identification annotation information, correcting the annotation information of the object in the non-sampled media resource by using the identification annotation information to obtain the corrected annotation information of the object in the non-sampled media resource.

[0013] Optionally, before taking all the target annotation information and the annotation information of the object in the non-sampled media resource as the annotation information of the object in the media resource to be annotated, the method further includes: obtaining the identification annotation information of the object in at least two sampled media resource frames; in a case where the initial annotation information of the object in the at least two sampled media resource frames is inconsistent with the identification annotation information, correcting the identification annotation information by using the initial annotation information to obtain the corrected identification annotation information; and in a case where the corrected identification annotation information is consistent with the initial annotation information, correcting the annotation information of the object in the non-sampled media resource by using the corrected identification annotation information to obtain the corrected annotation information of the object in the non-sampled media resource.

[0014] According to a second aspect of the present disclosure, a media resource labeling apparatus is provided, which comprises an acquisition module and a processing module. The acquisition module is configured to acquire initial labeling information of at least two sampled media resource frames in a media resource to be labeled, the initial labeling information being used to label an object in the sampled media resource frame; the processing module is configured to determine, for each sampled media resource frame, a similar sampled media resource frame from the at least two sampled media resource frames; the processing module is further configured to perform assimilation processing on the initial labeling information of the object in each sampled media resource frame and the initial labeling information of the object in the respective corresponding similar sampled media resource frame, to obtain target labeling information of the object in each sampled media resource frame; the processing module is further configured to determine, based on the target labeling information of the object in each sampled media resource frame, labeling information of the object in a non-sampled media resource between every two adjacent sampled media resource frames; the non-sampled media resource is a media resource other than the sampled media resource frames in the media resource to be labeled; and the processing module is further configured to take all the target labeling information and the labeling information of the object in the non-sampled media resource as the labeling information of the object in the media resource to be labeled.

[0015] Optionally, the processing module is further configured to determine a frame number of all non-sampled media resource frames in the non-sampled media resource; the processing module is further configured to determine, according to the frame number, position information in the target labeling information of the target object in a front sampled media resource frame and position information in the target labeling information of the target object in a rear sampled media resource frame, position information of the target object in each non-sampled media resource frame; the front sampled media resource frame and the rear sampled media resource frame are two adjacent sampled media resource frames; the processing module is further configured to label the target object in each non-sampled media resource frame according to parameter information in the target labeling information of the target object in the front sampled media resource frame, to obtain parameter information of the target object in each non-sampled media resource frame; and the processing module is further configured to take the position information and the parameter information of the target object in each non-sampled media resource frame as the labeling information of the target object in the non-sampled media resource.

[0016] Optionally, when the front sampled media resource frame is the first media resource frame, the processing module is further configured to supplement the labeling information of the target object in each non-sampled media resource frame in the non-sampled media resource according to the target labeling information of the target object in the front sampled media resource frame, to obtain the labeling information of the target object in the non-sampled media resource; and the front sampled media resource frame and the rear sampled media resource frame are two adjacent sampled media resource frames.

[0017] Optionally, the obtaining module is further configured to, when the current sampling media resource frame is a non-first media resource frame, obtain position information of the target object in a target non-sampling media resource frame including the target object; the target non-sampling media resource frame is any media resource frame in the non-sampling media resources; the front sampling media resource frame and the rear sampling media resource frame are two adjacent sampling media resource frames; the processing module is further configured to, when the position information of the target object in the target non-sampling media resource frame is inconsistent with the position information in the target annotation information of the target object in the front sampling media resource frame, determine the number of frames of all the non-sampling media resource frames in the non-sampling media resources; the processing module is further configured to perform linear processing on the position information of the target object in the front sampling media resource frame, the position information of the target object in the target non-sampling media resource frame, and the number of frames, to obtain the position information of the target object in each non-sampling media resource frame; the processing module is further configured to supplement the parameter information of the target object in each non-sampling media resource frame based on the parameter information in the target annotation information of the target object in the front sampling media resource frame, to obtain the parameter information of the target object in each non-sampling media resource frame; and the processing module is further configured to take the position information and the parameter information of the target object in each non-sampling media resource frame as the annotation information of the target object in the non-sampling media resources.

[0018] Optionally, the obtaining module is further configured to obtain position information of the target object in a target non-sampling media resource frame including the target object; the target non-sampling media resource frame is any media resource frame in the non-sampling media resources; the front sampling media resource frame and the rear sampling media resource frame are two adjacent sampling media resource frames; and the processing module is further configured to, when the position information of the target object in the target non-sampling media resource frame is consistent with the position information of the target object in the rear sampling media resource frame, take the target annotation information of the target object in the rear sampling media resource frame as the annotation information of the target object in each non-sampling media resource frame.

[0019] Optionally, the processing module is further configured to determine the number of frames of the non-sampled media resource when the position information of the target object in the target non-sampled media resource frame and the position information of the target object in the post-sampled media resource frame are inconsistent; the processing module is further configured to perform linear processing on the position information of the target object in the post-sampled media resource frame, the position information of the target object in the target non-sampled media resource frame, and the number of frames to obtain the position information of the target object in each non-sampled media resource frame; the processing module is further configured to supplement the parameter information of the target object in each non-sampled media resource frame based on the parameter information in the target annotation information of the target object in the post-sampled media resource frame to obtain the parameter information of the target object in each non-sampled media resource frame; and the processing module is further configured to take the position information and the parameter information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resource.

[0020] Optionally, the similarity between the object in each sampled media resource frame and the object in the corresponding similar sampled media resource frame is greater than a threshold value; the similarity is the similarity between the annotation information of the object in each sampled media resource frame and the annotation information of the object in the similar sampled media resource frame, and / or the similarity between the position information of the object in each sampled media resource frame and the position information of the object in the similar sampled media resource frame.

[0021] Optionally, the obtaining module is further configured to obtain the identification annotation information of the object in the at least two sampled media resource frames; and the processing module is further configured to correct the annotation information of the object in the non-sampled media resource by using the identification annotation information to obtain the corrected annotation information of the object in the non-sampled media resource when the initial annotation information of the object in the at least two sampled media resource frames is consistent with the identification annotation information.

[0022] Optionally, the obtaining module is further configured to obtain the identification annotation information of the object in the at least two sampled media resource frames; and the processing module is further configured to correct the identification annotation information by using the initial annotation information to obtain the corrected identification annotation information when the initial annotation information of the object in the at least two sampled media resource frames is inconsistent with the identification annotation information; and the processing module is further configured to correct the annotation information of the object in the non-sampled media resource by using the corrected identification annotation information to obtain the corrected annotation information of the object in the non-sampled media resource when the corrected identification annotation information is consistent with the initial annotation information.

[0023] According to a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement any of the optional media resource annotation methods according to the first aspect.

[0024] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, and the computer-readable storage medium stores instructions, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any one of the optional media resource labeling methods in the first aspect.

[0025] According to a fifth aspect of the present disclosure, a computer program product is provided, and the computer program product contains instructions, when the instructions in the computer program product are executed by a processor of an electronic device, the instructions implement any one of the optional media resource labeling methods in the first aspect.

[0026] The technical solutions provided by the embodiments of the present disclosure at least have the following beneficial effects: in the above solutions, the electronic device finds similar sample media resource frames corresponding to each sample media resource frame, and assimilates the labeling information of each sample media resource frame and the labeling information of the similar sample media resource frame to obtain the target labeling information of the object in each sample media resource frame. According to the obtained target labeling information, the non-sample media resource is labeled, and finally the labeling information of the object in the media resource to be labeled is obtained. Compared with the prior art, all media resource frames are labeled frame by frame, and the present disclosure only needs to label the sample media resource frames frame by frame, and then supplements the non-sample media resource according to the initial labeling information of the sample media resource frame to obtain the labeling information of the object in the media resource to be labeled. In this way, the labor cost is greatly reduced, and the efficiency of media resource labeling is improved.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0028] The accompanying drawings, which are incorporated into and form part of the specification, illustrate an embodiment consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0029] Figure 1 is one of the flowcharts of a media resource labeling method according to an exemplary embodiment;

[0030] Figure 2 is the second flowchart of a media resource labeling method according to an exemplary embodiment;

[0031] Figure 3 is the third flowchart of a media resource labeling method according to an exemplary embodiment;

[0032] Figure 4 is the fourth flowchart of a media resource labeling method according to an exemplary embodiment;

[0033] Figure 5 Figure 5 is a flowchart of a media resource labeling method according to an example embodiment;

[0034] Figure 6 Figure 6 is a flowchart of a media resource labeling method according to an example embodiment;

[0035] Figure 7 Figure 7 is a structural block diagram of a media resource labeling device according to an example embodiment;

[0036] Figure 8 Figure 8 is a structural schematic diagram of a computer program product of a media resource labeling method according to an example embodiment. DETAILED DESCRIPTION

[0037] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings.

[0038] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The data involved in the present disclosure can be data authorized by the user or sufficiently authorized by the parties. The implementation described in the following example embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0039] It should also be understood that the term "comprising" indicates the presence of described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or components.

[0040] Based on the background art, the present disclosure provides a media resource labeling method. The method is to find similar sample media resource frames corresponding to each sample media resource frame, and to assimilate the labeling information of each sample media resource frame and the labeling information of the similar sample media resource frame to obtain the target labeling information of the object in each sample media resource frame. According to the obtained target labeling information, the non-sampling media resource is labeled, and finally the labeling information of the object in the media resource to be labeled is obtained. Compared with the prior art, the present disclosure does not need manual frame-by-frame labeling, which reduces the labor cost.

[0041] The media resource labeling method provided by the embodiments of the present disclosure is exemplarily described as follows.

[0042] The media resource labeling method provided by the present disclosure can be applied to an electronic device.

[0043] In some embodiments, the electronic device can be a server, or a terminal, or other electronic device for labeling media resources, and the present disclosure does not limit the same.

[0044] The server can be a single server, or a server cluster composed of multiple servers. In some embodiments, the server cluster can also be a distributed cluster. The present disclosure does not limit the specific implementation of the server.

[0045] The terminal can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) \ virtual reality (VR) device, and other devices that can install and use a content community application (such as Kuaishou), and the present disclosure does not specially limit the specific form of the electronic device. It can interact with the user through one or more ways such as a keyboard, a touchpad, a touch screen, a remote controller, voice interaction, or a handwriting device.

[0046] The technical solutions in the embodiments of the present disclosure will be described clearly and completely in combination with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present disclosure.

[0047] As shown in Figure 1 When the media resource labeling method is applied to an electronic device, the media resource labeling method can include:

[0048] Step 11: The electronic device acquires initial labeling information of at least two sampled media resource frames in the media resource to be labeled. The initial labeling information is used to label the objects in the sampled media resource frames.

[0049] In the embodiments of the present disclosure, the electronic device obtains a media resource to be labeled, samples the media resource according to a preset sampling frequency to obtain a plurality of sampled media resource frames, and manually labels an object in each sampled media resource frame to obtain initial labeling information of the object in each sampled media resource frame.

[0050] The media resource can be a video or any resource with labeling requirements. The object in the sampled media resource frame can be text or an image. The labeling information includes position information and parameter information, and the parameter information includes an object category, such as a foreground object or a background object, an identification content of the object, a labeling box of the object, a presentation mode of the labeling box of the object, a dotting order of the labeling box of the object, edge information of the object, and the like. The present disclosure does not limit this according to actual use.

[0051] Step 12: The electronic device determines a similar sampled media resource frame from at least two sampled media resource frames for each sampled media resource frame.

[0052] In the embodiments of the present disclosure, after obtaining the sampled media resource frames and the labeling information of the object in each sampled media resource frame, the electronic device finds a similar sampled media resource frame similar to each sampled media resource frame for each sampled media resource frame. The principle of finding can be that a sampled media resource frame satisfying a preset condition is a similar sampled media resource frame, or can be set according to the actual performance of each sampled media resource frame, and the present disclosure does not limit the comparison.

[0053] Optionally, a similarity between the object in each sampled media resource frame and the object in the corresponding similar sampled media resource frame is greater than a threshold value. The similarity is a similarity between the labeling information of the object in each sampled media resource frame and the labeling information of the object in the similar sampled media resource frame, and / or a similarity between the position information of the object in each sampled media resource frame and the position information of the object in the similar sampled media resource frame.

[0054] Specifically, the disclosure provides a method for determining a similar sampling media resource frame corresponding to each sampling media resource frame. For example, the determination method is whether the similarity between the objects in each sampling media resource frame and the objects in the corresponding similar sampling media resource frame is greater than a threshold. The specific determination process is as follows: the media resource to be labeled includes a first sampling media resource frame, a second sampling media resource frame, a third sampling media resource frame, and the like. When determining the similar sampling media resource frame corresponding to the first sampling media resource frame, first, the bounding box of the object in the first sampling media resource frame, the bounding box of the object in the second sampling media resource frame, the bounding box of the object in the third sampling media resource frame, the object category of the object in the first sampling media resource frame, the object category of the object in the second sampling media resource frame, and the object category of the object in the third sampling media resource frame are determined.

[0055] Then, the intersection over union and the content edit distance between the bounding box of the object in the first sampling media resource frame and the bounding box of the object in the second sampling media resource frame are determined, as well as the intersection over union and the content edit distance between the bounding box of the object in the first sampling media resource frame and the bounding box of the object in the third sampling media resource frame. It is found that the intersection over union and the content edit distance between the bounding box of the object in the first sampling media resource frame and the bounding box of the object in the second sampling media resource frame are greater than a first threshold, and the intersection over union and the content edit distance between the bounding box of the object in the first sampling media resource frame and the bounding box of the object in the third sampling media resource frame are less than the first threshold.

[0056] Next, it is found that the object category of the object in the first sampling media resource frame is the same as that of the object in the second sampling media resource frame, and the object category of the object in the first sampling media resource frame is different from that of the object in the third sampling media resource frame. Therefore, it is considered that the second sampling media resource frame is the similar sampling media resource frame of the first sampling media resource frame.

[0057] The above embodiment provides a technical solution with at least the following beneficial effects: by using the above technical features, a method for determining a similar sampling media resource frame is provided, which provides a data basis for subsequent labeling of sampling video frames.

[0058] Step 13: The electronic device performs assimilation processing on the initial labeling information of the object in each sampling media resource frame and the initial labeling information of the object in the corresponding similar sampling media resource frame, to obtain the target labeling information of the object in each sampling media resource frame.

[0059] In the embodiment of the disclosure, the electronic device performs assimilation processing on the initial labeling information of the object in each sampling media resource frame and the initial labeling information of the object in the corresponding similar sampling media resource frame, so that the transition of the labeling information of the object in each sampling media resource frame is smoother and more coherent. Specifically, the assimilation processing is to adjust the position information and parameter information in the initial labeling information.

[0060] For example, the assimilation processing includes: adjusting the bounding box of the object in the sample media resource frame from a rectangular box to a quadrilateral box, modifying or deleting the element with content mutation in the object, and determining the object attribute in the parameter information as the object attribute selected in the object in the most similar sample media resource frame, and the like.

[0061] Step 14: The electronic device determines the labeling information of the object in the non-sample media resource between any two adjacent sample media resource frames based on the target labeling information of the object in each sample media resource frame. The non-sample media resource is the media resource in the media resource to be labeled except the sample media resource frame.

[0062] In the embodiment of the present disclosure, the electronic device supplements the labeling information of the non-sample media resource between any two adjacent sample media resource frames according to the target labeling information of the object in each sample media resource frame. Specifically, the labeling information of the non-sample media resource is supplemented according to the target labeling information, and the supplement content includes the position of the object in the non-sample media resource, the identification content of the object, the attribute information of the object, and the like. For example, the non-sample media resource is the part of the original media resource containing the sample media resource, and the part of the sample media resource is removed.

[0063] Step 15: The electronic device takes all the target labeling information and the labeling information of the object in the non-sample media resource as the labeling information of the object in the media resource to be labeled.

[0064] In the embodiment of the present disclosure, after obtaining the target labeling information and the labeling information of the object in the non-sample media resource, the labeling object of the object in the media resource to be labeled can be considered to be obtained.

[0065] The technical scheme provided by the above embodiment has at least the following beneficial effects: by using the above technical features, the similar sample media resource frame corresponding to each sample media resource frame is found, the labeling information of each sample media resource frame and the labeling information of the similar sample media resource frame are assimilated to obtain the target labeling information of the object in each sample media resource frame, the non-sample media resource is labeled according to the obtained target labeling information, and finally the labeling information of the object in the media resource to be labeled is obtained. Compared with the prior art, all media resource frames are labeled frame by frame, and the present disclosure only needs to label the sample media resource frame frame by frame, and then supplements the non-sample media resource according to the initial labeling information of the sample media resource frame to obtain the labeling information of the object in the media resource to be labeled. In this way, the labor cost is greatly reduced, and the efficiency of media resource labeling is improved.

[0066] Specifically, in combination with Figure 1 For example, Figure 2As shown, when the target object exists in both of the two adjacent sampling media resource frames, step 14 determines the labeling information of the object in the non-sampling media resource between every two adjacent sampling media resource frames based on the target labeling information of the object in each sampling media resource frame, including:

[0067] Step 21: The electronic device determines the frame number of all non-sampling media resource frames in the non-sampling media resource.

[0068] In the embodiments of the present disclosure, since the sampling media resource frame is obtained by sampling the media resource, comparing the media resource with the sampling media resource frame can obtain the non-sampling media resource between any two adjacent sampling media resource frames and the frame number of the non-sampling media resource frames in the non-sampling media resource.

[0069] Step 22: The electronic device determines the position information of the target object in each non-sampling media resource frame according to the frame number, the position information of the target object in the target labeling information in the front sampling media resource frame, and the position information of the target object in the target labeling information in the rear sampling media resource frame; wherein the front sampling media resource frame and the rear sampling media resource frame are two adjacent sampling media resource frames.

[0070] In the embodiments of the present disclosure, since the target object exists in both of the two adjacent sampling media resource frames, the target object also exists in the non-sampling media resource between the two adjacent sampling media resource frames. Based on the target object, the position information of the target object in the front sampling media resource frame, the position of the target object in the rear sampling media resource frame, and the frame number of the non-sampling media resource frame, the electronic device can determine the position change of the target object in each non-sampling media resource frame, and thus determine the position information of the target object in each non-sampling media resource frame. For example, the position information of the target object in the target labeling information in the front sampling media resource frame can be determined by using the four-corner point average sampling estimation.

[0071] Step 23: The electronic device labels the target object in each non-sampling media resource frame according to the parameter information in the target labeling information of the target object in the front sampling media resource frame, to obtain the parameter information of the target object in each non-sampling media resource frame.

[0072] In the embodiment of the present disclosure, since the annotation information includes the position information and the parameter information, after the position information of the target object in each non-sampling media resource frame is determined, the parameter information of the target object in each non-sampling media resource frame needs to be annotated. Specifically, the annotation of the parameter information of the target object in each non-sampling media resource frame is annotated according to the parameter information in the target annotation information of the target object in the front-sampling media resource frame, until the edge information in the text area of the annotated object changes greatly, which proves that the target object has been annotated. The annotated content includes: object type, object recognition content, object annotation box, object annotation box presentation mode, object annotation box dotting order, object edge information, and the like.

[0073] Step 24: The electronic device takes the position information and the parameter information of the target object in each non-sampling media resource frame as the annotation information of the target object in the non-sampling media resource.

[0074] In the embodiment of the present disclosure, after the position information and the parameter information of the target object in each non-sampling media resource frame are obtained, the annotation information of the target object in the non-sampling media resource can be considered to be obtained.

[0075] The technical scheme provided by the above embodiment has at least the following beneficial effects: by using the above technical features, if two adjacent sampling media resource frames both include the target object, the annotation information of the target object in each non-sampling media resource frame is supplemented according to the target annotation information of the target object in the two adjacent sampling media resource frames, so as to obtain the annotation information of the target object in each non-sampling media resource frame after annotation. In this way, the unannotated annotation information is supplemented by using the existing target annotation information, which effectively reduces the analysis workload, thereby improving the annotation speed.

[0076] Specifically, in combination with Figure 1 As shown in Figure 3 When the target object appears in the front-sampling media resource frame and does not appear in the rear-sampling media resource frame, step 14 determines the annotation information of the object in the non-sampling media resource between every two adjacent sampling media resource frames based on the target annotation information of the object in each sampling media resource frame, including:

[0077] Step 31: When the current sampling media resource frame is the first media resource frame, the electronic device supplements the annotation information of the target object in each non-sampling media resource frame in the non-sampling media resource according to the target annotation information of the target object in the front-sampling media resource frame, so as to obtain the annotation information of the target object in each non-sampling media resource frame. The front-sampling media resource frame and the rear-sampling media resource frame are two adjacent sampling media resource frames.

[0078] In the embodiments of the present disclosure, when the target object appears in the front-sampling media resource frame and does not appear in the rear-sampling media resource frame, the electronic device needs to determine whether the front-sampling media resource frame is the first media resource frame. If the front-sampling media resource frame is the first media resource frame, it is considered that the position information of the target object in each non-sampling media resource frame is consistent with the position information of the target object in the front-sampling media resource frame, and the annotation information of the target object in each non-sampling media resource frame is supplemented according to the target annotation information of the target object in the front-sampling media resource frame. The supplemented parameter information is detailed in step 23. Until the edge information of the supplemented target object changes greatly, it is considered that the supplement of the target annotation information of the target object is completed.

[0079] The technical solutions provided by the above embodiments have at least the following beneficial effects: by using the above technical features, when the target object appears in the front-sampling media resource frame and does not appear in the rear-sampling media resource frame, and the front-sampling media resource frame is the first media resource frame, the annotation information of the object in the front-sampling media resource frame is used as the annotation information of the target object in each non-sampling media resource frame. In this way, for the repeated target object, its annotation information can be directly borrowed, so that the annotation is quickly completed.

[0080] Specifically, in combination with Figure 1 As shown in Figure 4 When the target object appears in the front-sampling media resource frame and does not appear in the rear-sampling media resource frame, step 14 determines the annotation information of the object in the non-sampling media resource between every two adjacent sampling media resource frames based on the target annotation information of the object in each sampling media resource frame, including:

[0081] Step 41: When the current sampling media resource frame is not the first media resource frame, the electronic device acquires the position information of the target object in the target non-sampling media resource frame including the target object. The target non-sampling media resource frame is any media resource frame in the non-sampling media resource, and the front-sampling media resource frame and the rear-sampling media resource frame are two adjacent sampling media resource frames.

[0082] In the embodiments of the present disclosure, since the target object appears in the front-sampling media resource frame and does not appear in the rear-sampling media resource frame, it indicates that the target object ends in a certain non-sampling media resource frame between the front-sampling media resource frame and the rear-sampling media resource frame. At this time, the electronic device needs to acquire any target non-sampling media resource frame including the target object, so as to determine the position information of the target object in other non-sampling media resource frames between the front-sampling media resource frame and the rear-sampling media resource frame according to the position information of the target object in the target non-sampling media resource frame.

[0083] Step 42: When the position information of the target object in the target non-sampled media resource frame and the position information of the target object in the previous sampled media resource frame are inconsistent, the electronic device determines the number of frames of all non-sampled media resource frames in the non-sampled media resource.

[0084] In the embodiments of the present disclosure, when the electronic device determines that the position information of the target object in the target non-sampled media resource frame and the position information of the target object in the previous sampled media resource frame are inconsistent, it indicates that the position of the target object has deviated over time. Therefore, it is necessary to determine the number of frames of non-sampled media resource frames in the non-sampled media resource between the previous sampled media resource frame and the subsequent sampled media resource frame, so as to calculate the position information of the target object in each non-sampled media resource frame according to the position deviation and the number of frames of non-sampled media resource frames.

[0085] Step 43: The electronic device linearly processes the position information of the target object in the previous sampled media resource frame, the position information of the target object in the target non-sampled media resource frame, and the number of frames to obtain the position information of the target object in each non-sampled media resource frame.

[0086] In the embodiments of the present disclosure, the electronic device can calculate the position deviation of the target object according to the position information of the target object in the previous sampled media resource frame and the position information of the target object in the target non-sampled media resource frame, and then combine the number of frames of non-sampled media resource frames to obtain the position information of the target object in each non-sampled media resource frame.

[0087] Step 44: The electronic device supplements the parameter information of the target object in each non-sampled media resource frame based on the parameter information in the target annotation information of the target object in the previous sampled media resource frame to obtain the parameter information of the target object in each non-sampled media resource frame.

[0088] In the embodiments of the present disclosure, after obtaining the position information of the target object in each non-sampled media resource frame, the electronic device supplements the parameter information of the target object in each non-sampled media resource frame according to the parameter information in the target annotation information of the target object in the previous sampled media resource frame to obtain the annotated annotation information of the target object in each non-sampled media resource frame. The supplemented parameter information is detailed in step 23, and until the edge information of the target object changes greatly, it is considered that the supplement of the target annotation information of the target object is completed.

[0089] Step 45: The electronic device takes the position information and the parameter information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resource.

[0090] In the embodiments of the present disclosure, after obtaining the position information and the parameter information of the target object in each non-sampling media resource frame, it is considered that the labeling information of the target object in the non-sampling media resource is obtained.

[0091] The technical solutions provided by the above embodiments have at least the following beneficial effects: by using the above technical features, when the target object appears in the front-sampling media resource frame, does not appear in the rear-sampling media resource frame, and the front-sampling media resource frame is not the first media resource frame, an implementation scheme for media resource labeling of the target object is provided.

[0092] Specifically, in combination with Figure 1 As shown in Figure 5 When the target object does not appear in the front-sampling media resource frame and appears in the rear-sampling media resource frame, step 14 determines the labeling information of the object in the non-sampling media resource between every two adjacent sampling media resource frames based on the target labeling information of the object in each sampling media resource frame, including:

[0093] Step 51: The electronic device obtains the position information of the target object in a target non-sampling media resource frame including the target object. The target non-sampling media resource frame is any media resource in the non-sampling media resource, and the front-sampling media resource and the rear-sampling media resource frame are two adjacent sampling media resource frames.

[0094] In the embodiments of the present disclosure, since the target object appears in the rear-sampling media resource frame and does not appear in the front-sampling media resource frame, it indicates that the target object starts in a certain non-sampling media resource frame between the front-sampling media resource frame and the rear-sampling media resource frame. At this time, the electronic device needs to obtain a target non-sampling media resource frame including the target object, so as to determine the position information of the target object in other non-sampling media resource frames between the front-sampling media resource frame and the rear-sampling media resource frame according to the position information of the target object in the target non-sampling media resource frame.

[0095] Step 52a: When the position information of the target object in the target non-sampling media resource frame is consistent with the position information of the target object in the rear-sampling media resource frame, the target labeling information of the target object in the rear-sampling media resource frame is taken as the labeling information of the target object in each non-sampling media resource frame.

[0096] In the embodiments of the present disclosure, when the electronic device determines that the position information of the target object in the target non-sampling media resource frame is consistent with the position information of the target object in the rear-sampling media resource frame, it indicates that the position of the target object in the media resource frame has not deviated with the change of time. Then, the target labeling information of the target object in the rear-sampling media resource frame can be used to supplement the labeling information of the target object in each non-sampling media resource frame, so as to obtain the labeling information of the target object in each non-sampling media resource frame.

[0097] The technical solutions provided by the above embodiments have at least the following beneficial effects: when the target object appears in the post-sampling media resource frame and does not appear in the pre-sampling media resource frame, and the position of the target object does not shift, the technical solutions provide an implementation scheme for media resource labeling for the target object.

[0098] Specifically, in combination with Figure 1 As shown in the figure, if the target object does not appear in the pre-sampling media resource frame and appears in the post-sampling media resource frame, after the step 51 of obtaining the position information of the target object in the target non-sampling media resource frame including the target object, the following steps are further included: Figure 5

[0099] Step 52b: When the position information of the target object in the target non-sampling media resource frame and the position information of the target object in the post-sampling media resource frame are inconsistent, the electronic device determines the frame number of all non-sampling media resource frames in the non-sampling media resource.

[0100] In the embodiments of the present disclosure, when the electronic device determines that the position information of the target object in the target non-sampling media resource frame and the position information of the target object in the post-sampling media resource frame are inconsistent, it indicates that the position of the target object has shifted over time. Therefore, it is necessary to determine the frame number of the non-sampling media resource frame between the pre-sampling media resource frame and the post-sampling media resource frame, so as to calculate the position information of the target object in each non-sampling media resource frame according to the position shift amount and the frame number of the non-sampling media resource frame.

[0101] Step 53: The electronic device linearly processes the position information of the target object in the post-sampling media resource frame, the position information of the target object in the target non-sampling media resource frame, and the frame number to obtain the position information of the target object in each non-sampling media resource frame.

[0102] In the embodiments of the present disclosure, the electronic device can calculate the position shift amount of the target object according to the position information of the target object in the post-sampling media resource frame and the position information of the target object in the target non-sampling media resource frame, and then combine the number of non-sampling media resource frames in the pre-sampling media resource frame and the post-sampling media resource frame to obtain the position information of the target object in each non-sampling media resource frame.

[0103] Step 54: Based on the parameter information in the target labeling information of the target object in the post-sampling media resource frame, the electronic device supplements the parameter information of the target object in each non-sampling media resource frame to obtain the parameter information of the target object in each non-sampling media resource frame.

[0104] ​In this embodiment of the disclosure, after obtaining the position information of the target object in each unsampled media resource frame, the electronic device supplements the parameter information of the target object in each unsampled media resource frame according to the parameter information in the target annotation information of the target object in the post-sampled media resource frame, so as to obtain the parameter information of the target object in each unsampled media resource frame. For details of the supplementary content, please refer to step 23.

[0105] Step 55: The electronic device uses the position information and parameter information of the target object in each unsampled media resource frame as the annotation information of the target object in the unsampled media resource.

[0106] In this embodiment of the disclosure, after the electronic device obtains the location information and parameter information of the target object in each unsampled media resource frame, it considers that it has obtained the annotation information of the target object in the unsampled media resource.

[0107] The technical solution provided by the above embodiments has at least the following beneficial effects: by adopting the above technical features, when the target object appears in the post-sampled media resource frame, does not appear in the pre-sampled media resource frame, and the position of the target object is offset, a method is provided for the media resource annotation of the target object.

[0108] Specifically, in combination Figure 1 ,like Figure 6 As shown, in this media resource annotation method, before step 15, which uses all target annotation information and the annotation information of objects in non-sampled media resources as the annotation information of objects in the media resource to be annotated, it also includes:

[0109] Step 61: The electronic device acquires the identification and annotation information of objects in at least two sampled media resource frames.

[0110] In this embodiment of the disclosure, the electronic device uses Optical Character Recognition (OCR) technology to identify objects in each sampled media resource frame, and finally obtains the identification annotation information of the objects in each sampled media resource frame. The specific identification annotation information includes the identification content, object attributes, etc.

[0111] Step 62a: If the initial annotation information and identification annotation information of objects in at least two sampled media resource frames are consistent, the identification annotation information is used to correct the annotation information of objects in non-sampled media resources to obtain the corrected annotation information of objects in non-sampled media resources.

[0112] In the embodiments of the present disclosure, in the case that the initial annotation information and the recognition annotation information of the object in each sampled media resource frame are consistent, the OCR algorithm is used to correct the annotation information of the object in the non-sampled media resource, so as to obtain the annotation information of the object in the non-sampled media resource with higher accuracy. Specifically, the correction of the annotation information of the object in the non-sampled media resource is not limited to the OCR algorithm, and other algorithms can also be used, and the present disclosure does not limit this.

[0113] The technical solutions provided by the above embodiments have at least the following beneficial effects: by using the above technical features, the annotation information obtained by the method of the present disclosure is corrected by the OCR algorithm, so as to obtain the annotation information of the object in the media resource to be annotated with higher accuracy. Therefore, the annotation information is corrected in multiple ways, which can improve the accuracy of annotation and provide a data basis for subsequent related work.

[0114] Optionally, in combination with Figure 6 Before step 15 in the media resource annotation method, the annotation information of the object in the non-sampled media resource is obtained by using the initial annotation information and the recognition annotation information of the object in all the sampled media resource frames as the annotation information of the object in the media resource to be annotated.

[0115] Step 61: The electronic device obtains the recognition annotation information of the object in at least two sampled media resource frames. For details, see the foregoing.

[0116] Step 62b: In the case that the initial annotation information and the recognition annotation information of the object in at least two sampled media resource frames are inconsistent, the initial annotation information is used to correct the recognition annotation information, so as to obtain the corrected recognition annotation information.

[0117] In the embodiments of the present disclosure, in the case that the initial annotation information and the recognition annotation information of the object in each sampled media resource frame are inconsistent, since the initial annotation information is artificially annotated information, in general, it is considered that the accuracy of the artificially annotated initial annotation information is higher. Therefore, the initial annotation information is used to correct the recognition annotation information of the object in the sampled media resource frame, so as to obtain the corrected recognition annotation information.

[0118] Step 63b: In the case that the corrected recognition annotation information and the initial annotation information are consistent, the corrected recognition annotation information is used to correct the annotation information of the object in the non-sampled media resource, so as to obtain the corrected annotation information of the object in the non-sampled media resource.

[0119] In the embodiments of the present disclosure, in the case that the initial annotation information and the corrected recognition annotation information of the object in each sampled media resource frame are consistent, the corrected recognition annotation information is used to correct the annotation information of the object in the non-sampled media resource, so as to obtain the annotation information of the object in the non-sampled media resource with higher accuracy.

[0120] The technical solutions provided by the above embodiments have at least the following beneficial effects: the technical features described above are used to judge the accuracy of the recognition annotation information, if the accuracy of the recognition annotation information is low, the recognition annotation information is first corrected, and the corrected recognition annotation information is used to adjust the annotation information obtained by the method of the present disclosure, so as to obtain annotation information of the object in the non-sampling media resource with higher precision. Therefore, by bidirectional correction of the annotation information, the accuracy of the annotation can be improved.

[0121] The above Figures 2-6 The method provided by the embodiments of the present disclosure is described in detail. In order to achieve the above functions, the media resource annotation device includes hardware structures and / or software modules corresponding to each function, and these hardware structures and / or software modules corresponding to each function can constitute an electronic device. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.

[0122] The embodiments of the present disclosure can divide the functional modules of the electronic device according to the above method examples. For example, the electronic device can include a media resource annotation device, which can divide each functional module corresponding to each function, or integrate two or more functions in one processing module. The above integrated module can be realized in the form of hardware or software functional module. It should be noted that the division of modules in the embodiments of the present disclosure is illustrative, and is only a logical functional division. Actual implementation can have another division manner.

[0123] In the following, the embodiments of the present disclosure are described in combination with Figure 7 The media resource annotation device provided by the embodiments of the present disclosure is described in detail. It should be understood that the description of the device embodiments and the description of the method embodiments correspond to each other, therefore, the content not described in detail can be referred to the above method embodiments, and for brevity, the details are not described here.

[0124] Figure 7 is a structural schematic diagram of a media resource annotation device according to an exemplary embodiment, applied to an electronic device, see Figure 7As shown, the media resource labeling apparatus comprises an obtaining module 71 and a processing module 72. The obtaining module 71 is configured to obtain initial labeling information of at least two sampled media resource frames in a media resource to be labeled, the initial labeling information being used to label an object in the sampled media resource frame. For example, refer to Figure 1 As shown, the obtaining module 71 is configured to perform step 11. The processing module 72 is configured to determine, for each sampled media resource frame, a similar sampled media resource frame from the at least two sampled media resource frames. For example, refer to Figure 1 As shown, the processing module 72 is configured to perform step 12. The processing module 72 is further configured to perform assimilation processing on the labeling information of the object in each sampled media resource frame and the similar labeling information, to obtain target labeling information of the object in each sampled media resource frame. For example, refer to Figure 1 As shown, the processing module 72 is configured to perform step 13. The processing module 72 is further configured to determine, based on the target labeling information of the object in each sampled media resource frame, labeling information of the object in a non-sampled media resource between every two adjacent sampled media resource frames; the non-sampled media resource being a media resource other than the sampled media resource frames in the media resource to be labeled. For example, refer to Figure 1 As shown, the processing module 72 is configured to perform step 14. The processing module 72 is further configured to take all the target labeling information and the labeling information of the object in the non-sampled media resource as the labeling information of the object in the media resource to be labeled. For example, refer to Figure 1 As shown, the processing module 72 is configured to perform step 15.

[0125] Optionally, the processing module 72 is further configured to determine the number of frames of all non-sampled media resource frames in the non-sampled media resource. For example, refer to Figure 2 As shown, the processing module 72 is configured to perform step 21. The processing module 72 is further configured to determine, according to the number of frames, position information in the target labeling information of the target object in the front sampled media resource frame and position information in the target labeling information of the target object in the rear sampled media resource frame, position information of the target object in each non-sampled media resource frame; the front sampled media resource frame and the rear sampled media resource frame being two adjacent sampled media resource frames. For example, refer to Figure 2 As shown, the processing module 72 is configured to perform step 22. The processing module 72 is further configured to label the target object in each non-sampled media resource frame according to parameter information in the target labeling information of the target object in the front sampled media resource frame, to obtain parameter information of the target object in each non-sampled media resource frame. For example, refer to Figure 2As shown in FIG. 7, the processing module 72 is configured to perform step 23. The processing module 72 is further configured to take the position information and the parameter information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resource. For example, refer to Figure 2 As shown in FIG. 8, the processing module 72 is configured to perform step 24.

[0126] Optionally, the processing module 72 is further configured to take the annotation information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resource, when the current sampled media resource frame is the first media resource frame. For example, refer to Figure 3 As shown in FIG. 9, the processing module 72 is configured to perform step 31.

[0127] Optionally, the acquisition module 71 is further configured to acquire the position information of the target object in the target non-sampled media resource frame including the target object, when the current sampled media resource frame is a non-first media resource frame. The target non-sampled media resource frame is any media resource frame in the non-sampled media resource. The previous sampled media resource frame and the next sampled media resource frame are two adjacent sampled media resource frames. For example, refer to Figure 4 As shown in FIG. 10, the acquisition module 71 is configured to perform step 41. The processing module 72 is further configured to determine the frame number of all non-sampled media resource frames in the non-sampled media resource, when the position information of the target object in the target non-sampled media resource frame and the position information in the target annotation information of the target object in the previous sampled media resource frame are inconsistent. For example, refer to Figure 4 As shown in FIG. 11, the processing module 72 is configured to perform step 42. The processing module 72 is further configured to perform linear processing on the position information of the target object in the previous sampled media resource frame, the position information of the target object in the target non-sampled media resource frame, and the frame number, to obtain the position information of the target object in each non-sampled media resource frame. For example, refer to Figure 4 As shown in FIG. 12, the processing module 72 is configured to perform step 43. The processing module 72 is further configured to supplement the parameter information of the target object in each non-sampled media resource frame based on the parameter information in the target annotation information of the target object in the previous sampled media resource frame, to obtain the parameter information of the target object in each non-sampled media resource frame. For example, refer to Figure 4 As shown in FIG. 13, the processing module 72 is configured to perform step 44. The processing module 72 is further configured to take the position information and the parameter information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resource. For example, refer to Figure 4 As shown in FIG. 14, the processing module 72 is configured to perform step 45.

[0128] Optionally, the obtaining module 71 is further configured to obtain position information of the target object in a target non-sampled media resource frame including the target object; the target non-sampled media resource frame is any media resource frame in the non-sampled media resources; the front-sampled media resource frame and the back-sampled media resource frame are two adjacent sampled media resource frames. For example, referring to Figure 5 The processing module 72 is configured to perform step 52a. For example, referring to Figure 5 Optionally, the processing module 72 is further configured to determine the number of frames of all non-sampled media resource frames in the non-sampled media resources when the position information of the target object in the target non-sampled media resource frame and the position information of the target object in the back-sampled media resource frame are inconsistent. For example, referring to

[0129] The processing module 72 is configured to perform step 52b. The processing module 72 is further configured to perform linear processing on the position information of the target object in the back-sampled media resource frame, the position information of the target object in the target non-sampled media resource frame, and the number of frames to obtain the position information of the target object in each non-sampled media resource frame. For example, referring to Figure 5 The processing module 72 is configured to perform step 53. The processing module 72 is further configured to supplement the parameter information of the target object in each non-sampled media resource frame based on the parameter information in the target annotation information of the target object in the back-sampled media resource frame to obtain the parameter information of the target object in each non-sampled media resource frame. For example, referring to Figure 5 The processing module 72 is configured to perform step 54. The processing module 72 is further configured to take the position information and the parameter information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resources. For example, referring to Figure 2 The processing module 72 is configured to perform step 55. Figure 2 Optionally, the similarity between the object in each sampled media resource frame and the object in the corresponding similar sampled media resource frame is greater than a threshold; the similarity is the similarity between the annotation information of the object in each sampled media resource frame and the annotation information of the object in the similar sampled media resource frame, and / or the similarity between the position information of the object in each sampled media resource frame and the position information of the object in the similar sampled media resource frame.

[0130]

[0131] ​Optionally, the obtaining module 71 is further configured to obtain the recognition annotation information of the object in the at least two sample media resource frames. For example, refer to Figure 6 As shown in the figure, the obtaining module 71 is configured to perform step 61. The processing module 72 is further configured to, in the case that the initial annotation information of the object in the at least two sample media resource frames is consistent with the recognition annotation information, correct the annotation information of the object in the non-sample media resource by using the recognition annotation information to obtain the corrected annotation information of the object in the non-sample media resource. For example, refer to Figure 6 As shown in the figure, the processing module 72 is configured to perform step 62a.

[0132] Optionally, the obtaining module 71 is further configured to obtain the recognition annotation information of the object in the at least two sample media resource frames. For example, refer to Figure 6 As shown in the figure, the obtaining module 71 is configured to perform step 61. The processing module 72 is further configured to, in the case that the initial annotation information of the object in the at least two sample media resource frames is inconsistent with the recognition annotation information, correct the recognition annotation information by using the initial annotation information to obtain the corrected recognition annotation information. For example, refer to Figure 6 As shown in the figure, the processing module 72 is configured to perform step 62b. The processing module 72 is further configured to, in the case that the corrected recognition annotation information is consistent with the initial annotation information, correct the annotation information of the object in the non-sample media resource by using the corrected recognition annotation information to obtain the corrected annotation information of the object in the non-sample media resource. For example, refer to Figure 6 As shown in the figure, the processing module 72 is configured to perform step 63b.

[0133] Of course, the media resource annotation apparatus provided by the embodiments of the present disclosure includes but is not limited to the above-mentioned modules. For example, the media resource annotation apparatus can further include a storage module 73. The storage module 73 can be used to store the program code of the media resource annotation apparatus, and can also be used to store the data generated by the media resource annotation apparatus during the running process, such as the data in the write request.

[0134] In actual implementation, the obtaining module 71 and the processing module 72 can be realized by the data generation time apparatus shown in the figure. Figure 1 The specific execution process can refer to the description of any one of the media resource annotation methods shown in the figure. Figures 1-6

[0135] Another embodiment of the present disclosure also provides a computer readable storage medium, which stores instructions, when the instructions run on an electronic device, the electronic device performs any one of the media resource annotation methods shown in the above Figures 1-6

[0136] ​​Figure 8 A conceptual partial view of a computer program product provided by embodiments of the present disclosure is schematically shown, the computer program product including a computer program for executing a computer process on a computing device.

[0137] In one embodiment, the computer program product is provided using a signal bearing medium 810. The signal bearing medium 810 can include one or more program instructions which, when executed by one or more processors, can provide the functionality or some portion thereof described above with respect to Figure 1 the embodiments shown in FIG. 11. Thus, for example, one or more features of steps 11-14 can be undertaken by one or more instructions associated with the signal bearing medium 810. Further, the program instructions of the signal bearing medium 810 also describe example instructions. Figure 1 Figure 8 In some examples, the signal bearing medium 810 can include a computer readable medium 811 such as, but not limited to, a hard disk drive, a compact disc (CD), a digital versatile disc (DVD), a digital tape, memory, read-only memory (ROM), random access memory (RAM), etc.

[0138] In some examples, the signal bearing medium 810 can include a computer readable medium 811 such as, but not limited to, a hard disk drive, a compact disc (CD), a digital versatile disc (DVD), a digital tape, memory, read-only memory (ROM), random access memory (RAM), etc.

[0139] In some examples, the signal bearing medium 810 can include a computer readable medium 811 such as, but not limited to, a hard disk drive, a compact disc (CD), a digital versatile disc (DVD), a digital tape, memory, read-only memory (ROM), random access memory (RAM), etc.

[0140] In some examples, the signal bearing medium 810 can include a computer readable medium 811 such as, but not limited to, a hard disk drive, a compact disc (CD), a digital versatile disc (DVD), a digital tape, memory, read-only memory (ROM), random access memory (RAM), etc.

[0141] The signal bearing medium 810 can be conveyed by a wireless form of the communication medium 813. The one or more program instructions can be, for example, computer executable instructions or logic-implemented instructions.

[0142] In some examples, such as for the media resource annotation apparatus described above with respect to Figure 7 In some examples, such as for the media resource annotation apparatus described above with respect to

[0143] ​Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of functional modules is taken as an example, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete the whole classification or part of the functions described above.

[0144] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other manners. For example, the above-described device embodiments are merely illustrative. For example, the division of the modules or units can be different, and each can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0145] The units described as separate components can or can not be physically separate, and the components displayed as units can be one physical unit or multiple physical units, that is, can be located in one place, or can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0146] In addition, each functional unit in the various embodiments of the present disclosure can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.

[0147] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present disclosure essentially or say the part that makes contributions to the prior art or the whole classification or part of the technical solutions can be embodied in the form of a software product. The software product is stored in a storage medium, and includes several instructions to make a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the methods of the various embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, and various media that can store program codes.

[0148] The above merely describes specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and any changes or replacements within the technical scope disclosed by the present disclosure shall be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A media resource labeling method, characterized in that, The method comprises the following steps: acquiring initial labeling information of at least two sample media resource frames in a media resource to be labeled, the initial labeling information being used for labeling an object in the sample media resource frame; for each sample media resource frame, determining a similar sample media resource frame from the at least two sample media resource frames; performing assimilation processing on the initial labeling information of the object in each sample media resource frame and the initial labeling information of the object in the respective corresponding similar sample media resource frame, to obtain target labeling information of the object in each sample media resource frame; based on the target labeling information of the object in each sample media resource frame, determining labeling information of the object in a non-sample media resource between every two adjacent sample media resource frames; the non-sample media resource being a media resource other than the sample media resource frame in the media resource to be labeled; when a target object appears in a front sample media resource frame and does not appear in a rear sample media resource frame, the step of determining, based on the target labeling information of the object in each sample media resource frame, the labeling information of the object in the non-sample media resource between every two adjacent sample media resource frames, comprises: when the front sample media resource frame is the first media resource frame, supplementing, according to the target labeling information of the target object in the front sample media resource frame, the labeling information of the target object in each non-sample media resource frame in the non-sample media resource, to obtain the labeling information of the target object in the non-sample media resource; the front sample media resource frame and the rear sample media resource frame being two adjacent sample media resource frames; using all the target labeling information and the labeling information of the object in the non-sample media resource as the labeling information of the object in the media resource to be labeled.

2. The method of claim 1, wherein, when the target object exists in both of the two adjacent sample media resource frames, the step of determining, based on the target labeling information of the object in each sample media resource frame, the labeling information of the object in the non-sample media resource between every two adjacent sample media resource frames, comprises: determining the number of frames of all non-sample media resource frames in the non-sample media resource; determining the position information of the target object in each non-sample media resource frame according to the number of frames, the position information in the target labeling information of the target object in the front sample media resource frame, and the position information in the target labeling information of the target object in the rear sample media resource frame; the front sample media resource frame and the rear sample media resource frame being two adjacent sample media resource frames; labeling the target object in each non-sample media resource frame according to the parameter information in the target labeling information of the target object in the front sample media resource frame, to obtain the parameter information of the target object in each non-sample media resource frame; using the position information and the parameter information of the target object in each non-sample media resource frame as the labeling information of the target object in the non-sample media resource.

3. The method of claim 1, wherein, When the target object appears in the front-sampling media resource frame and does not appear in the rear-sampling media resource frame, the determining of the annotation information of the object in the non-sampling media resource between every two adjacent sampling media resource frames based on the target annotation information of the object in each sampling media resource frame comprises: When the front-sampling media resource frame is not the first media resource frame, the position information of the target object in a target non-sampling media resource frame including the target object is acquired; the target non-sampling media resource frame is any media resource frame in the non-sampling media resource, and the front-sampling media resource frame and the rear-sampling media resource frame are two adjacent sampling media resource frames; When the position information of the target object in the target non-sampling media resource frame and the position information in the target annotation information of the target object in the front-sampling media resource frame are inconsistent, the frame number of all non-sampling media resource frames in the non-sampling media resource is determined; The position information of the target object in the front-sampling media resource frame, the position information of the target object in the target non-sampling media resource frame and the frame number are linearly processed to obtain the position information of the target object in each non-sampling media resource frame; Based on the parameter information in the target annotation information of the target object in the front-sampling media resource frame, the parameter information of the target object in each non-sampling media resource frame is supplemented to obtain the parameter information of the target object in each non-sampling media resource frame; The position information and the parameter information of the target object in each non-sampling media resource frame are taken as the annotation information of the target object in the non-sampling media resource.

4. The method of claim 1, wherein, When the target object does not appear in the front-sampling media resource frame and appears in the rear-sampling media resource frame, the determining of the annotation information of the object in the non-sampling media resource between every two adjacent sampling media resource frames based on the target annotation information of the object in each sampling media resource frame comprises: The position information of the target object in a target non-sampling media resource frame including the target object is acquired; the target non-sampling media resource frame is any media resource frame in the non-sampling media resource, and the front-sampling media resource frame and the rear-sampling media resource frame are two adjacent sampling media resource frames; When the position information of the target object in the target non-sampling media resource frame and the position information of the target object in the rear-sampling media resource frame are consistent, the target annotation information of the target object in the rear-sampling media resource frame is taken as the annotation information of the target object in each non-sampling media resource frame.

5. The method of claim 4, wherein, The acquisition of the position information of the target object in the target non-sampling media resource frame including the target object further comprises: When the position information of the target object in the target non-sampling media resource frame and the position information of the target object in the rear-sampling media resource frame are inconsistent, the frame number of all non-sampling media resource frames in the non-sampling media resource is determined; linearly processing the position information of the target object in the post-sampling media resource frame, the position information of the target object in the target non-sampling media resource frame, and the frame number to obtain the position information of the target object in each non-sampling media resource frame; complementing the parameter information of the target object in each non-sampling media resource frame based on the parameter information in the target annotation information of the target object in the post-sampling media resource frame to obtain the parameter information of the target object in each non-sampling media resource frame; taking the position information and the parameter information of the target object in each non-sampling media resource frame as the annotation information of the target object in the non-sampling media resource.

6. The method of claim 1, wherein a similarity between the object in each sampling media resource frame and the object in the corresponding similar sampling media resource frame is greater than a threshold value; the similarity is a similarity between the annotation information of the object in each sampling media resource frame and the annotation information of the object in the similar sampling media resource frame, and / or a similarity between the position information of the object in each sampling media resource frame and the position information of the object in the similar sampling media resource frame. Before the step of taking all the target annotation information and the annotation information of the object in the non-sampling media resource as the annotation information of the object in the media resource to be annotated, the method further comprises:

7. The method according to any one of claims 1 to 6, characterized in that, obtaining identification annotation information of the object in the at least two sampling media resource frames; in a case where the initial annotation information and the identification annotation information of the object in the at least two sampling media resource frames are consistent, correcting the annotation information of the object in the non-sampling media resource by using the identification annotation information to obtain the annotation information of the object in the non-sampling media resource after correction. Before the step of taking all the target annotation information and the annotation information of the object in the non-sampling media resource as the annotation information of the object in the media resource to be annotated, the method further comprises:

8. The method according to any one of claims 1-6, characterized in that, obtaining identification annotation information of the object in the at least two sampling media resource frames; in a case where the initial annotation information and the identification annotation information of the object in the at least two sampling media resource frames are inconsistent, correcting the identification annotation information by using the initial annotation information to obtain identification annotation information after correction; in a case where the identification annotation information after correction and the initial annotation information are consistent, correcting the annotation information of the object in the non-sampling media resource by using the identification annotation information after correction to obtain the annotation information of the object in the non-sampling media resource after correction. comprises:

9. A media resource labeling apparatus, comprising: an obtaining module configured to obtain initial annotation information of at least two sampling media resource frames in a media resource to be annotated, the initial annotation information being used to annotate an object in the sampling media resource frame; a processing module configured to, for each sampling media resource frame, determine a similar sampling media resource frame from the at least two sampling media resource frames; ​ The processing module is further configured to perform assimilation processing on the initial annotation information of the object in each sampled media resource frame and the initial annotation information of the object in the respective corresponding similar sampled media resource frame, to obtain target annotation information of the object in each sampled media resource frame; The processing module is further configured to determine annotation information of the object in the non-sampled media resource between every two adjacent sampled media resource frames based on the target annotation information of the object in each sampled media resource frame; the non-sampled media resource is a media resource in the media resource to be annotated other than the sampled media resource frames; When the target object appears in a front sampled media resource frame and does not appear in a rear sampled media resource frame, the processing module is further configured to, when the front sampled media resource frame is the first media resource frame, supplement the annotation information of the target object in each non-sampled media resource frame in the non-sampled media resource according to the target annotation information of the target object in the front sampled media resource frame, to obtain the annotation information of the target object in the non-sampled media resource, the front sampled media resource frame and the rear sampled media resource frame being two adjacent sampled media resource frames; The processing module is further configured to take all the target annotation information and the annotation information of the object in the non-sampled media resource as the annotation information of the object in the media resource to be annotated.

10. The apparatus of claim 9, wherein The processing module is further configured to determine the number of frames of all non-sampled media resource frames in the non-sampled media resource; The processing module is further configured to determine the position information of the target object in each non-sampled media resource frame according to the number of frames, the position information in the target annotation information of the target object in the front sampled media resource frame, and the position information in the target annotation information of the target object in the rear sampled media resource frame, the front sampled media resource frame and the rear sampled media resource frame being two adjacent sampled media resource frames; The processing module is further configured to annotate the target object in each non-sampled media resource frame according to the parameter information in the target annotation information of the target object in the front sampled media resource frame, to obtain the parameter information of the target object in each non-sampled media resource frame; The processing module is further configured to take the position information and the parameter information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resource.

11. The apparatus of claim 9, wherein The acquisition module is further configured to, when the front sampled media resource frame is a non-first media resource frame, acquire the position information of the target object in a target non-sampled media resource frame including the target object; The target non-sampled media resource frame is any media resource frame in the non-sampled media resource, and the front sampled media resource frame and the rear sampled media resource frame are two adjacent sampled media resource frames. The processing module is further configured to determine the number of frames of all non-sampled media resource frames in the non-sampled media resource when the position information of the target object in the target non-sampled media resource frame and the position information in the target annotation information of the target object in the front-sampled media resource frame are inconsistent. The processing module is further configured to perform linear processing on the position information of the target object in the front-sampled media resource frame, the position information of the target object in the target non-sampled media resource frame, and the number of frames to obtain the position information of the target object in each non-sampled media resource frame. The processing module is further configured to supplement the parameter information of the target object in each non-sampled media resource frame based on the parameter information in the target annotation information of the target object in the front-sampled media resource frame to obtain the parameter information of the target object in each non-sampled media resource frame. The processing module is further configured to take the position information and the parameter information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resource.

12. The apparatus of claim 9, wherein The obtaining module is further configured to obtain position information of a target object in a target non-sampled media resource frame including the target object, the target non-sampled media resource frame being any media resource frame in the non-sampled media resource, and the front-sampled media resource frame and the rear-sampled media resource frame being two adjacent sampled media resource frames. The processing module is further configured to take the target annotation information of the target object in the rear-sampled media resource frame as the annotation information of the target object in each non-sampled media resource frame when the position information of the target object in the target non-sampled media resource frame is consistent with the position information of the target object in the rear-sampled media resource frame.

13. The apparatus of claim 12, wherein The processing module is further configured to determine the number of frames of all non-sampled media resource frames in the non-sampled media resource when the position information of the target object in the target non-sampled media resource frame and the position information in the target annotation information of the target object in the rear-sampled media resource frame are inconsistent. The processing module is further configured to perform linear processing on the position information of the target object in the rear-sampled media resource frame, the position information of the target object in the target non-sampled media resource frame, and the number of frames to obtain the position information of the target object in each non-sampled media resource frame. The processing module is further configured to supplement the parameter information of the target object in each non-sampled media resource frame based on the parameter information in the target annotation information of the target object in the rear-sampled media resource frame to obtain the parameter information of the target object in each non-sampled media resource frame. The processing module is further configured to take the position information and the parameter information of the target object in each non-sampled media resource frame as the annotation information of the target object in the non-sampled media resource.

14. The apparatus of claim 9, wherein, the similarity between the object in each of the sampled media resource frames and the object in the corresponding similar sampled media resource frame is greater than a threshold value; the similarity is a similarity between annotation information of the object in each of the sampled media resource frames and annotation information of the object in the similar sampled media resource frame, and / or a similarity between position information of the object in each of the sampled media resource frames and position information of the object in the similar sampled media resource frame.

15. The apparatus of any one of claims 9-14, wherein, the obtaining module is further configured to obtain identification annotation information of the object in the at least two sampled media resource frames; the processing module is further configured to, in a case where the initial annotation information of the object in the at least two sampled media resource frames and the identification annotation information are consistent, correct the annotation information of the object in the non-sampled media resource using the identification annotation information to obtain corrected annotation information of the object in the non-sampled media resource.

16. The apparatus of any one of claims 9-14, wherein, the obtaining module is further configured to obtain identification annotation information of the object in the at least two sampled media resource frames; the processing module is further configured to, in a case where the initial annotation information of the object in the at least two sampled media resource frames and the identification annotation information are inconsistent, correct the identification annotation information using the initial annotation information to obtain corrected identification annotation information; the processing module is further configured to, in a case where the corrected identification annotation information and the initial annotation information are consistent, correct the annotation information of the object in the non-sampled media resource using the corrected identification annotation information to obtain corrected annotation information of the object in the non-sampled media resource.

17. An electronic device, comprising: The electronic device includes: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the media resource annotation method of any one of claims 1-8.

18. A computer-readable storage medium having stored thereon instructions, the instructions comprising, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the media resource annotation method of any one of claims 1-8.

19. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the electronic device, the media resource annotation method of any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Image content marking method and apparatus thereof

    CN106682595A

  • Image annotation method, device and equipment and storage medium

    CN110991491A

  • Text labeling method

    CN113033380A