Data annotation methods, apparatus, readable media and electronic devices
By filtering and labeling target data in video data, and utilizing semantic description text sets and preset labeling rates, the problems of high labor costs and difficulty in ensuring quality in the data labeling process are solved, achieving efficient and low-cost data labeling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2026-03-13
AI Technical Summary
In the field of machine learning, data annotation is costly in terms of manpower and difficult to guarantee in terms of annotation quality. This is especially true when a large amount of video data needs to be annotated, particularly when the target category accounts for a small proportion. Existing technologies are not able to complete data annotation efficiently.
By acquiring multiple sets of text descriptions, target data is selected from the data to be labeled according to a preset labeling rate, and the target data is categorized. Semantic description text is used for data filtering and labeling, reducing the amount of data while ensuring the proportion of data in the specified category.
It reduces the manpower required for data annotation, saves annotation costs, and improves the quality and efficiency of data annotation, while ensuring the proportion of the specified category in the target data.
Smart Images

Figure CN116127373B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing, and more specifically, to a method, apparatus, readable medium, and electronic device for data annotation. Background Technology
[0002] As we all know, in some machine learning applications, it is necessary to first train the machine learning model based on labeled training data, and then complete tasks such as classification and prediction based on the trained model. Therefore, data labeling is a very important task in the field of machine learning, and the quality and quantity of labeled data largely determine the performance of the model itself and the business metrics that can be achieved.
[0003] In specific business scenarios, in order to obtain a certain amount of sample data for the target category, a large amount of data needs to be labeled, which places a great demand on labeling manpower and makes the manpower cost of labeling data high. Summary of the Invention
[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, this disclosure provides a data annotation method, the method comprising:
[0006] Multiple text description sets are obtained, with different text description sets corresponding to different annotation categories. Each text description set includes one or more semantic description texts, and one or more semantic description texts in the same text description set correspond to the same annotation category.
[0007] Target data is selected from the data to be labeled based on multiple sets of text descriptions and a preset labeling rate, wherein the preset labeling rate is the expected proportion of data of a specified category in the target data, and the labeling category includes the specified category;
[0008] The target data is categorized.
[0009] Secondly, a data annotation apparatus is provided, the apparatus comprising:
[0010] The acquisition module is used to acquire multiple text description sets. Different text description sets correspond to different annotation categories. Each text description set includes one or more semantic description texts. One or more semantic description texts in the same text description set correspond to the same annotation category.
[0011] A data filtering module is used to filter target data from the data to be labeled based on multiple sets of text descriptions and a preset labeling rate, wherein the preset labeling rate is the expected proportion of data of a specified category in the target data, and the labeling category includes the specified category;
[0012] The annotation module is used to label the target data by category.
[0013] Thirdly, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect of this disclosure.
[0014] Fourthly, an electronic device is provided, comprising:
[0015] A storage device on which computer programs are stored;
[0016] A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect of this disclosure.
[0017] Through the above technical solution, multiple text description sets are obtained, with different text description sets corresponding to different annotation categories. Each text description set includes one or more semantic description texts, and one or more semantic description texts in the same text description set correspond to the same annotation category. Target data is filtered from the data to be annotated based on the multiple text description sets and a preset annotation rate. The preset annotation rate is the expected proportion of data of a specified category in the target data, and the annotation category includes the specified category. The target data is then annotated by category. Since the amount of target data after filtering is less than the amount of data to be annotated, annotating the smaller amount of target data by category can reduce the workload of annotation and alleviate the manpower required for data annotation, thereby saving manpower costs. In addition, filtering data according to the preset annotation rate of the specified category can ensure that the data of the specified category has a certain proportion in the target data, thus guaranteeing the quality of data annotation.
[0018] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0020] Figure 1This is a flowchart illustrating a data annotation method according to an exemplary embodiment.
[0021] Figure 2 It is based on Figure 1 The illustrated embodiment shows a flowchart of a data annotation method.
[0022] Figure 3 It is based on Figure 1 The illustrated embodiment shows a flowchart of the method for step S102.
[0023] Figure 4 It is based on Figure 1 The illustrated embodiment shows a flowchart of a data annotation method.
[0024] Figure 5 This is a block diagram illustrating a data annotation apparatus according to an exemplary embodiment.
[0025] Figure 6 It is based on Figure 5 The illustrated embodiment shows a block diagram of a data annotation apparatus.
[0026] Figure 7 It is based on Figure 5 The illustrated embodiment shows a block diagram of a data annotation apparatus.
[0027] Figure 8 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0028] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0029] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0030] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0031] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0032] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0033] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0034] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0035] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0036] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0037] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0038] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0039] This disclosure is primarily applied to scenarios where sample data is labeled before model training. In specific business scenarios, due to differences in data distribution, many business categories that need to be distinguished account for a small proportion of the overall raw data. In this case, in order to obtain a certain number of sample data for the target category, a large amount of raw data needs to be labeled, which places a significant demand on labeling manpower. For example, if a positive sample accounts for 5% of the overall data, and model training requires 10,000 positive samples, directly labeling the raw data would require labeling 10,000 / 0.05 = 200,000 data points, resulting in high labeling manpower costs.
[0040] Specifically, because videos contain rich modal information, such as video frames (i.e., images), text, and audio, the annotation of video data is only truly completed after the information of each modality is labeled. Furthermore, if the video data of the target category accounts for a small proportion of the overall original video data, a large amount of manpower is needed to complete the annotation of the video data in a timely manner, which will result in higher annotation manpower costs.
[0041] When annotating video data, the usual method is to inject the entire video and title into the annotation queue according to the original information and then manually annotate them. However, since annotators need to pay attention to the information of each modality at the same time, the annotation efficiency of annotators will be reduced, and some modalities are more likely to be missed. For example, annotators may only pay attention to the information of video frames and not pay attention to the text information on the video frames, making it difficult to guarantee the overall annotation quality.
[0042] To address the aforementioned problems, this disclosure provides a data annotation method, apparatus, readable medium, and electronic device. This method can acquire multiple sets of text descriptions and filter target data from the data to be annotated based on these multiple sets of text descriptions and a preset annotation rate. The preset annotation rate is the expected proportion of data of a specified category in the target data. The annotation category includes the specified category. Then, the target data is annotated by category. Since the amount of target data after filtering is less than the amount of data to be annotated, annotating the smaller amount of target data by category can reduce the workload and manpower requirements for data annotation, thereby saving manpower costs. Furthermore, filtering data based on the preset annotation rate of the specified category ensures that data of the specified category has a certain proportion in the target data, thus guaranteeing the quality of data annotation.
[0043] The specific embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0044] Figure 1 This is a flowchart illustrating a data annotation method according to an exemplary embodiment, such as... Figure 1 As shown, the method includes the following steps:
[0045] In step S101, multiple text description sets are obtained. Different text description sets correspond to different annotation categories. Each text description set includes one or more semantic description texts. One or more semantic description texts in the same text description set correspond to the same annotation category.
[0046] The text description set is a pre-set set of texts used to represent the labeling categories of the data. Different text description sets represent different labeling categories, and a text description set includes one or more semantic description texts.
[0047] For example, suppose we want to classify the sentiment of video titles, and the sentiment categories can be divided into positive and negative. Specifically, titles with more positive text content can be labeled as positive, and titles with more negative text content can be labeled as negative. Based on the data annotation method provided in this disclosure, different text description sets can be established for different sentiment categories, that is, the text description sets are used to represent the sentiment categories. For example, the text description set 1 corresponding to the positive sentiment category can be {this short video is very good, this short video is very excellent, this short video is very amazing}, and the text description set 2 corresponding to the negative sentiment category can be {this short video is very horrible, this short video is very bad, this short video is very shocking}. Here, text description set 1 includes three semantic description texts: 'this short video is very good', 'this short video is very excellent', and 'this short video is very amazing'; text description set 2 includes 'this short video is very horrible', 'this short video is very bad', and 'this short video is very...'. The three semantic descriptions of 'shocking' are merely illustrative examples and are not intended to limit the scope of this disclosure.
[0048] It should be noted that in related technologies, data annotation methods often use numbers without semantic information for sentiment category labeling; for example, positive and negative video data are encoded as 0 / 1 numbers respectively. This disclosure uses a set of text descriptions with semantic information to label data categories, ensuring that both feature text and category label text in the training sample set contain semantic information. Compared to 0 / 1 labels that do not contain semantic information, this can lead to greater performance improvements.
[0049] In step S102, target data is selected from the data to be labeled based on multiple sets of text descriptions and a preset labeling rate. The preset labeling rate is the expected proportion of data of a specified category in the target data, and the labeling category includes the specified category.
[0050] For example, suppose the annotation categories to be labeled include two categories: positive samples and negative samples. If 10,000 positive samples need to be labeled, and the preset labeling rate for positive samples is 5%, meaning that positive sample data accounts for 5% of the data to be labeled, then the amount of data to be labeled would be 10,000 / 0.05 = 200,000. If the preset labeling rate for positive samples is increased to 20%, then the amount of data to be labeled would be 10,000 / 0.2 = 50,000. Obviously, increasing the preset labeling rate for a specific category of data can reduce the amount of data to be labeled. Therefore, in this step, target data can be filtered from the data to be labeled based on multiple text description sets and preset labeling rates, thereby reducing the amount of data to be labeled and saving annotation manpower costs.
[0051] In step S103, the target data is labeled with categories.
[0052] In this step, the annotators can label the target data with smaller volumes after filtering, saving manpower costs for annotation.
[0053] Using the above method, since the amount of target data after filtering is smaller than the amount of data to be labeled, class labeling of target data with a smaller amount of data can reduce the workload of labeling and reduce the manpower required for data labeling, thereby saving manpower costs for data labeling. In addition, data filtering based on the preset labeling rate of the specified category can also ensure that the data of the specified category has a certain proportion in the target data, thereby ensuring the quality of data labeling.
[0054] Figure 2 It is based on Figure 1 The illustrated embodiment presents a flowchart of a data annotation method, as shown in the figure. Figure 2 As shown, this set of text descriptions can be preset using the following steps:
[0055] In step S201, a preset description file corresponding to the data to be labeled is obtained. The preset description file is used to record the keyword information corresponding to different labeling categories of the data to be labeled.
[0056] In real-world data annotation scenarios, there are usually pre-set annotation documents (i.e., preset description files) that specifically describe the data annotation rules and processes. Annotators can use these preset description files to annotate data for various scenarios.
[0057] For example, taking the labeling of video titles with sentiment categories as an example, the preset description file corresponding to the video title to be labeled describes the keyword information corresponding to each sentiment category that needs to be labeled. For example, the keyword information corresponding to positive sentiment categories includes keywords such as good, excellent, and amazing, while the keyword information corresponding to negative sentiment categories includes keywords such as horrible, bad, and shocking.
[0058] In step S202, for each annotation category, one or more semantic description texts corresponding to the annotation category are generated based on the keyword information and preset description template, and the one or more semantic description texts are used as a text description set.
[0059] The preset description template can be a pre-defined fixed text description format.
[0060] For example, for the sentiment classification of video titles, the preset description template can be: "this short video is very A", or "this short video is very B". For the text description set corresponding to a positive sentiment classification, each keyword corresponding to the positive sentiment category can be used to replace "A" in the preset description template to obtain one or more semantic description texts, thus generating the text description set corresponding to the positive sentiment category. Similarly, for the text description set corresponding to a negative sentiment classification, each keyword corresponding to the negative sentiment category can be used to replace "B" in the preset description template to obtain one or more semantic description texts, thus generating the text description set corresponding to the negative sentiment category.
[0061] After generating a set of text descriptions that correspond one-to-one with the labeled categories, the set of text descriptions can be stored offline so that it can be directly read for data labeling before model training.
[0062] Figure 3 It is based on Figure 1 The flowchart of step S102 shown in the embodiment is as follows: Figure 3 As shown, step S102 includes the following sub-steps:
[0063] In step S1021, for each piece of data to be labeled, the data to be labeled is combined with each semantic description text in multiple text description sets to obtain multiple data pairs.
[0064] For example, suppose there are three data points to be labeled: A1, A2, and A3. If we perform a binary classification on these three data points, then there are two text description sets: a and b. Text description set a includes two semantic description texts, a1 and a2, and text description set b includes two semantic description texts, b1 and b2. (The symbols in this paragraph are merely symbolic representations of the corresponding data and are not considered specific limitations on the corresponding data.) Based on the method in this step, for each of the three data points to be labeled, the data point is combined with each semantic description text to obtain the following multiple data pairs: (A1, a1), (A1, a2), (A1, b1), (A1, b2), (A2, a1), (A2, a2), (A2, b1), (A2, b2), (A3, a1), (A3, a2), (A3, b1), (A3, b2). The above example is merely illustrative, and this disclosure does not limit its scope.
[0065] In addition, the data to be labeled includes videos to be labeled. Since video data includes multimodal information such as video frames, text data on each video frame, and audio data, when the data to be labeled is a video, the data pairs obtained based on the combination method in this step include multiple types of data pairs.
[0066] When the data to be labeled is a video, the data pair may include a first data pair, a second data pair, a third data pair, and a fourth data pair. Thus, before performing this step, multiple video frames of the video to be labeled, the first text data corresponding to each video frame, the second text data corresponding to the video to be labeled, and the audio corresponding to the video to be labeled can be obtained. The audio may include any audio in the video to be labeled (such as songs, background music, etc.), and the audio is converted into the third text data. The first text data may be, for example, text data such as bullet comments or subtitles on the video frame, and the second text data may be text data corresponding to the title and / or video description of the video to be labeled.
[0067] In one possible implementation, the acquired first and second text data can be converted into editable text using OCR (optical character recognition), and the acquired audio can be converted into third text data using ASR (Automatic Speech Recognition).
[0068] Thus, for the video to be labeled, during this step, for each video frame in the multiple video frames, the video frame is combined with each semantic description text to obtain multiple first data pairs; for each video frame, the first text data corresponding to the video frame is combined with each semantic description text to obtain multiple second data pairs; the second text data is combined with each semantic description text to obtain multiple third data pairs; and the third text data is combined with each semantic description text to obtain multiple fourth data pairs.
[0069] In step S1022, for each data pair, the first matching degree between the data to be labeled and the semantic description text in the data pair is calculated.
[0070] It is understandable that different types of data to be labeled will have different methods for calculating the matching degree. When the data to be labeled is a video, the first matching degree can include image-text matching degree and text matching degree. The image-text matching degree refers to the matching degree between the video frame in each first data pair and the corresponding semantic description text. The text matching degree refers to the matching degree between the text data and the corresponding semantic description text in each of the second, third, and fourth data pairs.
[0071] In this step, if the data to be labeled is a video, for each first data pair, the image-text matching degree between the video frame and the semantic description text in the first data pair can be determined by a pre-trained image-text matching model; for each of the second, third, and fourth data pairs, the text matching degree between the target text data and the semantic description text in the data pair can be determined by a pre-trained text similarity detection model, wherein the target text data refers to the text data in any one of the second, third, and fourth data pairs.
[0072] The image-text matching model can include, for example, any one of CLIP (Contrastive Language-Image Pre-Training) or ALIGN (A Large-scale Image and Noisy-text Embedding). The text similarity detection model can include, for example, the SBERT (Sentence-BERT) model.
[0073] Here, the specific implementation of calculating the image-text matching degree between the video frame and the semantic description text in the first data pair based on the image-text matching model, and the specific calculation method of calculating the text matching degree between the target text data and the semantic description text based on the text similarity detection model, can be found in the descriptions in related technologies, and will not be repeated here.
[0074] In step S1023, the target matching degree threshold is determined based on the data to be labeled and the preset labeling rate.
[0075] The preset labeling rate is the expected proportion of data of a specified category in the target data, and the target matching threshold is the matching threshold corresponding to the specified category. Therefore, when the specified category is different, the preset labeling rate and the target matching threshold need to be determined according to the corresponding specified category. That is, different labeling categories correspond to different preset labeling rates, and different labeling categories correspond to different target matching thresholds.
[0076] In specific application scenarios, the preset labeling rate can be determined based on the actual number of labeling personnel and the amount of data to be labeled. For example, suppose the current labeling task is to label 5,000 positive samples per day, and there are currently 5 labeling personnel, each of whom can label 2,000 samples per day. The original positive sample labeling rate is 5%. If no data filtering is performed, the original amount of data to be labeled is 5,000 / 0.05 = 100,000 samples. However, based on the current labeling personnel, a maximum of 5 * 2,000 = 10,000 samples can be labeled per day. Under the existing labeling personnel, in order to complete the labeling task of labeling 5,000 positive samples per day, the positive sample labeling rate needs to be increased to 50% so that 5,000 samples will be labeled as positive samples after labeling 10,000 samples. Therefore, the preset positive sample labeling rate can be set to 50%. This is just an example, and this disclosure does not limit it.
[0077] After determining the preset labeling rate corresponding to the specified category, the target matching degree threshold can be determined in the following way: randomly sample the data to be labeled to obtain a preset number of target sample data; for each target sample data, combine the target sample data with each semantic description text to obtain multiple data sample pairs; for each data sample pair, calculate the second matching degree between the target sample data and the semantic description text in the data sample pair; sort the data sample pairs according to the second matching degree, and determine the target matching degree threshold according to the sorting result and the preset labeling rate.
[0078] For example, 100 (i.e., a preset number) data points can be randomly sampled from 10,000 original data points to be labeled as the target sample data. Assuming that the preset labeling rate corresponding to the required positive samples is 20%, then at least 20 positive sample data points should be labeled from the 100 target sample data points. In this example, for each target sample data point, the target sample data point can be combined with each semantic description text in the target text description set to obtain multiple data sample pairs (the specific combination method can be referred to the relevant description in step 1021). The target text description set refers to the text description set corresponding to the specified category with the preset labeling rate. Then, based on the pre-set similarity matching model, the second matching degree between the target sample data point and the semantic description text in each data sample pair is calculated. After sorting the target sample data points from largest to smallest according to the second matching degree, the second matching degree corresponding to the 20th target sample data point can be selected as the target matching degree threshold. The above example is only for illustration, and this disclosure does not limit it.
[0079] In step S1024, data pairs with a first matching degree greater than or equal to the target matching degree threshold are taken as target data pairs.
[0080] For example, when the data to be labeled is a video, the calculated first matching degree includes image-text matching degree and text matching degree. Therefore, the target matching degree threshold may include the first matching degree threshold corresponding to the image-text matching degree and the second matching degree threshold corresponding to the text matching degree.
[0081] In this step, the first data pair with a text-image matching degree greater than or equal to the first matching degree threshold can be used as the target data pair, and the second, third, and fourth data pairs with a text matching degree greater than or equal to the second matching degree threshold can be used as the target data pairs.
[0082] In step S1025, the data to be labeled in each target data pair is taken as the target data.
[0083] Based on the above method, since the amount of target data after filtering is smaller than the amount of data to be labeled, class labeling of target data with a smaller amount of data can reduce the workload of labeling and reduce the manpower required for data labeling, thereby saving manpower costs for data labeling. In addition, data filtering based on the preset labeling rate of the specified category can also ensure that the data of the specified category has a certain proportion in the target data, thereby ensuring the quality of data labeling.
[0084] Figure 4 It is based on Figure 1 The illustrated embodiment presents a flowchart of a data annotation method, as shown in the figure. Figure 4As shown, the method also includes the following steps:
[0085] In step S104, for each target data pair, the recommendation category of the target data in the target data pair is determined based on the text description set corresponding to the target data pair.
[0086] Here, the text description set corresponding to the target data pair refers to the text description set corresponding to the semantic description text in the target data pair. Therefore, in this step, the recommendation category of the target data in the target data pair can be determined based on the semantic description text in the target data pair for each target data pair.
[0087] For example, suppose the labeled data to be labeled includes two categories, category 1 and category 2. Category 1 corresponds to text description set a, and category 2 corresponds to text description set b. Text description set a includes two semantic description texts, a1 and a2, and text description set b includes two semantic description texts, b1 and b2. Suppose the current target data pair is (A2, b2), where A2 is one of the target data and b2 is the semantic description text combined with the target data A2. Since the text description set containing b2 is text description set b, and the labeled category corresponding to text description set b is category 2, the recommended category corresponding to the target data A2 in the target data pair (A2, b2) is category 2. The above example is only for illustration and this disclosure does not limit it.
[0088] In step S105, the recommendation category corresponding to each target data is stored.
[0089] Thus, during the execution of step S103, each target data can be labeled according to the recommended category. In the process of labeling each target data according to the recommended category, the terminal can automatically label the target data as the recommended category, or it can recommend the recommended category to the labeler so that the labeler can use the recommended category as a reference for data labeling, thereby improving the efficiency of data labeling.
[0090] It should be noted that, when the data to be labeled is a video, for the target data pair determined from the first data pair, the video frame, the identification information of the video to be labeled, and the recommended category corresponding to the video frame can be pushed to the video frame labeling queue; for the target data pair determined from the second and third data pairs, the text data and their respective recommended categories can be pushed to the text labeling queue; since the third text data in the third data pair is converted from audio data, for the target data pair determined from the third data pair, the third text data and its corresponding recommended category can be pushed to the audio labeling queue. In this way, when labeling video data, the information of the three modalities of image, text, and audio can be distinguished and divided into different labeling queues for labeling, improving the efficiency of data labeling while ensuring the quality of data labeling.
[0091] Figure 5 This is a block diagram illustrating a data annotation apparatus according to an exemplary embodiment, such as... Figure 5 As shown, the device includes:
[0092] The acquisition module 501 is used to acquire multiple text description sets, different text description sets correspond to different annotation categories, and the text description set includes one or more semantic description texts, and one or more semantic description texts in the same text description set correspond to the same annotation category;
[0093] Data filtering module 502 is used to filter target data from the data to be labeled based on multiple sets of text descriptions and a preset labeling rate, wherein the preset labeling rate is the expected proportion of data of a specified category in the target data, and the labeling category includes the specified category;
[0094] The annotation module 503 is used to perform category annotation on the target data.
[0095] Optionally, Figure 6 It is based on Figure 5 The illustrated embodiment shows a block diagram of a data annotation apparatus, such as Figure 6 As shown, the device also includes:
[0096] The text description set generation module 504 is used to pre-set the text description set in the following ways:
[0097] Obtain a preset description file corresponding to the data to be labeled. The preset description file is used to record the keyword information corresponding to different labeling categories of the data to be labeled. For each labeling category, generate one or more semantic description texts corresponding to the labeling category based on the keyword information and preset description template, and use the one or more semantic description texts as the text description set.
[0098] Optionally, the data filtering module 502 is configured to, for each piece of data to be labeled, combine the data to be labeled with each semantic description text in the plurality of text description sets to obtain a plurality of data pairs; for each data pair, calculate a first matching degree between the data to be labeled and the semantic description text in the data pair; determine a target matching degree threshold based on the data to be labeled and the preset labeling rate; take data pairs with the first matching degree greater than or equal to the target matching degree threshold as target data pairs; and take the data to be labeled in each target data pair as the target data.
[0099] Optionally, the data to be labeled includes a video to be labeled, and the data pairs include a first data pair, a second data pair, a third data pair, and a fourth data pair. The data filtering module 502 is used to obtain multiple video frames of the video to be labeled, first text data corresponding to each video frame, second text data corresponding to the video to be labeled, and audio corresponding to the video to be labeled; convert the audio into third text data; for each of the multiple video frames, combine the video frame with each semantic description text to obtain multiple first data pairs; for each video frame, combine the first text data corresponding to the video frame with each semantic description text to obtain multiple second data pairs; combine the second text data with each semantic description text to obtain multiple third data pairs; and combine the third text data with each semantic description text to obtain multiple fourth data pairs.
[0100] Optionally, the first matching degree includes image-text matching degree and text matching degree. The data filtering module 502 is used to determine the image-text matching degree between the video frame and the semantic description text in each first data pair using a pre-trained image-text matching model; and to determine the text matching degree between the target text data and the semantic description text in each of the second, third, and fourth data pairs using a pre-trained text similarity detection model.
[0101] Optionally, the data filtering module 502 is used to randomly sample the data to be labeled to obtain a preset number of target sample data; for each target sample data, combine the target sample data with each semantic description text to obtain multiple data sample pairs; for each data sample pair, calculate the second matching degree between the target sample data and the semantic description text in the data sample pair; sort the data sample pairs according to the second matching degree, and determine the target matching degree threshold according to the sorting result and the preset labeling rate.
[0102] Optionally, Figure 7 It is based on Figure 5 The illustrated embodiment shows a block diagram of a data annotation apparatus, such as Figure 7 As shown, the device also includes:
[0103] The category recommendation module 505 is used to determine the recommended category of the target data in each target data pair based on the text description set corresponding to the target data pair; and to store the recommended category corresponding to each target data pair.
[0104] The data annotation module 503 is used to annotate each of the target data according to the recommended category.
[0105] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0106] Using the above-mentioned device, since the amount of data in the target data after filtering is smaller than the amount of data to be labeled, the workload of labeling can be reduced and the manpower required for data labeling can be reduced by labeling the target data with a smaller amount of data. This can save manpower costs for data labeling. In addition, data filtering based on the preset labeling rate of the specified category can also ensure that the data of the specified category has a certain proportion in the target data, thereby ensuring the quality of data labeling.
[0107] The following is for reference. Figure 8 This illustration shows a structural schematic of an electronic device 800 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0108] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0109] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0110] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.
[0111] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0112] In some implementations, the terminal can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0113] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0114] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the electronic device to: acquire multiple sets of text descriptions, different sets of text descriptions corresponding to different annotation categories, each set of text descriptions including one or more semantic description texts, and one or more semantic description texts in the same set of text descriptions corresponding to the same annotation category; filter target data from the data to be annotated according to the multiple sets of text descriptions and a preset annotation rate, the preset annotation rate being the expected proportion of data of a specified category in the target data, the annotation category including the specified category; and annotate the target data by category.
[0115] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0117] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a module does not necessarily limit the module itself; for example, an acquisition module can also be described as "a module for acquiring a set of text descriptions".
[0118] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0119] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0120] According to one or more embodiments of this disclosure, Example 1 provides a data annotation method, including:
[0121] Multiple text description sets are obtained, with different text description sets corresponding to different annotation categories. Each text description set includes one or more semantic description texts, and one or more semantic description texts in the same text description set correspond to the same annotation category.
[0122] Target data is selected from the data to be labeled based on multiple sets of text descriptions and a preset labeling rate, wherein the preset labeling rate is the expected proportion of data of a specified category in the target data, and the labeling category includes the specified category;
[0123] The target data is categorized.
[0124] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the text description set is pre-set in the following manner:
[0125] Obtain a preset description file corresponding to the data to be labeled, the preset description file being used to record keyword information corresponding to different labeling categories of the data to be labeled;
[0126] For each of the labeled categories, one or more semantic description texts corresponding to the labeled category are generated based on the keyword information and preset description template, and the one or more semantic description texts are used as the text description set.
[0127] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 1, wherein filtering target data from the data to be labeled based on a plurality of text description sets and a preset labeling rate includes:
[0128] For each of the data to be labeled, the data to be labeled is combined with each of the semantic description texts in the multiple sets of text descriptions to obtain multiple data pairs;
[0129] For each data pair, calculate the first matching degree between the data to be labeled in the data pair and the semantic description text;
[0130] The target matching degree threshold is determined based on the data to be labeled and the preset labeling rate;
[0131] Data pairs whose first matching degree is greater than or equal to the target matching degree threshold are taken as target data pairs;
[0132] The data to be labeled in each of the target data pairs is taken as the target data.
[0133] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 3, wherein the data to be labeled includes a video to be labeled, and the data pairs include a first data pair, a second data pair, a third data pair, and a fourth data pair. Before combining each piece of data to be labeled with each semantic description text in a plurality of text description sets to obtain a plurality of data pairs, the method further includes:
[0134] Obtain multiple video frames of the video to be labeled, first text data corresponding to each video frame, second text data corresponding to the video to be labeled, and audio corresponding to the video to be labeled;
[0135] Convert the audio into third-party text data;
[0136] For each piece of data to be labeled, the data to be labeled is combined with each semantic description text in the plurality of text description sets to obtain a plurality of data pairs, including:
[0137] For each of the plurality of video frames, the video frame is combined with each of the semantic description texts to obtain a plurality of first data pairs;
[0138] For each video frame, the first text data corresponding to the video frame is combined with each semantic description text to obtain multiple second data pairs;
[0139] The second text data is combined with each of the semantic description texts to obtain multiple third data pairs;
[0140] The third text data is combined with each of the semantic description texts to obtain multiple fourth data pairs.
[0141] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 4, wherein the first matching degree includes image-text matching degree and text matching degree, and the step of calculating the first matching degree between the data to be labeled and the semantic description text in each data pair includes:
[0142] For each of the first data pairs, the image-text matching degree between the video frames and the semantic description text in the first data pair is determined by a pre-trained image-text matching model;
[0143] For each of the second, third, and fourth data pairs, the text matching degree between the target text data and the semantic description text in that data pair is determined by a pre-trained text similarity detection model.
[0144] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 3, wherein determining the target matching degree threshold based on the data to be labeled and the preset labeling rate includes:
[0145] Randomly sample the data to be labeled to obtain a preset number of target sample data;
[0146] For each target sample data, the target sample data is combined with each semantic description text to obtain multiple data sample pairs;
[0147] For each data sample pair, calculate the second matching degree between the target sample data and the semantic description text in that data sample pair;
[0148] The data sample pairs are sorted according to the second matching degree, and the target matching degree threshold is determined according to the sorting result and the preset labeling rate.
[0149] According to one or more embodiments of this disclosure, Example 7 provides a method of any one of Examples 3-6, the method further comprising:
[0150] For each target data pair, the recommendation category of the target data in the target data pair is determined based on the set of text descriptions corresponding to the target data pair;
[0151] Store the recommendation category corresponding to each target data item;
[0152] The category labeling of the target data includes:
[0153] Each of the target data points is labeled according to the recommended category.
[0154] According to one or more embodiments of this disclosure, Example 8 provides a data annotation apparatus, the apparatus comprising:
[0155] The acquisition module is used to acquire multiple text description sets. Different text description sets correspond to different annotation categories. Each text description set includes one or more semantic description texts. One or more semantic description texts in the same text description set correspond to the same annotation category.
[0156] The data filtering module is used to filter target data from the data to be labeled based on multiple sets of text descriptions and a preset labeling rate, wherein the preset labeling rate is the expected proportion of data of a specified category in the target data, and the labeling category includes the specified category;
[0157] The annotation module is used to label the target data by category.
[0158] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-7.
[0159] According to one or more embodiments of this disclosure, Example 10 provides an electronic device, including:
[0160] A storage device having a computer program stored thereon; a processing device for executing the computer program in the storage device to implement the steps of any of the methods described in Examples 1-7.
[0161] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0162] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0163] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A data annotation method, characterized in that, The method includes: Multiple text description sets are obtained, with different text description sets corresponding to different annotation categories. Each text description set includes one or more semantic description texts, and one or more semantic description texts in the same text description set correspond to the same annotation category. Target data is selected from the data to be labeled based on multiple sets of text descriptions and a preset labeling rate, wherein the preset labeling rate is the expected proportion of data of a specified category in the target data, and the labeling category includes the specified category; The target data is categorized. The step of filtering target data from the data to be labeled based on multiple sets of text descriptions and a preset labeling rate includes: For each of the data to be labeled, the data to be labeled is combined with each of the semantic description texts in the multiple sets of text descriptions to obtain multiple data pairs; For each data pair, calculate the first matching degree between the data to be labeled in the data pair and the semantic description text; The target matching degree threshold is determined based on the data to be labeled and the preset labeling rate; Data pairs whose first matching degree is greater than or equal to the target matching degree threshold are taken as target data pairs; The data to be labeled in each of the target data pairs is taken as the target data. The step of determining the target matching threshold based on the data to be labeled and the preset labeling rate includes: Randomly sample the data to be labeled to obtain a preset number of target sample data; For each target sample data, the target sample data is combined with each semantic description text to obtain multiple data sample pairs; For each data sample pair, calculate the second matching degree between the target sample data and the semantic description text in that data sample pair; The data sample pairs are sorted according to the second matching degree, and the target matching degree threshold is determined according to the sorting result and the preset labeling rate.
2. The method according to claim 1, characterized in that, The set of text descriptions is pre-set in the following manner: Obtain a preset description file corresponding to the data to be labeled, the preset description file being used to record keyword information corresponding to different labeling categories of the data to be labeled; For each of the labeled categories, one or more semantic description texts corresponding to the labeled category are generated based on the keyword information and preset description template, and the one or more semantic description texts are used as the text description set.
3. The method according to claim 1, characterized in that, The data to be labeled includes videos to be labeled, and the data pairs include a first data pair, a second data pair, a third data pair, and a fourth data pair. Before combining each piece of data to be labeled with each semantic description text in the plurality of text description sets to obtain multiple data pairs, the method further includes: Obtain multiple video frames of the video to be labeled, first text data corresponding to each video frame, second text data corresponding to the video to be labeled, and audio corresponding to the video to be labeled; Convert the audio into third-party text data; For each piece of data to be labeled, the data to be labeled is combined with each semantic description text in the plurality of text description sets to obtain a plurality of data pairs, including: For each of the plurality of video frames, the video frame is combined with each of the semantic description texts to obtain a plurality of first data pairs; For each video frame, the first text data corresponding to the video frame is combined with each semantic description text to obtain multiple second data pairs; The second text data is combined with each of the semantic description texts to obtain multiple third data pairs; The third text data is combined with each of the semantic description texts to obtain multiple fourth data pairs.
4. The method according to claim 3, characterized in that, The first matching degree includes image-text matching degree and text matching degree. For each data pair, calculating the first matching degree between the data to be labeled and the semantic description text in that data pair includes: For each of the first data pairs, the image-text matching degree between the video frames and the semantic description text in the first data pair is determined by a pre-trained image-text matching model; For each of the second, third, and fourth data pairs, the text matching degree between the target text data and the semantic description text in that data pair is determined by a pre-trained text similarity detection model.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: For each target data pair, the recommendation category of the target data in the target data pair is determined based on the text description set corresponding to the target data pair; Store the recommendation category corresponding to each target data item; The category labeling of the target data includes: Each of the target data points is labeled according to the recommended category.
6. A data annotation device, characterized in that, The device includes: The acquisition module is configured to acquire multiple text description sets, with different text description sets corresponding to different annotation categories. Each text description set includes one or more semantic description texts, and one or more semantic description texts in the same text description set correspond to the same annotation category. The data filtering module is configured to: filter target data from the data to be labeled based on multiple sets of text descriptions and a preset labeling rate, wherein the preset labeling rate is the expected proportion of data of a specified category in the target data, and the labeling category includes the specified category; The annotation module is configured to: perform category annotation on the target data. The data filtering module is further configured as follows: For each of the data to be labeled, the data to be labeled is combined with each of the semantic description texts in the multiple sets of text descriptions to obtain multiple data pairs; For each data pair, calculate the first matching degree between the data to be labeled in the data pair and the semantic description text; Randomly sample the data to be labeled to obtain a preset number of target sample data; For each target sample data, the target sample data is combined with each semantic description text to obtain multiple data sample pairs; For each data sample pair, calculate the second matching degree between the target sample data and the semantic description text in that data sample pair; The data sample pairs are sorted according to the second matching degree, and the target matching degree threshold is determined according to the sorting result and the preset labeling rate. Data pairs whose first matching degree is greater than or equal to the target matching degree threshold are taken as target data pairs; The data to be labeled in each of the target data pairs is taken as the target data.
7. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processing device, it implements the steps of the method described in any one of claims 1-5.
8. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Semantic model tuning method and system
CN110347786A
Sample screening method and device and electronic equipment
CN114003724A
Content tag generation method and device, electronic equipment and storage medium
CN114021577A
Image labeling method and device, terminal equipment and storage medium
CN114169381A