Data labeling method, system and apparatus
By decoupling the storage methods of labeled data and attributes, multiple annotations can be performed on the same image, solving the problem of high coupling of labeled information and improving the training efficiency of deep learning models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU HIKROBOT TECH CO LTD
- Filing Date
- 2023-07-19
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, the high coupling between annotation information and the annotated images limits the training effect of deep learning models and makes it difficult to quickly select efficient versions of annotated data.
By completely decoupling the annotation data and annotation attributes and storing them separately in the visual dictionary platform, multiple annotations of the same image can be achieved. The earlier annotation data versions can be quickly filtered out using attributes such as structured annotation labels and annotation batches.
It improves the training efficiency of deep learning models, enabling the rapid selection of early labeled data versions when the later training results are poor, thereby improving the model training effect.
Smart Images

Figure CN116958970B_ABST
Abstract
Description
Technical Field
[0001] This application relates to image processing technology, and in particular to data annotation methods, systems and apparatus. Background Technology
[0002] In artificial intelligence applications, providing high-quality labeled data services for deep learning algorithms has become a crucial condition for successful AI applications. Currently, the coupling between labeled information and labeled images is very high; multiple versions of labeled data cannot exist for the same image, which limits the training of deep learning models. For example, during model training, if it is found that the training effect of a deep learning model using later labeled versions of a certain batch of training images (label information and labeled images) is not as good as the training effect using earlier labeled versions of the same batch of training images (label information and labeled images), the high coupling between the label information and labeled images makes it difficult to quickly select the earlier labeled versions, thus affecting the training of the deep learning model. Summary of the Invention
[0003] This application provides a data annotation method, system, and apparatus to achieve complete decoupling between images and annotation data, thereby improving the training efficiency of deep learning models.
[0004] This application provides a data annotation method, which is applied to a first client and includes:
[0005] Based on the pushed annotation task, an annotation pipeline is generated; the annotation pipeline is used to annotate the data in the dataset to be annotated; the annotation task is used to instruct the annotation of the data in the dataset searched by the second client on the visual dictionary platform; the annotation pipeline includes at least: annotation information; the annotation information includes at least: the object designated to perform the annotation task, the data in the dataset that the object is responsible for annotating, and the annotation features used by the object when annotating the data it is responsible for annotating;
[0006] The system obtains labeled data and corresponding labeled attributes obtained by labeling data in the dataset based on the labeling pipeline. The obtained labeled data and corresponding labeled attributes are stored separately in the visual dictionary platform. The labeled attributes are used to search for labeled data so as to automatically train the model using the searched labeled data. The labeled attributes corresponding to any labeled data include at least: structured label, label batch, and storage location identifier of the labeled data. The structured label is determined based on the labeling features used to label the data, and the same labeling pipeline corresponds to the same label batch.
[0007] A data annotation method, applied on a server-side platform, wherein the server-side platform deploys a visual dictionary platform, the method comprising:
[0008] The system receives raw data uploaded by any client; the raw data uploaded by any client has corresponding data attributes, which include at least: upload batch and / or original tag; the data uploaded by any client has corresponding original tag and upload batch, which is allocated by the visual dictionary platform based on the client's request before the client uploads the raw data; the raw data uploaded by any client and the data attributes of the raw data are stored separately on the visual dictionary platform.
[0009] The raw data uploaded by any client and the corresponding data attributes are stored separately on the visual dictionary platform;
[0010] Based on the search request from the second client, a dataset that meets the search request is found from the visual dictionary platform, so that the second client saves the dataset to its local favorites; the dataset includes: at least one original data uploaded by a client, and / or labeled data that has been labeled at least once; the search request includes at least specified data attributes or specified label attributes;
[0011] When the second client determines that the first client needs to annotate the dataset, it pushes an annotation task to the first client, enabling the first client to generate an annotation pipeline based on the pushed annotation task. The first client then obtains annotated data obtained by annotating the data in the dataset using the annotation pipeline, along with the corresponding annotation attributes. The obtained annotated data and its corresponding annotation attributes are stored separately on the visual dictionary platform. The annotation pipeline is used to annotate the data in the dataset. The annotation task is used to instruct the annotation of the data in the dataset. The annotation pipeline includes at least: annotation information; the annotation information includes at least: the object designated to perform the annotation task, the data in the dataset that the object is responsible for annotating, and the annotation features used by the object when annotating the data; the annotation attributes are used to search for annotated data to automatically train the model using the searched annotated data; the annotation attributes corresponding to any annotated data include at least: structured annotation labels, annotation batch, and the storage location identifier of the annotated data; the structured annotation labels are determined based on the annotation features used to annotate the data, and the same annotation pipeline corresponds to the same annotation batch.
[0012] A data annotation system, comprising: at least one client and one server;
[0013] Either of the at least one client is used to perform the steps in the first method described above;
[0014] The server is used to execute the steps in the second method above.
[0015] A data annotation device, applied to a first client, includes:
[0016] A pipeline task unit is used to generate a labeling pipeline based on a pushed labeling task; the labeling pipeline is used to label data in the dataset to be labeled; the labeling task is used to instruct the labeling of data in the dataset searched by the second client on the visual dictionary platform; the labeling pipeline includes at least: labeling information; the labeling information includes at least: the object designated to perform the labeling task, the data in the dataset that the object is responsible for labeling, and the labeling features used by the object when labeling the data it is responsible for labeling;
[0017] The annotation processing unit is used to obtain annotated data and corresponding annotation attributes obtained by annotating the data in the dataset based on the annotation pipeline, and to store the obtained annotated data and corresponding annotation attributes separately in the visual dictionary platform; the annotation attributes are used to search for annotated data so as to automatically train the model using the searched annotated data; the annotation attributes corresponding to any annotated data include at least: structured annotation label, annotation batch, and storage location identifier of the annotation data; the structured annotation label is determined based on the annotation features used to annotate the data, and the same annotation pipeline corresponds to the same annotation batch.
[0018] A data annotation device is applied to a server, wherein the server deploys a visual dictionary platform, and the device includes:
[0019] A receiving unit is configured to receive raw data uploaded by any client, and store the raw data uploaded by the client and its corresponding data attributes separately in the visual dictionary platform; the raw data uploaded by any client has corresponding data attributes, the data attributes including at least: upload batch and / or original tag; the data uploaded by any client has corresponding original tag and upload batch, the upload batch being allocated by the visual dictionary platform based on the client's request before the client uploads the raw data; the raw data uploaded by any client and its data attributes are stored separately in the visual dictionary platform.
[0020] The retrieval unit is configured to find a dataset that satisfies the retrieval request from the visual dictionary platform based on the retrieval request of the second client, so that the second client can store the dataset in its local favorites; the dataset includes: at least one original data uploaded by a client, and / or labeled data that has been labeled at least once; the retrieval request includes at least specified data attributes or specified label attributes;
[0021] A push unit is used to push a labeling task to the first client when the second client determines that the first client needs to label the dataset. This allows the first client to generate a labeling pipeline based on the pushed task and obtain labeled data and corresponding labeling attributes obtained by labeling the data in the dataset using the pipeline. The obtained labeled data and corresponding labeling attributes are stored separately in the visual dictionary platform. The labeling pipeline is used to label the data in the dataset. The labeling task is used to instruct the labeling of the data in the dataset. The labeling pipeline includes at least: labeling information; the labeling information includes at least: the object designated to perform the labeling task, the data in the dataset that the object is responsible for labeling, and the labeling features used by the object when labeling the data; the labeling attributes are used to search for labeled data to automatically train the model using the searched labeled data; the labeling attributes corresponding to any labeled data include at least: structured label, labeling batch, and storage location identifier of the labeled data; the structured label is determined based on the labeling features used to label the data, and the same labeling pipeline corresponds to the same labeling batch.
[0022] An electronic device comprising: a processor and a machine-readable storage medium;
[0023] The machine-readable storage medium stores machine-executable instructions that can be executed by the processor;
[0024] The processor is used to execute machine-executable instructions to implement the steps in any of the methods described above.
[0025] As can be seen from the above technical solutions, in this embodiment, the annotation data, such as the annotated image, and the corresponding annotation attributes, such as structured annotation labels and annotation batches, are completely decoupled during storage. The two are stored separately without any mandatory binding relationship. This decouples the existing annotation information and the annotated image, allowing the same image to be annotated multiple times and for the same image to have multiple versions of annotation data. When the deep learning model training effect achieved by the later annotated version (annotation information and annotated image) of a certain batch of training images is not as good as the training effect achieved by the earlier annotated version (annotation information and annotated image) of the same batch of training images, the earlier annotated version of the annotation data can be quickly selected based on the above annotation attributes, such as structured annotation labels and annotation batches, thereby improving the training efficiency of the deep learning model. Attached Figure Description
[0026] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0027] Figure 1 A flowchart illustrating the method provided in this application embodiment;
[0028] Figure 2a This is a schematic diagram illustrating the separate storage of raw data and data attributes provided in the embodiments of this application;
[0029] Figure 2b This is a schematic diagram illustrating the separate storage of annotation data and annotation attributes provided in the embodiments of this application;
[0030] Figure 3 This is a retrieval diagram provided for an embodiment of this application;
[0031] Figure 4 This is a schematic diagram of the labeled pipeline provided in the embodiments of this application;
[0032] Figure 5 Another method flowchart provided for embodiments of this application;
[0033] Figure 6 This is a schematic diagram of the device structure provided in the embodiments of this application;
[0034] Figure 7 This is a schematic diagram of another device structure provided in an embodiment of this application;
[0035] Figure 8 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0036] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0037] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0038] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, and to make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0039] See Figure 1 , Figure 1 This is a flowchart of a method provided in an embodiment of this application. This method can be applied to any client. For ease of description, the client here can be referred to as the first client.
[0040] like Figure 1 As shown, the process may include the following steps:
[0041] Step 101: Generate a labeling pipeline based on the pushed labeling task; the labeling pipeline is used to label the data in the dataset to be labeled.
[0042] As an example, the aforementioned dataset consists of raw data, such as images, uploaded by at least one client and retrieved by the second client through a distributed search engine on a visual dictionary platform. Here, the second client refers to any client other than the first client.
[0043] In this embodiment, the raw data uploaded by any client (including the first client and the second client mentioned above), such as an image, has corresponding data attributes. These data attributes include at least: upload batch and / or original tag. The upload batch is assigned by the visual dictionary platform based on the client's request before the client uploads the raw data. The original tag is used to characterize the raw data, such as the industry, process, camera type of the camera that captured the raw data (such as the image), and data type (e.g., image) of the scene it displays.
[0044] In this embodiment, the raw data uploaded by any client and its data attributes are stored separately on the visual dictionary platform. That is, the raw data and its data attributes are decoupled and stored separately, without any binding relationship. Figure 2a An example of a separately stored structure is shown. For example... Figure 2a As shown, the original data, such as images, is stored in the distributed file storage server of the visual dictionary platform in a distributed file storage manner, while the data attributes are stored in the relational database of the visual dictionary platform.
[0045] based on Figure 2a Given the storage structure shown, in step 101 above, the second client can quickly search for the dataset in the distributed file storage server using a distributed search engine and based on specified data attributes such as a certain upload batch and / or a certain original tag.
[0046] As another embodiment, the data in the above dataset can also be non-original data searched by the second client on the visual dictionary platform through a distributed search engine, such as data that has been labeled at least once according to step 101 to step 102 below, such as images. See step 102 for details on how to search for data, which will not be elaborated here.
[0047] Optionally, in this embodiment, after retrieving the dataset, the second client adds the attribute set corresponding to the dataset to its personal favorites. Here, the attribute set refers to the collection of data attributes / standard attributes corresponding to each data point in the dataset.
[0048] In practical applications, if, based on requirements, it is determined that the first client will annotate the attribute set saved in the personal bookmarks of the second client, then as described in step 101, an annotation task can be pushed to the first client, so that the first client can generate an annotation pipeline based on the pushed annotation task. Optionally, the above-mentioned requirements may include, for example, the second client specifying that the first client should annotate the above-mentioned dataset, or other actual business requirements, etc., which are not specifically limited in this embodiment.
[0049] Alternatively, in this embodiment, the annotation task is used to instruct the second client to annotate the data in the dataset to be annotated found by the visual dictionary platform. This task can be pushed to the first client by the deployed annotation platform or by the second client; this embodiment is not specifically limited. As an example, the annotation task can carry the attribute set, so that the first client can download the dataset from the visual dictionary platform based on the attribute set.
[0050] As an example, the above-described annotation pipeline is used to annotate data in the dataset to be annotated. In a specific implementation, the annotation pipeline includes at least: annotation information. Here, the annotation information includes at least: the object designated to perform the annotation task, such as a user; the data in the dataset that the object is responsible for annotating; and the annotation features used by the object when annotating the data. These annotation features may include, for example, the shape used for annotation (such as a circle, triangle, etc.) and the label used for annotation, etc., etc., which are not specifically limited in this embodiment. Corresponding structured annotation labels can be generated through these annotation features. For example, if the annotation feature is a shape used for annotation, such as a circle, then the structured annotation label at least indicates that circle.
[0051] After generating the above annotation pipeline, annotations can be performed based on this pipeline. Then proceed to step 102.
[0052] Step 102: Obtain the labeled data and the labeled attributes corresponding to the labeled data obtained by labeling the data in the dataset based on the labeling pipeline, and store the obtained labeled data and the labeled attributes corresponding to the labeled data separately in the above-mentioned visual dictionary platform.
[0053] In this embodiment, the annotation attribute corresponding to any annotation data is used to search for that annotation data, so as to automatically train the model using the searched annotation data.
[0054] In a practical implementation, the annotation attributes corresponding to any annotation data may include at least: structured annotation labels and / or annotation batches.
[0055] In this embodiment, the structured annotation labels are determined based on the annotation features used to annotate the data, which will not be elaborated here.
[0056] In this embodiment, the same annotation pipeline corresponds to the same annotation batch. In other words, the annotation data obtained by annotating the data under the same annotation pipeline belongs to the same calibration batch.
[0057] In this embodiment, the annotation data and the corresponding annotation attributes are stored separately in the visual dictionary platform. For example, the annotation data is stored in the distributed file storage server of the visual dictionary platform in a distributed file storage manner, while the annotation attributes corresponding to the annotation data are stored in the relational database of the visual dictionary platform. Figure 2b An example of a storage structure is provided. Based on... Figure 2b The storage structure shown can be used by a distributed search engine to quickly search for the above-mentioned labeled data in the distributed file storage server based on the structured tags and / or labeling batches in the above-mentioned labeled attributes, so as to automatically train the model using the searched labeled data. Figure 3 The image search structure is illustrated using this example.
[0058] It should be noted that in this embodiment, if the data in the dataset to be labeled has been labeled at least once, the labeling applied in step 102 will not overwrite the previous labeling information. This allows the same data, such as an image, to have multiple versions of labeled data simultaneously. Furthermore, to facilitate differentiation between batches of labeled images, an additional backup of each labeled image can be made for each batch, enabling quick retrieval of images labeled in different batches later.
[0059] This concludes the process. Figure 1 The process is shown below.
[0060] pass Figure 1As can be seen from the process shown, in this embodiment, the annotation data, such as the annotated image, and the corresponding annotation attributes, such as structured annotation labels and annotation batches, are completely decoupled during storage. The two are stored separately without any mandatory binding relationship. This decouples the existing annotation information and the annotated image, allowing the same image to be annotated multiple times and for the same image to have multiple versions of annotation data. When the deep learning model training effect achieved by the later annotated version (annotation information and annotated image) of a certain batch of training images is not as good as the training effect achieved by the earlier annotated version (annotation information and annotated image) of the same batch of training images, the earlier annotated version of the annotation data can be quickly selected based on the above annotation attributes, such as structured annotation labels and annotation batches, thereby improving the training efficiency of the deep learning model.
[0061] In this embodiment, the above-mentioned annotation pipeline further includes: review information and quality inspection information. Here, the review information is used to indicate that the annotation information on the labeled data is reviewed, the quality inspection information is used to indicate that the annotation information on the labeled data is inspected, and the acceptance information is used to indicate that the annotation information on the labeled data is accepted. Figure 3 An example is shown illustrating the structure of a labeled pipeline.
[0062] Specifically, the verification method, quality inspection method, and acceptance method are not specifically limited in this embodiment; they can be set according to actual needs to ensure that the annotation information of any data conforms to the regulations. Of course, if, based on the above verification information, and / or quality inspection information, it is determined that the annotation information of a data is abnormal, such as the annotated graphic not meeting the requirements, then the annotation of that data will be retried to ensure that the final annotation information conforms to the regulations.
[0063] In addition, in this embodiment, the following can be further performed: Figure 4 The steps shown are as follows:
[0064] Step 401: Based on the target annotation attributes, search the visual dictionary platform for the first target annotation dataset corresponding to the first target annotation attributes using a distributed search engine.
[0065] Optionally, the first target annotation dataset includes: annotation data obtained by performing N annotations on the original data, such as images uploaded in a certain upload batch; where N is greater than 1.
[0066] Step 402: Train the model based on the first target labeled dataset to train the algorithm model. If the algorithm model does not meet the set requirements, proceed to step 403.
[0067] In this embodiment, if it is found that the model trained by the original data with labeled data obtained through N annotations is not as good as the model trained by the original data with labeled data obtained through M annotations, where M is a positive integer less than N, then step 403 can be executed.
[0068] Step 403: Based on the first target annotation dataset, search for the second target annotation dataset to train an algorithm model that meets the set requirements based on the second target annotation dataset.
[0069] The second target annotation dataset includes: annotated data obtained by performing M annotation operations on the original data; where M is a positive integer less than N. Specifically, in step 403, the annotation attributes of the second target annotation dataset can be used as keywords to search for the second target annotation dataset on the visual dictionary platform using a distributed search engine.
[0070] pass Figure 4 As shown in the process, once it is found that the training effect of the deep learning model achieved by using the later labeled data version (labeling information and labeled images) of a certain batch of training images is not as good as the training effect achieved by using the earlier labeled data version (labeling information and labeled images) of the same batch of training images, the earlier labeled data version can be quickly filtered out directly through the visual dictionary platform by means of the complete decoupling between the labeling information and the labeled images, thus improving the training efficiency of the deep learning model.
[0071] The methods provided in the embodiments of this application have been described above from the client's perspective. The methods provided in the embodiments of this application will now be described from the server's perspective:
[0072] See Figure 5 , Figure 5 Another method flowchart provided for an embodiment of this application. This method is applied to a server, where a visual dictionary platform is deployed. The method includes the following steps:
[0073] Step 501: Receive raw data uploaded by any client, and store the raw data uploaded by any client and the corresponding data attributes separately in the visual dictionary platform.
[0074] The raw data uploaded by any client has corresponding data attributes, which include at least: upload batch and / or original tag; the upload batch is allocated by the visual dictionary platform based on the client's request before the client uploads the raw data.
[0075] Step 502: Based on the retrieval request of the second client, find a dataset that satisfies the retrieval request from the visual dictionary platform, so that the second client stores the dataset in its local favorites; the dataset includes: at least one original data uploaded by a client, and / or labeled data that has been labeled at least once; the retrieval request includes at least specified data attributes or specified label attributes.
[0076] Step 503: When the second client determines that the first client needs to annotate the dataset, it pushes the annotation task to the first client.
[0077] When the first client receives the annotation task, it executes as follows: Figure 1 The process is shown below.
[0078] This concludes the process. Figure 5 The process is shown below.
[0079] pass Figure 5 As can be seen from the process shown, in this embodiment, the raw data uploaded by the client, such as images, and the corresponding data attributes, such as original labels and upload batches, are completely decoupled during storage. The two are stored separately without any mandatory binding relationship. This decouples the existing annotation information and the annotated images, allowing the same image to be annotated multiple times and for the same image to have multiple versions of annotation data. When the deep learning model training effect achieved by the annotated data version (annotation information and annotated images) of a certain batch of training images in a later period is not as good as the training effect achieved by the annotated data version (annotation information and annotated images) of the same batch of training images in an earlier period, the earlier annotated data version can be quickly selected based on the above annotation attributes, such as structured annotation labels and annotation batches, thereby improving the training efficiency of the deep learning model.
[0080] The methods provided in the embodiments of this application have been described above. The systems provided in the embodiments of this application are described below:
[0081] The data annotation system provided in this application includes: at least one client and one server;
[0082] Wherein, any one of the at least one clients is used to perform, such as Figure 1 The steps in the method shown; the server is used to execute, as follows Figure 5 The steps in the method shown.
[0083] As an example, this application also provides a data annotation device, which is applied to a first client, such as... Figure 6 As shown, the device may include:
[0084] A task unit is used to generate a labeling pipeline based on a pushed labeling task; the labeling pipeline is used to label data in the dataset to be labeled; the labeling task is used to instruct the labeling of data in the dataset searched by the second client on the visual dictionary platform; the labeling pipeline includes at least: labeling information; the labeling information includes at least: the object designated to perform the labeling task, the data in the dataset that the object is responsible for labeling, and the labeling features used by the object when labeling the data it is responsible for labeling;
[0085] The processing unit is used to obtain labeled data and corresponding labeled attributes obtained by labeling data in the dataset based on the labeling pipeline, and to store the obtained labeled data and corresponding labeled attributes separately in the visual dictionary platform; the labeled attributes are used to search for labeled data so as to automatically train the model using the searched labeled data; the labeled attributes corresponding to any labeled data include at least: structured label, label batch, and storage location identifier of the labeled data; the structured label is determined based on the labeling features used to label the data, and the same labeling pipeline corresponds to the same label batch.
[0086] Optionally, the labeling pipeline may further include: verification information, quality inspection information, and acceptance information;
[0087] The review information is used to indicate that the annotation information of the labeled data should be reviewed;
[0088] The quality inspection information is used to indicate whether the labeling information of the labeled data is subject to quality inspection.
[0089] The acceptance information is used to indicate the acceptance of the annotation information on the labeled data;
[0090] The method further includes: when it is determined, based on the review information, and / or the quality inspection information, and / or the acceptance information, that the annotation information of a data being annotated is abnormal, then the annotation of the data is retried.
[0091] Optionally, the dataset includes: raw data uploaded by at least one client, and / or labeled data obtained by performing annotation at least once;
[0092] The raw data uploaded by at least one client is retrieved by the second client from the visual dictionary platform using a distributed search engine based on specified data attributes; the visual dictionary platform stores the raw data uploaded by each client in a distributed file storage manner; the raw data uploaded by any client has corresponding data attributes, which include at least: upload batch and / or original tag; the data uploaded by any client has corresponding original tag and upload batch, which is allocated by the visual dictionary platform based on the client's request before the client uploads the raw data; the raw data uploaded by any client and the data attributes of the raw data are stored separately on the visual dictionary platform;
[0093] The labeled data is retrieved by the second client based on specified label attributes and searched on the visual dictionary platform using a distributed search engine.
[0094] Optionally, the processing unit further searches for a first target annotation dataset corresponding to the target annotation attributes on a visual dictionary platform using a distributed search engine, based on the target annotation attributes; the first target annotation dataset includes: annotation data obtained by performing N annotations on the original data; N is greater than 1; model training is performed based on the first target annotation dataset to train an algorithm model; when the algorithm model does not meet the set requirements, a second target annotation dataset is searched based on the first target annotation dataset to train an algorithm model that meets the set requirements based on the second target annotation dataset; the second target annotation dataset includes: annotation data obtained by performing M annotations on the original data; M is a positive integer less than N.
[0095] Optionally, the attribute set corresponding to the dataset is stored in the local favorites of the second client; the attribute set refers to the collection of data attributes or annotation attributes corresponding to each data in the dataset.
[0096] The task unit further receives the attribute set stored in the local favorites sent by the second client, so as to download the corresponding dataset based on the attribute set.
[0097] This concludes the process. Figure 6 Structural description of the device shown.
[0098] This application embodiment also provides another data annotation device, which is applied to a server, wherein the server deploys a visual dictionary platform, such as... Figure 7 As shown, the device includes:
[0099] A receiving unit is configured to receive raw data uploaded by any client, and store the raw data uploaded by the client and its corresponding data attributes separately in the visual dictionary platform; the raw data uploaded by any client has corresponding data attributes, the data attributes including at least: upload batch and / or original tag; the data uploaded by any client has corresponding original tag and upload batch, the upload batch being allocated by the visual dictionary platform based on the client's request before the client uploads the raw data; the raw data uploaded by any client and its data attributes are stored separately in the visual dictionary platform.
[0100] The retrieval unit is configured to find a dataset that satisfies the retrieval request from the visual dictionary platform based on the retrieval request of the second client, so that the second client can store the dataset in its local favorites; the dataset includes: at least one original data uploaded by a client, and / or labeled data that has been labeled at least once; the retrieval request includes at least specified data attributes or specified label attributes;
[0101] A push unit is used to push a labeling task to the first client when the second client determines that the first client needs to label the dataset. This allows the first client to generate a labeling pipeline based on the pushed task and obtain labeled data and corresponding labeling attributes obtained by labeling the data in the dataset using the pipeline. The obtained labeled data and corresponding labeling attributes are stored separately in the visual dictionary platform. The labeling pipeline is used to label the data in the dataset. The labeling task is used to instruct the labeling of the data in the dataset. The labeling pipeline includes at least: labeling information; the labeling information includes at least: the object designated to perform the labeling task, the data in the dataset that the object is responsible for labeling, and the labeling features used by the object when labeling the data; the labeling attributes are used to search for labeled data to automatically train the model using the searched labeled data; the labeling attributes corresponding to any labeled data include at least: structured label, labeling batch, and storage location identifier of the labeled data; the structured label is determined based on the labeling features used to label the data, and the same labeling pipeline corresponds to the same labeling batch.
[0102] This concludes the process. Figure 7 Structural description of the device shown.
[0103] Based on the same application concept as the method described above, embodiments of this application also provide an electronic device, such as... Figure 8As shown, the electronic device includes: a processor and a machine-readable storage medium; the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example.
[0104] Based on the same application concept as the above method, this application embodiment also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the method disclosed in the above examples of this application.
[0105] For example, the aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For instance, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0106] The systems, devices, modules, or units described in the above embodiments can be implemented by a computer or entity, or by a product with a certain function. A typical implementation device is a computer, which can be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.
[0107] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0108] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0109] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0110] Furthermore, these computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0111] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0112] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A data annotation method, characterized in that, This method is applied to the first client and includes: Based on the pushed annotation task, an annotation pipeline is generated; the annotation pipeline is used to annotate the data in the dataset to be annotated; the annotation task is used to instruct the annotation of the data in the dataset searched by the second client on the visual dictionary platform; the annotation pipeline includes at least: annotation information; the annotation information includes at least: the object designated to perform the annotation task, the data in the dataset that the object is responsible for annotating, and the annotation features used by the object when annotating the data it is responsible for annotating; The system obtains labeled data and corresponding labeling attributes obtained by labeling data in the dataset based on the labeling pipeline. This labeled data does not overwrite previously labeled information. The obtained labeled data and corresponding labeling attributes are stored separately in the visual dictionary platform. The labeled data is stored in a distributed file storage server within the visual dictionary platform using a distributed file storage method, and the corresponding labeling attributes are stored in a relational database within the visual dictionary platform. The labeling attributes are used to search for labeled data through a distributed search engine to automatically train the model using the searched labeled data. The labeling attributes corresponding to any labeled data include at least: structured label and / or label batch. The structured label is determined based on the labeling features used to label the data, and the same labeling pipeline corresponds to the same label batch.
2. The method according to claim 1, characterized in that, The labeling pipeline also includes: verification information, quality inspection information, and acceptance information; The review information is used to indicate that the annotation information of the labeled data should be reviewed; The quality inspection information is used to indicate whether the labeling information of the labeled data is subject to quality inspection. The acceptance information is used to indicate the acceptance of the annotation information on the labeled data; The method further includes: when it is determined, based on the review information, and / or the quality inspection information, and / or the acceptance information, that the annotation information of a data being annotated is abnormal, then the annotation of the data is retried.
3. The method according to claim 1, characterized in that, The dataset includes: raw data uploaded by at least one client, and / or labeled data obtained by performing annotation at least once; The raw data uploaded by at least one client is retrieved by the second client from the visual dictionary platform using a distributed search engine based on specified data attributes; the visual dictionary platform stores the raw data uploaded by each client in a distributed file storage manner; the raw data uploaded by any client has corresponding data attributes, which include at least: upload batch and / or original tag; the upload batch is allocated by the visual dictionary platform based on the client's request before the client uploads the raw data; the raw data uploaded by any client and the data attributes of the raw data are stored separately on the visual dictionary platform; The labeled data is retrieved by the second client based on specified label attributes and searched on the visual dictionary platform using a distributed search engine.
4. The method according to claim 1, characterized in that, The method further includes: Based on the target annotation attributes, a first target annotation dataset corresponding to the target annotation attributes is searched on a visual dictionary platform using a distributed search engine; the first target annotation dataset includes: annotated data obtained by performing N annotations on the original data; N is greater than 1; The algorithm model is trained based on the first target labeled dataset. When the algorithm model does not meet the set requirements, a second target annotation dataset is searched based on the first target annotation dataset to train an algorithm model that meets the set requirements based on the second target annotation dataset; the second target annotation dataset includes: annotation data obtained by performing annotation on the original data M times; M is a positive integer less than N.
5. The method according to claim 1, characterized in that, The attribute set corresponding to the dataset is stored in the local favorites of the second client. The attribute set refers to the collection of data attributes or annotation attributes corresponding to each data in the above dataset; The method further includes: The system receives the attribute set stored in the local favorites from the second client, and downloads the corresponding dataset based on the attribute set.
6. A data annotation method, characterized in that, This method is applied to a server-side application, where a visual dictionary platform is deployed. The method includes: The system receives raw data uploaded by any client and stores the raw data and its corresponding data attributes separately on the visual dictionary platform. The raw data uploaded by any client has corresponding data attributes, which include at least: upload batch and / or original tag. The upload batch is allocated by the visual dictionary platform based on the client's request before the client uploads the raw data. Based on the search request from the second client, a dataset that meets the search request is found from the visual dictionary platform, so that the second client saves the dataset to its local favorites; the dataset includes: at least one original data uploaded by a client, and / or labeled data that has been labeled at least once; the search request includes at least specified data attributes or specified label attributes; When the second client determines that the first client needs to annotate the dataset, it pushes an annotation task to the first client. This allows the first client to generate an annotation pipeline based on the pushed task and obtain annotated data from the dataset annotated using the pipeline, along with the corresponding annotation attributes. This annotated data does not overwrite any previously annotated data. The obtained annotated data and its corresponding annotation attributes are stored separately on the visual dictionary platform. The annotated data is stored in a distributed file storage server within the visual dictionary platform using a distributed file storage method, while the annotation attributes are stored in a relational database within the visual dictionary platform. The annotation pipeline is used to annotate the dataset... The data in the dataset is labeled; the labeling task is used to instruct the data in the dataset to be labeled; the labeling pipeline includes at least: labeling information; the labeling information includes at least: the object designated to perform the labeling task, the data in the dataset that the object is responsible for labeling, and the labeling features used by the object when labeling the data it is responsible for labeling; the labeling attributes are used to search for labeled data through a distributed search engine, so as to automatically train the model using the searched labeled data; the labeling attributes corresponding to any labeled data include at least: structured label, labeling batch, and storage location identifier of the labeled data; the structured label is determined based on the labeling features used to label the data, and the same labeling pipeline corresponds to the same labeling batch.
7. A data annotation system, characterized in that, The system includes: at least one client and one server; Any of the at least one client is used to perform the steps in the method as claimed in any one of claims 1 to 5; The server is used to perform the steps in the method as described in claim 6.
8. A data annotation device, characterized in that, The device is applied to the first client and includes: A task unit is used to generate a labeling pipeline based on a pushed labeling task; the labeling pipeline is used to label data in the dataset to be labeled; the labeling task is used to instruct the labeling of data in the dataset searched by the second client on the visual dictionary platform; the labeling pipeline includes at least: labeling information; the labeling information includes at least: the object designated to perform the labeling task, the data in the dataset that the object is responsible for labeling, and the labeling features used by the object when labeling the data it is responsible for labeling; The processing unit is configured to obtain labeled data obtained by labeling data in the dataset based on the labeling pipeline, and the labeling attributes corresponding to the labeled data, wherein the labeled data does not overwrite the labeling information previously labeled in the data; the obtained labeled data and the labeling attributes corresponding to the labeled data are stored separately in the visual dictionary platform; the labeled data is stored in a distributed file storage server in the visual dictionary platform in a distributed file storage manner, and the labeling attributes corresponding to the labeled data are stored in a relational database in the visual dictionary platform; the labeling attributes are used to search for labeled data through a distributed search engine, so as to automatically train the model using the searched labeled data; the labeling attributes corresponding to any labeled data include at least: structured label, labeling batch, and storage location identifier of the labeled data; the structured label is determined based on the labeling features used to label the data, and the same labeling pipeline corresponds to the same labeling batch.
9. The apparatus according to claim 8, characterized in that, The dataset includes: raw data uploaded by at least one client, and / or labeled data obtained by performing annotation at least once; The raw data uploaded by at least one client is retrieved by the second client from the visual dictionary platform using a distributed search engine based on specified data attributes; the visual dictionary platform stores the raw data uploaded by each client in a distributed file storage manner; the raw data uploaded by any client has corresponding data attributes, which include at least: upload batch and / or original tag; the data uploaded by any client has corresponding original tag and upload batch, which is allocated by the visual dictionary platform based on the client's request before the client uploads the raw data; the raw data uploaded by any client and the data attributes of the raw data are stored separately on the visual dictionary platform; The labeled data is retrieved by the second client based on specified label attributes and searched on the visual dictionary platform using a distributed search engine.
10. A data annotation device, characterized in that, This device is used on a server-side application, where a visual dictionary platform is deployed. The device includes: A receiving unit is configured to receive raw data uploaded by any client, and store the raw data uploaded by any client and its corresponding data attributes separately in the visual dictionary platform; the raw data uploaded by any client has corresponding data attributes, the data attributes including at least: upload batch and / or original tag; the upload batch is allocated by the visual dictionary platform based on the client's request before the client uploads the raw data; The retrieval unit is configured to find a dataset that satisfies the retrieval request from the visual dictionary platform based on the retrieval request of the second client, so that the second client can store the dataset in its local favorites; the dataset includes: at least one original data uploaded by a client, and / or labeled data that has been labeled at least once; the retrieval request includes at least specified data attributes or specified label attributes; A push unit is used to push a labeling task to the first client when the second client determines that the first client needs to label the dataset. This allows the first client to generate a labeling pipeline based on the pushed task and obtain labeled data from the dataset labeled using the pipeline, along with the corresponding labeling attributes. This labeled data does not overwrite any previously labeled information. The obtained labeled data and its corresponding labeling attributes are stored separately on the visual dictionary platform. The labeled data is stored in a distributed file storage server within the visual dictionary platform, and the corresponding labeling attributes are stored in a relational database within the platform. The labeling pipeline is used... The annotation pipeline is used to annotate the data in the dataset; the annotation task is used to instruct the data in the dataset to be annotated; the annotation pipeline includes at least: annotation information; the annotation information includes at least: an object designated to perform the annotation task, the data in the dataset that the object is responsible for annotating, and the annotation features used by the object when annotating the data it is responsible for annotating; the annotation attributes are used to search for annotated data through a distributed search engine, so as to automatically train the model using the searched annotated data; the annotation attributes corresponding to any annotated data include at least: structured annotation label, annotation batch, and storage location identifier of the annotated data; the structured annotation label is determined based on the annotation features used to annotate the data, and the same annotation pipeline corresponds to the same annotation batch.
11. An electronic device, characterized in that, The electronic device includes: a processor and a machine-readable storage medium; The machine-readable storage medium stores machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method steps of any one of claims 1-6.