Label data synchronization method and device, electronic equipment and storage medium
By receiving and processing data annotation tasks in the AI platform and using the application program interface for HTTP requests for data synchronization, the problem of high development and maintenance of Label Studio annotation data synchronization to the AI platform is solved, and efficient and accurate data synchronization is achieved, ensuring the accuracy and consistency of model training.
Patent Information
- Application Number
- CN202510712860.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The method of synchronizing data marked by Label Studio to the AI platform is difficult to develop and maintain, and the synchronization efficiency is inefficient and the data accuracy is low.
By receiving data annotation tasks submitted by users, calling the storage interface to store tasks in the storage database, allocating computing resources to the task, starting the built-in data annotation tool, creating a mount path for the storage database, setting read and write permissions and execution permissions, running the data annotation task, uploading the to-named data to be marked to the mount path for labeling, and uploading the annotation data to the mount path, and finally obtaining the annotation data from the mount path for model training, and using the HTTP requested application program interface for data synchronization.
It reduces the development and maintenance difficulty of Label Studio labeled data synchronization to the AI platform, improves synchronization efficiency and data accuracy, ensures data consistency and accuracy, reduces dependence on specific programming languages and libraries, and simplifies the maintenance process.
Smart Images

Figure CN120256169A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device, and storage medium for synchronizing labeled data. Background Art
[0002] With the development of artificial intelligence (AI) technology, the demand for high-quality labeled data is increasing. As an open-source data annotation tool, Label Studio supports the annotation of multi-modal data, and the data annotated by it can be synchronized to the AI platform for model training. In related technologies, the data annotated by Label Studio is synchronized to the AI platform through the Python Software Development Kit (SDK) of Label Studio and the Python SDK of the AI platform.
[0003] However, in related technologies, the method of synchronizing the data annotated by Label Studio to the AI platform is difficult to develop and maintain. Summary of the Invention
[0004] This application provides a method, device, electronic device, and storage medium for synchronizing labeled data, so as to at least solve the problem that the method of synchronizing the data annotated by Label Studio to the AI platform in related technologies is difficult to develop and maintain.
[0005] This application provides a method for synchronizing labeled data, including: Receiving a data annotation task submitted by a user, calling a storage interface, storing the data annotation task in a storage database, and allocating computing resources for the data annotation task, where the data annotation task is encapsulated through an application programming interface; Starting the built-in data annotation tool; Creating a mount path for the data annotation task in the storage database, and setting the read-write permission and execution permission of the mount path for the artificial intelligence platform and the data annotation tool; Running the data annotation task, uploading the data to be annotated of the data annotation task to the mount path, so that the data annotation tool obtains the data to be annotated from the mount path, annotates the data to be annotated, and uploads the annotated data to the mount path; Obtaining the annotated data from the mount path and performing model training based on the annotated data.
[0006] This application also provides a device for synchronizing labeled data, including: A receiving module, configured to receive a data annotation task submitted by a user, call a storage interface, store the data annotation task in a storage database, and allocate computing resources for the data annotation task, wherein the data annotation task is encapsulated through an application programming interface; A starting module, configured to start a built-in data annotation tool; A creating module, configured to create a mounting path for the data annotation task in the storage database, and set read-write permissions and execution permissions for the artificial intelligence platform and the data annotation tool on the mounting path; An operating module, configured to run the data annotation task, upload the data to be annotated in the data annotation task to the mounting path, so that the data annotation tool obtains the data to be annotated from the mounting path, annotates the data to be annotated, and uploads the annotated data to the mounting path; A obtaining module, configured to obtain the annotated data from the mounting path and perform model training based on the annotated data.
[0007] The present application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above-mentioned annotation data synchronization methods when executing the computer program.
[0008] The present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program implements the steps of any one of the above-mentioned annotation data synchronization methods when executed by a processor.
[0009] The present application further provides a computer program product, including a computer program, wherein the computer program implements the steps of any one of the above-mentioned annotation data synchronization methods when executed by a processor.
[0010] With this application, since the data annotation task submitted by the user is received, the storage interface is called to store the data annotation task in the storage database, and computing resources are allocated for the data annotation task, where the data annotation task is encapsulated through the application programming interface; the built-in data annotation tool is started; a mounting path is created for the data annotation task in the storage database, and the artificial intelligence platform and the data annotation tool are set with read-write permissions and execution permissions for the mounting path; the data annotation task is run, and the data to be annotated in the data annotation task is uploaded to the mounting path, so that the data annotation tool can obtain the data to be annotated from the mounting path, annotate the data to be annotated, and upload the annotated data to the mounting path; the annotated data is obtained from the mounting path, and model training is performed based on the annotated data. By encapsulating the data annotation task using the application programming interface and running the data annotation task, the synchronization of the annotated data can be achieved. Since the application programming interface only requires a tool that can send HyperText Transfer Protocol (HTTP) requests, the dependence on specific programming languages and libraries is reduced. Therefore, the technical problem of the high development and maintenance difficulty of the method for synchronizing the data annotated by Label Studio to the AI platform in the related technology can be solved, and the technical effect of reducing the development and maintenance difficulty of synchronizing the data annotated by Label Studio to the AI platform is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 FIG. is a schematic structural diagram of a system for synchronizing annotated data provided by an embodiment of the present application; Figure 2 FIG. is a schematic flowchart of a method for synchronizing annotated data provided by an embodiment of the present application; Figure 3 FIG. is a flowchart for creating a data annotation task provided by an embodiment of the present application; Figure 4 FIG. is a schematic flowchart of another method for synchronizing annotated data provided by an embodiment of the present application; Figure 5 FIG. is an interaction diagram for synchronizing the data annotated by the data annotation tool to the AI platform provided by an embodiment of the present application; Figure 6 FIG. is a schematic structural diagram of a device for synchronizing annotated data provided by an embodiment of the present application; Figure 7Schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0015] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0016] With the wide application of artificial intelligence technology in multiple fields such as autonomous driving, medical imaging, and industrial quality inspection, the demand for high-quality labeled data is growing at an exponential rate. When enterprises and research institutions develop AI projects, they increasingly attach importance to the quality and efficiency of data labeling, as well as the seamless connection between the labeled data and subsequent model training, evaluation, and other links.
[0017] As a comprehensive tool for artificial intelligence development, the AI platform provides a one-stop working environment from data management to model training, model evaluation, and model deployment. In this process, high-quality labeled data is the basis for model training and directly determines the success or failure of AI projects.
[0018] Label Studio is an open-source data annotation tool that supports data annotation of various types such as text, images, and audio. It provides rich annotation functions and a friendly user interface, enabling users to complete data annotation tasks quickly and efficiently. However, the annotation data generated by Label Studio needs to be combined with the AI platform for model training. Therefore, synchronizing the annotation data generated by Label Studio to the AI platform for model training is the basis for ensuring the accuracy and efficiency of model training.
[0019] In related technologies, when synchronizing Label Studio annotated data to the AI platform, the LabelStudio Python SDK is generally used. As a powerful open source data annotation tool, Label Studio provides developers with the Label Studio Python SDK, with which developers can easily establish a connection between the AI platform and the Label Studio server. Specifically, by configuring the Uniform Resource Locator (URL) and Application Programming Interface (API) key and other necessary information of the Label Studio server in the Python script or application, the connection process is completed to obtain the annotated data.
[0020] When the AI platform is successfully connected to the Label Studio server, developers can obtain various types of labeled data objects by calling the rich related methods in the Label StudioPython SDK. These labeled data objects contain detailed information about the task, such as the task identifier (ID), annotator information, and labeling results. Based on the obtained labeled data objects, developers can perform a series of corresponding operations and processing, such as filtering the labeling results of a specific task, converting the format of the labeling results according to the requirements of the AI platform, or extracting key information.
[0021] However, simply acquiring and processing the labeled data of Label Studio is not enough to synchronize the data to the AI platform. It is also necessary to combine the Python SDK or third-party library of the AI platform. Different AI platforms provide their own Python SDKs to facilitate developers' integration. Developers can use the Python SDKs or third-party libraries of these AI platforms to upload the data obtained from Label Studio to the target AI platform. It is understandable that during the upload process, the data format and interface specifications specified by the target AI platform need to be followed.
[0022] It can be seen that in the related technology, by writing scripts in the Python environment and connecting the Label Studio Python SDK and the AI platform's Python SDK, the data annotated by Label Studio can be synchronized to the AI platform.
[0023] However, this method requires developers to master the Label Studio Python SDK and the Python SDK of the target platform, and be familiar with their API documentation and usage methods. This involves different concepts, data structures, and operation processes. For beginners, the learning curve is relatively steep, and it may take a lot of time and effort to master. Moreover, it is necessary to manually export data files from Label Studio and then import them to the AI platform, resulting in low efficiency and accuracy of annotation data synchronization.
[0024] In the process of synchronizing the data annotated in Label Studio to the AI platform, it is necessary to write code to complete multiple steps such as data acquisition, processing, conversion, and uploading. To meet the requirements of different platforms, it may also be necessary to handle complex data format conversion and error handling logic. As the functions increase and the requirements change, the code will become more and more complex, and the difficulty of maintenance and extension will also become higher.
[0025] The Python SDK usually depends on multiple third-party libraries, and there may be version compatibility issues with these libraries. In different development environments and production environments, the code may run incorrectly due to inconsistent library versions. In addition, when the dependent libraries are updated, the code also needs to be adjusted and tested accordingly, increasing the workload of development and maintenance.
[0026] To solve the above technical problems, an embodiment of the present application provides a method, device, electronic device, and storage medium for synchronizing annotation data. The method for synchronizing annotation data includes: receiving a data annotation task submitted by a user, invoking a storage interface, storing the data annotation task in a storage database, and allocating computing resources for the data annotation task, where the data annotation task is encapsulated through an application programming interface; starting a built-in data annotation tool; creating a mounting path for the data annotation task in the storage database, and setting the artificial intelligence platform and the data annotation tool with read-write permissions and execution permissions for the mounting path; running the data annotation task, uploading the data to be annotated of the data annotation task to the mounting path, so that the data annotation tool can obtain the data to be annotated from the mounting path, annotate the data to be annotated, and upload the annotated data to the mounting path; obtaining the annotated data from the mounting path, and performing model training based on the annotated data. The method provided by the above solution realizes the synchronization of annotation data by encapsulating the data annotation task through an application programming interface and running the data annotation task. Since the application programming interface only requires a tool that can send HTTP requests, the dependence on specific programming languages and libraries is reduced. Therefore, the technical problem of the high development and maintenance difficulty of synchronizing the data annotated by Label Studio to the AI platform in the related technology can be solved, and the technical effect of reducing the development and maintenance difficulty of synchronizing the data annotated by Label Studio to the AI platform is achieved.
[0027] Moreover, the application programming interface usually maintains a certain stability and backward compatibility. Therefore, when the application programming interface is updated, developers only need to make corresponding adjustments to the application programming interface according to the new API documentation, without the need to wait for the update of the Python SDK and perform complex version upgrades as when using the Python SDK. This makes it easier to maintain and upgrade when the application programming interface is updated, where the application programming interface is the API.
[0028] Furthermore, the application programming interface call is a lightweight operation based on the HTTP protocol and does not require loading the entire SDK library as when using the Python SDK. When processing a large number of annotation data synchronization tasks, the response speed of the application programming interface call is faster, which can reduce unnecessary resource consumption and improve the efficiency of annotation data synchronization.
[0029] Since the application programming interface call is based on HTTP requests, developers can conveniently implement parallel processing, send multiple requests simultaneously to obtain and synchronize annotation data, and thus make full use of the multi-core processors and network bandwidth of the system to further improve the performance of annotation data synchronization.
[0030] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the annotation data synchronization method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0031] The annotation data synchronization method, device, electronic device, and storage medium provided by the embodiments of this application are applicable to synchronizing the data annotated by Label Studio to the AI platform. As Figure 1 shown, it is a schematic structural diagram of the annotation data synchronization system based on this application. The annotation data synchronization system includes an artificial intelligence platform and a data annotation tool. Among them, the artificial intelligence platform is used to receive the data annotation tasks submitted by users, call the storage interface, store the data annotation tasks in the storage database, and allocate computing resources for the data annotation tasks. Among them, the data annotation tasks are encapsulated through application programming interfaces; start the built-in data annotation tool, create a mounting path for the data annotation tasks in the storage database, and set the read-write permissions and execution permissions for the mounting path for the artificial intelligence platform and the data annotation tool; run the data annotation tasks, upload the data to be annotated of the data annotation tasks to the mounting path, so that the data annotation tool can obtain the data to be annotated from the mounting path, annotate the data to be annotated, and upload the annotated data to the mounting path; obtain the annotated data from the mounting path and perform model training based on the annotated data. It can be understood that the data annotation tool is built into the artificial intelligence platform in the form of a data annotation tool image, and the built-in data annotation tool is started by starting the built-in data annotation tool image.
[0032] The embodiments of this application provide an annotation data synchronization method, which is applied to the AI platform. Figure 2 It is a flowchart of the annotation data synchronization method provided by the embodiments of this application. As Figure 2 shown, this process includes: Step S201, receive the data annotation tasks submitted by users, call the storage interface, store the data annotation tasks in the storage database, and allocate computing resources for the data annotation tasks. Among them, the data annotation tasks are encapsulated through application programming interfaces.
[0033] Figure 3 It is a creation flowchart of the data annotation tasks provided by the embodiments of this application. As Figure 3 shown, when a user submits a data annotation task to the AI platform, the data annotation task needs to be created in the AI platform. First, call the storage interface, store the data annotation task in the storage database, and allocate computing resources for the data annotation task. Figure 3 Not shown in
[0034] It should be noted that the storage interface is the istorage interface. Istorage is a storage service. The data annotation task includes multiple items. Storing the data annotation task in the storage database means storing the data annotation task and the items in the task in the istorage database table. The computing resources include Central Processing Unit (CPU) resources, Graphics Processing Unit (GPU) resources, etc.
[0035] Step S202, start the built-in data annotation tool.
[0036] Then, start the mirror image of the built-in data annotation tool of the AI platform to start the built-in data annotation tool of the AI platform. Among them, the data annotation tool is Label Studio, Figure 3 which is not shown in the figure.
[0037] It can be understood that the user selects the Label Studio mirror image to start in the development environment.
[0038] Step S203, create a mount path for the data annotation task in the storage database, and set the read-write and execution permissions of the artificial intelligence platform and the data annotation tool for the mount path.
[0039] Finally, create a mount path for the data annotation task in the storage database, and set the read-write and execution permissions of the AI platform and LabelStudio for the mount path, that is, set the permission of the mount path to 775. Among them, the first 7 in 775 represents the permission of the file owner (User), which means having read, write, and execution permissions. The second 7 represents the permission of the file's group (Group), which means that users within the file's group also have read, write, and execution permissions. The third 5 represents the permission of others (Others), which means having read and execution permissions but no write permission.
[0040] After the above steps S201 to S203 are executed, the creation of the data annotation task is completed on the AI platform by calling the iresource interface.
[0041] Step S204, run the data annotation task, upload the data to be annotated of the data annotation task to the mount path, so that the data annotation tool can obtain the data to be annotated from the mount path, annotate the data to be annotated, and upload the annotated data to the mount path.
[0042] It should be noted that after the creation of the data annotation task is completed on the AI platform, the data annotation task is run to complete the connection between the AI platform and the data annotation tool and the synchronization of the data annotated by the data annotation tool to the AI platform.
[0043] It can be understood that the data to be annotated includes the data to be annotated in each item of the data annotation task. The AI platform creates a data synchronization form by calling the underlying API interface of the data annotation tool to inform the data annotation tool which data needs to be processed. The istorage service of the AI platform calls the local storage data synchronization API to synchronize the data to be annotated of the data annotation task to the mounted path according to the project to which it belongs. At the same time, a local data synchronization function is created in Label Studio. After starting the local data synchronization function, Label Studio scans the data to be annotated under the mounted path and lists all the data to be annotated that can be synchronized to Label Studio, that is, the synchronization list. Based on the synchronization list, the data to be annotated is obtained from the mounted path and synchronized to the corresponding project in Label Studio.
[0044] The istorage service of the AI platform calls the data synchronization API interface of the data to be annotated to track the processing progress of LabelStudio for the data to be annotated in real time.
[0045] The AI platform calls the export annotation result API of the data annotation tool at preset time intervals to enable the data annotation tool to upload the incremental annotation data generated within the preset time interval to the mounted path. Among them, the preset time interval is set by technical personnel and is not specifically limited here. Exemplarily, the preset time interval is 10 seconds. The data annotation task includes the export annotation result API interface of the data annotation tool.
[0046] It can be understood that the AI platform can export the annotation data corresponding to the data to be annotated in one or more projects in the data annotation task to the mounted path by calling the export annotation result API of the data annotation tool. That is to say, the parameters and rules of data synchronization can be flexibly configured through the API, such as the synchronization frequency, data filtering conditions, etc. The administrator can adjust these configurations at any time according to actual needs to meet the business needs at different stages.
[0047] Step S205, obtain the annotation data from the mounted path and perform model training based on the annotation data.
[0048] Among them, the annotation data is obtained from the mounted path at preset time intervals to obtain the latest annotation data, and model training is performed based on the latest annotation data.
[0049] The AI platform calls the import annotation data API of the AI platform to obtain annotation data from the mounted path, and performs model training based on the annotation data. Among them, the data annotation task includes the import annotation data API.
[0050] The annotation data synchronization method provided by the embodiments of the present application builds a high-speed data channel between Label Studio and the AI platform by means of API interfaces, including the import annotation data API and the export annotation result API interface, automatically synchronizes the annotation data to the AI platform, avoids the manual import / export link, saves a lot of time and effort, especially when dealing with large-scale annotation data, and improves the efficiency. And after the data annotation tool completes new annotation data or modifies the existing annotation data, the updated annotation data can be immediately synchronized to the mounted path through the API interface, and then synchronized to the AI platform, so that the AI platform always uses the latest annotation data for model training and optimization, avoiding problems such as model training errors or inaccurate evaluations caused by inconsistent data, ensuring the consistency and accuracy of the annotation data at different stages and different platforms, and accelerating the AI development process.
[0051] And since the application programming interface only requires a tool that can send HTTP requests, reducing the dependence on specific programming languages and libraries, therefore, the technical problem of the high development and maintenance difficulty of the method for synchronizing the data annotated by Label Studio to the AI platform in the related art can be solved, and the technical effect of reducing the development and maintenance difficulty of synchronizing the data annotated by Label Studio to the AI platform is achieved.
[0052] And the API interface can ensure that no information is lost during the transmission of the annotation data, and completely synchronize the annotation data in LabelStudio to the AI platform. At the same time, annotation data verification and error correction can be performed during the transmission process, further ensuring the quality of the data. Real-time annotation data synchronization and accurate data quality ensure that the model of the AI platform can be trained and optimized more efficiently, improve the performance and effect of the model, and thus reduce the additional costs brought by inaccurate models.
[0053] The embodiments of the present application provide an annotation data synchronization method, which is applied to the AI platform. Figure 4 It is a schematic flowchart of the annotation data synchronization method provided by the embodiments of the present application. As Figure 4 shown, the process includes: Step S401, receive the data annotation task submitted by the user, call the storage interface, store the data annotation task in the storage database, and allocate computing resources for the data annotation task. Among them, the data annotation task is encapsulated through the application programming interface. For details, please refer to Figure 2Step S201 of the illustrated embodiment will not be elaborated herein.
[0054] Step S402, start the built-in data annotation tool. For details, please refer to Figure 2 Step S202 of the illustrated embodiment will not be elaborated herein.
[0055] Step S403, create a target mount path for the data annotation task in the storage database, and set the read / write and execution permissions of the artificial intelligence platform and the data annotation tool for the target mount path. For details, please refer to Figure 2 Step S203 of the illustrated embodiment will not be elaborated herein.
[0056] Step S404, run the data annotation task, and upload the data to be annotated in the data annotation task to the target mount path, so that the data annotation tool can obtain the data to be annotated from the mount path, annotate the data to be annotated, and upload the annotated data to the mount path.
[0057] Specifically, the above step S404 includes: Step S4041, run the data annotation task, upload the data to be annotated in the data annotation task to the mount path, and redirect the data annotation task to the data annotation tool address, so that the data annotation tool can automatically execute the data annotation task, obtain the data to be annotated from the mount path, annotate the data to be annotated, and upload the annotated data to the mount path.
[0058] It should be noted that after running the data annotation task, the underlying layer will automatically redirect the data annotation task to the Label Studio address, and the address redirection is performed by istorage. At the same time, the redirected address is the project list of Label Studio.
[0059] Step S405, obtain the annotated data from the mount path, and perform model training based on the annotated data. For details, please refer to Figure 2 Step S205 of the illustrated embodiment will not be elaborated herein.
[0060] For the annotation data synchronization method provided in the embodiments of the present application, the data annotation tool can automatically obtain the data to be annotated from the mount path without manual intervention, reducing the time and errors of manual operations. After annotation, the annotated data will be automatically uploaded to the mount path, realizing the automation of the whole process and improving the overall efficiency.
[0061] In some alternative embodiments, when starting the built-in data annotation tool, the above annotation data synchronization method further includes: Step a1: Inject the user authentication information of the artificial intelligence platform into the data annotation tool as an environment variable, so that the permissions of the users of the artificial intelligence platform in the data annotation tool are the same as their permissions in the artificial intelligence platform.
[0062] Among them, the user authentication information includes the username, password, user authentication information (token), etc. The AI platform can use this token to call the API of the data annotation tool to implement functions such as creating projects and synchronizing storage.
[0063] The annotation data synchronization method provided by the embodiments of this application reduces the need for repeated configuration of user permissions by injecting user authentication information into the data annotation tool as an environment variable. The design of consistent permissions can effectively prevent unauthorized access and ensure that users can only access and operate the data and functions they are authorized to.
[0064] In some optional embodiments, before model training based on the annotation data, the above annotation data synchronization method further includes: Step b1: Map the annotation data to a data format that meets the requirements of the artificial intelligence platform.
[0065] Among them, the annotation data exported by the data annotation tool is JSON data. Parse the annotation data and map the annotation data to the data format required by the AI platform.
[0066] Label Studio and various AI platforms may adopt different data formats and storage methods. The annotation data synchronization method provided by the embodiments of this application ensures the smooth flow of data between Label Studio and the AI platform by converting and adapting the annotation data exported by the data annotation tool, improves the compatibility and integration of the system, and solves the problems of data format and communication protocol differences between Label Studio and different AI platforms.
[0067] In some optional embodiments, before uploading the data to be annotated of the data annotation task to the target mount path, the above annotation data synchronization method further includes: Step c1: Determine whether the file extension of the data to be annotated meets the requirements of the preset file extension.
[0068] Among them, the requirements of the preset file extension are set by technical personnel and include multiple file extensions. The data to be annotated exists in the form of a file.
[0069] Step c2: If the file extension of the data to be annotated meets the requirements of the preset file extension, then determine whether the naming of the data to be annotated meets the preset naming specification.
[0070] It is understandable that if the file extension of the data to be annotated does not meet the requirements of the preset file extension, the files in the data to be annotated that do not meet the requirements of the preset file extension will be deleted.
[0071] Step c3, if the naming of the data to be annotated does not conform to the preset naming specification, rename the data to be annotated based on the preset naming specification.
[0072] Among them, the preset naming specification is set by the technical staff.
[0073] Step c4, obtain the hash value corresponding to each file in the data to be annotated that conforms to the preset naming specification or the renamed data to be annotated.
[0074] Step c5, determine whether there are at least two files with the same corresponding hash value.
[0075] Step c6, if there are at least two files with the same corresponding hash value, retain one of the at least two files with the same hash value, and delete the remaining files among the at least two files with the same hash value.
[0076] Among them, if there are at least two files with the same corresponding hash value, it means that there are duplicate files in the data to be annotated. Then, only one of the duplicate files needs to be retained, and the redundant files among the duplicate files are deleted.
[0077] Step c7, perform integrity verification on each file in the data to be annotated that conforms to the preset naming specification or the renamed data to be annotated.
[0078] Step c8, for any file in the data to be annotated that conforms to the preset naming specification or the renamed data to be annotated, if the file fails the integrity verification, determine that the file is a damaged file.
[0079] Step c9, delete the damaged files in the data to be annotated that conforms to the preset naming specification or the renamed data to be annotated.
[0080] The annotation data synchronization method provided by the embodiments of the present application ensures the correctness of the data format of the data to be annotated, improves the naming consistency of the data to be annotated, reduces the redundant data of the data to be annotated, guarantees the data integrity of the data to be annotated, and improves the accuracy of annotating the data to be annotated.
[0081] In some optional implementation manners, the above step S201 includes: Step d1, obtain the amount of data to be annotated for the data annotation task and the annotation complexity of the data annotation task.
[0082] Among them, based on the type of data to be annotated for the data annotation task, the annotation complexity of the data annotation task is determined. The corresponding relationship between the type of data to be annotated and the annotation complexity is preset in advance.
[0083] Step d2: Based on the amount of data to be annotated for the data annotation task and the annotation complexity of the data annotation task, allocate computing resources for the data annotation task; Among them, the more the amount of data to be annotated and the higher the annotation complexity, the more computing resources are allocated.
[0084] Specifically, based on the amount of data to be annotated for the data annotation task, determine the first computing resource allocated for the data annotation task. Based on the annotation complexity of the data annotation task, determine the second computing resource allocated for the data annotation task. Based on the first computing resource and the second computing resource, determine the total computing resource allocated for the data annotation task.
[0085] In the case where the amount of data to be annotated for the data annotation task exceeds the first data volume threshold, determine that the first computing resource allocated for the data annotation task is the third computing resource.
[0086] In the case where the amount of data to be annotated for the data annotation task exceeds the second data volume threshold, determine that the first computing resource allocated for the data annotation task is the fourth computing resource. Among them, the first data volume threshold is less than the second data volume threshold, and the third computing resource is less than the fourth computing resource.
[0087] In the case where the annotation complexity of the data annotation task exceeds the preset complexity, determine that the second computing resource allocated for the data annotation task is the fifth computing resource.
[0088] In the case where the annotation complexity of the data annotation task does not exceed the preset complexity, determine that the second computing resource allocated for the data annotation task is the sixth computing resource. Among them, the fifth computing resource is more than the sixth computing resource.
[0089] The annotation data synchronization method provided by the embodiments of the present application realizes resource allocation for the data annotation task by based on the amount of data to be annotated for the data annotation task and the annotation complexity of the data annotation task, improves resource utilization, and enhances the execution efficiency of the data annotation task.
[0090] In some optional implementation manners, the above step S204 includes: Step e1: Determine whether the amount of data to be annotated for the data annotation task is greater than the preset data volume threshold.
[0091] Among them, the preset data volume threshold is set by technical personnel and is not specifically limited here.
[0092] Step e2, if the data volume of the data to be annotated in the data annotation task is greater than the preset data volume threshold, divide the data to be annotated into multiple data blocks to be annotated.
[0093] It can be understood that the data to be annotated can be evenly divided into multiple data blocks to be annotated.
[0094] Step e3, perform encryption processing on each data block to be annotated respectively.
[0095] Step e4, perform compression processing on multiple encrypted data blocks to be annotated to obtain the compressed data blocks to be annotated.
[0096] Step e5, upload the compressed data blocks to be annotated to the mount path.
[0097] It can be understood that the data annotation tool obtains the compressed data blocks to be annotated, performs decompression and decryption processing, and then annotates the data to be annotated.
[0098] To make the annotation data synchronization method provided by the embodiments of the present application clearer, it is described in combination with Figure 5 as follows. Among them, Figure 5 This is an interaction diagram for synchronizing the data annotated by the data annotation tool to the AI platform provided by the embodiments of the present application. It should be noted that the process of creating a data annotation task is not shown in this diagram, and the interaction process after running the data annotation task is described. As Figure 5 shown, this interaction process includes: The user clicks on the data annotation task name on the interface of the AI platform to enter the task details interface, where the detailed information of the data annotation task can be viewed. Among them, the interface of the AI platform provides an operation guide for the user to quickly get started with using the data annotation tool for data annotation, helping new users understand the basic operation process.
[0099] The task details interface includes a list of all items under the task, which is used to display all specific items included in the current data annotation task. The list of all items under the task is obtained by calling the istorage project interface in the istorage service. Among them, the istorage project interface obtains the list of all items under the task by calling the project list interface in the data annotation tool. The user can understand the specific work content included in the data annotation task through this list and select operations.
[0100] The task details interface also includes a local storage synchronization interface. The user initiates a request to synchronize the locally stored data, i.e., the data to be annotated, to the mounted path at this interface. By calling the local storage list and synchronization interface in the istorage service, the data to be annotated is synchronized to the mounted path. Among them, the local storage list and synchronization interface in the istorage service call the local storage list and synchronization interface in the data annotation tool to synchronize the data to be annotated to the data annotation tool, that is, the data annotation tool obtains the data to be annotated from the mounted path.
[0101] The task details interface also includes an interface for exporting annotated data. The user initiates a request to export annotated data at this interface. By calling the annotated data export interface in the istorage service, the annotated data is obtained from the mounted path. Among them, the annotated data export interface in the istorage service calls the annotated data export interface in the data annotation tool to obtain the annotated data from the mounted path.
[0102] In the embodiment of the present application, the annotated data of a specified project is obtained through the API interface of Label Studio, and the data obtained from Label Studio is converted to meet the requirements of the AI platform. The converted data is uploaded to the AI platform using the import API of the AI platform, completing the operations of data acquisition, conversion, and upload, thereby achieving the goal of synchronizing the annotated data generated in Label Studio to the AI platform, realizing the direct transmission of the annotated data, and avoiding the intermediate links of manual export / import.
[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0104] The embodiment of the present application also provides an annotated data synchronization device, as Figure 6 shown, including: A receiving module 601, configured to receive a data annotation task submitted by a user, call a storage interface, store the data annotation task in a storage database, and allocate computing resources for the data annotation task, where the data annotation task is encapsulated through an application programming interface.
[0105] A starting module 602, configured to start a built-in data annotation tool.
[0106] A creating module 603, configured to create a mounted path for the data annotation task in the storage database, and set the read / write permissions and execution permissions for the mounted path for the artificial intelligence platform and the data annotation tool.
[0107] The running module 604 is used to run the data annotation task, upload the data to be annotated in the data annotation task to the mount path, so that the data annotation tool can obtain the data to be annotated from the mount path, annotate the data to be annotated, and upload the annotated data to the mount path.
[0108] The obtaining module 605 is used to obtain the annotated data from the mount path and perform model training based on the annotated data.
[0109] In some alternative embodiments, the running module 604 includes: The running unit is used to run the data annotation task, upload the data to be annotated in the data annotation task to the mount path, redirect the data annotation task to the data annotation tool address, so that the data annotation tool can automatically execute the data annotation task, obtain the data to be annotated from the mount path, annotate the data to be annotated, and upload the annotated data to the mount path.
[0110] In some alternative embodiments, the annotated data synchronization device further includes: The injection unit is used to inject the user authentication information of the artificial intelligence platform into the data annotation tool as an environment variable, so that the permissions of the user of the artificial intelligence platform in the data annotation tool are consistent with their permissions in the artificial intelligence platform.
[0111] In some alternative embodiments, the annotated data synchronization device further includes: The mapping unit is used to map the annotated data into a data format that meets the requirements of the artificial intelligence platform.
[0112] In some alternative embodiments, the annotated data synchronization device further includes: The first judgment unit is used to judge whether the file extension of the data to be annotated meets the requirements of the preset file extension.
[0113] The second judgment unit is used to judge whether the naming of the data to be annotated meets the preset naming specification if the file extension of the data to be annotated meets the requirements of the preset file extension.
[0114] The renaming unit is used to rename the data to be annotated based on the preset naming specification if the naming of the data to be annotated does not meet the preset naming specification.
[0115] The first obtaining unit is used to obtain the hash value corresponding to each file in the data to be annotated that meets the preset naming specification or the renamed data to be annotated.
[0116] The third judgment unit is used to judge whether there are at least two files with the same corresponding hash value.
[0117] The first deletion unit is configured to, if there are at least two files with the same hash value, retain one of the at least two files with the same hash value and delete the remaining files among the at least two files with the same hash value.
[0118] The verification unit is configured to perform integrity verification on each file in the data to be labeled that conforms to the preset naming specification or the data to be labeled after being renamed.
[0119] The determination unit is configured to, for any file in the data to be labeled that conforms to the preset naming specification or the data to be labeled after being renamed, if the file fails the integrity verification, determine that the file is a damaged file.
[0120] The second deletion unit is configured to delete the damaged files in the data to be labeled that conforms to the preset naming specification or the data to be labeled after being renamed.
[0121] In some optional embodiments, the receiving module 601 includes: The second acquisition unit is configured to acquire the amount of data to be labeled and the labeling complexity of the data labeling task.
[0122] The allocation unit is configured to allocate computing resources for the data labeling task based on the amount of data to be labeled and the labeling complexity of the data labeling task.
[0123] Wherein, the more the amount of data to be labeled and the higher the labeling complexity, the more computing resources are allocated.
[0124] In some optional embodiments, the running module 604 includes: The fourth judgment unit is configured to judge whether the amount of data of the data to be labeled in the data labeling task is greater than a preset data amount threshold.
[0125] The partitioning unit is configured to, if the amount of data of the data to be labeled in the data labeling task is greater than the preset data amount threshold, partition the data to be labeled into multiple data blocks to be labeled.
[0126] The encryption unit is configured to perform encryption processing on each data block to be labeled respectively.
[0127] The compression unit is configured to perform compression processing on multiple encrypted data blocks to be labeled to obtain compressed data blocks to be labeled.
[0128] The uploading unit is configured to upload the compressed data blocks to be labeled to the mounting path.
[0129] For the description of the features in the embodiments corresponding to the labeled data synchronization device, reference can be made to the relevant descriptions in the embodiments corresponding to the labeled data synchronization method, which will not be elaborated here one by one.
[0130] Embodiments of the present application further provide an electronic device, such as Figure 7 as shown, including a processor 701 and a memory 702. A computer program is stored in the memory 702, and the processor 701 is configured to run the computer program to execute the steps in any of the above-described embodiments of the annotation data synchronization method.
[0131] Embodiments of the present application further provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the annotation data synchronization method when running.
[0132] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media capable of storing computer programs such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs.
[0133] Embodiments of the present application further provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the annotation data synchronization method are implemented.
[0134] Embodiments of the present application further provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the annotation data synchronization method are implemented.
[0135] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0136] The above has introduced in detail a method, device, electronic device, and storage medium for synchronizing labeled data provided in this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for synchronizing labeled data, characterized in that, Including: Receiving a data annotation task submitted by a user, calling a storage interface, storing the data annotation task in a storage database, and allocating computing resources for the data annotation task, wherein the data annotation task is encapsulated through an application programming interface; Starting a built-in data annotation tool; Creating a mounting path for the data annotation task in the storage database, and setting the read-write and execution permissions of the artificial intelligence platform and the data annotation tool for the mounting path; Running the data annotation task, uploading the data to be annotated of the data annotation task to the mounting path, so that the data annotation tool obtains the data to be annotated from the mounting path, annotates the data to be annotated, and uploads the annotated data to the mounting path; Obtaining the annotated data from the mounting path and performing model training based on the annotated data.
2. The method according to claim 1, characterized in that The running the data annotation task, uploading the data to be annotated of the data annotation task to the mounting path, so that the data annotation tool obtains the data to be annotated from the mounting path, annotates the data to be annotated, and uploads the annotated data to the mounting path includes: Running the data annotation task, uploading the data to be annotated of the data annotation task to the mounting path, and redirecting the data annotation task to the data annotation tool address, so that the data annotation tool automatically executes the data annotation task, obtains the data to be annotated from the mounting path, annotates the data to be annotated, and uploads the annotated data to the mounting path.
3. The method according to claim 1, characterized in that, When starting the built-in data annotation tool, the method further includes: Injecting the user authentication information of the artificial intelligence platform into the data annotation tool as an environment variable, so that the permissions of the user of the artificial intelligence platform in the data annotation tool are the same as those in the artificial intelligence platform.
4. The method according to claim 1, wherein Before performing model training based on the annotated data, the method further includes: Mapping the annotated data to a data format that meets the requirements of the artificial intelligence platform.
5. The method according to claim 1, characterized in that, Before uploading the data to be annotated of the data annotation task to the mounting path, the method further includes: Judging whether the file extension of the data to be annotated meets the requirements of a preset file extension; If the file extension of the data to be annotated meets the requirements of the preset file extension, judging whether the naming of the data to be annotated meets the preset naming specification; If the naming of the data to be annotated does not meet the preset naming specification, renaming the data to be annotated based on the preset naming specification; Obtaining the hash value corresponding to each file in the data to be annotated that meets the preset naming specification or the renamed data to be annotated; Judging whether there are at least two files with the same corresponding hash value; If there are at least two files with the same corresponding hash value, retaining one of the at least two files with the same hash value and deleting the remaining files among the at least two files with the same hash value; Performing integrity verification on each file in the data to be annotated that meets the preset naming specification or the renamed data to be annotated; For any file in the to-be-annotated data that conforms to the preset naming specification or the to-be-annotated data after renaming, if the file fails the integrity check, determine that the file is a damaged file; Delete the damaged files in the to-be-annotated data that conforms to the preset naming specification or the to-be-annotated data after renaming.
6. The method according to claim 1, characterized in that, The allocating computing resources for the data annotation task includes: Obtain the amount of to-be-annotated data of the data annotation task and the annotation complexity of the data annotation task; Based on the amount of to-be-annotated data of the data annotation task and the annotation complexity of the data annotation task, allocate computing resources for the data annotation task; Wherein, the more the amount of to-be-annotated data and the higher the annotation complexity, the more computing resources are allocated.
7. The method according to claim 1, wherein The uploading the to-be-annotated data of the data annotation task to the mounting path includes: Judge whether the amount of data of the to-be-annotated data of the data annotation task is greater than a preset data volume threshold; If the amount of data of the to-be-annotated data of the data annotation task is greater than the preset data volume threshold, divide the to-be-annotated data into multiple to-be-annotated data blocks; Perform encryption processing on each of the to-be-annotated data blocks respectively; Perform compression processing on multiple encrypted to-be-annotated data blocks to obtain compressed to-be-annotated data blocks; Upload the compressed to-be-annotated data blocks to the mounting path.
8. A labeled data synchronization device, characterized in that, It includes: A receiving module, configured to receive a data annotation task submitted by a user, call a storage interface, store the data annotation task in a storage database, and allocate computing resources for the data annotation task, wherein the data annotation task is encapsulated through an application programming interface; A starting module, configured to start a built-in data annotation tool; A creating module, configured to create a mounting path for the data annotation task in the storage database, and set the read-write permission and execution permission of the artificial intelligence platform and the data annotation tool for the mounting path; An operating module, configured to operate the data annotation task, upload the to-be-annotated data of the data annotation task to the mounting path, so that the data annotation tool obtains the to-be-annotated data from the mounting path, annotates the to-be-annotated data, and uploads the annotated data to the mounting path; An obtaining module, configured to obtain the annotated data from the mounting path and perform model training based on the annotated data.
9. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the annotation data synchronization method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the annotation data synchronization method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Patent Citations
Data labeling method and device
CN114090534A
Data fast uploading engine implementation method based on DataX
CN115729938A
Task processing method, system and platform and automatic question answering method
CN116431316A
Artificial intelligence labeling platform, method and device and storage medium
CN118863094A
Task processing methods, system and platform, and automated question answering method
WO2024253580A1