Method, apparatus, device, and storage medium for data annotation
Through mutual reference and frequency analysis between multiple models, the accuracy and consistency of data annotation in machine learning training is solved, and more efficient data annotation results are achieved.
Patent Information
- Application Number
- CN202210598807.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-30
AI Technical Summary
In machine learning training, obtaining high-quality labeled data is a challenge, and it is difficult for the prior art to effectively utilize the labeling results of multiple models for accurate and consistent data annotation.
The target data is marked separately through multiple models, and the final annotation result is determined by the frequency of occurrence of multiple annotation results. The models refer to each other and adjust the annotation result to improve accuracy.
A more accurate and consistent annotation results of the target data are achieved, which avoids the dispersion and inconsistency of the annotation results and improves the convergence of the annotation process.
Smart Images

Figure CN114970724B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular to the fields of artificial intelligence, speech technology, and deep learning technology. Background Art
[0002] With the advent of the era of artificial intelligence, machine learning has been applied to more and more fields. In the machine learning training of supervised learning, the first problem to be solved is the acquisition of high-quality training data. Only with the labeled data can the subsequent machine learning training process be carried out. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, and storage medium for data annotation.
[0004] According to one aspect of the present disclosure, a method for data annotation is provided, including:
[0005] Using multiple models to respectively annotate target data to obtain multiple annotation results, wherein the annotation result of at least one model among the multiple models is obtained according to the annotation results of at least one other model among the multiple models. And
[0006] Determining the final annotation result of the target data according to the occurrence frequencies of the multiple annotation results.
[0007] According to another aspect of the present disclosure, a device for data annotation is provided, including:
[0008] An annotation module for using multiple models to respectively annotate target data to obtain multiple annotation results, wherein the annotation result of at least one model among the multiple models is obtained according to the annotation results of at least one other model among the multiple models. And
[0009] A determination module for determining the final annotation result of the target data according to the occurrence frequencies of the multiple annotation results.
[0010] According to another aspect of the present disclosure, an electronic device is provided, including:
[0011] At least one processor; and
[0012] A memory communicatively connected to the at least one processor; wherein,
[0013] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of any embodiment in the present disclosure.
[0014] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of the embodiments in the present disclosure.
[0015] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method according to any one of the embodiments in the present disclosure.
[0016] According to the solution of the present disclosure, a more accurate annotation result of the target data can be obtained.
[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0019] Figure 1 is a flowchart of the method for data annotation according to an embodiment of the present disclosure;
[0020] Figure 2 is a schematic diagram of an application scenario of the method for data annotation according to an embodiment of the present disclosure;
[0021] Figure 3 is a flowchart of step S101 of the method for data annotation according to an embodiment of the present disclosure;
[0022] Figure 4 is a flowchart of step S101 of the method for data annotation according to another embodiment of the present disclosure;
[0023] Figure 5 is a schematic structural diagram of the device for data annotation according to an embodiment of the present disclosure;
[0024] Figure 6 is a block diagram of an electronic device for implementing the method for data annotation according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] Embodiments of the present disclosure provide a method for data annotation, as Figure 1 shown, which is a flowchart of the method for data annotation in this embodiment. The method may include the following steps:
[0027] S101: Use multiple models to respectively annotate the target data to obtain multiple annotation results, where the annotation result of at least one model among the multiple models is obtained based on the annotation results of at least one other model among the multiple models. And
[0028] S102: Determine the final annotation result of the target data according to the occurrence frequencies of the multiple annotation results.
[0029] It should be noted that the target data can be voice data, image data, text data, etc. When the target data is voice data, through the annotation result, it can be determined what the text content corresponding to the voice data is. When the target data is image data, through the annotation result, it can be determined what image contents are included in the image data, and what specific people, things, etc. are in the image. When the target data is text data, through the annotation result, it can be determined what field the text data describes, the type of the text data, etc.
[0030] The models used can be understood as software or systems with annotation functions, or can be understood as pre-trained neural network models with annotation functions. After each model annotates the target data, at least one annotation result can be obtained.
[0031] The multiple models may include Model A. After Model A obtains a preliminary annotation result based on the target data, it can then adjust the preliminary annotation result according to the annotation results of one or more other models, and thus obtain the actual annotation result of Model A for the target data. The multiple models may also include Model B. Model B can directly obtain an annotation result based on the target data without referring to the annotation results of certain others. The multiple models may include at least one Model A, and the multiple models may not include Model B.
[0032] The occurrence frequencies of the multiple annotation results can be understood as: the number of occurrences of each annotation result. For example, four models output four annotation results, where the annotation results of the first model and the second model are both a, the annotation result of the third model is b, and the annotation result of the fourth model is c. Then the occurrence frequency of the a annotation result is 2, and the occurrence frequencies of the b and c annotation results are 1.
[0033] The determined final annotation result can be a single unique annotation result or multiple different annotation results.
[0034] According to the solution of the present disclosure, since the annotation process of some models refers to the annotation results of other models and adjusts its own annotation results based on this, the annotation results obtained by each model can be made more accurate, thereby further ensuring that the final annotation result obtained based on the annotation results of all models is more accurate. At the same time, since some models refer to the annotation results of other models, the problem that the annotation results of each model are too scattered and inconsistent is also avoided. After using multiple models to annotate the target data, the final annotation result can be quickly converged.
[0035] The data annotation method provided by the embodiments of the present disclosure can be applied to a scenario framework as Figure 2 shown. In Figure 2 , 10 represents a terminal, 20 represents a server, and 30 represents a distributed computer system. The data annotation method of the present disclosure can be jointly executed by the terminal 10, the server 20, or the distributed computer system 30, or can be executed by one or more of the terminal 10, the server 20, or the distributed computer system 30. The terminal 10 can be used to report / send the target data to the server 20 or the distributed computer system 30. After the server 20 or the distributed computer system 30 completes the data annotation method of the present disclosure through multiple models, the final annotation result can be fed back to the terminal 10. The terminal 10, the server 20, or the distributed computer system 30 can also independently complete the data annotation method of the present disclosure using multiple models according to the target data.
[0036] In one embodiment, the data annotation method provided by the embodiments of the present disclosure includes steps S101 and S102, where step S101: using multiple models to respectively annotate the target data to obtain multiple annotation results, where the annotation result of at least one model among the multiple models is obtained according to the annotation results of at least one other model among the multiple models, and it can further include:
[0037] Using multiple models to sequentially and respectively annotate the target data to obtain multiple annotation results, where the annotation result of the model that comes later among the multiple models is obtained according to the annotation results of at least one model that comes earlier.
[0038] It should be noted that using multiple models to sequentially and respectively annotate the target data can be understood as: after the previous model annotates the target data to obtain an annotation result, the subsequent model then annotates the target data.
[0039] The model that comes earlier can be understood as the previous model of the model that comes later, or can also be understood as any one or more models before the model that comes later.
[0040] The annotation result of the subsequent model is obtained based on the annotation results of at least one prior model. It can be understood that: after the subsequent Model A obtains a preliminary annotation result based on the target data, it then adjusts the preliminary annotation result according to the annotation result of the prior Model B (i.e., the model that has already annotated the target data before Model A performs annotation), and thus obtains the actual annotation result of Model A for the target data. It can also be understood that: after the subsequent Model A obtains a preliminary annotation result based on the target data, it then adjusts the preliminary annotation result according to the annotation results of the prior Model B and Model C, and thus obtains the actual annotation result of Model A for the target data.
[0041] According to the solution of the present disclosure, since the annotation process of the subsequent model refers to the annotation results of other prior models and adjusts its own annotation result based on this, the annotation result obtained by the subsequent model can be made more accurate, thereby further ensuring that the final annotation result obtained based on the annotation results of all models is more accurate. At the same time, since the subsequent model refers to the annotation results of the prior models, the problem of the annotation results of each model being too scattered and inconsistent is also avoided, so that after multiple models are used to annotate the target data, the final annotation result can be quickly converged.
[0042] In one example, affected by the audio quality of the voice, the accent of the voice speaker, or the ambient noise of the voice, each model may annotate a sentence of voice with different annotation results. In order to avoid the situation where the same annotation result cannot be obtained, it is necessary to use the data annotation method of the embodiments of the present disclosure to annotate a sentence of voice. As Figure 3 shown, the multiple models include a first model, a second model, a third model, a fourth model, and a fifth model, and the voice data (target data) is annotated by the five models. After the first model annotates the voice data, a first annotation result is obtained. When the second model annotates the voice data, it adjusts the preliminary annotation result obtained by itself with reference to the first annotation result, and thus obtains a second annotation result. When the third model annotates the voice data, it adjusts the preliminary annotation result obtained by itself with reference to the second annotation result, and thus obtains a third annotation result. When the fourth model annotates the voice data, it adjusts the preliminary annotation result obtained by itself with reference to the third annotation result, and thus obtains a fourth annotation result. When the fifth model annotates the voice data, it adjusts the preliminary annotation result obtained by itself with reference to the fourth annotation result, and thus obtains a fifth annotation result.
[0043] In a specific example, the first annotation result obtained by the above-mentioned first model for the content of the voice data includes five features, where A1: today, A2: day, A3: sky, A4: weather, A5: sunny. The preliminary annotation result obtained by the second model includes six features, where B1: today, B2: day, B3: sky, B4: weather, B5: clear, B6: cool. On this basis, the preliminary annotation result is adjusted according to the five features of the first annotation result to obtain the second annotation result of the second model, B1: today, B2: day, B3: sky, B4: weather, B5: sunny, B6: clear.
[0044] In one implementation manner, the data annotation method provided by the implementation manner of the present disclosure includes steps S101 and S102, where step S101: using multiple models to separately annotate target data to obtain multiple annotation results, where the annotation result of at least one model among the multiple models is obtained according to the annotation results of at least one other model among the multiple models, and it may further include:
[0045] Using multiple models to sequentially and separately annotate target data to obtain multiple annotation results, where starting from the third model among the multiple models, the annotation result is obtained according to the annotation results of two preceding models.
[0046] It should be noted that the two preceding models can be understood as any two preceding models before the current model. It can also be understood as the two preceding models arranged in sequence before the current model. For example, including five models A, B, C, D, and E, the two preceding models based on which the D model is determined can be the B and C models, or the A and B models, or the A and C models.
[0047] According to the solution of the present disclosure, since starting from the third model, the annotation process of the model refers to the annotation results of two preceding models and adjusts its own annotation result based on this, the annotation results obtained by the third model and subsequent models can be made more accurate, thereby further ensuring that the final annotation result obtained based on the annotation results of all models is more accurate. At the same time, since the subsequent models refer to the annotation results of the preceding models, the problem of the annotation results of each model being too scattered and inconsistent is also avoided, so that after using multiple models to annotate the target data, the final annotation result can be quickly converged.
[0048] In one implementation manner, the data annotation method provided by the implementation manner of the present disclosure includes steps S101 and S102, where step S101: using multiple models to separately annotate target data to obtain multiple annotation results, where the annotation result of at least one model among the multiple models is obtained according to the annotation results of at least one other model among the multiple models, and it may further include:
[0049] The target data is labeled sequentially by multiple models to obtain multiple labeling results. Among them, starting from the third model among the multiple models, the labeling result is obtained based on the labeling results of two previous models. Starting from the fourth model among the multiple models, one of the labeling results of the two previous models used is the labeling result of the model immediately preceding the current model.
[0050] According to the solution of the present disclosure, since starting from the third model, the labeling process of the model refers to the labeling results of two previous models and adjusts its own labeling result based on this, the labeling results obtained by the third model and subsequent models can be made more accurate, thereby further ensuring that the final labeling result obtained based on the labeling results of all models is more accurate. At the same time, since the subsequent models refer to the labeling results of the previous models, the problem of the labeling results of each model being too scattered and inconsistent is also avoided, so that after labeling the target data with multiple models, the final labeling result can be quickly converged.
[0051] In one example, affected by the audio quality of the speech, the accent of the speaker, or the environmental noise of the speech, each model may label a sentence of speech with different labeling results. In order to avoid the situation where the same labeling result cannot be obtained, it is necessary to use the data labeling method of the embodiments of the present disclosure to label a sentence of speech. As Figure 4 shown, the multiple models include a first model, a second model, a third model, a fourth model, and a fifth model. The speech data (target data) is labeled by the five models. After the first model labels the speech data, a first labeling result is obtained. After the second model labels the speech data, a second labeling result is obtained. When the third model labels the speech data, it refers to the first labeling result and the second labeling result to adjust the preliminary labeling result obtained by itself, and then obtains a third labeling result. When the fourth model labels the speech data, it refers to the third labeling result and the second labeling result to adjust the preliminary labeling result obtained by itself, and then obtains a fourth labeling result. When the fifth model labels the speech data, it refers to the third labeling result and the fourth labeling result to adjust the preliminary labeling result obtained by itself, and then obtains a fifth labeling result.
[0052] In a specific example, the content of the first annotation result obtained by the above first model for the voice data includes five features, where A1: today, A2: day, A3: sky, A4: weather, A5: sunny. The content of the second annotation result obtained by the second model for the voice data includes six features, where B1: today, B2: day, B3: sky, B4: weather, B5: sunny, B6: clear. The preliminary annotation result obtained by the third model based on the voice data includes six features, where C1: Beijing, C2: day, C3: field, C4: weather, C5: sunny, C6: ah. On this basis, the preliminary annotation result is adjusted according to the first annotation result and the second annotation result to obtain the third annotation result of the third model, where C1: today, C2: day, C3: sky, C4: weather, C5: sunny, C6: ah.
[0053] In an example, starting from the fourth model among multiple models, the annotation result of the fourth model is obtained according to the annotation results (annotation result A and annotation result B) of two prior models, including:
[0054] Obtain the third annotation result of the third model, which is the previous prior model of the fourth model, and use it as annotation result A; and
[0055] Obtain the first annotation result of the first model and the second annotation result of the second model, and determine the annotation result with a relatively large difference from the annotation result of the third model as annotation result B (for example, the first annotation result);
[0056] According to annotation result A (i.e., the third annotation result) and annotation result B (i.e., the first annotation result), adjust the preliminary annotation result obtained by the fourth model based on the target data to obtain the fourth annotation result of the fourth model.
[0057] It should be noted that when there are a fifth model, a sixth model, and even more models among multiple models, when these models select the annotation results of prior models, they can all refer to the way the fourth model selects the annotation results of prior models (annotation result A and annotation result B).
[0058] In an implementation manner, the method for data annotation provided by the implementation manner of the present disclosure includes steps S101 and S102, where step S102: determining the final annotation result of the target data according to the occurrence frequencies of multiple annotation results may further include:
[0059] According to the occurrence frequencies of multiple annotation results, determine the annotation result with the highest occurrence frequency as the final annotation result of the target data.
[0060] It should be noted that, according to the occurrence frequencies of multiple annotation results, determining the final annotation result of the target data can be understood as follows: after all models have obtained annotation results, the annotation result with the highest occurrence frequency is selected as the final annotation result. It can also be understood as follows: when only some models have obtained annotation results and the remaining models have not obtained annotation results yet, if the occurrence frequency of the same annotation result in the part of the models that have obtained annotation results exceeds half of the total number of models, the annotation result can also be determined as the final annotation result at this time.
[0061] According to the solution of the present disclosure, selecting the annotation result with the highest occurrence frequency as the final annotation result of the target data can ensure the accuracy of the obtained final annotation result.
[0062] In one example, if the model is a software or system with an annotation function, when the subsequent model adjusts its preliminary annotation result based on the annotation result of the previous model, it can modify the calculation parameters of the software or system based on the annotation result of the previous model, so as to obtain the annotation result again.
[0063] In one example, if the model is a pre-trained neural network model with an annotation function, when the subsequent model adjusts its preliminary annotation result based on the annotation result of the previous model, it can modify the parameters or weights of the neural network model based on the annotation result of the previous model, so as to obtain the annotation result again.
[0064] An embodiment of the present disclosure provides a data annotation device, as Figure 5 shown, which is a structural block diagram of the data annotation device of this embodiment. The device may include:
[0065] An annotation module 510, configured to use multiple models to respectively annotate target data to obtain multiple annotation results, where the annotation result of at least one model among the multiple models is obtained according to the annotation result of at least one other model among the multiple models. And
[0066] A determination module 520, configured to determine the final annotation result of the target data according to the occurrence frequencies of the multiple annotation results.
[0067] In one implementation manner, the annotation module 510 is configured to use multiple models to sequentially and respectively annotate target data to obtain multiple annotation results, where the annotation result of the subsequent model among the multiple models is obtained according to the annotation result of at least one previous model.
[0068] In one implementation manner, the annotation module 510 is configured to use multiple models to sequentially and respectively annotate target data to obtain multiple annotation results, where starting from the third model among the multiple models, the annotation result is obtained according to the annotation results of two previous models.
[0069] In one embodiment, the annotation module 510 is configured to sequentially annotate the target data using multiple models respectively to obtain multiple annotation results. Among them, starting from the third model among the multiple models, the annotation result is obtained based on the annotation results of two previous models. Starting from the fourth model among the multiple models, one of the annotation results of the two previous models on which the current model is based is the annotation result of the previous model of the current model.
[0070] In one embodiment, the determination module 520 is configured to determine, according to the occurrence frequency of the multiple annotation results, that the annotation result with the highest occurrence frequency is the final annotation result of the target data.
[0071] In one embodiment, the target data is any one of voice data, image data, or text data.
[0072] For the specific functions and examples of each module and sub-module of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0073] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0074] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0075] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0076] As Figure 6As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 602 or computer programs loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0077] Multiple components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0078] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method of data annotation. For example, in some embodiments, the method of data annotation can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method of data annotation described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method of data annotation in any other appropriate way (e.g., by means of firmware).
[0079] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0080] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0081] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0082] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0083] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0084] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0085] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0086] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for data annotation, comprising: Using multiple models to respectively annotate target data to obtain multiple annotation results, wherein the annotation result of at least one model among the multiple models is obtained based on the annotation results of at least one other model among the multiple models; and Determining a final annotation result of the target data according to the occurrence frequencies of the multiple annotation results; Wherein, the step of using multiple models to respectively annotate target data to obtain multiple annotation results includes: using the multiple models to sequentially and respectively annotate the target data to obtain the multiple annotation results, wherein the annotation result of a later model among the multiple models is obtained by adjusting a preliminary annotation result obtained by the later model based on the target data according to the annotation results of at least one earlier model; Wherein, the annotation result of a later model among the multiple models is obtained by adjusting a preliminary annotation result obtained by the later model based on the target data according to the annotation results of at least one earlier model, including: starting from the fourth model among the multiple models, the annotation result of the later model is obtained by adjusting a preliminary annotation result obtained by the later model based on the target data according to the annotation result of the previous model of the later model and the annotation results of other earlier models among the multiple earlier models that have a relatively large difference from the annotation result of the previous model; 2. The method according to claim 1, wherein The step of using multiple models to respectively annotate target data to obtain multiple annotation results, wherein the annotation result of at least one model among the multiple models is obtained based on the annotation results of at least one other model among the multiple models, includes: Using multiple models to sequentially and respectively annotate target data to obtain multiple annotation results, wherein starting from the third model among the multiple models, the annotation result is obtained based on the annotation results of two earlier models; 3. The method according to claim 1 or 2, wherein The step of determining a final annotation result of the target data according to the occurrence frequencies of the multiple annotation results includes: Determining the annotation result with the highest occurrence frequency as the final annotation result of the target data according to the occurrence frequencies of the multiple annotation results; 4. The method according to claim 1 or 2, wherein The target data is any one of voice data, image data or text data; 5. A device for data annotation, comprising: An annotation module, configured to use multiple models to respectively annotate target data to obtain multiple annotation results, wherein the annotation result of at least one model among the multiple models is obtained based on the annotation results of at least one other model among the multiple models; and A determination module, configured to determine a final annotation result of the target data according to the occurrence frequencies of the multiple annotation results; The annotation module is configured to sequentially and respectively annotate the target data by using the multiple models to obtain the multiple annotation results, wherein the annotation result of the subsequent model among the multiple models is obtained by adjusting the preliminary annotation result obtained by the subsequent model according to the target data based on the annotation results of at least one preceding model; wherein the annotation result of the subsequent model among the multiple models is obtained by adjusting the preliminary annotation result obtained by the subsequent model according to the target data based on the annotation results of at least one preceding model, including: starting from the fourth model among the multiple models, the annotation result of the subsequent model is obtained by adjusting the preliminary annotation result obtained by the subsequent model according to the target data based on the annotation result of the model preceding the subsequent model and the annotation results of other preceding models among the multiple preceding models that have a relatively large difference from the annotation result of the preceding model.
6. The device according to claim 5, wherein, The annotation module is configured to sequentially and respectively annotate the target data by using the multiple models to obtain the multiple annotation results, wherein starting from the third model among the multiple models, the annotation result is obtained based on the annotation results of two preceding models.
7. The device according to claim 5 or 6, wherein, The determination module is configured to determine, according to the occurrence frequencies of the multiple annotation results, that the annotation result with the highest occurrence frequency is the final annotation result of the target data.
8. The device according to claim 5 or 6, wherein The target data is any one of voice data, image data, or text data.
9. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 4.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 4.
11. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Data annotation correction method and device, computer readable medium and electronic equipment
CN110399933A
Image annotation method and device, electronic equipment and computer readable storage medium
CN113299373A