Multi-channel data-based labeling method and labeling system

Through the multi-channel data annotation method and system, the 2.5D photometric stereo camera and unsupervised learning model, combined with the Q-learning algorithm, the automation and efficient collaboration of image annotation are achieved, and the problems of time-consuming and labor-intensive manual annotation and data inconsistency in the existing technology are solved, and the labeling efficiency and accuracy are improved.

CN120340033AActive Publication Date: 2025-07-18ZHEJIANG JIN FEI MASCH CO LTD

Patent Information

Application Number
CN202510821815.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

In the prior art, image annotation methods rely on manual labor, which leads to time-consuming and labor-intensive, and have problems such as data inconsistency and low collaboration efficiency.

Method used

Using a multi-channel data-based annotation method, a 2.5D photometric stereo camera is used to acquire images, combined with an unsupervised learning model and Q-learning algorithm, the pre-labeling and manual re-check of multi-channel images can be automated and efficiently collaborated.

Benefits of technology

It significantly reduces the labor amount of manual labeling, improves the efficiency and accuracy of labeling, ensures the quality of labeling, and reduces errors through the multi-channel image checksum consensus mechanism, improving the collaboration efficiency of the overall labeling system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340033A_ABST
    Figure CN120340033A_ABST
Patent Text Reader

Abstract

The invention aims to solve the problem that an existing image annotation model needs more workers to participate in annotation and is time-consuming and labor-consuming. According to the multi-channel data-based labeling method and labeling system provided by the invention, multi-person online collaborative labeling, the pre-labeling capability of an unsupervised model and a mapping strategy are fused, so that comprehensive optimization of a multi-channel image data labeling task is realized, the labor intensity of manual labeling is remarkably reduced, and the labeling efficiency is improved. The overall efficiency of the labeling work is improved, and the labeling quality is ensured; and meanwhile, by dynamically adjusting task allocation, real-time data synchronization and intelligent labeling quality evaluation, labeling efficiency and data quality are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of training image annotation, and particularly to an annotation method and an annotation system based on multi-channel data. Background Art

[0002] In the process of training deep learning models, high-quality data annotation is a key factor in improving model performance. Especially in the fields of industrial defect detection, object detection, medical image analysis, etc., accurate image annotation data can directly affect the generalization ability of the model and the final application effect. However, current image annotation methods mainly rely on manual annotation, and this process usually has the following problems: 1) Due to the large amount of data, manual annotation often requires a large amount of manpower and time. In the process of developing enterprise-level high-precision customized models, the time for data annotation may account for more than 50% of the entire project cycle, greatly slowing down the progress of algorithm development.

[0003] 2) Due to the different experiences of annotators, it is easy to cause data inconsistency, especially in scenarios that require high-precision annotation, such as industrial defect detection, small target detection, etc. Annotation errors or data noise will directly affect the accuracy of the trained model.

[0004] 3) Traditional annotation methods lack efficient multi-person collaboration tools. It is difficult to synchronize the data annotated by different personnel, and problems such as data conflicts and chaotic version management are likely to occur. In addition, the assignment and quality review of annotation tasks need to be manually operated, increasing the communication cost.

[0005] To solve the above-mentioned annotation problems, a method for automatically annotating images based on an algorithm model has been proposed, such as the content shown in the patent with the application publication number CN118429279A. However, the annotation images used in this patent still need to be manually annotated. Although the finally trained model can help complete image annotation, a large amount of manual annotation is still required during the training of this model, and the aforementioned problems still exist, which is time-consuming and laborious. Therefore, an annotation method and an annotation system that can reduce the workload of annotators are needed. Summary of the Invention

[0006] The purpose of the present invention is to solve the deficiencies of the prior art and provide an annotation method and an annotation system based on multi-channel data.

[0007] To solve the above problems, the present invention adopts the following technical solutions: An annotation method based on multi-channel data includes the following steps: Step 1: Collect multi-channel images of multiple target objects for training through a 2.5D photometric stereo camera; Step 2: According to whether the target has defects, the collected multi-channel images are divided into two major categories. One is the image data of the target with defects, and the other is the image data of the target without defects; Step 3: Input the image data of the target without defects into the unsupervised learning model for training to obtain an unsupervised training model; Step 4: Input the image data of the target with defects into the trained unsupervised training model for pre-annotation; Step 5: Assign the pre-annotated defective target images to different annotators for manual recheck; Step 6: Map the annotation information in the rechecked target images to the other channel image data of the multi-channel image corresponding to the image; Step 7: Obtain the multi-channel images of several targets with annotations for training the deep learning model, and end the steps.

[0008] Further, in the step 3, the images of the same channel in the multi-channel images of multiple targets are input into the unsupervised learning model for training to obtain an unsupervised training model; corresponding to the multi-channel images, unsupervised training models corresponding to the number of channels are obtained.

[0009] Further, in the step 4, the multi-channel target image data with defects are respectively input into the unsupervised training models corresponding to their channels for pre-annotation.

[0010] Further, in the step 5, when assigning images to the annotators, it is necessary to complete the image assignment based on the Q-learning algorithm, using the reward mechanism, combined with the historical annotation data of the annotators and the image annotation requirements corresponding to the targets.

[0011] Further, the historical annotation data of the annotators include annotation speed, annotation accuracy, the number of historical incorrect annotation times, and the total annotation duration; the image annotation requirements corresponding to the targets include the difficulty of image annotation and the accuracy of annotation requirements.

[0012] Further, in the recheck process of the step 6, if the pre-annotation is correct, the correct pre-annotation information is mapped to the image data of other channels; if the annotation is incorrect or not annotated, the incorrect annotation is manually modified or manually added, and the result is mapped to the image data of other channels.

[0013] Further, the multi-channel images obtained in the step 1 are set with labels of image acquisition time and channel name.

[0014] A multi-channel data annotation system. Based on the above-mentioned annotation method, the annotation system includes a server and at least one client, where the client is connected to the server to form a distributed collaboration system; the server is used to store data and assign image annotation tasks; the client is used to complete annotation recheck.

[0015] Furthermore, the server includes a database and a task image allocation module; the database uses a MySql database, which is responsible for storing and managing the annotation data of images of all channels and the historical annotation information of annotators; the task image allocation module dynamically allocates the annotation recheck tasks of images according to the historical annotation information of annotators and the image annotation requirements.

[0016] Furthermore, the client includes a UI interface, a server interaction module, an annotator collaboration module, an annotation quality evaluation module, and a monitoring module; Among them, the UI interface is developed using C++ combined with the Qt framework, and the UI interface provides functions such as image loading, box selection, annotation modification, and area addition of annotations; The server interaction module is responsible for data communication between the client and the server. The server interaction module conducts data interaction with the server based on the TCP / IP protocol. The operations of the UI interface of the client will be synchronized to the server in real time through the server interaction module, and the server will also issue the image annotation recheck tasks through the server interaction module; The annotator collaboration module is based on the long polling technology to synchronize the annotation modifications of multiple clients on the same image; for the annotation content being modified, it will also be "locked" based on the optimistic lock mechanism; The annotation quality evaluation module uses a consensus mechanism to evaluate the annotated images rechecked by annotators, and measures the annotation differences between different annotators for the same annotated image through the Kappa coefficient; if the annotation differences exceed the threshold, it is considered that the differences are too large and there are annotation errors, and the administrator is reminded to review; The monitoring module includes an administrator interface, which is developed using the Qt framework. The administrator interface displays the task completion status, annotation speed, and annotation quality of annotators; the administrator interface can also manually adjust the task allocation of annotated images.

[0017] The beneficial effects of the present invention are as follows: By using the defect-free target object image data to train an unsupervised training model for completing the preliminary image pre-annotation task, the labor intensity of manual annotation is reduced, and the original annotation work of annotators is converted into recheck work, improving efficiency and accuracy; By using multi-channel images for annotation, on the one hand, various types of defects of the target object can be displayed more clearly, and on the other hand, the multi-channel images can be used for verification with each other, improving the accuracy of annotation; By applying the Q-learning algorithm and setting up a reward mechanism, the annotation recheck task is allocated, and the image content to be annotated is reasonably distributed, improving the overall efficiency while ensuring the quality of annotation; By applying the long polling technique and the optimistic locking mechanism, the modification of the same image by multiple annotators is synchronized, improving the collaboration efficiency while avoiding the generation of incorrect data; Through a consensus mechanism, multiple annotators conduct a consistency evaluation of the annotation results of the image, reducing the possibility of incorrect annotation. Brief Description of the Drawings

[0018] Figure 1 It is a flowchart of the annotation method in Embodiment 1. Detailed Embodiments

[0019] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0020] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0021] Embodiment 1: As Figure 1 shown, a multi-channel data-based annotation method includes the following steps: Step 1: Use a 2.5D photometric stereo camera to collect multi-channel images of multiple objects for training; Step 2: According to whether the object has defects, the collected multi-channel images are divided into two major categories. One is the image data of the object with defects, and the other is the image data of the object without defects; Step 3: Input the image data of the object without defects into an unsupervised learning model for training to obtain an unsupervised training model; Step 4: Input the image data of the object with defects into the trained unsupervised training model for pre-annotation; Step 5: Assign the pre-annotated defective object images to different annotators for manual recheck; Step 6: Map the annotation information in the target object image after re-inspection to the image data of other channels of the multi-channel image corresponding to this image; Step 7: Obtain multi-channel images of several target objects with annotations for training a deep learning model, and end the steps.

[0022] Since traditional optical cameras only reflect the two-dimensional image information of the target object at one angle, it is easy to cause some defects in the target object to not be clearly displayed. Therefore, a 2.5D photometric stereo camera is used in this application; the 2.5D photometric stereo camera is constructed based on the principle of photometric stereo. Compared with traditional optical cameras, different processing methods are adopted, such as calculating the minimum value, median value, maximum value, and normal vector for each pixel, and projecting them in different directions of x, y, and z to obtain rich multi-channel image data. On the one hand, it can show various types of defects of the target object more clearly, and on the other hand, it can obtain more training image data. It should be noted that for the convenience of data tracing and the completion of the annotation work, in this example, the multi-channel images collected will also be labeled with "acquisition time + channel name". In this way, it is convenient to classify the images of different target objects collected in the same channel and train them in the same model in the subsequent steps.

[0023] In step 2, the image data is divided into target object image data with defects and target object image data without defects by means of manual screening.

[0024] In step 3, the images of the same channel in the multi-channel images of multiple target objects are input into an unsupervised learning model for training to obtain an unsupervised training model; corresponding to the multi-channel images, an unsupervised training model corresponding to the number of channels is obtained. The specific process of training includes the following steps: Step 31: Preprocess the image, including normalizing the pixel values of the image; Step 32: Construct a training model, including a pre-trained convolutional neural network and distribution modeling based on a Gaussian mixture model; Step 33: Train the model by inputting the preprocessed image into the completed training model; where the image training objective is set to minimize the reconstruction error of normal images or feature distribution matching; the training parameters are set as Epochs: 100, learning rate adjustment: reduce to 1 / 10 of the original when the loss stagnates; use a variational autoencoder to constrain the latent space distribution; it should be noted that a denoising autoencoder is also required to denoise the input when inputting into the model; Step 34: Obtain the completed unsupervised training model, and end the steps.

[0025] In step 4, the multi-channel target object image data with defects are respectively input into the unsupervised training models corresponding to their channels for pre-annotation.

[0026] In step 5, when allocating images to annotators, based on the Q-learning algorithm, a reward mechanism is used, combined with the historical annotation data of the annotators and the image annotation requirements corresponding to the target object, to complete the image allocation; the reward mechanism includes rewards for correct annotation and rewards for annotation speed, etc. It should be noted that the annotated images that need to be rechecked will be allocated to several annotators, so that mutual verification can be carried out according to the recheck results of each annotator, reducing the possibility of annotation errors. In addition, it is also convenient for the recheck task of complex annotated images to achieve collaborative annotation.

[0027] The historical annotation data of the annotators include annotation speed, annotation accuracy, the number of historical incorrect annotations, and the total annotation duration; the image annotation requirements corresponding to the target object include the difficulty of image annotation and the accuracy of annotation requirements.

[0028] In the recheck process of step 6, if the pre-annotation is correct, the correct pre-annotation information is mapped to the image data of other channels; if the annotation is incorrect or not annotated, the incorrect annotation is manually modified or manually added, and the result is mapped to the image data of other channels. In this way, the annotation efficiency of multi-channel images of the same target object can be improved. In addition, combined with the annotation content displayed on the images of other channels, or the supplementary annotation content corresponds to the image content, the accuracy of the annotation can be better verified to ensure the annotation quality.

[0029] It should be noted that before the multi-channel images in step 7 are used for training, manual review is also required to prevent situations such as duplicate annotations.

[0030] A multi-channel data annotation system, based on the above annotation method, the annotation system includes a server and at least one client, where the client is connected to the server to form a distributed collaboration system; the server is used to store data and allocate image annotation tasks; the client is used to complete annotation recheck.

[0031] The server includes a database and a task image allocation module; the database uses a MySql database, which is responsible for storing and managing the annotation data of images of all channels and the historical annotation information of annotators; the task image allocation module dynamically allocates the annotation recheck tasks of images according to the historical annotation information of annotators and the image annotation requirements.

[0032] The client includes a UI interface, a server interaction module, an annotator collaboration module, an annotation quality evaluation module, and a monitoring module.

[0033] The UI interface is developed using C++ in combination with the Qt framework. The UI interface provides functions such as image loading, box selection, annotation modification, and adding annotations to regions. Through the UI interface, annotators can conveniently view and modify the content of the image, and perform operations such as box selection, annotation modification, and adding annotations to regions on this interface.

[0034] The server interaction module is responsible for data communication between the client and the server. The server interaction module conducts data interaction with the server based on the TCP / IP protocol. The operations of the UI interface on the client will be synchronized to the server in real time through the server interaction module. The server will also issue image annotation recheck tasks through the server interaction module. For example, the box selection regions added and the annotations added or modified by the annotator during the annotation operation on the client will be transmitted to the server in real time.

[0035] The annotator collaboration module is based on the long polling technology to synchronize the annotation modifications of multiple clients for the same image, which is used to achieve the task collaboration of multiple annotators. Among them, for the annotation content being modified on a certain client, an "optimistic lock" mechanism is also used to ensure that only one annotator can perform annotation modification operations at the same time, preventing chaos caused by multiple modifications of the same data. However, there is no "locking" function for the annotations added to regions. Therefore, in the subsequent process, it is necessary to conduct a review in the monitoring module to prevent the situation of repeated addition of annotations.

[0036] The annotation quality assessment module uses a consensus mechanism to evaluate the annotated images rechecked by the annotators, and measures the annotation differences among different annotators for the same annotated image through the Kappa coefficient. If the annotation difference exceeds the threshold, it is considered that the difference is too large and there is an annotation error, and the administrator is reminded to conduct a review.

[0037] The monitoring module includes an administrator interface, which is developed using the Qt framework. The administrator interface displays the task completion status, annotation speed, and annotation quality of the annotators. When necessary, the administrator interface can also manually adjust the task allocation of the annotated images, such as when the annotator has difficulty continuing to complete the assigned annotation tasks. In addition, the monitoring module also provides data monitoring and statistical functions. Among them, data monitoring represents the work information data of the annotators, and the statistical function is used to count the work tasks completed by the annotators.

[0038] In the implementation process, an unsupervised training model is obtained through training with defect-free target object image data, which is used to complete the preliminary image pre-annotation task, reduce the labor intensity of manual annotation, and convert the original annotation work of annotators into recheck work, improving efficiency and accuracy; by using multi-channel images for annotation, on the one hand, various types of defects of the target object can be more clearly displayed, and on the other hand, the multi-channel images can be used for verification with each other, improving the accuracy of annotation; by using the Q-learning algorithm and setting a reward mechanism, the annotation recheck task is allocated, the image content to be annotated is reasonably allocated, improving the overall efficiency while ensuring the quality of annotation; by using the long polling technology and the optimistic lock mechanism, the modification of the same image by multiple annotators is synchronized, improving the collaboration efficiency while avoiding the generation of incorrect data; through the consensus mechanism, multiple annotators conduct a consistency assessment of the annotation results of the image, reducing the possibility of incorrect annotation.

[0039] The above description is only a specific example of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these corrections and changes based on the idea of the present invention are still within the protection scope of the claims of the present invention.

Claims

1. A labeling method based on multi-channel data, characterized in that, It includes the following steps: Step 1: Collect multi-channel images of multiple target objects for training through a 2.5D photometric stereo camera; Step 2: According to whether the target object has defects, divide the collected multi-channel images into two major categories. One is the image data of the target object with defects, and the other is the image data of the target object without defects; Step 3: Input the image data of the target object without defects into an unsupervised learning model for training to obtain an unsupervised training model; Step 4: Input the image data of the target object with defects into the trained unsupervised training model for pre-labeling; Step 5: Assign the pre-labeled defective target object images to different annotators for manual recheck; Step 6: Map the annotation information in the rechecked target object images to the other channel image data of the multi-channel images corresponding to the images; Step 7: Obtain multi-channel images of several target objects with annotations for training a deep learning model, and end the steps.

2. The annotation method based on multi-channel data according to claim 1, wherein In the said Step 3, input the images of the same channel in the multi-channel images of multiple target objects into the unsupervised learning model for training to obtain an unsupervised training model; corresponding to the multi-channel images, obtain unsupervised training models corresponding to the number of channels.

3. The annotation method based on multi-channel data according to claim 2, wherein In the said Step 4, input the image data of the multi-channel target object with defects into the unsupervised training model corresponding to its channel for pre-labeling respectively.

4. A labeling method based on multi-channel data according to claim 1, characterized in that In the said Step 5, when assigning images to annotators, it is necessary to complete the image assignment based on the Q-learning algorithm, using a reward mechanism, in combination with the historical annotation data of the annotator and the image annotation requirements corresponding to the target object.

5. A labeling method based on multi-channel data according to claim 4, characterized in that, The historical annotation data of the said annotator includes annotation speed, annotation accuracy, the number of historical incorrect annotations, and the total annotation duration; the image annotation requirements corresponding to the target object include the difficulty of image annotation and the accuracy of annotation requirements.

6. A labeling method based on multi-channel data according to claim 1, characterized in that In the recheck process of the said Step 6, if the pre-labeling is correct, map the correct pre-labeling information to the image data of other channels; if the annotation is incorrect or not annotated, manually modify the incorrect annotation or manually add an annotation, and map the result to the image data of other channels.

7. A labeling method based on multi-channel data according to claim 1, characterized in that, The multi-channel images obtained in the said Step 1 are set with labels of image acquisition time and channel name.

8. A labeling system for multi-channel data, characterized in that Based on the annotation method according to any one of claims 1 to 7, the annotation system includes a server and at least one client, where the client is connected to the server to form a distributed cooperation system; the server is used to store data and assign image annotation tasks; The client is used to complete annotation recheck.

9. The annotation system based on multi-channel data according to claim 8, wherein, The said server includes a database and a task image assignment module; among them, the database uses a MySql database, which is responsible for storing and managing the annotation data of images of all channels and the historical annotation information of annotators; the task image assignment module dynamically assigns the annotation recheck tasks of images according to the historical annotation information of annotators and the image annotation requirements.

10. A labeling system based on multi-channel data according to claim 8, characterized in that, The said client includes a UI interface, a server interaction module, an annotator cooperation module, an annotation quality evaluation module, and a monitoring module; The UI interface is developed using C++ combined with the Qt framework. The UI interface provides functions such as image loading, box selection, annotation modification, and adding annotations to regions. The server interaction module is responsible for data communication between the client and the server. The server interaction module conducts data interaction with the server based on the TCP / IP protocol. Operations on the UI interface of the client will be synchronized to the server in real time through the server interaction module, and the server will also issue image annotation recheck tasks through the server interaction module. The annotator collaboration module is based on the long-polling technology to synchronize annotation modifications of multiple clients on the same image. For the annotation content being modified, it will also be "locked" based on the optimistic lock mechanism. The annotation quality assessment module uses a consensus mechanism to evaluate the annotated images rechecked by annotators, and measures the annotation differences between different annotators for the same annotated image through the Kappa coefficient. If the annotation difference exceeds the threshold, it is considered that the difference is too large and there is an annotation error, and the administrator is reminded to review. The monitoring module includes an administrator interface, which is developed using the Qt framework. The administrator interface displays the task completion status, annotation speed, and annotation quality of annotators. The administrator interface can also manually adjust the task allocation of annotated images.

Citation Information

Patent Citations

  • Intelligent labeling method and system for welding defects

    CN118429279A

  • Model-assisted data annotation system and annotation method

    CN110880021A

  • Joint multi-channel product defect classification method based on deep learning

    CN111709918A

  • Three-dimensional ground penetrating radar image underground pipeline identification method based on 2.5D-CNN algorithm

    CN113780361A

  • Defect detection method and system based on photometric stereo

    CN114998308A

Cited By

  • Multi-source data-based labeling task adaptive optimization method

    CN121543065A