Intelligent labeling method based on instance segmentation
Through the intelligent annotation method based on instance segmentation, the FastSAM model is used to segment and label the target instance, which solves the problem of time-consuming, labor-intensive and error-prone traditional annotation methods, and achieves efficient and accurate intelligent annotation.
Patent Information
- Application Number
- CN202510099658.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional intelligent annotation methods rely on manual operations, are time-consuming and labor-intensive, and are prone to human errors, especially when processing large amounts of data.
Using an intelligent annotation method based on instance segmentation, we create a data set for training an instance segmentation model, train the FastSAM model, perform instance segmentation of the target, and convert the results into intelligent annotation results, and finally fine-tune them.
It realizes the rapid identification of specific features in the data, automatically generates high-quality labels, reduces the workload of manual labels, improves the labeling speed and accuracy, and is suitable for fields such as autonomous driving and medical images.
Smart Images

Figure CN120107971A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent labeling method based on instance segmentation. Background Art
[0002] Intelligent labeling is a technology based on artificial intelligence and machine learning, which aims to label images, videos, text and other data through automated means. Traditional labeling work relies on manual operations, which is often time-consuming and laborious. Especially in scenarios where large amounts of data need to be processed, manual labeling not only consumes resources, but is also prone to human errors. Summary of the invention
[0003] In view of the technical defects and technical drawbacks in the prior art, the embodiments of the present invention provide an intelligent labeling method based on instance segmentation to overcome the above problems or at least partially solve the above problems. The specific scheme is as follows:
[0004] An intelligent labeling method based on instance segmentation, the method comprising:
[0005] Step 1: Create a dataset for training the instance segmentation model;
[0006] Step 2: training an instance segmentation model as the basis for intelligent annotation;
[0007] Step 3: Based on the trained instance segmentation model, perform instance segmentation on the targets that need intelligent labeling;
[0008] Step 4: Convert the instance segmentation result into an intelligent labeling result;
[0009] Step 5: Fine-tune the results of intelligent labeling.
[0010] Furthermore, in step 1, images and videos that meet the task requirements are taken by a camera, mobile phone or drone as the data set, wherein the taken images and videos cover different angles, lighting conditions and environmental backgrounds of the target object to increase the generalization ability of the model.
[0011] Furthermore, step 1 also includes: preprocessing the data in the data set, wherein the preprocessing includes image size standardization, grayscale adjustment, and denoising, and making the data format input to the instance segmentation model unified and compliant with the standard.
[0012] Furthermore, in step 1, LabelMe is used to annotate the data in the dataset, wherein a polygon tool is used to annotate the target objects, and each target object is named, edited, and the annotation processing is adjusted, and an annotation file in JSON format is automatically generated. Each annotation file includes the category label of the target object and the corresponding bounding box and segmentation mask information. After the annotation is completed, the annotation file is converted into a format that adapts to different instance segmentation model training requirements.
[0013] Furthermore, in step 2, the FastSAM model is used as the instance segmentation model.
[0014] Furthermore, the process of training the FastSAM model includes standard deep learning steps, including:
[0015] Prepare a dataset for instance segmentation tasks, which includes images and their corresponding object segmentation mask labels;
[0016] Initialize the model architecture, define the loss function and optimizer to build the training process, where the loss function uses cross entropy loss or Dice loss to measure the difference between the segmentation mask predicted by the model and the actual segmentation mask; the optimizer uses Adam or SGD to continuously update the model weights through back propagation to optimize the model performance;
[0017] The images and their corresponding segmentation masks are input into the model in small batches. The model performs instance segmentation on each image, predicts whether each pixel belongs to the target object, and adjusts the weights by calculating the loss function so that the prediction results gradually approach the true value.
[0018] In the process of training the model, cross-validation technology is used to evaluate the performance of the model. A part of the data is selected as the validation set, and the performance of the model on the validation set is regularly evaluated. By observing the changes in the validation error of the model, the learning rate and regularization parameters are dynamically adjusted to avoid overfitting or underfitting of the model. In order to accelerate the training process and improve the performance of the model, a transfer learning strategy of the pre-trained model is adopted. By using the weights pre-trained on a large-scale data set, the model converges faster and obtains better results on a small data set. After the training is completed, the model is deployed on the target device.
[0019] Furthermore, in order to improve the generalization ability of the model, data augmentation technology is used in the training process to make the model adapt to different image changes and avoid over-reliance on specific image features.
[0020] Furthermore, the instance segmentation in step 3 includes four methods: full segmentation, point segmentation, box segmentation, and text segmentation;
[0021] Among them: full segmentation is to perform instance segmentation on the entire image; point segmentation is to mark the objects to be segmented in the form of points on the image, and mark the background that does not need to be segmented with background points, so that the algorithm can intelligently segment the objects to be segmented; frame segmentation is to mark a rectangular frame on the image, and the algorithm segments part of the graphics in the frame, and only the target with the largest area in the frame is retained; text segmentation is to describe the objects to be segmented in the form of text descriptions, identify the targets in the image that are closest to the corresponding descriptions through the algorithm, and segment the corresponding targets.
[0022] Furthermore, in the instance segmentation task, the instance segmentation result obtained is a segmentation mask for each detected object, and the segmentation mask is a pixel-level binary image that represents the position of the target object in the image;
[0023] In step 4, converting the instance segmentation result into an intelligent annotation result includes: converting the segmentation mask obtained by the instance segmentation into an annotation file in VOC format, specifically including:
[0024] Step 401, generating a bounding box and a corresponding class label for each object, and each segmentation mask determines the bounding box of the object by finding the circumscribed rectangle of the target area;
[0025] Step 402, in accordance with the requirements of the VOC format, organize the information obtained in step 401 into an XML file to describe the label, bounding box coordinates, and related information of each object; the segmentation mask itself is saved in a separate PNG file as a segmented image, and the PNG file path is referenced in the XML file, thereby converting the instance segmentation mask predicted by the model into a standard VOC format annotation file.
[0026] Further, step 5 includes: after the intelligent annotation is completed, the generated annotation file information is sent to the front-end annotation interface, and the annotation information is intuitively presented on the contour of the target object in the form of feature points, and the target boundary to be annotated is outlined by the feature points;
[0027] Drag and move the corresponding feature points to precisely adjust the contour shape of the target;
[0028] Add or delete feature points as needed to achieve the required annotation accuracy;
[0029] Every adjustment made during the annotation process will be reflected in real time on the interface, allowing users to intuitively evaluate the effects of the modifications.
[0030] The present invention has the following beneficial effects:
[0031] The present invention provides an intelligent labeling method based on instance segmentation. By training a well-trained model, the system can quickly identify specific features in the data and automatically generate high-quality labels. This process includes model pre-training, preliminary labeling of data and manual verification, forming an efficient labeling workflow. Intelligent labeling can not only reduce the workload of manual labeling, but also greatly improve the speed of labeling, especially in the fields of autonomous driving, medical imaging, e-commerce and video surveillance. It has a wide range of applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A flow chart of an intelligent labeling method based on instance segmentation provided by an embodiment of the present invention.
[0033] Figure 2 A flowchart of the use provided by the embodiment of the present invention. DETAILED DESCRIPTION
[0034] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0035] See also Figure 1 As shown, an intelligent labeling method based on instance segmentation provided by an embodiment of the present invention mainly includes:
[0036] Step 1: Create a dataset for training the instance segmentation model;
[0037] Step 2: training an instance segmentation model as the basis for intelligent annotation;
[0038] Step 3: Based on the trained instance segmentation model, perform instance segmentation on the targets that need intelligent labeling;
[0039] Step 4: Convert the instance segmentation result into an intelligent labeling result;
[0040] Step 5: Fine-tune the results of intelligent labeling.
[0041] The present invention provides an intelligent labeling method based on instance segmentation. By training a well-trained model, the system can quickly identify specific features in the data and automatically generate high-quality labels. Intelligent labeling can not only reduce the workload of manual labeling, but also greatly improve the labeling speed. It has a wide range of applications, especially in the fields of autonomous driving, medical imaging, e-commerce and video surveillance.
[0042] In some embodiments, in step 1, a large number of high-quality image data sets for training the instance segmentation model are collected through data acquisition equipment. These data sets should cover a variety of scenes, including different lighting conditions, angles, background complexity, etc., to ensure the generalization ability of the model. Data collection is carried out in a variety of ways, such as using professional cameras, drones, surveillance cameras and other equipment to shoot at multiple angles and in multiple environments. The collected images are screened to remove low-quality images that are blurred, out of focus or have too much noise to ensure the accuracy and representativeness of the model training data. In this process, the resolution of the image should be kept high enough to capture the details of the target in the subsequent annotation process.
[0043] After data collection is completed, the data set will be preprocessed. The preprocessing steps include image size standardization, grayscale adjustment, denoising, etc., to ensure that the data format input to the model is unified and meets the standards. In addition, in order to further enhance the diversity of the data set, data enhancement techniques will be applied, such as random rotation, translation, scaling, cropping, and color transformation. These operations increase the diversity of the data, improve the robustness of the model in practical applications, and avoid overfitting the model to specific image features during training.
[0044] In some embodiments, step 1 further includes labeling the data in the dataset using LabelMe.
[0045] The data labeling process is also crucial. For instance segmentation tasks, it is necessary to accurately label the target objects in each image at the pixel level. The labeling tool used in this embodiment is LabelMe, which supports the use of polygon tools to outline the contours of each target object and generates corresponding labeling files (such as JSON format or COCO format). Each labeling file includes the category label of the target object and the corresponding bounding box, segmentation mask and other detailed information. In this process, it is crucial to maintain the consistency and accuracy of the labeling. In order to ensure the quality of the data set, further manual review is performed after the labeling is completed to confirm whether the labeling accurately covers the boundaries and details of the target.
[0046] In some embodiments, in step 2, the FastSAM model is used as the instance segmentation model.
[0047] The FastSAM model is used in this embodiment. It is a lightweight and efficient instance segmentation model, which is particularly suitable for scenarios with high real-time requirements and limited computing resources, such as embedded systems and edge computing devices. The model structure of FastSAM has been optimized and can significantly improve the inference speed while maintaining good segmentation accuracy. The training process first needs to use a carefully annotated data set to ensure that the image and the corresponding segmentation mask data provide the model with sufficient information for learning the segmentation boundaries of the target.
[0048] In the training phase of the present invention, the model architecture of FastSAM will be initialized. Defining the loss function and optimizer is the key first step. The loss function uses cross entropy loss or Dice loss to measure the difference between the segmentation mask predicted by the model and the actual segmentation mask. For instance segmentation tasks, the loss function not only needs to measure the accuracy of target classification, but also needs to pay attention to whether each pixel is correctly assigned to the target instance. The choice of optimizer may include Adam, SGD, etc., which can continuously update the model weights through back propagation and optimize the model performance.
[0049] During the training process, images and their corresponding segmentation masks are input into the model in small batches. The model performs instance segmentation on each image and predicts whether each pixel belongs to the target object. By calculating the loss function, the model adjusts its weights so that the prediction results are closer and closer to the true values. To improve the generalization ability of the model, data enhancement techniques can be used during training, such as random rotation, scaling, cropping, mirroring, etc. These techniques allow the model to better adapt to different image changes and avoid over-reliance on specific image features.
[0050] In the process of training the model, cross-validation technology is used to evaluate the performance of the model. A part of the data is selected as the validation set, and the performance of the model on the validation set is regularly evaluated. By observing the changes in the validation error of the model, hyperparameters such as the learning rate and regularization parameters are dynamically adjusted to avoid overfitting or underfitting of the model. In order to further accelerate the training process and improve the performance of the model, a transfer learning strategy of the pre-trained model is adopted. By using weights pre-trained on large-scale data sets, the model can converge more quickly and achieve better results on small data sets. After training is completed, the model will be saved as a weight file and can be used for inference and deployment on the target device.
[0051] As a lightweight model, FastSAM can be easily deployed on devices with limited computing resources, such as mobile devices or edge computing nodes. In order to meet the needs of real-time processing, the model is optimized for inference so that it can generate accurate segmentation masks with low latency, ensuring its application in scenarios such as autonomous driving and intelligent monitoring.
[0052] After completing the training and development of the instance segmentation model for intelligent labeling, the model is deployed in the intelligent labeling platform as the basis of the intelligent labeling system. The overall process of users using the intelligent labeling system for labeling is as follows: Figure 2As shown, first of all, it is necessary to clearly identify and confirm the annotated objects in the image. The instance segmentation model trained by the present invention provides four flexible and powerful segmentation methods, namely full segmentation, point segmentation, frame segmentation and text segmentation. These methods can adapt to the annotation requirements in different scenarios and greatly improve the efficiency and accuracy of annotation. First, full segmentation is suitable for scenes where all objects in the entire picture need to be fully annotated. The system will perform instance segmentation on the entire image, automatically identify each target object in the picture, and generate a corresponding segmentation mask for each target. This method does not require manual intervention by the user, and can quickly obtain the segmentation results of all targets. It is suitable for complex scenes or tasks that require global annotation; secondly, point segmentation provides more precise user control. The user specifies the target object of interest by clicking on the picture. The specific operation is that the user puts a point on the area of each object to be segmented, and puts a background point on the background area of the image. The system will intelligently segment the objects that the user wants to annotate based on the location information of these points and the context of the image. The advantage of point segmentation is that it not only quickly locks the target of user attention and avoids the system's incorrect segmentation of irrelevant areas, but also improves the flexibility of annotation through easy user interaction, which is particularly suitable for tasks that require rapid segmentation of specific objects; the frame segmentation method further simplifies the target annotation of local areas. The user only needs to draw a rectangular box on the image and select the area of interest. The system will perform instance segmentation on part of the image in the box and intelligently retain the target with the largest area in the box; this method is particularly suitable for annotating the main objects in certain specific areas of the image. It can avoid redundant data caused by full image segmentation, and quickly lock the main target in the box, reducing unnecessary operations. It is suitable for segmentation tasks with relatively concentrated targets or clear regional limitations. By providing four methods: full segmentation, point segmentation, frame segmentation and text segmentation, the FastSAM model provides users with a comprehensive segmentation solution. Whether it is necessary to quickly and comprehensively annotate, accurately control the annotation area, or rely on text descriptions for complex target recognition, users can flexibly choose the appropriate segmentation method according to actual needs, making the annotation process more efficient, convenient and highly controllable. This flexible annotation method greatly improves the user experience and ensures the adaptability and efficiency of the intelligent annotation system in various scenarios.
[0053] After the intelligent annotation is completed, the system will send the generated annotation information to the front-end annotation interface through a dedicated interface. In the annotation interface, the user can see the segmentation results of each target object, where the contour boundary of each target is displayed on the image in the form of a series of feature points. These feature points outline the precise boundary of the target in a polygonal manner. The user can intuitively view the effect of the system annotation. If the user is dissatisfied with the intelligent annotation results, he can make manual adjustments. The adjustment method includes directly dragging and moving the feature points to accurately define the specific shape of the target contour. In addition, the user can freely add or delete feature points according to needs to deal with those details that the automatic annotation fails to accurately cover. Through this operation method, the user can accurately fine-tune the target to ensure the high accuracy of the annotation results.
[0054] Each time the user adjusts a feature point, the system will immediately provide feedback and present the adjusted effect in the annotation interface, allowing the user to observe the changes in the annotation in real time. This interactive adjustment method greatly improves the user's operating experience while retaining the efficiency of intelligent annotation. The system also provides undo and redo functions, allowing users to flexibly adjust annotations without affecting the overall progress. After the adjustment is completed, the annotation results will be stored in the system for subsequent calls and further processing. Through this efficient feedback mechanism, users can quickly complete accurate annotation of images.
[0055] Finally, the results of manual adjustment and refined annotation are exported to a standard annotation format according to user needs, such as the commonly used Pascal VOC or COCO format. These annotation files contain the bounding box coordinates, category labels, and corresponding segmentation mask information of each target object. The generated annotation files can be used to train new deep learning models or as benchmark data for model evaluation. The exported annotation files can be used for multiple downstream tasks, such as image classification, target detection, instance segmentation, etc., and are widely used in application scenarios such as autonomous driving, medical image analysis, video surveillance, and e-commerce. The intelligent annotation system proposed in the present invention can significantly improve the annotation efficiency while ensuring the annotation accuracy, and provide users with flexible and intuitive adjustment methods.
[0056] The present invention significantly improves the efficiency and accuracy of instance segmentation in the intelligent annotation process through a series of innovative designs. First, by utilizing the lightweight structure and efficient reasoning ability of the FastSAM model, the system can achieve rapid target segmentation in a resource-constrained environment and adapt to a variety of real-time application scenarios. Through various methods such as full segmentation, point segmentation, frame segmentation and text segmentation, users can flexibly select segmentation methods according to the needs of different scenarios. The system can not only accurately mark the target object, but also greatly simplify the manual intervention in the annotation process. At the same time, the results after the annotation are adjusted in real time, and the user can make fine modifications to the annotation results through the interactive interface to ensure that the final annotation has a high degree of accuracy and consistency. The annotation file finally generated can also be exported to a standardized format, which is suitable for a wide range of application scenarios such as subsequent model training and data analysis. The combination of these technologies has enabled the intelligent annotation system of the present invention to achieve significant improvements in accuracy, efficiency and adaptability.
[0057] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An intelligent labeling method based on instance segmentation, characterized in that: The method comprises: Step 1: Create a dataset for training the instance segmentation model; Step 2: training an instance segmentation model as the basis for intelligent annotation; Step 3: Based on the trained instance segmentation model, perform instance segmentation on the targets that need intelligent labeling; Step 4: Convert the instance segmentation result into an intelligent labeling result; Step 5: Fine-tune the results of intelligent labeling.
2. The intelligent labeling method based on instance segmentation according to claim 1, characterized in that: In step 1, images and videos that meet the task requirements are taken by a camera, mobile phone or drone as the data set, wherein the taken images and videos cover different angles, lighting conditions and environmental backgrounds of the target object to increase the generalization ability of the model.
3. The intelligent labeling method based on instance segmentation according to claim 1, characterized in that: Step 1 also includes: preprocessing the data in the data set, which includes image size standardization, grayscale adjustment and denoising, and making the data format input to the instance segmentation model unified and in compliance with the standard.
4. The intelligent labeling method based on instance segmentation according to claim 1, characterized in that: In step 1, use LabelMe to label the data in the dataset. Use the polygon tool to label the target objects, name, edit and adjust the labeling for each target object, and automatically generate a labeling file in JSON format. Each labeling file includes the category label of the target object and the corresponding bounding box and segmentation mask information. After the labeling is completed, choose to convert the labeling file into a format that adapts to different instance segmentation model training requirements.
5. The intelligent labeling method based on instance segmentation according to claim 1, characterized in that: In step 2, the FastSAM model is used as the instance segmentation model.
6. The intelligent labeling method based on instance segmentation according to claim 5, characterized in that: The process of training the FastSAM model consists of standard deep learning steps, including: Prepare a dataset for instance segmentation tasks, which includes images and their corresponding object segmentation mask labels; Initialize the model architecture, define the loss function and optimizer to build the training process, where the loss function uses cross entropy loss or Dice loss to measure the difference between the segmentation mask predicted by the model and the actual segmentation mask; the optimizer uses Adam or SGD to continuously update the model weights through back propagation to optimize the model performance; The images and their corresponding segmentation masks are input into the model in small batches. The model performs instance segmentation on each image, predicts whether each pixel belongs to the target object, and adjusts the weights by calculating the loss function so that the prediction results gradually approach the true value. In the process of training the model, cross-validation technology is used to evaluate the performance of the model. A part of the data is selected as the validation set, and the performance of the model on the validation set is regularly evaluated. By observing the changes in the validation error of the model, the learning rate and regularization parameters are dynamically adjusted to avoid overfitting or underfitting of the model. In order to accelerate the training process and improve the performance of the model, a transfer learning strategy of the pre-trained model is adopted. By using the weights pre-trained on a large-scale data set, the model converges faster and obtains better results on a small data set. After the training is completed, the model is deployed on the target device.
7. The intelligent labeling method based on instance segmentation according to claim 6, characterized in that: The method also includes: in order to improve the generalization ability of the model, using data enhancement technology during the training process, so that the model can adapt to different image changes and avoid over-reliance on specific image features.
8. The intelligent labeling method based on instance segmentation according to claim 1, characterized in that: The instance segmentation in step 3 includes four methods: full segmentation, point segmentation, box segmentation, and text segmentation; Among them: full segmentation is to perform instance segmentation on the entire image; point segmentation is to mark the objects to be segmented in the form of points on the image, and mark the background that does not need to be segmented with background points, so that the algorithm can intelligently segment the objects to be segmented; frame segmentation is to mark a rectangular frame on the image, and the algorithm segments part of the graphics in the frame, and only the target with the largest area in the frame is retained; text segmentation is to describe the objects to be segmented in the form of text descriptions, identify the targets in the image that are closest to the corresponding descriptions through the algorithm, and segment the corresponding targets.
9. The intelligent labeling method based on instance segmentation according to claim 1, characterized in that: In the instance segmentation task, the instance segmentation result obtained is a segmentation mask for each detected object. The segmentation mask is a pixel-level binary image that indicates the position of the target object in the image. In step 4, converting the instance segmentation result into an intelligent annotation result includes: converting the segmentation mask obtained by the instance segmentation into an annotation file in VOC format, specifically including: Step 401, generating a bounding box and a corresponding class label for each object, and each segmentation mask determines the bounding box of the object by finding the circumscribed rectangle of the target area; Step 402, in accordance with the requirements of the VOC format, organize the information obtained in step 401 into an XML file to describe the label, bounding box coordinates, and related information of each object; the segmentation mask itself is saved in a separate PNG file as a segmented image, and the PNG file path is referenced in the XML file, thereby converting the instance segmentation mask predicted by the model into a standard VOC format annotation file.
10. The intelligent labeling method based on instance segmentation according to claim 1, characterized in that: Step 5 includes: after the intelligent annotation is completed, the generated annotation file information is sent to the front-end annotation interface, and the annotation information is intuitively presented on the contour of the target object in the form of feature points, and the target boundary to be annotated is outlined by the feature points; Drag and move the corresponding feature points to precisely adjust the contour shape of the target; Add or delete feature points as needed to achieve the required annotation accuracy; Every adjustment made during the annotation process will be reflected in real time on the interface, allowing users to intuitively evaluate the effects of the modifications.