Image labeling method, computer device and storage medium
Through the combination of image segmentation model and neural radiation field technology, the autonomous driving system is fused with multiple perspective angles and timing, solving the problem of high cost and long-term pixel-level annotation, and achieving more efficient and accurate annotation results.
Patent Information
- Application Number
- PCT/CN2024/126354
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-22
- Publication Date
- 2025-05-08
AI Technical Summary
In the autonomous driving system, the pixel-level labeling of the image segmentation model is expensive and time-consuming, and cannot support large-scale data volume, which limits the improvement of the accuracy of the perception function.
The image segmentation model is used to segment multi-view images, and the neural radiation field is used to perform multi-view fusion and timing fusion, correct single-view and single-frame errors, and generate more accurate labeling results.
It reduces the cost and time of labeling, improves the labeling accuracy and the accuracy of the perception function of the autonomous driving system, and supports the labeling of large-scale data volumes.
Smart Images

Figure CN2024126354_08052025_PF_FP_ABST
Abstract
Description
Image annotation method, computer device and storage medium
[0001] This application claims priority to Chinese patent application No. 202311422619.7, filed on October 31, 2023, entitled “Image Annotation Method, Computer Device and Storage Medium,” the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of automatic data annotation, and more specifically to an image annotation method, a computer device for implementing the method, and a computer storage medium for implementing the method. Background Art
[0003] Object perception in autonomous driving systems relies on model iteration, which is fundamentally based on large amounts of labeled data. Image segmentation models, crucial perception and recognition models in autonomous driving, differ from detection and classification models in that they require pixel-level annotation. Pixel-level annotation requires extremely high accuracy, requiring perfect fit of object edges. This results in extremely high segmentation and annotation costs for a single image, and a significant time commitment. Due to its high complexity and cost, this type of pixel-level annotation cannot support large amounts of data. Consequently, the perception capabilities of autonomous driving systems are limited by data volume, hindering further accuracy improvements.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field.
[0005] Summary of the Invention
[0006] In order to solve or at least alleviate one or more of the above problems, the following technical solutions are provided. Embodiments of the present application provide an image annotation method, a computer device for implementing the method, and a computer storage medium for implementing the method, which can reduce annotation costs and improve annotation efficiency, thereby improving the iteration efficiency of the autonomous driving model.
[0007] According to a first aspect of the present application, an image annotation method is provided, which includes the following steps: A. using an image segmentation model to segment an original image under multiple perspectives to output a segmentation mask map of the original image under each perspective, wherein the segmentation mask map contains annotation information indicating the contour of the target object; B. using a neural radiation field to perform multi-perspective fusion on the segmentation mask maps under each perspective to correct single-perspective errors and generate a corrected first mask map; and C. performing multi-frame temporal fusion on the first mask map according to a temporal relationship to correct single-frame errors and generate a second mask map from a bird's-eye view perspective.
[0008] As an alternative or supplement to the above solution, in an image annotation method according to an embodiment of the present application, the method further includes: performing frame extraction processing on the video data captured by the multi-view video capture unit, and generating a series of original images with timestamps.
[0009] As an alternative or supplement to the above scheme, in an image annotation method according to an embodiment of the present application, step A includes: performing pixel-level recognition on the original image using the image segmentation model, wherein each pixel in the original image is labeled with category information; and drawing the outline of the target object based on the labeled category information.
[0010] As an alternative or supplement to the above scheme, in an image annotation method according to an embodiment of the present application, the image segmentation model is constructed based on a training data set including sample images and mask information of the sample images, the mask information is generated based on the target recognition result of the target object in the sample image, and the image segmentation model includes one or more of the following items: an instance segmentation model for rigid objects, a semantic segmentation model for non-rigid objects, and a panoramic segmentation model for rigid objects and non-rigid objects.
[0011] As an alternative or supplement to the above scheme, in an image annotation method according to an embodiment of the present application, step B includes: using the neural radiation field to perform three-dimensional reconstruction of the segmentation mask map to project the two-dimensional mask in the segmentation mask map under multiple perspectives to a three-dimensional mask grid; and back-projecting the reconstructed three-dimensional mask map to the two-dimensional mask grid to generate a corrected first mask map under multiple perspectives.
[0012] As an alternative or supplement to the above scheme, in an image annotation method according to an embodiment of the present application, step C includes: for the first mask image under multiple perspectives with the same timestamp, using affine transformation to project it to the bird's-eye view perspective in the world coordinate system to generate a third mask image under the bird's-eye view perspective at each moment; and performing multi-frame temporal fusion of the third mask image at each moment according to the timing relationship to generate the second mask image, wherein the size of the second mask image is larger than the size of the third mask image.
[0013] As an alternative or supplement to the above scheme, in an image annotation method according to an embodiment of the present application, the method further includes: calculating the correction error of each target object during the multi-frame temporal fusion process; and back-projecting the second mask image to a specific perspective to generate a fourth mask image at the specific perspective.
[0014] As an alternative or supplement to the above scheme, in an image annotation method according to an embodiment of the present application, the method further includes: if the correction errors of each target object are less than or equal to the first threshold, outputting the fourth mask image as the final annotation result, wherein the specific perspective includes multiple original perspectives of the original image.
[0015] As an alternative or supplement to the above scheme, in an image annotation method according to an embodiment of the present application, the method further includes: if the corrected error of the first target object is greater than the first threshold, performing single target aggregation on the first target object in the fourth mask image to repair the error of the first target object in the fourth mask image, and generating a corrected fifth mask image, wherein the specific perspective includes one or more of the multiple original perspectives of the original image; and using the fifth mask image to update the segmentation mask image, and re-executing steps B and C.
[0016] As an alternative or supplement to the above scheme, in an image annotation method according to an embodiment of the present application, calculating the correction error of the target object includes calculating the pixel difference of the target object in adjacent frames, and / or the single target aggregation is performed based on an interactive segmentation model.
[0017] According to a second aspect of the present application, a computer device is provided, comprising: a memory; a processor; and a computer program stored on the memory and executable on the processor, wherein the execution of the computer program enables any one of the image annotation methods described in the first aspect of the present application to be executed.
[0018] According to a third aspect of the present application, a computer storage medium is provided, wherein the computer storage medium includes instructions, and the instructions, when run, execute any one of the image annotation methods according to the first aspect of the present application.
[0019] The image annotation scheme according to one or more embodiments of the present application is implemented based on image segmentation and temporal dynamic repair technology. This scheme introduces the temporal characteristics in autonomous driving. It not only has the ability to annotate a single frame, but also can use the neural radiation field to correct the single-view error of the image, and automatically correct the single-frame error of the image according to the temporal relationship between the front and back. Compared with the existing image segmentation and annotation methods, this scheme can reduce the annotation errors caused by problems such as jitter, holes, and errors in single-frame images, thereby improving the annotation accuracy and improving the perception function accuracy of the autonomous driving system. In addition, this scheme does not require manual annotation, which greatly reduces the annotation cost and reduces invalid redundant annotations. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and / or other aspects and advantages of the present application will become clearer and easier to understand through the following description of various aspects in conjunction with the accompanying drawings, in which the same or similar elements are represented by the same reference numerals. In the drawings:
[0021] FIG1 is a schematic flowchart of an image annotation method 10 according to one or more embodiments of the present application;
[0022] FIG2 is a schematic block diagram of a computer device 20 according to one or more embodiments of the present application. DETAILED DESCRIPTION
[0023] The description of the following specific embodiments is merely exemplary in nature and is not intended to limit the disclosed technology or the application and use of the disclosed technology. In addition, there is no intention to be bound by any express or implied theory presented in the foregoing technical field, background technology or the following specific embodiments.
[0024] In the following detailed description of the embodiments, numerous specific details are set forth to provide a more thorough understanding of the disclosed technology. However, it will be apparent to one of ordinary skill in the art that the disclosed technology can be practiced without these specific details. In other instances, well-known features are not described in detail to avoid unnecessarily complicating the description.
[0025] Terms such as "comprising" and "including" indicate that in addition to the units and steps directly and explicitly stated in the specification, the technical solution of the present application does not exclude the presence of other units and steps that are not directly or explicitly stated. Terms such as "first" and "second" do not indicate the order of units in terms of time, space, size, etc., but are only used to distinguish between the units. The technology of the present application is generally used in electric vehicles, including but not limited to pure electric vehicles (BEVs), hybrid electric vehicles (HEVs), fuel cell vehicles (FCEVs), etc.
[0026] Some image segmentation and manual annotation methods primarily achieve image annotation by drawing a circumscribed polygon around the target object. With this annotation method, the circumscribed polygon required for a single target object typically consists of 10-20 points, and the time consumed ranges from approximately 10-30 seconds. Therefore, the image annotation time for a complex scene with hundreds of target objects often exceeds 30 minutes. To improve annotation efficiency, semi-automatic interactive segmentation can also be used as an annotation method. Semi-automatic interactive segmentation typically involves extracting object features at the mouse click location, calculating the similarity of surrounding pixels, and clustering pixels exceeding a threshold to output the object outline. By using an image algorithm for automatic feature extraction, the model can automatically identify target objects to a certain extent, significantly reducing annotation speed and cost. However, the annotation results of semi-automatic interactive segmentation are often inaccurate and still contain certain errors, necessitating manual correction to compensate. Therefore, this application proposes an image annotation solution based on image segmentation and time-series dynamic repair technology to achieve higher annotation accuracy while reducing annotation cost and time, thereby improving the perception capabilities of autonomous driving.
[0027] Hereinafter, various exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings.
[0028] Referring to the drawings below, FIG1 is a schematic flowchart of an image annotation method 10 according to one or more embodiments of the present application.
[0029] As shown in FIG. 1 , in step S110 , an image segmentation model is used to perform segmentation processing on the original image under multiple viewing angles to output a segmentation mask image of the original image under each viewing angle.
[0030] According to one or more embodiments of the present invention, multiple video capture units are used to capture information about the vehicle's surroundings from multiple perspectives and output corresponding multi-perspective video data. The multiple video capture units can be positioned at different preset locations on the vehicle, each corresponding to a specific perspective. To avoid blind spots, the fields of view of the multiple video capture units may overlap.
[0031] Exemplarily, the above-mentioned multiple video acquisition units include one or more front-view cameras and / or one or more side-view cameras, which can be installed on the high side of the roof, on the vehicle fender, or on the light strip above the rearview mirror, or at the connection between the rearview mirror and the door panel to increase visual redundancy and reduce perception blind spots. Exemplarily, before performing image segmentation processing, the video data collected by the above-mentioned multi-view video acquisition units (e.g., left front camera, front camera, right front camera) can also be subjected to frame extraction processing, and a series of original images with timestamps are generated (e.g., original images under the left front perspective, original images under the front perspective, original images under the right front perspective). In this article, cameras, lenses, video cameras, cameras, etc. all refer to devices that can obtain images or images within the coverage area. They have similar meanings and can be interchangeable. The present invention does not impose any restrictions on this.
[0032] Image segmentation involves dividing an image into several non-overlapping regions based on features such as grayscale, color, spatial texture, and geometric shape, so that these features appear consistent or similar within the same region, but distinctly different across different regions. For example, the image segmentation model can be one of SegNet, DeepLab, Mask R-CNN, U-Net, and Gated SCNN.
[0033] Optionally, the image segmentation model is constructed based on a training data set containing sample images and mask information of the sample images, wherein the mask information is generated based on the target recognition results for the target objects in the sample images. In some embodiments of the present application, in order to avoid labeling the sample images by manual labeling, one or more target detection models can be pre-trained to identify the area where the target object in the sample image is located and the category of the target object, and determine whether the pixel in the sample image is a pixel in the target object, thereby generating corresponding mask information. The image segmentation model is then trained using the training data set containing the sample image and the mask information of the sample image. In step S110, the image segmentation model trained in the above manner can be used to recognize the original image under each perspective (e.g., left front perspective, front perspective, and right front perspective), so that each pixel in the original image under each perspective is labeled with category information (e.g., vehicle, pedestrian, lane line), so as to meet the pixel-level labeling requirements. In addition, in step S110, the outline of the target object can be drawn based on the labeled category information and a segmentation mask map containing the labeling information indicating the outline of the target object can be output. For example, in the output segmentation mask image, pixels with the same category information can be marked with the same color, and pixels with different category information can be marked with different colors.
[0034] Optionally, image segmentation models with different segmentation frameworks can be used based on the characteristics of the data itself, the objects that need to be labeled, and other features. Exemplarily, the image segmentation model in step S110 can be a combination of one or more of an instance segmentation (IS) model, a semantic segmentation (SS) model, and a panoptic segmentation (PS) model. Semantic segmentation is based on pixel-level segmentation without distinguishing instances within a category separately. It is mainly applicable to non-rigid objects, such as road surfaces, lane lines, ground signs, etc. Exemplarily, all lane lines in the image can be labeled yellow and all road surfaces can be labeled black. Instance segmentation is a combination of object detection and semantic segmentation, that is, the target object is detected in the image and then a label is assigned to each pixel. Instance segmentation is usually limited by the fact that the target object must be a regular frameable feature, and is therefore suitable for segmenting rigid objects, such as vehicles, pedestrians, obstacles, traffic lights, etc. It should be noted that the semantic segmentation results do not distinguish between different instances belonging to the same category (for example, all pedestrians are marked in red), while the instance segmentation results distinguish between different instances of the same category (for example, different pedestrians are distinguished by different colors). Panoramic segmentation is a combination of instance segmentation and panoramic segmentation capabilities. It is necessary to detect all target objects and distinguish different instances in the same category to output the pixel outlines of all objects in the image. It should be noted that instance segmentation only detects and segments the target objects (for example, pedestrians) in the image by pixel, and uses different colors to distinguish different instances, while panoramic segmentation detects and segments all objects in the image (including the background), and uses different colors to distinguish different instances. The image segmentation results (i.e., segmentation masks) output by the above three image segmentation models can all be used as input for subsequent steps, where the image segmentation results are masks with the same size ratio as the original image.
[0035] Next, in step S120 , a neural radiance field (NeRF) is used to perform multi-view fusion on the segmentation mask images under each view to correct single view errors, and a corrected first mask image is generated.
[0036] NeRF is a neural network-based three-dimensional (3D) reconstruction technology. Unlike traditional 3D reconstruction methods that explicitly represent scenes as point clouds, meshes, voxels, etc., NeRF models the scene as a continuous radiation field and implicitly stores it in the neural network. Simply by inputting two-dimensional (2D) images from multiple angles, NeRF can be trained to generate NeRF, which can then be used to render clear photos from any perspective. Step S120 uses NeRF and multi-view information for 3D reconstruction, fusing image segmentation results from different viewpoints to generate high-quality new 2D views from any perspective.
[0037] Optionally, in step S120, the trained NeRF can be used to first project the 2D mask in the image segmentation result (i.e., the segmentation mask map under each viewing angle) to a three-dimensional mask grid for 3D reconstruction (for example, a 3D mask map is generated based on the 2D segmentation mask map under the left front viewing angle, the 2D segmentation mask map under the front viewing angle, and the 2D segmentation mask map under the right front viewing angle), and the reconstructed 3D mask map is back-projected to the 2D mask grid to generate a corrected first mask map under multiple viewing angles (for example, a first mask map under the left front viewing angle, a first mask map under the front viewing angle, and a first mask map under the right front viewing angle are generated, which are respectively corrected versions of the aforementioned 2D segmentation mask map under the left front viewing angle, the 2D segmentation mask map under the front viewing angle, and the 2D segmentation mask map under the right front viewing angle). It should be noted that the image segmentation results under a single perspective are prone to defects such as incomplete segmentation and category errors. Therefore, compared with the segmentation mask images under each perspective output by the image segmentation module in step S110 (for example, the 2D segmentation mask image under the left front perspective, the 2D segmentation mask image under the front perspective, and the 2D segmentation mask image under the right front perspective), the first mask image output in step S120 (for example, the first mask image under the left front perspective, the first mask image under the front perspective, and the first mask image under the right front perspective) integrates the information of each perspective, thereby correcting the defects under the single perspective through information complementarity and improving the accuracy of data labeling.
[0038] In step S130, the first mask images generated from each perspective in step S120 are fused in a time-sequential manner to correct single-frame errors, generating a second mask image from a bird's-eye view. Because some samples cannot be corrected after multi-perspective fusion, step S130 aims to further improve the problem of inaccurate details in the image segmentation results by combining perspective fusion with time-sequential fusion.
[0039] Optionally, in step S130, first, for the first mask images under multiple perspectives with the same timestamp, an affine transformation is used to project them to the bird's eye view (BEV) perspective in the world coordinate system to generate a third mask image under the BEV perspective at each moment, and then the third mask images at each moment are temporally fused in multiple frames according to the temporal relationship to generate a second mask image. For example, the first mask images under the left front perspective, the front perspective, and the right front perspective obtained based on the raw data collected by the left front camera, the front camera, and the right front camera at the same time can be spatially fused to obtain a third mask image under the BEV perspective, wherein the size of the third mask image under the BEV perspective is larger than the size of the first mask image under the original perspective. Next, according to the temporal relationship between the previous and next frames, the third mask images at multiple moments are fused to obtain a second mask image after temporal fusion, wherein the size of the second mask image is larger than the size of the third mask image. It can be understood that due to the fusion of multi-view information and multi-frame information, the second mask image has repaired the single-frame errors compared to the first mask image, thereby reducing single-frame jitter, holes and other problems, and further improving the accuracy of data labeling.
[0040] The above-mentioned image annotation method 10 utilizes a combination of image segmentation and temporal dynamic repair technology. Compared with existing segmentation schemes, it introduces temporal features in autonomous driving. It not only has the ability to annotate single frames, but can also use NeRF to correct single-view errors in images and automatically correct single-frame errors in images based on the temporal relationship between the previous and next frames. While reducing the annotation cost, it also reduces the annotation errors caused by jitter, holes, errors, and other problems in single-frame images, improves the annotation accuracy, and improves the perception function accuracy of the autonomous driving system. At the same time, the present invention recognizes that in practice, there may still be some long-term annotation errors that cannot be resolved through temporal dynamic repair. Therefore, a single-target aggregation step can be further added to perform fixed-point repair on erroneous samples.
[0041] Optionally, the method 10 may further include step S140: calculating the correction error of each target object during the multi-frame temporal fusion process, and back-projecting the second mask image to a specific viewing angle to generate a fourth mask image under the specific viewing angle.
[0042] In one example, the correction error can be determined by calculating the pixel difference of the target object in adjacent frames. In another example, the correction error can be determined by calculating the position and posture changes of the target object in the mask image before and after time series fusion, such as the translation amount, rotation angle, or scale change.
[0043] Exemplarily, if the correction errors of each target object are less than or equal to the first threshold, it is determined that there is no need to perform single target aggregation, and the second mask image is back-projected to a fourth mask image under a specific perspective as the final labeling result, where the specific perspective is multiple original perspectives of the original image (for example, the left front perspective, the front perspective, and the right front perspective).
[0044] Exemplarily, if the correction error of a first target object among multiple target objects is greater than a first threshold, it is determined that single-target aggregation needs to be performed for the first target object. For example, single-target aggregation can be performed for the first target object in the fourth mask image under a specific perspective to repair the error of the first target object in the fourth mask image and generate a corrected fifth mask image, where the specific perspective includes one or more of the multiple original perspectives of the original image. Through single-target aggregation, the annotation of the first target object under the original perspective has been corrected. To this end, the corrected fifth mask image can be used to update the segmentation mask image generated in step S110. That is, the fifth mask image is used as input in step S120, and multi-view fusion is performed on the fifth mask image, and multi-frame temporal fusion in step S130 is performed to further improve the accuracy of the annotation. Exemplarily, the above-mentioned single-target aggregation can be performed based on an interactive segmentation model. That is, through the three steps of inputting the target mask image, extracting features of the corresponding area, and finding the nearest neighbor pixels of the features, erroneous and missing pixels are further corrected and clustered, thereby obtaining a more accurately annotated image.
[0045] FIG2 is a schematic block diagram of a computer device 20 according to one or more embodiments of the present application. The computer device 20 includes a memory 210, a processor 220, and a computer program 230 stored in the memory 210 and executable on the processor 220. The execution of the computer program 230 enables the image annotation method 10 shown in FIG1 to be executed. In addition, as described above, the present application can also be implemented as a computer storage medium, in which a program for causing a computer to execute the image annotation method 10 shown in FIG1 is stored. Here, as a computer storage medium, various computer storage media such as disks (e.g., magnetic disks, optical disks, etc.), cards (e.g., memory cards, optical cards, etc.), semiconductor memories (e.g., ROMs, non-volatile memories, etc.), and tapes (e.g., magnetic tapes, cassettes, etc.) can be used.
[0046] In the applicable situation, the combination of hardware, software or hardware and software can be used to realize the various embodiments provided by the application. Moreover, in the applicable situation, without departing from the scope of the application, the various hardware components and / or software components set forth herein can be combined into a composite component comprising software, hardware and / or both. In the applicable situation, without departing from the scope of the application, the various hardware components and / or software components set forth herein can be divided into a subcomponent comprising software, hardware or both. In addition, in the applicable situation, it is contemplated that the software component can be implemented as a hardware component, and vice versa.
[0047] Software according to the present application (such as program code and / or data) can be stored on one or more computer storage media. It is also contemplated that the software identified herein can be implemented using one or more general or special computers and / or computer systems, networked and / or otherwise. Where applicable, the order of the various steps described herein can be changed, combined into composite steps and / or divided into sub-steps to provide the features described herein.
[0048] The embodiments and examples set forth herein are provided to best illustrate embodiments according to the present application and its specific applications, and thereby enable those skilled in the art to make and use the present application. However, those skilled in the art will appreciate that the above description and examples are provided for ease of illustration and example only. The descriptions set forth are not intended to be exhaustive of all aspects of the present application or to limit the present application to the precise forms disclosed.
Claims
1. An image annotation method, characterized in that: The method comprises the following steps: A. performing segmentation processing on the original image under multiple viewing angles using an image segmentation model to output a segmentation mask map of the original image under each viewing angle, wherein the segmentation mask map includes annotation information indicating the contour of the target object; B. using the neural radiation field to perform multi-view fusion on the segmentation mask images under each view to correct the single view error, and generate a corrected first mask image; and C. Perform multi-frame temporal fusion on the first mask image according to the temporal relationship to correct single-frame errors, and generate a second mask image from a bird's-eye view perspective.
2. The image annotation method according to claim 1, wherein: The method further comprises: The video data collected by the multi-view video acquisition unit is subjected to frame extraction processing, and a series of original images with time stamps are generated.
3. The image annotation method according to claim 1, wherein: Step A includes: Using the image segmentation model to perform pixel-level recognition on the original image, wherein each pixel in the original image is labeled with category information; and The contour of the target object is drawn based on the labeled category information.
4. The image annotation method according to claim 1, wherein: The image segmentation model is constructed based on a training data set including sample images and mask information of the sample images, wherein the mask information is generated based on a target recognition result for the target object in the sample image, and The image segmentation model includes one or more of the following: an instance segmentation model for rigid objects, a semantic segmentation model for non-rigid objects, and a panoramic segmentation model for rigid objects and non-rigid objects.
5. The image annotation method according to claim 1, wherein: Step B includes: Reconstructing the segmentation mask map in three dimensions using the neural radiation field to project the two-dimensional mask in the segmentation mask map under multiple viewing angles onto a three-dimensional mask grid; and The reconstructed three-dimensional mask map is back-projected onto a two-dimensional mask grid to generate a corrected first mask map under multiple views.
6. The image annotation method according to claim 1, wherein: Step C includes: For the first mask images under multiple perspectives with the same timestamp, project them to the bird's-eye view perspective in the world coordinate system using affine transformation to generate a third mask image under the bird's-eye view perspective at each moment; and The third mask images at each moment are subjected to multi-frame time-series fusion according to the time-series relationship to generate the second mask image, wherein the size of the second mask image is larger than the size of the third mask image.
7. The image annotation method according to claim 1, wherein: The method further comprises: Calculating the correction error of each target object during the multi-frame temporal fusion process; and The second mask image is back-projected to a specific viewing angle to generate a fourth mask image at the specific viewing angle.
8. The image annotation method according to claim 7, wherein: The method further comprises: If the correction errors of each target object are less than or equal to the first threshold, the fourth mask image is output as a final labeling result, wherein the specific viewing angle includes multiple original viewing angles of the original image.
9. The image annotation method according to claim 7, wherein: The method further comprises: If the corrected error of the first target object is greater than the first threshold, performing single target aggregation on the first target object in the fourth mask image to repair the error of the first target object in the fourth mask image, and generating a corrected fifth mask image, wherein the specific viewing angle includes one or more of the multiple original viewing angles of the original image; and The segmentation mask map is updated using the fifth mask map, and step B and step C are re-executed.
10. The image annotation method according to claim 9, wherein: Calculating the correction error of the target object includes calculating the pixel difference of the target object in adjacent frames, and / or The single target aggregation is performed based on an interactive segmentation model.
11. A computer device, characterized in that: The invention comprises: a memory; a processor; and a computer program stored in the memory and executable on the processor, wherein the execution of the computer program enables the image annotation method according to any one of claims 1 to 10 to be executed.
12. A computer storage medium, characterized in that: The computer storage medium comprises instructions, which, when executed, execute the image annotation method according to any one of claims 1-10.
Citation Information
Patent Citations
Image annotation method, computer device and storage medium
CN117152753B
Multi-view input aerial view semantic segmentation method and device
CN115131787A
Image segmentation method and device, electronic equipment and storage medium
CN116091768A
Automatic SUV calculation method and system based on neural network time sequence image segmentation
CN116309644A
Image labeling method, computer equipment and storage medium
CN117152753A
Cited By
Image enhancement method, system and equipment for automatic driving vision task
CN120563789A
An image enhancement method, system, and device for autonomous driving vision tasks
CN120563789B
Method, device and equipment for detecting target in open world game based on YOLO
CN120612471A
Method for evaluating orientation degree of sintered neodymium-iron-boron magnet
CN120971469A